A Concept Project for Segmentologists

As I noted in my last post (Drinking Through A Fire Hose), I have over 10,000 DNA Matches with pretty solid genealogy paths back to our Common Ancestors. I’m about 1/3 of the way through entering and Tagging these Matches and their paths back to our CAs in my main Tree at Ancestor.

Let’s look at an example. My Ancestor [A0856] John HIGGINBOTHAM b 1695; married 1713 in Amherst Co, VA to Frances RILEY. I have identified 827 DNA Match/cousins who descend from them.

Note 1: most of my Matches do not have a Tree back to John – I and ThruLines determined most of the paths.

Note 2: search Public Member Trees for John [drumroll….]: 10,323 Trees! WOW! Take a guess at how many Ancestry members actually descend from John and Francis, but don’t have Trees that reach 9 generations back…  – 100 thousand? A million?

Note 3: take a guess at how many DNA test takers there are in addition to the 827 folks I have already documented back to John and Francis. I’m sure it’s a LOT!!

Suppose we were all working within one Tree…

Over the years there have been several attempts to establish one family tree: OneWorldTree (Ancestry); World Family Tree (Geni); World Connect (RootsTech); WikiTree; FamilySearch Family Tree. IMO, there are a lot of issues within these attempts, as individuals interpret records differently, or worse, enter names and relationships without any documentation, etc, etc. Many NPEs are never discovered…

Concept: suppose we started Tagging ourselves, our DNA Matches, the DNA Connections, and the DNA Common Ancestors in WikiTree or FamilySearch.

As I wrote about in “Advanced Genetic Genealogy, Techniques and Case Studies”, I had identified three separate Triangulated Groups from John HIGGINBOTHAM and Frances RILEY [my 7XG grandparents – Ahnentafel 856].  In other words there were three finite segments in my maternal DNA that were in each of my Ancestors going back to John or Frances; and a shared/overlapping  DNA segment [part of my segment] in each of my 827 Matches, and in their Ancestors in a path of descendants from John or Frances down to and including each Match. [Remember each Match overlapped some of my DNA segment in the full Triangulated Group.]

Note that, on average, each of my Matches probably also had about 3 segments that went back to John and Frances. And I am not the center of the universe – all the other DNA test takers have their own, independent, experience. Certainly, there were many other descendants of John and Frances who had different DNA segments.

So what benefits would accrue to this concept…

1. Our accumulated Tags would provide a consensus that the paths we shared back to an Ancestor, passed through true genetic genealogy paths. This would be evidence that was independent of the genealogy analysis.

2. Perhaps a TG segment was actually from a different Ancestor. In the grand scheme of DNA, there would be certain distinct DNA segments that would be passed down from each Ancestor to various living test takers. It seems to me that, in general, these would wind up in multiple descendants. Note: we know that a given DNA segment could possibly be from a range of Ancestors, but when a number of DNA test takers all have the same [overlapping] DNA segment, it must surely be from one Common Ancestor.

3. So, this concept would help weed out which DNA segments came from which Ancestors.

4. Also, we’d start to accumulate specific DNA segments that came from specific Ancestors. We’d have the accumulated data to “paint” some of the Ancestor’s DNA.

I’m looking for feedback on this post. Pros and cons… Additional ideas… Is this already being tried? Shouldn’t we Segmentologists be working on a way to share and combine our DNA data to benefit each other? I’m looking for better language to articulate the possibilities and/or drawbacks of this concept.

[22DL] Segment-ology: A Concept Project for Segmentologists; by Jim Bartlett 20260622

24 thoughts on “A Concept Project for Segmentologists

  1. Jim i have a question , on myheritage i have a dna match where in his familytree my ancestor appear two times in his paternal side but my ancestor appear in grandfather and grandmother side from his father , so our unique common segment size will double or not? I share with him only 1 segment

    Like

    • swagf – in theory, yes – but the DNA often does not correspond with the theory. The way atDNA goes, we are lucky to get one shared segment with a Match – many true cousins do not share any segments with us. It is not unusual to get one segment instead of two, or one segment about half the size we might expect. I’m grateful for the one… Jim

      Like

      • Thanks for the reply. I also have the feeling that our only shared segment is inflated—larger than it should be. I asked an AI if the segment could be inflated, and it said no, but I still have the feeling that it is.

        Like

  2. Jim, thank you for your post. It is a relief to me that another genealogist, working with DNA, is thinking of a way to share information. I’ve worked on DNA for the last 12 or so years, have identified the ancestor(s) for probably 60% of my segments, but have no effective means of sharing that information publicly. I have put short pieces of information on my Ancestry tree, for each ancestor or marital pair, designating the chromosome and segment, in a DNA field, along with my Gedmatch kit number, but there is no information that I can publish about the other DNA matches who share that ancestor along with their chromosome/segment information, without invading their privacy.

    None of the DNA websites have taken this obvious next step, connecting DNA segments to ancestors, so I have suggested to Gedmatch that they do so. There are so many errors on unproved family trees,, especially on Ancestry, that this step would greatly help in resolving issues, plus assist in resolving NPEs. I would love to see a means of saving this information on FamilySearch and other websites, too, as a means of proof.

    Documents can be wrong, so even the best documented family tree can be incorrect. Putting collaborative DNA on trees would be a huge step forward for genealogists.

    Liked by 1 person

    • Sandra, Tim Janzen and I discussed this years ago. I believe each of my 372 TGs represents a very specific autosomal DNA Haplotype – a very unique string of tested ACGTs that only exist in descendants of specific Ancestors (just like Y-DNA and mtDNA). My point was that these atDNA haplogoups would very hard to easily document and catalog. However, for those of us interested in this concept, I think it could be done at WikiTree or FamilySearch (or both). I can think of 3 ways to identify these segments: Chr# and string of ACGTs (usually over 1,000 long) [too much data, too hard to use]; or TGID (like 17:56-93Mbp) tied to an Ancestor (Name, dates) [also huge, in theory a file for every human who over lived]; or TGID (like 17:46-93) that is added to Ancestors (and their lines of descent to Matches). In this 3rd case we’d use an already extant, and growing, Tree of mankind; and only the folks interested in this would do it. An M or P could be added to the TGID at each Ancestor on the path – it’s not all one gender as it is for Y or mt). And the folks adding them would also include their own Wike or FS ID; and we’d try to build a consensus. It cannot be a “one and done” process – the point is we need the consensus.
      Personally, I try to avoid the word “proof” – it is such a lightning rod and comes with a lot of baggage. To my thinking a DNA icon posted with an Ancestor is NOT the answer…
      Jim

      Like

  3. Pingback: Friday’s Family History Finds | Empty Branches on the Family Tree

  4. Hello Jim, I have followed you for years. A turn of phrase I use is convergence. I have curated the DNA match lines of 88 related participants and have generated a dataset of 2,000 unique genetic lines similat to your description.

    What the Study Is Testing-The DNA Study Frame is the working model used here to evaluate network-centered autosomal lineage evidence. Instead of relying on one tester or one DNA match, the study asks whether multiple descendants, across represented branches, repeatedly point toward the same ancestral structure.

    The focus is not whether any single person carries enough DNA to answer a lineage question alone. The focus is whether the active descendant network represented in the current study dataset shows repeated and reviewable convergence.

    You are welcome to review all of my work to determine if there exists anything that might be helpful. Because of the burden of heavy data, I produced a python script engine with the help of AI. If you were interested I would try to tweak the engine within my skill level; I would gift the use of this engine to you at no cost. I run the engine on a Google Colab notebook.

    https://yatesville.net/vault_yates_v2/vault-public-facing/about.shtml

    Regards, Ron Yates

    Liked by 1 person

    • Ron, Thanks very much for your input. I haven’t coded since college over 60 years ago. I’m a hobby genealogist, and always looking for ways to make this hobby easier – DNA has provided a new dimension to this hobby. Although I have DNA for myself, father, brother, uncle, and a few close cousins, I am laser-focused on getting the most out of my one kit. I’m retired and I put a lot of hours into this hobby. I see others who manage 5 or 10 or 20 or more kits. I cannot keep up with my one kit (at 6 companies); and have not done much with the other kits. My personal goals include verifying my own Ancestors (and pushing through Brick Walls); and then developing a Chromosome Map that links my DNA back to the Ancestors who passed it down (I think such a Map will soon be and important medical tool).

      I agree with you about developing networks (a single person has a very limited view of their own Ancestors). My desire is some easy way for genetic genealogists to provide their own segment data to an existing Tree so we could all see where the accumulating data points – “convergence” may be a good term for the objective. I was hoping this could be done in a “wiki” fashion without the “burden of heavy data”.

      These days, we can look up a lot of family trees – and there are several attempts to form a One-Tree – I’m looking for a easy way for folks currently working on, say, FamilySearch or WikiTree, to add in their segment data. Then we could all see it (like the Trees we now see) and/or programmers like you, or DNAPainter or Borland or DNAGedCom could help us analyze it.

      Jim

      Like

      • Jim,

        I admire your tenacity and your commitment to the plan. Your idea of accumulating segment evidence toward convergence makes sense to me, especially if it can be done in a way that does not bury hobby genealogists under heavy data-management demands.

        As for me, I am not really a programmer. My need for technical help arrived at almost the same time consumer AI became useful, and I have become fairly adept at asking for very specific technical assistance and then shaping the results into tools that serve the genealogy problem.

        As you continue working through this, you are certainly welcome to connect with me as useful. I may not be the coder you are looking for, but I do understand the genealogy problem and the need for practical, evidence-based tools.

        Ron

        Liked by 1 person

      • Ron, Thank you very much for the kind offer. I may well take you up on it. I like your wording “technical help” and don’t “bury hobby genealogists” – there is a balancing act here. I’m looking for fairly simple things we can do to make the DNA work for us. Jim

        Like

  5. hello Jim. I am hoping a platform will introduce back propagating segments on a DB scale for 10 years.

    Perhaps 23andMe is closest with their reconstructed ancestors feature. Even though it currently only works on parent and grandparent level. And they do not provide segmental information on the reconstructed kits.

    Practically, Borland Tools is surely the closest. But it is all community based. Not platform orchestrated. I am not paying my Borland subscription most months but I understand in paid membership, there is an option of working across individual accounts. I deeply appreciate being able to reconstruct my ancestors several generations back but sadly lack access to most of my cousin raw data that would enable achieving much better results.

    On yet another note, I am now heavily using LLM agents to work on my genetic genealogical data from visual phasing at scale, through segment mapping and allocation to combining genetic data with genealogy.
    The technology we now need is now vibe-codeable so it is only a matter of combining deep expertise, entrepreneurial spirit and becoming viral enough.

    OTOH, any third party site will never have the DB size of the big 3 (4).

    I hope your health (body and mind) and sound for many more years.

    Kuba Krchak

    Liked by 1 person

    • Hi Kuba, Would you share your workflow? I’m working with EOL ancestors 6+ generations back and having to manually collect triangulating segments (MyHeritage, 23andMe) for fishing in GEDmatch, and shared matches (Ancestry) for cluster analysis. I’m a user, not a coder, but I definitely need automation. Thank you, zeppley (Gmail)

      Like

      • Hello Zeppley, the beautiful thing about LLM agents is you do not have to be a coder at all, just know what you want.

        If you have no LLM subscription, I suggest you pick Claude Code for your experiments, it is currently a little bit more advanced, but if you have GPT/Gemini subscription, just pick Codex or Antigravity.

        Just keep telling your agent(s) what you need and they will work you through it.
        Naturally, I am still using my pre-23andMe mass-login-exploit downloads from all platforms.

        I suggest you try things out, play with it. You will see today’s LLMs (Opus 4.8 / GPT 5.5 / Gemini 3.5) agents can get your asks done. Once the current ban on stronger models is solved, we will likely see Fable-strength model from all labs for even stronger capabilities.

        Good luck and have fun.

        Like

      • Why am I saying I think we need this on the main DB sites? Because doing it by hand will be a proverbial drop in the ocean and get us nowhere. This is a “big data” project by its nature.

        I can at least imagine doing Borland projects, were all people work together on a village / parish or e.g. known colonial figures.

        I cannot imagine a successful crowdsourcing of this on WikiTree level. WikiTree is already minuscule as it is compared to FS. And any hand imputed segment information will be completely insignificant.

        I can imagine AI skills deployed at Borland and DNA Painter, that would automatically (when requested by e.g. a button) populate these segments across all available generations. But e.g. in DNA Painter, working across many generations is notoriously repetitive. And in Borland, the reconstructed ancestors drop of really quickly and could cover only few first missing generations and hardly do a deep dive. Unfortunately the computation cost of this might not be coverable with existing fees.

        I believe the only realistic way is education the top DBs via questionnaires and letters.

        I am consistently writing this as a wish to any questionnaire or personal letter to Ancestry / 23andMe / MH over most of the last 10 years. But realistically, this means 2-5 suggestions of this per company from me. We need to request this as thousand(s) strong.

        Also building some version in the external space and showing the value to community and thus DB companies might be a prerequisite.

        Let’s keep pushing…

        Liked by 1 person

      • Kuba

        You definitely have a big-picture viewpoint. I love my genealogy hobby, but part of that love comes from personally putting the jigsaw puzzle pieces in. Pushing a button to arrange the puzzle for me… I guess I’d do it a few time, and then move on. I do yearn for automation to download my Match lists (and all the Notes I’ve entered) as well as my segment data for each Match. But then, I want to analyze that data. I am good for pushing a button to gather data; I’m wary about pushing a button to connect the dots – my fear of GIGO… For instance, I have a lot of my ThruLines that were wrong, but they gave me just enough information that I was able to piece together to get a “correct” lineage. I guess what I’m searching for is a way to foster collaboration among cousins… Jim

        Like

      • It’s because thinking big and thinking automation will be so much more value, than few select people doing this manually…

        Like

    • Kuba, Thanks for your feedback. My version of propagating segments back is imbedded in Triangulated Group (TG) segments. Each of my TGs is in reality one of my DNA segments from an Ancestor. Say one of my TGs is paternal Ch17: 35 to 87 Mbp – that’s a very specific location in my DNA which (behind the scene) has say 3,500 SNPs. By comparing this area of my raw data with several other DNA Matches, I could probably determine the precise string of ACGTs. And this TG segment (and those specific ACTGs) had to exist in each of my Ancestors going back to, and including, the Common Ancestor (CA). The trick is: who is the real CA? I have mostly Colonial Virginia ancestry, so the possibilities are numerous. Usually we rely on a consensus of genealogy with other Matches in the TG. However with this concept, at say WikiTree, we could track many DNA segments. In the ultimate usage, we could take an adoptee and find matching segments in Ancestors…. Jim

      Like

      • Hello Jim, I am a reader since your first blogs.

        Sorry I generalize your term of walking segments back without properly distinguishing / explaining that I meant it in the context of multiple testers (ideally on DB level), not just on personal level.

        And I understand your colonial ancestry troubles. It is similar for us with Czech ancestry. I have people like four times 4C etc. Tough to narrow the segments down when there are multiple CA pairs between people.

        Some time ago, I was publishing a small study how to narrow down which of the multiple CAs does a segment belong to in our FB group for educating Czech genetic genealogists.

        https://www.facebook.com/groups/geneticka.genealogie/posts/1607638406661087/

        Here is an LLM translation into English:

        A short case study: Using triangulation and time frames to assign segments in the traditional, mildly endogamous environment of a Czech parish, including assigning a subsegment to an older generation.

        Most of you understand that all the DNA we have can be divided into small pieces inherited from our individual ancestors. Using current autosomal methods, each of us has about 300–400 identifiable pieces of DNA that can be mapped to our ancestors.

        A necessary condition for this mapping work is comparing the family trees of the people involved, ideally as complete as possible. This typically requires high-quality processing of the relevant families in your parish or parishes.

        What are the prerequisites for this kind of exercise?

        We have two autosomal matches for whom we have been able to find a common ancestor, more often several common ancestors, whose area of origin overlaps in some way, meaning that the common ancestors come from the same parish, and who triangulate on a DNA segment that we all share together. See the illustrative image.

        We will look at a situation where, on my side (M), we have about 10 related kits: one in the oldest generation, three in the second-oldest generation, and the others supporting the validity of the segment.

        On the second side (D), we have two siblings in the oldest generation, equivalent to the second generation on my side, and one child. (D) is represented by a very experienced genealogist, who herself helped me extend some of the shared lines further back. The largest overlap between kits from (M) and (D) is 54 cM across 4 segments, the largest of which is 22 cM.

        On the third side (T), we have a father and son in the oldest and second generation. (T) has a somewhat less developed tree, which I was able to connect to common ancestors after about three hours of searching in Acta Publica, the online archive for South Moravia.

        The largest overlap between kits from (M) and (T) is 66 cM across 3 segments, the largest of which is 27 cM.

        Common representatives of all three groups, (M), (D), and (T), triangulate on an 8 cM segment on chromosome 16. In the image, you can see:

        the longer beige segment, 22 cM, shared by (M) and (D);

        the shorter red segment, 12 cM, shared by (M) and (T);

        and the rounded frame marking the triangulation, meaning an 8.5 cM segment shared by (M), (D), and (T): that is, by (M) and (D), by (M) and (T), and by (D) and (T).

        Based on this, it is biologically necessary that this shared framed segment comes from a common ancestor of all the people involved. In other words, we have found an intersection, or triangulation, in the DNA, which means that there must also be an intersection, or triangulation, in the family trees. The task now is to find it: that is, the closest common ancestral couple of all three participants.

        In this example, we can see that serious genealogists, often with reasonably developed or even extensive family trees, seriously test multiple family members. This helps us (a) place segments as far back as possible using tested living, or recently deceased, cousins and ancestors before (b) switching to tracing through genealogy and distant DNA matches.

        If you are already involved in genetic genealogy, you have surely encountered multiple relationships in the already mentioned Czech endogamous environment. (M) and (D) have 21 identified common ancestors, based on well-developed trees. The three youngest independent couples encompass those 21 ancestors, with some deeper couples appearing multiple times because of different descendant lines and pedigree collapse.

        In my first archive session, I was able to connect (M) and (T) through one branch of common ancestors. That did not leave me satisfied, and after further effort I found two more branches of common ancestors. It is possible that there are still additional common ancestral couples, but in a quick search back to the 1680/1690 generation, the main ones appear to have been found.

        Together, (M), (D), and (T) have two independent common ancestral couples.

        From which of these independent common couples does the 8 cM segment come?

        The two common couples are chronologically placed at 1692 = 1694 and 1655 = about 1660.

        Because (M) and (D) share the longer beige segment, it is somewhat more likely that their segment comes from the same descendant, or closest common ancestral couple, and that this descendant is different from the descendant who is the ancestor of (T). This does not have to be the case; all three may have inherited the segments from three different descendants of the closest common ancestral couple, with a shorter piece surviving to the present in one case and a longer piece in another. But let us look at how, in this case, the genealogical evidence supports the most likely scenario.

        (M) and (D) really do share a common couple, 1718 = 1724, where the person born in 1718 is the son of the triangulating closest common ancestral couple, 1692 = 1694. (T) is descended from another child of the couple 1692 = 1694.

        So for our overlapping 22 cM and 8 cM segments, we have two possibilities.

        The 22 cM segment comes from the couple 1718 = 1724, and its triangulating subsegment comes from the closest common ancestral couple 1692 = 1694.

        Both segments, 22 cM and 8 cM, come from the closest common ancestral couple 1655 = about 1660.

        Of course, based on the information we have, both options are possible. But it is probably clear to everyone who has managed to follow me this far that option 1, in which the 22 cM segment is two generations younger and its subsegment one generation younger, is statistically much more likely than both segments coming from a couple two generations older.

        Based on research into the triangulating segment and the individual triangulating matches, including the identification of two different triangulating closest common ancestral couples in the family tree, their context in relation to the individual shared ancestral couples of each triangulating match, and their time frames, we were able to choose the most likely option.

        Thanks to this, we assigned the 22 cM segment to the couple 1718 = 1724, from at least five possible candidates, and assigned its 8 cM triangulating subsegment to the man from 1718 and his parents, 1692 = 1694, the closer of the two closest-common-ancestral-couple options for the shared segment.

        I hope you were able to follow this typical case study of work in a Czech endogamous parish, and that it will help you make similar use of your DNA matches and map your segments to specific ancestors among multiple cousins.

        I look forward to your questions and comments. Based on them, I will try to improve the original text before possibly posting it on gengen.cz.

        Liked by 1 person

      • Kuba – a great analysis – you are truly a Segmentologist! I think I followed it all the way – and agree.
        In general, the shared DNA is redeuced by about 1/4 for each generation (1C average 880cM; 2C average 220cM; 3C 55cM; etc. So the reverse engineering of this is that given a TG and two possibilites for the CA source, the best option is the closest CA.
        Also, I quibble – just a little – over a TG being the “framed segment”. In my thinking, one person is the *base* (usually yourself), and all of the shared DNA segments basically “paint” the *base’s* segment. So the TG is the maximum painted area of the shared segment. Each of the other Matches in the TG could do their own Triangulation (or you could do it at GEDmatch) and they would almost always get their own TG which is different from your TG – sometimes only slightly, sometimes significantly (always looking out for a large TG segment representing a closer relationship that may span two, or more, of your TGs.
        With my Concept we’d be able to collect all the TGs from multiple Matches and paint a larger segment in our CA.

        I am envious of your facebook group that works together – we need more sharing like that. Jim

        Like

  6. Si se ho 18 corrispondenze dna su mhyeritage che sta crescendo solo pochi tra loro sono cugini tra 1e terzo grado e tutti si sovrappongono stesso segmento che va da 7,1-9cM da quante figli del antenato possono discendere?

    Liked by 1 person

    • Kevin, The number of children of an Ancestor who descend to a living person today can vary a LOT. Clearly you descend from one of the children; and each of your Matches descends from a different child. But in most cases most of the Matches descend from only one or two children of the Ancestor – they are more closely related to each other than they are to you. So I would look to their family groups. It is rare to find Matches who share the same DNA segment, to descend from more than two children. Jim

      Like

  7. Great idea. This is what genealogy today is about.

    I like Wikitree – they do their best to encourage you to include sources, AND there is provision for adding DNA data. If you don’t explain the connection, you get a nice message asking you to do so. I have added all my known ancestors with sources and basic DNA match information. 

    Frankly, I think Familysearch’s world tree is a mess. It is rare for me to look at someone on it and not find a hodgepodge of information. 

    I will have a go at adding segment information to my Wikitree ancestors, but not everyone uploads to gedmatch or myheritage. We have to start somewhere. 

    Liked by 1 person

    • Jean,
      I agree with you. I’m not LDS, but I’ve been with the local Family History (now Search) Center since the mid 1990s. I love what they do, and support them all I can, but I, too, have trouble with their Tree. I subscribe to WikiTree and do a little more there since they added some DNA tweeks. But as I thought about the power of many people adding their DNA segment data, the light bulb came on… We need something quick and easy. One thing WikiTree might do to help us is to create a side-bar box for DNA segments – available for each person in their Tree. We can help by coming up with a standard way to add segment data each of us, individually can add in. Perhaps Ch17: 47-73 Jim Bartlett (or my WikiTree Code). As the data grows (in the box), each of us can see a consensus (perhaps), as well as the cousins we should have… Jim

      Liked by 1 person

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.