Success With Small Segments

Time out for an update on small segments. I have a little over 100,000 Matches at AncestryDNA; and 9441 rows in my spreadsheet of Ancestry Matches with Common Ancestors – it’s roughly 9%, so far. Of the 9441, 4281 are in the 6cM to 10cM range; and another 2,434 are in the 11cM to15cM range. So… depending on how you characterize “small” segments, from 45% to 71% of my “success” has been with cousins in the small segment range.

These cousins have been “proved” by me using traditional genealogical methods – not, necessarily GPS standards, but my 50 years of genealogy experience. It matters NOT to me that the small segments may be real or poison – in each case, they led me to a cousin. And in my spreadsheet, I also note most are closely related to other Matches per Pro Tools.

If your hair is on fire about this, just pretend that I found and linked these folks (Matches) to my Tree without any DNA evidence… It’s OK, I did that for 36 years before atDNA came along.

I am *still* drinking through a firehose at Ancestry and MyHeritage with Matches of Shared Matches showing relationships to known cousins. And I am more convinced than ever that a large share of our DNA Matches (at all the companies) are true genealogy cousins within a genealogy time frame, say back into the early 1700s in Colonial America. The limiting factor is not the DNA, it is the genealogy records. Pro Tools is demonstrating that many of our Matches are closely related – they are NOT 10 to 20 (or more) generations away. Focus on the positive viewpoint, and dig in on the genealogy research!

That’s my take.,,,

[06I] Segment-ology: Success With Small Segments; by Jim Bartlett 20270719

18 thoughts on “Success With Small Segments

  1. Jim, I am very interested in your statement: “Pro Tools is demonstrating that many of our Matches are closely related – they are NOT 10 to 20 (or more) generations away.” What range of generations would you say these small matches might be? – Louis

    Liked by 1 person

    • Louis, You make a good point. Although they the Matches are closely related I don’t really know, from that, how closely related they are to me. I was thinking of the times I have a known Match – say a 4C or 5C or 8C. And the shared Match is 7cM to me, but 1,558cM/niece to my known Match – in that scenario I know exactly how distant the shared Match is to me – just as much in a genealogy timeframe as the known Match was. Does that make sense? Jim

      Like

      • In those situations, i.e. that your 7 cM match is a 1558 cM niece of a known match, what is the cM that you match with the known match? Theoretically, the niece should share on average half of what her uncle/aunt shares with you, but I know the amounts can be disproportionate to each other.

        Like

    • Louis,

      I am comfortable with 9 generations back (8C level). For me, and, I think, many others, the genealogy just gets too fuzzy and it’s hard to find more distant relatives. Historical note: Ancestry used to find and tell us Circles, out to 8C. Even at the 6C level of ThruLines, I see a big jump in the number of total Matches reported with each generation back. I would that curve would be much larger if ThruLines went out to 9 generations (8C). I just did a quick sort and analysis of my spreadsheet Ancestry Matches with Common Ancestors – 10,400 (there are some dups related to me multiple ways, but I think the numbers will tell the tale. I had one 12C (11cM); 13 11C (6-21cM) 34 10C (6-34cM); 32 9C L6-26cM); 259 8C (6-40cM); 245 7C (6-41cM); 2,258 6C (6-62cM); 510 5C (6-58cM) and tapers off down to 1C2C…There were many more in the 1R and 2R and Half categories (at my age I tend to get 1R cousins… haha). The point is the big number at 6C which reflects the power of ThruLines. I’ve been going through this spreadsheet (starting with closest Matches and woring down to more distant. At each generation I’ve been able to about double the number of cousins using ProTools. I’ve just started on the 6C level (I did Ahnentafel 128 a few weeks ago, and shifted to Ahnentafe 142 to help some cousin-friends work on a puzzle and we needed to investigate the all-female lines). In both of those I was able to “swell the ranks” with more cousin “recruits” by using ProTools and just looking at each shared match with an Unlinked Tree (which Ancestry skips). I’m drinking through a fire hose… Which prompts me to predict that there are *many* more genetic cousins in my Match list that I have uet tp discover and document. There probably are some cousins in my list in the 10 to 20 generation category, but I think there are many more than I can find who are closer. I’ll do an update on stats when I get through my 8C list…
      Now, that is my take. You dig hard into genetic genealogy, too; and I very much value your insights, input, or even a WAG. What percentage of our current Ancestry Match list would you estimate whould be within the 9 generations back (8C) range? I’m going to stick my neck out and say over 50% (which means I have about 40,000 to go…) Jim

      Like

      • Jim: Very interesting stats on your number of matches and cM range by 12C down to 5C. If you could estimate the percent of your relatives that you believe you’ve identifed at each of 12C down to 5C and divide your number of matches by that, you’ll get an estimate of the true number of matches that you have at each cousin level. Your actual number has its maximum at 6C. Will the true number have its maximum at a higher cousin level?

        Liked by 1 person

      • Louis,
        IMO, YES, I think so. That is why I’m working down my spreadsheet. My experience at the 2C to 5C levels shows an increase beyond the mostly ThruLines Matches, as I investigate with ProTools. I often find parent, child, sibling, AUNN, Matches with no/private/scrawny Trees, but I can usually easily check and add them. I’m even adding some 1C based on their shared Match lists indicating a clear grouping (I put them in my Tree with a parent labeled as sibling of the known Match parent – these real people are listed as living, so they are hidden). My 6C list will grow. My 7C list will also grow, but I’m hampered there as ThruLines has not peeked into private Trees for me. I don’t have Ancestry’s master “world tree” in the background to pull from as ThruLines does – I’ve got to build those Trees back by hand – not sure if I’ll have the time/patience for that.
        As we know, there are competing forces. The actual number of cousins we have increases a lot for each generation going back. I’ve got to believe the percent of living cousins who took a DNA test is a relatively static number, and so the number of my distant couusins who have tested also increases greatly. On the other hand the amount of shared DNA decreases by a factor of 4 with each generation going back (the simple math is 880cM for 1C; 220cM for 2C; 55cM for 3C; 16cM for 4C; 4cM for 5C; etc. By the 5C level the average is below the threshold for DNA Matching. However, the distribution curve results in a lot of true 5C having more than, say, an 8cM threshold. So the two curves are in competition: a rising number of cousins vs. a smaller amount of shared cM. My statistics show that the number of cousins wins this “battle” (with significantly large numbers) out to at least the 6C level. Beyond that, the genealogy records “factor” starts to kick in, and the true 7C and 8C get harder and harder to identify with a Common Ancestor. however, my experience with Walking the Clusters Back showed an increasing number of Clusters as the upper and lower cMs were adjusted downward. When I got down to including Matches below 20cM the number of Clusters was still growing – and I “timed out” on that project. However, the WTCB project demonstrated to me that many of those Matches were still in the 8C range.
        Estimating all of this gets mushy. However, the genealogists can easily look at the shared Matches of their low cM Matches. Most of them still have shared Matches (I’ve not tried to scroll through thousands of 8cM Matches to see what percentage have no shared Matches with me – I see some infrequently).
        Anyway – all of these factors lead me to believe we still have a lot of Matches who are true cousins in a genealogy time frame.
        And this is important. In my Matches with Common Ancestors spreadsheet – now bulked up – I can see several things: some children of my Ancestors are not included (some of them probably would be included in similar spreadsheets of close relatives); some show up with birth years clearly out of whack; some Ancestors had large families, some small (one of my Ancestors had only one child, so no Match cousins). The spreadsheet confinues to be a great Quality Control tool.

        Like

  2. Ciao Jim io su mhyeritage anche se sono sud italiano ho molte corrispondenze chi fa parte di gruppi si che sono isolate ma ho molte corrispondenze da paesi come Polonia Ucraina ex Jugoslavia e ecc insomma molte corrispondenze da paesi diciamo est Europa.

    Like

  3. /It matters NOT to me that the small segments may be real or poison – in each case, they led me to a cousin./

    I have always had this same approach. I would often wonder about the advice given in many DNA groups which basically discouraged the use of ‘poisoned’ segments despite the fact that they are actually evidence.

    Liked by 1 person

    • Bobby – Spot on! Thanks! The main point is that small DNA segments should not be used as the “only” evidence. I had a case recently where a known DNA Match had a Pro Tools match of 3,466cM – WOW – except that Match was his mother, and the known DNA Match was a 4C to me through his father. He and his mother had to be related to me a different way – which I couldn’t find. So, instead of “evidence”, I tend to think of small segments as clues – just like ThruLines (which turn out to be about 90% correct for me) The key is, always, to build the case. Jim

      Like

  4. 14 cM is my normal cutoff for Ancestry matches, but I investigate plenty of segments down to 6 cM, especially for matches which are found as part of a cluster with larger matches.

    I am missing something though. With Ancestry, you can’t plot your segments on a chromosome map, so you can’t use those segments to find a common ancestor, can you?

    Liked by 1 person

    • Greg; You are correct! However, I have over 1,000 AncestryDNA Matches for whom I know our shared segments and Triangulated Groups. At least once a week I go to GEDMatch and run One-to-Many and sort on age (days) to get the new once on top. I then focus on the ones with an Ancestry test. In most cases the name is different, but close enough to find the Match at Ancestry (I usually also send a message to them at Ancestry to confirm and also pass along what I know about the TG group (some my TGs have robust genealogy, some not). Sometimes I get a clue in the email. Sometimes I run One-to-Many on *their* kit to see if they have close kits I already know.. When all else fails, I’ll send them an email, asking for their Ancestry name (often doesn’t work). I am trying to “Note” my GEDmatch Matches with the TGID, so I can see which ones I’ve already worked out.

      Like

  5. Hi Jim,

    I just did a quick count on my Common Ancestor matches. The results were much like yours. My total group (38,000 matches) resulted in 1225 Common Ancestors. Of these, 39% scored 16 or higher; 33% scored from 11-15; and 30% scored 6-10 cM — this last group including about 100 Common Ancestor matches that I saved during the Ancestry switchover.

    So your point is well taken that lower scores really help.

    Jim Baker

    Liked by 1 person

    • Jim,
      Thanks for your feedback. And I can virtually *guarantee* you that there are many more in the small segment category. How do I know this? Two ways – one: ThruLines has found many cousins/paths that I would never have found; and each one has, on average, at least another Match through Pro Tools; and two: I’ve been walking down my spreadsheet – I’m just now using Pro Tools near the beginning of my 6C Matches (with a lot of 7C and 8C Matches in the on-deck circle) – the farther out I review, the smaller the segments will be. That percentage will go up as I press on.

      My thesis here is that too many real genealogists have been lured into genetic genealogy. That’s OK, particularly if you’ve already spent a lot of pure genealogy time finding and documenting your Ancestors and cousins. But a DNA-ONLY list of your true cousins is severely truncated – maybe all of your true 2C, but missing a larger and larger percentage of your more distant cousins. But your distant cousins who don’t test, or test and don’t match, could also have key insights to your full Tree….

      Just sayin’
      Jim

      Like

  6. Ciao Jim allora corrispondenze dna con triangolazione tra me è loro che condividono diciamo tra 8cM fino a un massimo di 23,3cM con 1 segmento 2 segmenti e uno con 3 segmenti con triangolazione del segmento tra 9cM fino ha 7,1cM e in un arco temporale genealogico? E poi il segmento triangolazione ca da 826-846 ma alcune, anno uguale fine e inizio altri inizio uguale fino ni ma alcune sono tra 819-853 cosa può significare? Loro est Europa io sud Italia.

    Like

    • Kevin, it is almost impossible to understand more from small segments beyond the fact that they Triangulate and are probably distant cousins. There is a very wide ranged of possibilities – your next steps have to be with genealogy records. Jim

      Like

      • Si ma più o meno la triangolazione massimo che generazione può arrivare? E se queste corrispondenze dna nel tg sono da paesi diversi est Europa e tra loro alcuni rom alcuni misti rom alcuni balcani ma sempre collegati alle corrispondenze rom e alle corrispondenze miste rom da dove può essere il collegamento?

        Like

  7. That’s reassuring… the power of small segment matches! I am intrigued by some ancestry that probably goes back to the 18th or 17th centuries… maybe earlier. I evidently have a Moulin ancestor who came to East Frisia as a Pastor. That was the surname of one of my Frisian great grandmothers (maternal)! Once in awhile I get what I call French French matches, often involving the X chromosome. And I also am curious about how I am related to a lot of people who evidently are NOT German like my parents, but nevertheless are related like me to people who are clearly only Jewish… either Askenazi or Sephardic. (In the latter case, the shared matches are often in or from countries like Brazil, Cuba, Mexico, Spain, Turkey, etc). The biggest relevant lists of shared matches (involving usually chromosome 9, 10 and/or 13) average about 10cM, with the “higher” matches around 15-20 cM. – I recently upgraded my membership at 23andMe and discovered that with quite little bits of DNA, am linked to about 5 Danes or “Vikings” who are identified as St Brice’s Day Massacre victims. They use very small bits of what I guess they regard as distinctive DNA. I think even under 6 cM. Of course those people would would have been killed in November 1002 CE! I am female incidentally… these Viking matches were male and related to each other I’d assume. – It’s all very interesting. Your observations suggest that all of these connections may just be legitimate after all.

    Liked by 1 person

    • H.A.K. – There was an old saying “Witches have to make friends where they can.” It’s the same for genealogists. I try to be “friends” with every clue (particulary in Trianguated Segmentsand Cluster which point toward a Common Ancestor). And, in your case, where the segment might lead to a particular part of your ancestry which is of interest to you. And to your last line – to me the shared segment may or may not be legitimate, I need to work the genealogy. Jim

      Like

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.