Unknown's avatar

About Jim Bartlett

I've been a genealogist since 1974; and started my first Y-DNA surname project in 2002. Autosomal DNA is a powerful tool, and I encourage all genealogists to take a DNA test.

A New Cluster on the Block

Featured

AncestryDNA has rolled out an “auto” Cluster program. I tried it and got 8 Clusters, ranging from 3 to 9 Matches in each one. A total of 40 of my 60 Matches above 65cM. The other 20 Matches were not included because they didn’t form a Cluster of at least 3 Matches. I know the Common Ancestors for each of the 40 Matches and the program clustered them 100% correctly. I’d give AncestryDNA an A+ for this new program. I’m impressed and anxious to have the ability to adjust the cM ranges downward to get more Clusters.

Some additional input on auto-Clustering.

It began in late 2018, with Genetic Affairs (by EJ Blom), and soon we also had Shared Clustering (by Jonathan Brecher) and DNAGedcom Client (by Rob Warthen). I tried all three. I had already done segment Triangulation on all my Matches at FamilyTreeDNA, and I worked with Johathan Brecher and we Clustered those same Matches. There was over 90% concurrence between the hundreds of Clusters and the hundreds of Triangulated Groups. Not enough to say the two processes were equivalent (they are not), but certainly this analysis showed a strong tendency of Clusters to point to a Common Ancestor between me and all the Matches in each Cluster. A very strong clue in each case.

I then Clustered all of my Matches at AncestryDNA – down to about 18cM. Many of the Clusters had a Common Ancestor consensus (easily seen in the Match Notes I had previously entered – many from ThruLines). So, I imputed that Common Ancestor to the rest of the Matches in each Cluster. I used Ahnentafel numbers to represent my Ancestors and developed a tagging code: e.g. #A0020. The #A means a confirmed Common Ancestor with a Match, and 20 is Ahnentafel for William MITCHELL 1824-1895. This code is the first thing in the Notes field. When I impute a Common Ancestor to a Match from a Cluster consensus, I use #L0020 – which means the Match is highly Likely to have that Common Ancestor with me. With a #A or a #L, I tagged almost all my Ancestry Matches over 20cM and many below that. This was in the 2019-21 time frame.

Recently, with ProTools, I’ve been able to determine how many more Matches fit into my Tree – and thus our Common Ancestor. For well over 90% of all these new Match cousins, the #L tag turned out to be correct – I only needed to change the L to A.

Bottom line 1: I am a big fan of Clustering at AncestryDNA and really look forward to expanding the coverage to more Matches.

Bottom line 2: Use ProTools with Clustered Matches to really nail down Common Ancestors to Matches.

[22DI] Segment-ology: A New Cluster on the Block by Jim Bartlett 20250725

Segment Triangulation Insight

Featured

Your DNA segments are from your Ancestors. They are adjacent to each other and fill up (or “cover” or paint) each of your Chromosomes. You have shared DNA segments with your Matches. With a browser, you can see your shared DNA on a chromosome – visually as a bar and by the start and end points in the data. Segment Triangulation lets us group overlapping segments and identify your full segment from an Ancestor. It also places each Triangulated segment where it belongs on one of your 46 chromosomes. Genealogy helps you decide if each segment is on a maternal or paternal chromosome. Once you do that, it’s then relatively easy to “fit” the Triangulated segments along each chromosome.  

Three key elements of Segment Triangulation:

1. A browser to give you the data – where is each segment on a chromosome.

2. Determine the segments are on the same chromosome (you have two of each chromosome – one maternal and one paternal). Several ways to do this…

3. Determine where one of your segments stops and another starts – i.e. the crossover points. A judgment call based on the consensus of the data.

A fourth key element is determining the MRCA for the Triangulated segment, and the path the segment took from the MRCA down a line of your Ancestors to a parent to you. This is mainly a genealogy task, working with your Matches and their Trees to build a consensus.

I hope this “insight” provides a clearer picture of what Segment Triangulation is all about and why it is a worthwhile process – for specific segments or all of your DNA.

[08F] Segment-ology: Segment Triangulation Insight by Jim Bartlett 20250525

Half-Identical Region (HIR)

Featured

Your DNA segments (that make up the 23 Chromosomes passed down to you from a parent) are not the same as shared DNA segments with a Match (as described by a chromosome browser) aka a Half Identical Region (HIR). All of your DNA is real, down to any size you want to analyze. This is not necessarily so for a shared DNA segment (or HIR)!

From the ISOGG Wiki: A half-identical region (HIR) is a region of two paired chromosomes where at least one of the two alleles from one person’s pair of chromosomes matches at least one of the two alleles from a different person’s pair of chromosomes throughout the entire region. A half-identical region may be either identical by descent (IBD) or identical by state (IBS).

In my words, for genetic genealogy, a computer compares your DNA test to a potential Match’s DNA test. The computer compares the two raw DNA data files – about 600,000 SNPs with two values (alleles) for each SNP. The two values are one from the DNA passed down from the father and one from the mother. The computer is looking for a long string of matching SNPs, which are then reported as a shared DNA segment. This meets the HIR definition above – at least one value is the same at each SNP in the shared segment. The theory is that, although much of our DNA will be the same, there is some variation, and a long enough string of matching SNPs will indicate this segment of DNA is from a Common Ancestor. This also implies that the long string is on one side – on one chromosome from our mother OR our father. A lot of reported genetic data indicates that such an HIR is true when it’s at least 15cM.

But why aren’t all shared DNA segments true? Because the computer algorithm blindly looks at *both* values at each SNP for you and the potential Match. The computer may create a string of your SNPs that agree with your potential Match’s SNPs, but some are from your father and some from your mother. Clearly this “zig-zag” result, using SNPs from both your parents’ DNA, is not a representation of your DNA on one chromosome. It’s not a DNA segment passed down from one of your parents to you. It’s a false segment! Or this might have happened with your potential Match’s data, or with both of you. Bottom line: wherever the “zig-zag” occurred, the shared DNA segment is false.

The good news is that this “zig-zag” result doesn’t occur with long enough segments – over 15cM. And it occurs very infrequently with 14cM shared DNA segments. And there is a rough distribution curve – probably different for each of us – which drops down to about half of our 7cM segments are false. And most shared DNA segments are false below 7cM – which is why they are generally not used. Some of the companies use other, proprietary, algorithms to discard (not report) some of these false Matches. Also, as I’ve blogged before, Triangulated Groups are very good at culling out the false segments.

This also ties into the ISOGG terms: Identical By Descent (IBD) and Identical By State (IBS), noted above. IBD would apply to true shared DNA segments – you and your DNA Match got the shared DNA segment from a Common Ancestor. IBS means the computer found a “match”, but IBS is usually used in genetic genealogy to indicate the false segments. I usually just stick to “true” and “false” shared DNA segments (or HIRs).

Another quirk in this discussion is using the term HIR to refer to a shared DNA segment.  This is proper and OK. But, an HIR only refers to a shared DNA segment between you and one particular Match. We virtually never find exactly the same HIR with two Matches (although it’s possible with Matches who are closely related to each other.) When we look at segment Triangulation, the Triangulated Group is comprised of different HIRs. So HIR should not be used to refer to a TG. A TG represents a segment of your DNA (from a specific Ancestor) – there are many different HIRs in a TG. And each Match in a TG would have a different (but overlapping) segment from the Common Ancestor, with different HIRs. Because the whole process is so random, we just don’t get the same segments from our Common Ancestors that our Matches get.

Bottom Line: A shared DNA segment is also an HIR – formed by a computer by comparing raw DNA test data (about 600,000 SNPs) with two values (alleles) for each SNP. Shared DNA Over 15cM all are true segments (IBD); below 15cM some are false (IBS). A shared DNA segment (aka HIR) is usually unique to a specific Match.

[22DH] Segment-ology: Half Identical Region by Jim Bartlett 20250521

HAPPY 10TH ANNIVERSARY

Featured

10 years ago, I blogged: “What is a segment?”, and noted the difference between an ancestral segment (your DNA segment) – passed down from an Ancestor to you; and a shared segment (created by a computer algorithm) which usually indicates a Common Ancestor for both you and your Match.

This is still the fundamental concept that is key to genetic genealogy.

We’ve looked at a lot of twists and turns based on this concept…

– How segments are measured

– Why the data is a little fuzzy, but that doesn’t negate its power

– How our DNA is passed down in identifiable segments from our Ancestors

– How each generation of our Ancestors contributes two full genomes (46 Chr) to us

– Why some of our segments must be sticky (persistent) for multiple generations

– How we “see” our own segments through shared segments

– How we can map (or paint) our segments on our chromosomes

– How shared segment “size” predicts relationships

– How we can group Matches by segment Triangulation or shared Match Clusters

– How we can use groups to solve brick walls, NPEs, Bio-Ancestors, unknowns

– Which ancestors always, or sometimes, or never have shared Matches

– Why all of our shared segments (6cM and up) may be important to us

– How to Walk Ancestors, Clusters, Segments back in our genealogy

– How spreadsheets can help us collect, arrange, analyze, QC, and use data

– How to use new tools: autoClustering, DNA Painter, browsers, ProTools, etc.

You have all been part of this journey of learning – as in fact, we are all learning from each other. I very much value your feedback and suggestions.

As some of you know, I also host DNA Special Interest Group (SIG), through the Washington DC Family Search Center. It was in person/local until Covid. We are now international via Zoom – 2nd Wednesday of each month 7-9pm ET. This is now an Advanced DNA SIG, and members are encouraged to participate and/or present (learn from each other). If you’d like to join, please email me at jim4bartletts@verizon.net

Happy Anniversary – your suggestions/observations/comments are “gifts” to us all.

[99F] Segment-ology: Happy 10th Anniversary by Jim Bartlett 20250507

SPECIAL ANNIVERSARY COMING UP

Featured

My first real Segmentology blog post was on 7 May 2015 – so an anniversary is coming up soon. I’m looking to consolidate and re-package the approximately 200 posts in Segmentology. If you would like any new or revised topics included, please feel free to use the comments or email me at jim4bartletts@verizon.net. NOTE: The Table of Contents (Outline in the header bar) has been updated, and all the posts are hyperlinked.

[99E] Segment-ology: Special Anniversary Coming Up by Jim Bartlett 20250422

ProTools Part 26

Documenting a GUESS

Setup… A Match, with No Family Tree, is a 1C to a Known Match per ProTools. The Known Match is in my Tree with a specific line of descent from our MRCA; and a 1C estimate is very reliable. I want to put the new Match in my Tree and place them in my Common Ancestor spreadsheet – to “take care of” that Match by placing them almost certainly where they belong in my Tree.

As I’ve blogged before, there are only two options to place a 1C to a Known Match: 1. a grandchild of the Known Match’s paternal grandparents; or 2. a grandchild of the Known Match’s maternal grandparents. In other words, the new Match is a child of a sibling of the Known Match’s father or mother. A quick review of my Shared Match list with this new Match, clearly reveals the Match is on the same side (paternal or maternal) that I am on with the Known Match. In other words, I know the path from the Known Match back to our MRCA is through their father or mother. I can now see, through ProTools,  the new Match is related to me that way, too.

So I know the path from the MRCA down to the new Match – it’s the same path that I have with the Known Match down to, and including, the grandparent of the Known Match. What I don’t know is the name of the son or daughter of that grandparent = the parent of the new Match.

Up until recently, I’ve just named that son or daughter “block” as GUESS or Unknown in my Tree and in the “cell” of my spreadsheet. I’m now up to a dozen or so of these and can see many more on the horizon. My index of people in my Tree is filling up with GUESS and Unknown people…

I see four options for a name:

1. Continue with GUESS or Unknown [I usually reserve GUESS for iffy guesses]. I don’t like this – it’s not helpful to me or others reviewing my Tree – someday it may be very confusing.

2. Child of [name the grandparent]; ex: “Child of Bob JONES”

3. Parent of [the new Match]; ex: “Parent of Horatio Mitchell”

4. Sibling of [name the Known Match’s parent]; ex: “Sibling of Martha SMITH”

The Tree “box” and spreadsheet “cell” would have these entries and appear very close to other, known, boxes and cells. They would also be more specific in the Tree index, instead of a generic “GUESS” or “Unknown”.

I think I like (4) Sibling of Known Match’s parent the best because it specifically precludes the Known Match’s parent. In fact, I just did one new Match who was 1C to two different Matches so the description was: [sibling of John and Mary SURNAME] to rule them both out [after checking with ProTools].

I am interested in feedback on this topic – i.e. how to efficiently document Matches which clearly fit in a specific Tree branch. I am experimenting with 1C1R and even some 2C which clearly cannot fit anywhere else. Keyword here is “efficiently” – there is a LOT to do, and I don’t want to have to write a paragraph about each one. This is primarily for my own research. If I leave them as alive, no one else will see them; and if I mark them as deceased, the only people who will care will be close relatives to the new Match, and they may provide some feedback to me. I hope so…

[22DH] Segment-ology: ProTools 26 – Documenting a GUESS by Jim Bartlett 20250302

MITx Class on DNA is Free

Featured

MITx offers a wide range of free, on-line, self-paced semester-long courses to anyone in the world. Coming up next week is Introduction to Biology – The Secret of Life. I’ve taken this course (actually twice). It’s taught by Professor Eric Lander – the founding director of the BROAD Institute and a principle leader of the Human Genome Project – and a fantastic instructor (his course is fun). This course is targeted at non-biology students. This is not about genealogy, it’s about DNA. Anecdote:  I was about halfway through the course, and one night my wife called out: “Jim, what are you doing – it’s 3 AM.” My reply: “I’m in a lab, folding proteins to capture a virus”.  If you are into DNA and Segment-ology, this is a great opportunity to get a firm grounding.  As a side note, I think MITx is a great undertaking and am a regular donor to that program. Free, world-wide MIT classes…

Here is a link: https://www.edx.org/learn/biology/massachusetts-institute-of-technology-introduction-to-biology-the-secret-of-life

Click on the short YouTube video… Enjoy.

[99D] Segment-ology: MITx Class on DNA is Free by Jim Bartlett 20250128

ProTools Part 25

Featured

The Path Is Key

This may be an extension of my “genealogy sacrilege” outlook or rant.

But before I begin, to each their own – you get to choose your objectives.

My two main objectives are to get my genealogy right; and to get the Chromosome Map of segments from my Ancestors at each generation right. My objectives do not include finding all of the descendants of all of my Ancestors. However, I do think that documenting how my DNA Matches interrelate to me and each other is very helpful in achieving my two objectives – and this swells my Tree somewhat. I’m finding: Match paper trail paths (and ThruLines clues) that are impossible, given the DNA evidence; and DNA evidence that has revealed genealogy paths I never would have otherwise found (not just limited to breaking through brick walls).

So, a lot of work to do to document what will be over 10,000 Matches…  Time is precious…

When documenting DNA Matches and their line of descent from our MRCA to them, the “Path Is  Key”. Dotting all of the “i”s and crossing all the “t”s is NOT! The DNA segments do not “know” their hosts’ names (or dates, or places), just that the segments are passed along. We genealogists document what we can about each of these Match ancestor DNA hosts. It helps us keep track – in time and place. But how much effort do we need to put into documenting our Matches’ lines? My opinion is: not much! We need to be sure of the path. We don’t need to know the full names, or pet names, or titles. It’s nice to know the birth/death years, but how much digging should we do to find the complete birth date or place? What do we do when several different descendants insist on different given names … I could go on and on, but I’ve decided it’s not my job to adjudicate their family “wars” – my objective is to be clear of the path.

Therefore, I’m now using terms like Pvt, Unknown, GUESS, sibling of XYZ, etc. to describe Match Ancestors – particularly those close to the Match.I don’t really care about their parent’s or grandparent’s names or genealogy info – just the path that must exist for a DNA segment. [NB: proving a specific genealogy-DNA link is a separate issue; a potential path is not a proven path.]

I am still documenting the child and grandchild of the MRCA (given name and birth year at least). But, IMO, the further down the path from the MRCA to the Match, the less precise this info needs to be. The Key Is the Path. I don’t want to introduce incorrect info, so I’m introducing “other” terms in the name field when it is unclear, in debate, or might take days to research and resolve. I note the “path” that has to be and move on.This allows me to get as many DNA Matches as possible into the spreadsheet. Then the interrelationships can be better evaluated.

SUMMARY:  Don’t worry about “fully” documenting the MRCA-to-Match path; just that the path does exist, and no incorrect info is introduced (unless your Tree is private). And, of course, it’s up to your own judgment as to if/how much of this recommendation to follow. My plan is to get as many Matches as possible into MRCA family groups in a spreadsheet, and then study the interrelationships with ProTools. Get Matches in my Tree and my Common Ancestor spreadsheet, but “do no harm”.

[22DG] Segment-ology: ProTools 25 – The Path Is Key by Jim Bartlett 20250222

ProTools Part 24

Featured

Small Segment Stats

Ancestry DNA Matches who share 6-7cM and have a known MRCA with me: 1,160.

Total Ancestry DNA Matches at any cM level: 7450.

About 15% of my DNA Matches with a known MRCA share only 6-7cM.

This is NOT a statement linking DNA and Ancestors.

This IS a statement about the many true cousins we will not see in our Match lists because the current threshold at AncestryDNA is 8cM.

I’m glad I Dotted and saved some of my 6-7cM Matches when Ancestry made the threshold change – it was a fraction of the total. I wish I’d have saved them all…

To end on a higher note – I still have 2,600 other 6-7cM Matches to work with – many of them are being determined as close cousins to known MRCA Matches by using ProTools.

[22DF] Segment-ology: ProTools Part 24 – Small Segment Stats by Jim Bartlett 20250221

ProTools Part 23

Featured

Integrating With Genealogy

ProTools is a powerful tool. But it has it’s limits. 1C and closer relationships are very accurate, in my experience. Beyond that, the range of possibilities grows quickly as the cMs fall below the 1C range. But think about what that means… A 1C relationship takes us back to our grandparent level. Think of a 20 year old genealogist with a 50 year old parent, and 80 year old grandparents. Those grandparents would be in the 1950 census. And the census is a pretty good tool back to 1850 – another few generations. You might argue that the census is not rock solid in every case. There may be adoptions, NPEs, etc. That is true, but those individuals will not show up as DNA Matches – for the most part.

Yes, there are still a few situations that may slip through. But on the plus side, the census and ProTools will sort out a high percentage of false relationships, and/or incorrect genealogy “research”.

Used together, the census and ProTools can pretty accurately cover the past 175 years.

[22DE] Segment-ology: ProTools 23 – Integrating With Genealogy by Jim Bartlett 20250131