Showing posts with label Research Methodology. Show all posts
Showing posts with label Research Methodology. Show all posts

Saturday, 28 January 2023

Background to using DNA

 As a trained scientist I have no difficulty understanding the basics of the science and mathematical statistics underlying genetic genealogy, but I also do not believe in re-inventing the wheel, so I started looking for websites which demonstrated techniques to use. This blog post is a list of the sites I have found most useful and why. This is not intended to be a one-time blog post, but hopefully will grow as I come across more useful sites.


Diahan Southard's "Your DNA Guide"

 This really set me on the right track in organising my DNA matches, grouping them by "Most Recent Common Ancestor(s)" (MRCA). Yes, I bought the book. If you want to and don't want a solid copy I would suggest purchasing the PDF version from the web site. In particular I would suggest reading the following blog posts on the site:

  • DNA Triangulation: explains how to use shared matches (in Ancestry and other sites) to identify MRCAs without ploughing down to segment level,
  • What is a Genetic Network: explains the concepts underlying DNA Triangulation

DNA Painter

 Online visualisation tools for getting your head around your DNA matches (particularly for a very visual person like me). It has tools to help people with recent "non-parental events" locate themselves within a potential family tree and for those people who want to dig down to segment level in their analysis. The section I use most often is:

  • The Shared cM Tool: based on the work of a group of genetic statisticians, this allows to isolate the possible relationships a certain level of common cMs can cover.

The Genetic Genealogist

 One of the pioneers of genetic genealogy (and yes I've purchase one of his books), his blog ranges widely over the uses of DNA in genealogy. He was one of the founders of:

  • The Shared cM Project: a statistical analysis of relationships and the variations in shared cM which can occur. This is the data underlying "The Shared cM Tool" (above)

The Leeds Method

 A method for sorting out your high cM matches into genetic networks, and seeing if you have any recent pedigree collapse. Plenty of explanatory blog posts to help you.


DNA Explained

 I've only just come across this blog via one of the best coverages of ThruLinesTM I've seen:


 That's all for now, folks. Good luck with your DNA matches.

Tuesday, 9 August 2022

When Ancestry ThruLines™ gets it Wrong

 ThruLines is a marvellous tool for sorting out your more distant DNA connections, but it depends on the majority of people having accurate trees on ancestry (see "Getting the best out of an autosomal DNA test"). When there are a substantial number of inaccurate trees for a family ThruLines can be led astray.

In January 2021 I found a AncestryDNA© Match with a suggested ThruLines link1:

You and [DNA Match]
< 1% shared DNA | 6 cM across 1 segments
Unweighted shared DNA: 6 cM
Longest segment: 6 cM

Shortly after this Matches with such low cM values were cut from the results unless you had marked or noted them in some way. It was also relatively early days for ThruLines.

Figure 1 shows the relationship suggested by ThruLines with John Wickham (1737-1825) and Mary Baldwyn (1736-1825) as our Most Recent Common Ancestors (MRCAs).



Figure 1: The genetic link as proposed by ThruLines.

The late marriage is not unusual, but for a premarital child of a couple to keep their mother’s surname is. It usually indicates that the child is not the offspring of the groom. Also the middle name ‘Field’ is unexpected if the father is John Wickham.

Looking at Poor Law records:

East Sussex Bastardy Orders: Mayfield:
“1829 Dec 4 Maintenance order on Joseph FIELD, labourer tp for a bastard son of Phoebe DANN born 2 Jan 1828 (Par.422/34/2/127)
[Phoebe DANN married 27 June 1835 at Mayfield, John WICKHAM, widower].
Alfred Field son of Phoebe DANN, spinster, baptised 8 February 1829.”2

So according to Phoebe, Alfred Field Dann is her son by Joseph Field, which at least explains the middle name. But where does the Wickham link come from? Is it via Joseph Field or Phoebe Dann? It turns out there is a link via Phoebe herself3.

Phoebe Dann is the daughter of Thomas Dann and Jane Hobbs. Jane Hobbs is the daughter of William Hobbs and Elizabeth WICKHAM. Elizabeth Wickham is the daughter of Richard Wickham and Ann Colchin. She is also the sister of John Wickham, father of my 3xgreat-grandmother, Ruth Wickham, and grandfather of Phoebe's husband John Wickham, who is thus a second cousin of Phoebe Dann as well as her husband. So the MRCAs for my DNA Match and myself are (in the absence of any other link) Richard Wickham and Ann Colchin, and the DNA match and I are seventh cousins. Figure 2 shows the actual relationship.



Figure 2: An actual genetic link

I now have this information in my Ancestry tree, but the DNA Match does not appear as a ThruLine any more as it is eight generations back. I don’t mind, I value accuracy above ease.

This sort of multiple relationship can make sorting out DNA links complicated. Because of Phoebe’s Wickham ancestry, it would be easy to assume that all her children were fathered by her eventual husband, John Wickham. However any descendant of Phoebe Dann would show a Wickham link, as shown by the DNA Match of mine descended from Alfred Dann. Statistically though descendants of Phoebe by John Wickham would have the chance of a double dose of Wickham/Colchin DNA and average twice the amount of shared DNA compared to the descendants of Phoebe alone. The suggested link still appears on ThruLines for other DNA matches of mine who are descendants of Alfred Field Dann, hence the use of this example in a blog post.

Multiple marriages can also cause confusion, particularly for husbands who favour a particular given name in their wives. Another DNA match of mine traces their ancestry back to Jane Knight (1824-1904, daughter of Richard and Elizabeth Knight)1. In their tree they identify Elizabeth Knight with Elizabeth Rolph (1810-1855, married Richard Knight in 1830) and our MRCAs to be William Rolph (1787-1860) and Sarah Borders (1789-1869). (This match is part of a large genetic network of Rolph/Borders descendants.) However this identification would make Elizabeth Rolph only 14 years old when Jane was born (and occurs six years before the Knight-Rolph marriage). Of course the answer is that Richard Knight had a previous marriage (to Elizabeth Stapleton (1799-1828) and Jane is the child of this marriage. Jane cannot then carry the Rolph/Borders genes, so where is the link?

Looking at the family tree we find that Jane Knight's son, Daniel Filler (1851-1928) married Phoebe Philpott (1853-1902). Phoebe Philpott is the daughter of Eliza Knight (1832-1914), the granddaughter of Elizabeth Rolph (1810-1855) and thus half-cousin of Daniel Filler. So it is Eliza Knight (half sister of Jane Knight) who carries the Rolph/Borders genes and is the actual link back to William Rolph (1787-1860) and Sarah Borders (1789-1869)3.



Figure 3: How incorrect trees can confuse ThruLines

Figure 3 shows the situation with the red dashed lines showing the genetic link as currently suggested by ThruLines and the green lines showing the actual genetic link, confirming the Rolph-Borders marriage as the MRCAs. Of course there is also a possibility of other, as yet undiscovered, links particularly in the sort of small, country villages inhabited by both the above examples.

The moral of this post is to always, ALWAYS double check anything in an online tree or suggested tree. Check for supporting documentation, check for reasonableness, check for alternative possibilities.

Sources

  1. Ancestry ThruLines for Susan Law (https://www.ancestry.com.au/discoveryui-geneticfamily/thrulines/080AB020-82F7-4145-BC9C-3BF177830102?filterBy=all).
  2. Burchall, Michael J. East Sussex Bastardy Papers 1594-1845 (CD ROM). Edited by The Parish Register Transcription Society. Lewes, UK: Sussex Family History Group, 2009.
  3. Susan Law's Ancestry Tree (https://www.ancestry.com.au/family-tree/tree/52067358/family?cfpid=13296141854). Private tree, access by request.

Tuesday, 24 August 2021

Getting the best out of an autosomal DNA test

This blog post is based on my experience using AncestryDNA® for family history research.  AncestryDNA is an autosomal DNA test which (subject to statistical variation) is a “broad spectrum” test covering all antecedent lines.1 Much of what I say will, however, be applicable to any autosomal DNA test system, but Ancestry does have the biggest sample set of tests.

Once your DNA sample has been analysed Ancestry provides three sets of results:

  1. An Ethnicity Estimate,
  2. A list of all your DNA Matches with shared DNA equal to 8 cM or greater,
  3. ThruLinesTM, a suggested line of connection, based on analysis of Ancestry family trees (private and otherwise) if one can be found.

Further background on the AncestryDNA testing and analysis process is available on the Ancestry web site.2

1. Ethnicity

Unless you are primarily interested in anthopology, the Ethnicity Estimate doesn’t help much with family history. Due to the random nature of DNA inheritance, the whole of a particular ancestor’s DNA can have disappeared from your chromosome set. I have 4% European Jewish DNA. A remote cousin via my Jewish ancestors has none.

2. DNA matches

When I first did my test this was the only result useful for family history. Even so the usefulness depended on how much work the matches had put in their end and how accurate that work was (and whether, like me, they prefer to keep their tree private). I have three first cousins in the list whom I know and who were already in my tree. Two of them have online trees, helpful, but mainly containing information I already knew.

The next three with 93-109cM match values have no trees and minimal information on their profiles. By looking at the shared matches I can assign them to either my mother’s or my father’s side and define a rough branch. Two of them have surnames which do not appear in my tree. I don’t know where to start looking for them. The third has a family surname, but I still can’t track him down.

No tree, no personal info, no use (even with high level of DNA match).

The next match (91cM) had a family surname and a VERY small tree, most of which were living people, but she had linked her DNA to it and the one named person was in my tree. I was able to investigate her line and add it to my tree.

Even a small tree is useful as long as it is accurate.

A match at 58cM had a private tree. Even though it is linked, the only way I can directly make use of it is by contacting the owner and requesting access. I don’t blame them. When I first subscribed to Ancestry (15 years ago) I uploaded my careful research to a public tree and found large chunks of it taken and attached to totally unrelated trees. So I made my tree private and have happily shared it with anyone who requests access and can prove how they are related to me. This process can be tedious and this is where ThruLines comes in (see below).

Be prepared to respond to tree access requests in a reasonable and timely fashion if you want to keep your tree private.

A match at 24cM was linked to a public tree with 13,000+ people. Dead cert you might think? There were a lot of common names, but no matching people and no obvious common ancestor(s). The shared matches indicate my mother’s side, but that is as far as I can go.

Accuracy matters more than size in a tree.

3. ThruLines

Added in the last year or so, ThruLines is where Ancestry’s size comes in. Despite all the people who don’t bother with a tree, there are enough trees on Ancestry (with or without DNA links) for someone to have some of your lines in their tree somewhere – as long as what you have in your tree is accurate and you link yourself to it. And this applies to private trees too. You are no longer totally cutting someone off from your tree if you keep it private.

My 91cM match has a relationship suggested in ThruLines, because the linked tree is accurate, even though she only has 8 people in her tree.

My 58cM match with a private tree has a relationship suggested in ThruLines even though her tree is private. I can fill in my relationship with her without accessing information she prefers to keep private.

My 24cM match with a 13,000+ tree does not have a relationship suggested in ThruLines. Considering how many other relationships ThruLines suggests for my tree, I conclude that there is an error in the other tree.

To get the best out of ThruLines you need to link your DNA result to an accurate tree.

My experience with ThruLines is that it is ~75% reliable. It is dependent upon people having accurate trees. For most of the other 25%, you can work out what has gone wrong and still make use of the ThruLines information, but I have had one case where the suggested relationship was rubbish, even to the point of suggesting a relationship on my Father’s side when all the shared matches were on my Mother’s.

With ThruLines, don’t just copy, CHECK carefully.

The final thing I have to say is: what a waste it is to do a DNA test and then do nothing with it. Just a teensy tree with your result linked to it can help. Alternatively, find a relative you trust and let them manage your DNA result and link it to their tree.3

Sources

  1. International Society of Genetic Genealogy Wiki, ‘Autosomal DNA’, https://isogg.org/wiki/Autosomal_DNA, accessed 23 Aug 2021.
  2. Ancestry Australia, ‘Ancestry DNA White Papers’ https://support.ancestry.com.au/s/article/AncestryDNA-White-Papers3, accessed 23 Aug 2021.
  3. Ancestry Australia, ‘Assigning a Manager to Your AncestryDNA® Test’, https://support.ancestry.com.au/s/article/Assigning-a-Manager-to-Your-AncestryDNA-Test, accessed 24 Aug 2021.