<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">An updated dataset of early SARS-CoV-2 diversity supports a wildlife market origin</title>
				<funder ref="#_846bB26 #_fmRG7ea">
					<orgName type="full">unknown</orgName>
				</funder>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct status="extracted">
					<analytic>
						<author role="corresp">
							<persName><forename type="first">Zach</forename><surname>Hensel</surname></persName>
							<email>zach.hensel@itqb.unl.pt</email>
							<affiliation key="aff0">
								<orgName type="institution" key="instit1">ITQB NOVA</orgName>
								<orgName type="institution" key="instit2">Universidade NOVA de Lisboa</orgName>
								<address>
									<addrLine>Av. da República</addrLine>
									<postCode>2780-157</postCode>
									<settlement>Oeiras Lisbon</settlement>
									<country key="PT">Portugal</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Florence</forename><surname>Débarre</surname></persName>
							<affiliation key="aff1">
								<orgName type="department">Institut d&apos;Écologie et des Sciences de l&apos;Environnement (IEES-Paris</orgName>
								<orgName type="laboratory">UMR 7618)</orgName>
								<orgName type="institution" key="instit1">CNRS</orgName>
								<orgName type="institution" key="instit2">Sorbonne Université</orgName>
								<orgName type="institution" key="instit3">UPEC, IRD</orgName>
								<orgName type="institution" key="instit4">INRAE</orgName>
								<address>
									<settlement>Paris</settlement>
									<country key="FR">France</country>
								</address>
							</affiliation>
						</author>
						<title level="a" type="main">An updated dataset of early SARS-CoV-2 diversity supports a wildlife market origin</title>
					</analytic>
					<monogr>
						<imprint>
							<date/>
						</imprint>
					</monogr>
					<idno type="MD5">F619F6E8ED2AA9834FFAD528FDEC2E7A</idno>
					<idno type="DOI">10.1101/2025.04.05.647275</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.9.0" ident="GROBID" when="2026-07-12T15:53+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<label type="revision">0.9.0</label>
					<label type="parameters">startPage=-1, endPage=-1, consolidateCitations=0, consolidateHeader=0, consolidateFunders=0, includeRawAffiliations=false, includeRawCitations=true, includeRawCopyrights=false, generateTeiIds=false, generateTeiCoordinates=[], flavor=null</label>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>The origin of SARS-CoV-2 has been intensely scrutinized, and epidemiological and genomic evidence has consistently pointed to Wuhan's Huanan Seafood Wholesale Market as the epicenter of the COVID-19 pandemic. Early cases were associated with this market, and environmental sequencing placed the common ancestor of SARS-CoV-2 genomic diversity within the market. Phylogenetic analysis also suggested separate introductions of lineages A and B into the human population, a finding that can be tested with additional data. Here, we curated an expanded sequence dataset of early SARS-CoV-2 viral genomes, including newly available sequences from mid-January 2020. In this dataset, we found no additional support for previously proposed alternative progenitor sequences, or for any evolutionary intermediates between lineages A and B in the human population. Instead, we identified SARS-CoV-2 lineages that may have spread from the market, and additional samples of a sublineage of lineage A with three mutations, including one found in closely related bat coronaviruses. Although our analysis of early pandemic genomes suggests that this mutation is unlikely to characterize the immediate SARS-CoV-2 ancestor, it is more plausible than two previously proposed ancestral genomes.</p><p>These findings reinforce the proposed emergence of SARS-CoV-2 from the wildlife trade at the Huanan market, demonstrating how new data continues to both solidify and clarify our understanding of how the pandemic began.</p><p>An updated dataset of early SARS-CoV-2 diversity supports a wildlife market origin Z. Hensel &amp; F. Débarre Phylogenetic analysis and epidemic modeling suggested that the two clades of SARS-CoV-2 in the early pandemic (lineage A and lineage B, separated by two mutations) were founded by at least two introductions from an unsampled reservoir in proximal animal hosts <ref type="bibr" target="#b12">[13]</ref>. Identification of samples with transitional sequences in mid-January 2020 or earlier could challenge this interpretation of the data <ref type="bibr" target="#b13">[14]</ref>. However, most analyses to date have been limited to sequence datasets with few samples until mid-January 2020. Furthermore, the earliest samples are dominated by those collected from market-linked patients, or from a familial cluster in Guangdong province and linked to a Wuhan hospital <ref type="bibr" target="#b14">[15]</ref>.</p><p>Some have predicted that additional data could support an alternative scenario, according to which the pandemic arose from a single SARS-CoV-2 introduction carrying one additional mutation in addition to the two characterizing lineage A <ref type="bibr" target="#b15">[16,</ref><ref type="bibr" target="#b16">17]</ref>. In such a scenario, the likelihood that the pandemic originated in the market is reduced, because the inferred date of the primary infection would be earlier (further from the earliest known market-linked infections), and because the proposed ancestral mutations have not been reported in market-linked sequences. Conversely, additional data could further support the market as the pandemic epicenter if mutations found in market-linked sequences are also identified in sequences without a known link to the market. Analysis of early-pandemic sequencing data to date has largely focused on genome sequences available in the genetic sequence repository GISAID, in addition to sequences detailed in the 2021 joint WHO-China study on the origins of <ref type="bibr" target="#b17">18]</ref>. Motivated by the recent availability of early-pandemic sequences collected in Shanghai and Anyang, China <ref type="bibr" target="#b18">[19,</ref><ref type="bibr" target="#b19">20]</ref>, we searched for additional sequences not included in previously published phylogenetic analyses, including from raw sequencing data that we assembled. We included these new sequences in a curated and updated early-pandemic sequence dataset, from which we excluded sequences identified as duplicates or whose coverage was low. Importantly, new sequences in our dataset were collected in mid-January 2020, a period with few sequences in previous analyses. We did not find significant support for either of two previously proposed progenitor sequences <ref type="bibr" target="#b15">[16,</ref><ref type="bibr" target="#b16">17]</ref>, and we found no transitional sequences between lineage A and lineage B. Instead, we found significant support both inside and outside of Wuhan for a lineage associated with a market-linked cluster.</p><p>We also found evidence supporting earlier emergence of a clade separated from lineage A by three mutations, one of which is found in closely related bat coronavirus genomes. Empirical analysis of SARS-CoV-2 phylogenetic trees showed that a similarly divergent clade is an unlikely, but plausible, coincidence. A divergent clade carrying a potentially ancestral mutation could emerge from a spillover from an unsampled population in a proximal host. However, Bayesian phylodynamic analysis using a multi-type birth-death model indicated that the ancestral mutation in this clade is more likely to be derived from lineage A than present in the sequence of the proximal SARS-CoV-2 ancestor. Remaining uncertainty might be resolved through the acquisition, publication, and analysis of additional sequences from early-pandemic samples, painting a clearer picture of SARS-CoV-2 emergence.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Results</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Additional early-pandemic sequences do not increase support for alternative ancestral sequences or a single introduction</head><p>To construct our dataset, we started with a recently reported dataset of 863 sequences corresponding to samples collected between 24-Dec-2019 and 15-Feb-2020 [14]. In it, we identified 41 sequences to remove: 22 based on low reported coverage <ref type="bibr" target="#b20">[21]</ref> and 19 identified as duplicates based on matching sample metadata in public databases and corresponding</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Introduction</head><p>The origin of SARS-CoV-2 has been the subject of intense scrutiny and speculation since the virus was first identified in Wuhan, China in late 2019. The earliest reported COVID-19 cases were associated with Wuhan's Huanan Seafood Wholesale Market <ref type="bibr" target="#b0">[1,</ref><ref type="bibr" target="#b1">2]</ref>, immediately suggesting that SARS-CoV-2 emerged from its mammalian wildlife trade <ref type="bibr" target="#b2">[3,</ref><ref type="bibr" target="#b3">4]</ref>.</p><p>Subsequent evidence has consistently supported this hypothesis. Spatiotemporal epidemiological data unambiguously distinguished the area near the market as the epicenter of the outbreak <ref type="bibr" target="#b4">[5]</ref><ref type="bibr" target="#b5">[6]</ref><ref type="bibr" target="#b6">[7]</ref><ref type="bibr" target="#b7">[8]</ref>. Sequencing of environmental samples placed all known genomic diversity of SARS-CoV-2 at the end of 2019 in a small section of the market <ref type="bibr" target="#b8">[9]</ref>. A stone's throw away within the market, a stall was found to contain nucleic acid from both SARS-CoV-2 and animals capable of SARS-CoV-2 transmission <ref type="bibr" target="#b9">[10]</ref><ref type="bibr" target="#b10">[11]</ref><ref type="bibr" target="#b11">[12]</ref>.</p><p>An updated dataset of early SARS-CoV-2 diversity supports a wildlife market origin Z. Hensel &amp; F. Débarre papers <ref type="bibr" target="#b21">[22]</ref><ref type="bibr" target="#b22">[23]</ref><ref type="bibr" target="#b23">[24]</ref><ref type="bibr" target="#b24">[25]</ref>. We added 187 sequences from public databases and new assemblies from public sequencing reads. We identified published studies corresponding to many of the added sequences <ref type="bibr" target="#b25">[26]</ref><ref type="bibr" target="#b26">[27]</ref><ref type="bibr" target="#b27">[28]</ref><ref type="bibr" target="#b28">[29]</ref><ref type="bibr" target="#b29">[30]</ref><ref type="bibr" target="#b30">[31]</ref>. Some sequences from Wuhan were published at the same time and in the same BioProject as sequences described in a genomic epidemiological study of Anyang <ref type="bibr" target="#b19">[20]</ref>. Additional sequences from Beijing were added after deduplication based upon sample metadata in papers with overlapping groups of patients <ref type="bibr" target="#b31">[32]</ref><ref type="bibr" target="#b32">[33]</ref><ref type="bibr" target="#b33">[34]</ref><ref type="bibr" target="#b34">[35]</ref><ref type="bibr" target="#b35">[36]</ref>. Details on databases where data can be found and corresponding accession numbers are available in Supplemental Data 1.</p><p>The final dataset is composed of 1,009 sequences. Figure <ref type="figure" target="#fig_0">1</ref> shows their temporal distribution compared to the initial dataset (not including duplicates and low coverage sequences). Newly added data follows a similar distribution as the initial dataset starting in late January 2020. However, additional sequences collected until mid-January substantially add to the number of early samples, especially the number of samples without either a known market link or a link to a familial cluster in Guangdong province <ref type="bibr" target="#b14">[15]</ref>. The earliest added sequence, collected on 7-Jan-2020 (YS8011), has a reported market link <ref type="bibr" target="#b25">[26]</ref>. For a first comparison of the composition of the initial and updated datasets, we compared the fractions of sequences in lineage A (sequences with mutations C8782T and T28144C compared to the SARS-CoV-2 reference sequence). This addresses the possibility that lineage A was under-ascertained in early-pandemic sequence data, since market-linked patient sequences were in lineage B. Figure <ref type="figure" target="#fig_1">2</ref> shows that added data have an excess of lineage A sequences relative to the initial An updated dataset of early SARS-CoV-2 diversity supports a wildlife market origin Z. Hensel &amp; F. Débarre dataset starting in late January. However, this difference is driven by intensive sequencing in Anyang to characterize transmission chains <ref type="bibr" target="#b19">[20]</ref>. The frequencies of lineage A sequences are plainly indistinguishable after excluding Anyang-focused sampling (52 sequences collected between 24-Jan and 14-Feb-2020), suggesting that additional data has a similar overall composition to the initial dataset. In order to see whether additional data identifies any SARS-CoV-2 lineages as underrepresented in the initial dataset, we examined the distribution of newly added sequences in the context of a phylogenetic reconstruction of all 1,009 sequences (Figure <ref type="figure" target="#fig_3">3A</ref>). Added sequences are distributed throughout lineage A and lineage B. We focused on lineage A sequences since the most recent common ancestor (MRCA) of the pandemic is likely in lineage A <ref type="bibr" target="#b9">[10,</ref><ref type="bibr" target="#b12">13]</ref>. Two groups of added sequences in lineage A in late January and early February consist of samples from clusters in Anyang <ref type="bibr" target="#b19">[20]</ref>.</p><p>We first looked for potential intermediate sequences between lineage A and B (separated by the mutations C8782T and T28144C relative to the reference sequence Wuhan-Hu-1). We identified one sample, Guangzhou/ID098, collected on 31-Jan-2020, carrying C8782T and additional mutations G5062T, A9707G, and C29303T. However, this is not a sample from a transitional lineage, since it shares the mutation C29303T with other sequences in lineage A, including some collected in Guangzhou. We identified another sample, Guangzhou/ID106, collected on 31-Jan-2020, carrying T28144C. However, it similarly is not a transitional sequence since it shares C10604T, A15647G, and G29868A with other sequences in lineage B. Thus, this analysis of additional data, like previous analyses, found that intermediate sequences are either artifactual or derived from lineage A or B <ref type="bibr" target="#b12">[13,</ref><ref type="bibr" target="#b13">14]</ref>.</p><p>An updated dataset of early SARS-CoV-2 diversity supports a wildlife market origin Z. Hensel &amp; F. Débarre Next, we noted that previously proposed ancestral sequences in lineage A with mutations C18060T <ref type="bibr" target="#b16">[17]</ref> or C29095T <ref type="bibr" target="#b15">[16]</ref> were found in newly added sequences. These mutations are found in RaTG13 <ref type="bibr" target="#b36">[37]</ref> and other closely related bat coronaviruses. It is likely that the recent ancestors of SARS-CoV-2 prior to the first human infection harbored one of these or another similarly "ancestral" mutation. However, phylodynamics analyses so far have found that these genotypes were unlikely to represent the MRCA of SARS-CoV-2 <ref type="bibr" target="#b9">[10,</ref><ref type="bibr" target="#b12">13]</ref>.</p><p>We examined the updated dataset to see if it was likely to increase support for C29095T (10 new sequences; earliest 10-Jan-2020) or C18060T (7 new sequences; earliest 23-Jan-2020) as being ancestral. We also investigated an early pandemic clade that was previously highlighted as bearing a reversion, C24034T, to bat coronaviruses RaTG13, RpYN06, and RmYN02 <ref type="bibr" target="#b15">[16]</ref>. Although there were 13 newly added samples in the C24034T clade (earliest 10-Jan-2020), 9 of these An updated dataset of early SARS-CoV-2 diversity supports a wildlife market origin Z. Hensel &amp; F. Débarre were sampled in Anyang and share a lineage-defining mutation, T490A <ref type="bibr" target="#b19">[20]</ref>. Figure <ref type="figure" target="#fig_3">3B</ref> shows that these mutations are unique in characterizing relatively large clades in lineage A with potential ancestral mutations. Other reversions to sequences in related animal coronavirus mutations are also shown on Figure <ref type="figure" target="#fig_3">3B</ref>, annotated for reversions to the inferred ancestral sequence RecCA <ref type="bibr" target="#b12">[13]</ref>.</p><p>For each of these mutations, additional data did not significantly change the overall composition of our dataset (Figure <ref type="figure" target="#fig_4">4</ref>). Thus, additional data reinforces previous analyses finding it unlikely that the first human SARS-CoV-2 infection had C18060T or C29095T <ref type="bibr" target="#b9">[10,</ref><ref type="bibr" target="#b12">13]</ref>. The mutation C24034T was not in the inferred ancestral sequence RecCA used in those analyses. The C24034T clade stands out as bearing three mutations (Figure <ref type="figure" target="#fig_3">3B</ref>), and in the updated dataset its earliest sampling date moves from 15-Jan to 10-Jan-2020. The possibility that this divergent clade emerged from a spillover from an unsampled proximal host is discussed below. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Two early pandemic sublineages may have spread from the Huanan market</head><p>Sequences from market-linked patients were collected starting in late December 2019, several weeks after primary SARS-CoV-2 infection(s) seeded the outbreak <ref type="bibr" target="#b12">[13]</ref>. If the market played a significant role in the early spread of SARS-CoV-2, mutations in market-linked patient sequences could appear in unlinked samples if transmission chains leading to them came from the market. In contrast, it would be less likely to detect market-linked mutations in unlinked patients under the hypothesis of an earlier origin elsewhere, according to which apparent market centrality reflects ascertainment bias <ref type="bibr" target="#b15">[16]</ref>. To test this, we investigated the updated dataset for evidence of spread of mutations first identified in market-linked patients.</p><p>Mutations in common between market-linked and market-unlinked patients could signify transmission chains from the Huanan market. The report of the 2021 China-WHO joint mission in Wuhan <ref type="bibr" target="#b17">[18]</ref> includes as Table <ref type="table">6</ref> </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>a list of mutations</head><p>An updated dataset of early SARS-CoV-2 diversity supports a wildlife market origin Z. Hensel &amp; F. Débarre found in SARS-CoV-2 sequences from patients with Dec-2019 onset dates and a market link (they are all in lineage B sequences). Some initially identified mutations, still present in published consensus sequences (and included in our dataset), were not confirmed after reanalysis described in the WHO report. In our dataset, mutations, C6968A, T11764A, T13270C, and A24325G are found in sequences from market-linked patients with Dec-2019 onset. However, the WHO report found that C6968A and T11764A in sample WH01 could not be confirmed in reanalysis. In our dataset, T13270C only appears elsewhere in a clade unrelated to the market-linked sequence and sharing additional mutations C3885A, C12778T, and G29449T. Another market-linked mutation that could not be confirmed by reanalysis in the WHO report and is not in any market-linked sequence in our dataset, C28253T, is found in six sequences, but three of these sequences were clearly derived from other lineages without any market association, reflecting C28253T homoplasy.</p><p>Since C28253T was identified as an intrahost variant in a market-linked sample <ref type="bibr" target="#b37">[38]</ref>, we concluded that it was unlikely that C28253T sequences in our dataset reflect spread from the Huanan market.</p><p>In contrast, the market-linked mutation A24325G appears eight times in our dataset, including five times in newly added sequences. Furthermore, the A24325G mutation was the only mutation we could identify shared by a cluster of market-linked patients in December 2019 (Methods), so we investigated whether additional sequencing data supported its spread from the market. Most previous analyses considered datasets in which there was only one A24325G sequence collected prior to 15-Feb-2020 with a market link: a sequence collected in California, USA on 12-Feb-2020 (CA-CDC-8).</p><p>The recent publication of sequencing from Shanghai <ref type="bibr" target="#b18">[19]</ref> added another collected on 6-Feb-2020 (SH-P243-2-Shanghai), and A24325G was also detected in a partial sequence collected in Wuhan on 30-Jan-2020 <ref type="bibr" target="#b23">[24,</ref><ref type="bibr" target="#b38">39]</ref>. Newly added data considered in this study includes two A24325G sequences on 10-Jan-2020 and three others starting in late January. The earliest full-length sequence with A24325G without a known market link is now 27 days earlier, increasing the likelihood that these sequences reflect spread of this sublineage from the market.</p><p>Another market-linked patient was not included in the WHO report Table <ref type="table">6</ref> because symptom onset occurred after Dec-2020. This patient visited clinics for treatment in Wuhan between 1-Jan and 4-Jan-2020, but was only later diagnosed with COVID-19 after returning to his hometown of Jingzhou <ref type="bibr" target="#b39">[40]</ref>. The patient's sequence (HBCDC-HB-01/2020) includes the C26370T mutation <ref type="bibr" target="#b40">[41]</ref>, which we found in one newly added sequence collected on 10-Jan-2020 in Wuhan and four sequences collected in China and Thailand starting in late January. Also of note, given the absence of reported market-linked patient sequences in lineage A, this study reported that another sequence in lineage A (HBCDC-HB-03/2020) came from a patient residing about 2 km from Huanan market, making it similar to two other early sequences in lineage A <ref type="bibr" target="#b1">[2]</ref>. This is important because it adds additional evidence that early transmission of lineage A occurred near the Huanan market <ref type="bibr" target="#b4">[5]</ref>. Together, newly added sequences in our dataset increase the likelihood that SARS-CoV-2 lineages carrying A24325G or C26370T spread from the market, further supporting the Huanan market as the pandemic epicenter.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>A divergent early-pandemic sublineage carries a potentially ancestral mutation</head><p>We asked whether new data would support any additional, distinguishable introductions of SARS-CoV-2 (i.e. with sequences other than lineage A or lineage B). As noted above, a previous analysis had annotated an early pandemic lineage harboring one mutation, C24034T, towards the bat coronavirus RaTG13 as well as two other mutations, with no transitional samples <ref type="bibr" target="#b15">[16]</ref>. Multiple spillovers need not have high sequence diversity-for instance, the viral genomes RshSTT182 and RshSTT200 were collected from two different bats and differ by only 3 mutations <ref type="bibr" target="#b41">[42]</ref>. However, branches with an unusual degree of divergence could arise from additional introductions from an unsampled reservoir An updated dataset of early SARS-CoV-2 diversity supports a wildlife market origin Z. Hensel &amp; F. Débarre with some genomic diversity. Although the presence of a mutation found in the most closely related bat coronaviruses could indicate an ancestral sequence, such a reversion could also occur by chance, as has occurred throughout the pandemic <ref type="bibr" target="#b12">[13]</ref>.</p><p>Figure <ref type="figure" target="#fig_4">4</ref> shows that the earliest sample in the C24034T clade was collected on 10-Jan-2020, the same day as the earliest sample for the previously proposed progenitor C29095T and nine days earlier than the earliest sample for C18060T. This is five days earlier than the earliest C24034T sample in our initial dataset, supporting earlier emergence of this clade.</p><p>Phylogenetic reconstruction of the dataset of 1,009 samples revealed no evidence of samples with transitional sequences (lacking any of the three mutations characterizing this clade: C24034T, T26729C, or G28077C). Sample Hong_Kong/VM20001061-2 has ambiguous C24034Y (C or T) and shares C1663T and G22661T with another sequence from Hong Kong. Sample CHN/AY508 lacks G28077C, but shares T490A with 10 other sequences from China, South Korea, and the USA. Additionally, the C24034T clade was found in two samples in Wuhan in partial sequences from late January <ref type="bibr" target="#b38">[39]</ref>. Thus, the only two samples lacking all three mutations were clearly not transitional sequences. This identifies the C24034T clade as a potential introduction carrying an ancestral mutation, which we now test.</p><p>Figure <ref type="figure" target="#fig_5">5</ref> shows a phylogenetic reconstruction of closely related human, bat, and pangolin coronaviruses in the region surrounding the C24034T mutation. Although C24034T was not included in the inferred ancestral sequence RecCA in previous work <ref type="bibr" target="#b12">[13]</ref>, it subsequently was found in closely related bat coronaviruses <ref type="bibr" target="#b42">[43,</ref><ref type="bibr" target="#b43">44]</ref> and it now appears likely that a recent ancestor of SARS-CoV-2 carried C24034T. However, it is challenging to infer the recombinant history of this region of the genome; position 24034 falls in a variable and frequently recombinant genomic region within the sequence coding for the spike fusion peptide, following the S2′ proteolytic cleavage site. Although the C24034T clade stands out as a divergent clade within the early-pandemic phylogenetic tree (Figure <ref type="figure" target="#fig_3">3B</ref>), we first empirically characterized the rarity of a similar topology in SARS-CoV-2 phylogenetic trees. We noted that there were three potential reversions (C18060T, C24034T, and C29095T) that gave rise to small clades, and we estimated the likelihood that one of three would form a similarly divergent polytomy by chance in the GISAID Global Phylogeny (12.6 million sequences). We found that 1.83% of clades similarly sized or larger (4,728 of 169,599) were separated by three or more mutations from their parent node. Accounting for the fact that any of the three common early-pandemic reversions could have been in a divergent clade by chance, this results in a 5.49% chance that one clade out of three would be this divergent from lineage A.</p><p>A similar analysis can be applied to the early-pandemic split between lineages A and B, which were separated by two mutations and formed large, similarly sized polytomies. This tree topology appeared in 3.1% of simulated epidemics <ref type="bibr" target="#b12">[13]</ref>.</p><p>In the GISAID phylogeny, 2.99% of polytomies with at least 20 children had a similarly sized child node (size ratio no more than three) with two additional mutations (2,626 of 87,691). These empirical analyses confirm the rarity of these features of the SARS-CoV-2 phylogenetic tree under the hypothesis of a single spillover; they motivate considering an alternate model of multiple spillovers from a diverse proximal reservoir.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Phylodynamic analysis does not support any alternative ancestral genotype</head><p>Having shown that the divergence of the C24034T clade from lineage A was unlikely to happen by chance in the observed human SARS-CoV-2 tree, we next investigated the hypothesis that this clade was separately introduced from an unsampled proximal host. We also aimed to compare the relative likelihoods of the proposed ancestral sequences. We inferred parameters and phylogenetic structure in a multi-type birth-death model recently applied to MERS sequencing data <ref type="bibr" target="#b44">[45]</ref>. In this model, infections propagate through competing processes of reproduction (infection of one host by another of the same type, or by a different type in spillover events) and removal (end of infectious period). We aimed for model simplicity given phylogenetic uncertainty, examining sequences collected prior to 23-Jan-2020, a period in which the sample size exhibits exponential growth (Methods).</p><p>We posited the existence of a sample from the proximal reservoir of human SARS-CoV-2 infections with a date of 15-Nov-2019 (a plausible date of primary human infection(s) <ref type="bibr" target="#b12">[13]</ref>) and a lineage A sequence with unknown nucleotides (N) at positions 18060, 24034, and 29095. This choice essentially neglected scenarios in which the pandemic MRCA is lineage B or one of two possible transitional sequences, because we focused on the relative likelihoods of lineage A introduction or one of the three genotypes with one of C18060T, C24034T, or C29095T. This model is agnostic to the nature of the proximal reservoir other than being approximated by a birth-death process and plausibly being sampled on 15-Nov-2019. Exploration of prior distributions indicated that such a model was sensitive to assumptions about spillover rates (Figures <ref type="figure" target="#fig_0">S1</ref> and <ref type="figure" target="#fig_1">S2</ref>), so caution is required when considering absolute likelihoods that particular clades originated from spillover events. Furthermore, our site model did not account for irreversible substitution rates, an important consideration for SARS-CoV-2 phylogenetic inference <ref type="bibr" target="#b12">[13]</ref>. However, since all substitutions of interest were C→T, we were able to compare the relative support for different ancestral sequences.</p><p>We found that Lineage B was monophyletic in 99.3% of sampled trees (N=90,001 sampled trees after discarding burn-in), and that it was introduced via one or more spillover events in 69.1% of those trees (overall, 82.6% of trees had multiple spillovers including trees with multiple spillovers in lineage A, but no lineage B spillover). Although this is consistent with a previous analysis concluding that lineage B and lineage A likely emerged from multiple introductions <ref type="bibr" target="#b12">[13]</ref>, it will be  We compared the plausibility of proposed ancestral sequences in single-or multiple-spillover scenarios. We measured the likelihood in the posterior tree distribution that a clade of interest (C18060T, C24034T, or C29095T) was evolutionarily distinct from the MRCA of all other sequences (i.e. an outgroup to the clade containing lineage B and remaining lineage A sequences, as illustrated in Figure <ref type="figure" target="#fig_7">6B</ref>). Since this does not require that the ancestral state includes this mutation, we also included a clade defined by the mutation C15480T as an estimate of how frequently these topologies occur by chance; C15480T was chosen as a well-supported clade (N=4 sequences) with a single C→T mutation in lineage A.</p><p>Results are summarized in Table <ref type="table">1</ref>. Despite the limitations of very different methods in this work, results were comparable to an analysis constrained to RecCA (which lacks C24034T) that found 0.7% and 0.3% likelihoods for C29095T and C18060T pandemic MRCA sequences, respectively <ref type="bibr" target="#b9">[10]</ref>. That analysis also found low, but non-zero likelihoods for T26729C or G28077C MRCA sequences (the two mutations occurring together with C24034T). This showed that an unusually divergent lineage modestly increased support for an ancestral genotype even without an ancestral mutation.</p><p>Clade C24034 + 2 C29095T C18060T C15480T Total outgroup likelihood 8.75% 1.63% 0.74% 0.47% Multiple spillover 7.72% 1.35% 0.63% 0.40% Single spillover 1.04% 0.28% 0.11% 0.07% Clade spillover likelihood 38.07% 20.90% 14.64% 13.06% Table <ref type="table">1</ref>. Fraction of posterior tree distribution (N=90,001 after discarding burn-in) with a tree topology in which a specified clade is evolutionarily distinct from the MRCA of all other sequences; the total outgroup likelihood is sum of likelihoods of the three topologies shown in Figure <ref type="figure" target="#fig_7">6B</ref> and can occur in trees with multiple or single spillovers. The clade spillover likelihood is the likelihood that the parent node of each clade is the proximal host type regardless of the tree topology. The C24034T clade additionally has mutations T26729C and G28077C.</p><p>An updated dataset of early SARS-CoV-2 diversity supports a wildlife market origin Z. Hensel &amp; F. Débarre Previous work discussed the likelihoods of candidate progenitors in single-introduction scenarios, focusing on the mutations C18060T and C29095T <ref type="bibr" target="#b15">[16,</ref><ref type="bibr" target="#b16">17]</ref>. However, we found that trees with topologies consistent with such alternative ancestry were no less likely to have multiple introductions than other trees (Table <ref type="table">1</ref>). The clade with C24034T was much more likely to be an outgroup to other sequences than clades with C18060T or C29095T, showing that it is a more plausible SARS-CoV-2 progenitor genotype. However, it was most likely that none of these three mutations characterized the progenitor. Furthermore, since our model did not account for the high rate of C→T mutations relative to T→C mutations, outgroup and spillover likelihoods in Table <ref type="table">1</ref> should be conservatively considered as upper limits.</p><p>Thus, it is plausible, but unlikely, that the C24034T clade was an additional spillover, or that the MRCA of human SARS-CoV-2 had C24034T. The most likely scenario was that all clades were derived from lineage A.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Discussion</head><p>We assembled an updated dataset of early-pandemic SARS-CoV-2 sequences, identifying sequences for exclusion and adding additional sequences not considered in previous analyses. This provided an opportunity to see whether or not previous conclusions were robust to adding a substantial amount of data-approximately doubling the number of sequences collected by mid-January 2020 without a reported link to the Huanan market or an early familial cluster. We found no intermediate sequences between lineages A and B and no significant change in the relative composition of these two lineages. Finding either, especially in early samples, would reduce the likelihood that there were two or more introductions of SARS-CoV-2, but this was not the case. We found no indication that additional data significantly changed the fraction of C29095T or C18060T in our dataset. Finding an increased frequency for either would have increased the likelihood of an alternative ancestral sequence, but this was not the case.</p><p>Rather than overturning previous results, newly added sequences and annotation of sample metadata made it possible to identify increased support for two sublineages of lineage B spreading from the Huanan market. This would be unlikely to occur if the market outbreak started as a small part of a wider outbreak. This is in addition to the low likelihoods that, by mere coincidence, both lineage A and lineage B would be found in market environmental sampling, and that three of the earliest lineage A sequences would be from patients located near the market. Conversely, all of these observations are straightforward predictions from the hypothesis that SARS-CoV-2 emerged from the Huanan market wildlife trade, a hypothesis first posed at the end of 2019 on social media before almost any data existed to test it <ref type="bibr" target="#b2">[3,</ref><ref type="bibr" target="#b3">4]</ref>.</p><p>We also identified a clade separated by three mutations from lineage A including one potentially ancestral mutation.</p><p>Multi-type phylodynamic analysis suggested that this mutation characterizes a plausible, but unlikely, proximal ancestor genome. This part of our analysis could be improved by incorporating realistic spatiotemporal heterogeneities in transmission and sampling likelihoods as well as a more realistic substitution model, possibly utilizing empirical estimates of site-specific substitution rates <ref type="bibr" target="#b45">[46]</ref>. Although our model was agnostic about the nature of the proximal host, a market-specific model could additionally incorporate a location trait for market-linked sequences.</p><p>More complex analyses would benefit from more complete sample metadata, especially with respect to market exposure and focused sequencing of patient clusters. An important limitation in this work is that sequences without a known market link may, in fact, have unpublished, direct links to the Huanan market. Epidemiological metadata was key, for example, to attribute some features of our dataset to focused sampling in Anyang <ref type="bibr" target="#b19">[20]</ref>. The dataset reported here can also certainly be improved with additional scrutiny. For example, we have not removed duplicates from two studies that collected and sequenced independent samples from overlapping sets of patients <ref type="bibr" target="#b18">[19,</ref><ref type="bibr" target="#b46">47]</ref>; amending published metadata to</p><p>An updated dataset of early SARS-CoV-2 diversity supports a wildlife market origin Z. Hensel &amp; F. Débarre include cross-referenced patient ID numbers would make make it possible to construct a more accurate dataset.</p><p>Additionally, analysis of early pandemic genomic diversity will be improved by publishing additional sequences such as those from samples collected in January 2020 <ref type="bibr" target="#b6">[7]</ref>.</p><p>Our analysis demonstrates how evidence continues to accumulate and clarify the origin of SARS-CoV-2. This is not limited to human SARS-CoV-2 sequences, but also to sequences of animal viruses such as those supporting C24034T as a reversion to an ancestral sequence. In addition to viral genome sequencing data, we note that other sources of data are often overlooked in commentary questioning the strength of the evidence for SARS-CoV-2 emergence in the Huanan market. These other datasets include spatiotemporal data of healthcare worker infections showing spread from the area around the market <ref type="bibr" target="#b5">[6]</ref> and serosurveys of blood donor samples failing to find a broader early outbreak <ref type="bibr" target="#b47">[48,</ref><ref type="bibr" target="#b48">49]</ref>. A combination of additional data and models that can synthesize diverse sources of data will continue to refine our understanding of SARS-CoV-2 origin.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Methods</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Curation of sequence dataset</head><p>Our initial dataset was a recently reported dataset of 863 genomes collected between 24-Dec-2019 and 15-Feb-2020 <ref type="bibr" target="#b13">[14]</ref>;</p><p>we gathered assemblies corresponding to sample names in analysis configuration files available at <ref type="url" target="https://github.com/pekarj/SC2_intermediates">https://github.com/pekarj/SC2_intermediates</ref>. To identify some additional sequences, we searched the GISAID database for newly published sequences and searched the RCoV19 resource (<ref type="url" target="https://ngdc.cncb.ac.cn/ncov/">https://ngdc.cncb.ac.cn/ncov/</ref>) for sequences not in GISAID or GenBank databases. Other published assemblies were identified from publications as detailed in the Results</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>section.</head><p>Supplementary Data 1 includes information for the database, assembly accession, related BioProject accession (where available), corresponding manuscript (where one was identified), and sample collection date. Collection dates were obtained from linked BioSample metadata. Mid-January 2020 sequences from Wuhan were published in January 2024 to BioProject PRJCA002163 together with sequences from Anyang <ref type="bibr" target="#b19">[20]</ref>; BioSample data had been published one year earlier.</p><p>New assemblies in the dataset were constructed from sequencing reads associated with one or more of several related publications on samples from Beijing patients <ref type="bibr" target="#b31">[32]</ref><ref type="bibr" target="#b32">[33]</ref><ref type="bibr" target="#b33">[34]</ref><ref type="bibr" target="#b34">[35]</ref><ref type="bibr" target="#b35">[36]</ref>, some of which were made public at our request. Samples were deduplicated and collection dates were identified by cross-referencing experiment and BioSample metadata with supplementary tables in related manuscripts, retaining the earliest high-quality sequence for each patient. We assembled sequences from all samples from several related BioProjects and our dataset includes sequences from CRA002626 (15), HRA000181 (26), HRA000349 (15), and PRJNA667180 (1). Unmapped reads were mapped with Minimap2 <ref type="bibr" target="#b49">[50]</ref> (options: -ax sr). Consensus sequences were assembled using ViralConsensus <ref type="bibr" target="#b50">[51]</ref> (options: -q 20 -d 10 -f 0.5). Assemblies were subsequently refined with a 75% consensus threshold. Sequences with less than 95% coverage were excluded from the dataset for consistency with methods used to construct the initial dataset <ref type="bibr" target="#b12">[13]</ref>.</p><p>Additionally, some samples in the initial dataset were removed. Supplementary Data 1 describes the rationale for removal such as evidence supporting sequences as duplicates. Other samples were excluded because of coverage below 95% according to supplementary data in a related publication <ref type="bibr" target="#b20">[21]</ref>. We also made a modification for which market-linked sequences were included in our dataset. Although A24325G was reportedly found in consensus genomes for two market-linked patients (S04 and S12 in Table <ref type="table">6</ref> of the Joint WHO-China Study <ref type="bibr" target="#b17">[18]</ref>), 1 of 5 sequences for S04 lacks previously made for other sequence-patient pairs in the report <ref type="bibr" target="#b51">[52]</ref>. For patient S12, we opted for IME-WH02 over WIV06 in our dataset due to higher reported sequencing depth in the Joint WHO-China Study <ref type="bibr" target="#b17">[18]</ref>; both sequences contain no substitutions relative to lineage B. This left the sequence for patient S04 (WH19008) as the only market-linked sequence with A24325G in our dataset. Further inspection of sequencing and metadata <ref type="bibr" target="#b37">[38]</ref> identified another sequence, WH19003, containing A24325G and likely corresponding to a different patient in the same cluster as WH19008 (Cluster 2 in the Joint WHO-China Study). Although coverage is too low to include sample WH19003 in our dataset, together this identifies A24325G as the only mutation shared by a cluster of patients in Huanan market in late December 2019.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Phylogenetic analyses</head><p>The entire dataset of 1,009 sequences was analyzed using the Nextstrain ncov pipeline <ref type="bibr" target="#b52">[53]</ref> available at <ref type="url" target="https://github.com/nextstrain/ncov">https://github.com/nextstrain/ncov</ref>, rooting on lineage A to infer the phylogenetic tree and date samples collected between 24-Dec-2019 and 15-Feb-2020. Annotation was added based upon mutations in RecCA <ref type="bibr" target="#b12">[13]</ref> identified by Nextclade <ref type="bibr" target="#b53">[54]</ref>, with the addition of C24034T. Auspice was used to visualize trees for both Figure <ref type="figure" target="#fig_3">3</ref> and Figure <ref type="figure" target="#fig_5">5</ref>.</p><p>To investigate the C24034T mutation, we first used MAFFT <ref type="bibr" target="#b54">[55]</ref> to add BtSY2, also known as CX1 <ref type="bibr" target="#b42">[43]</ref>, to a multiple sequence alignment of 111 animal and human coronaviruses <ref type="bibr" target="#b43">[44]</ref>. The region of the alignment corresponding to the phylogenetic coloured genomic bootstrap (CGB) barcode around SARS-CoV-2 position 24034 was extracted from the alignment. IQTree <ref type="bibr" target="#b55">[56]</ref> (options: -st DNA -m GTR+F -czb --keep-ident -nt 4 -redo -seed 2 -asr) was used to construct the phylogenetic tree of this region, and only a portion of the tree is shown in Figure <ref type="figure" target="#fig_5">5</ref> for clarity.</p><p>We analyzed SARS-CoV-2 phylogenetic trees with Python scripts using the ETE Toolkit <ref type="bibr" target="#b56">[57]</ref> to estimate the likelihood that one of three potentially ancestral mutations would form a similar divergent polytomy. Each node we examined had at least 26 descendants, matching the number of C18060T samples in our dataset. In the GISAID Global Phylogeny <ref type="bibr" target="#b57">[58]</ref>, 1.83% of these nodes formed polytomies at least three mutations diverged from their parent with 10 or more children.</p><p>This was similar to 1.81% in the public UShER tree <ref type="bibr" target="#b58">[59]</ref> and 2.69% for an early-pandemic phylogenetic tree <ref type="bibr" target="#b59">[60]</ref>. Little change in likelihood was observed as a function of distance from the root of the tree. The likelihood that two similarly sized polytomies would be separated by 2 mutations was analyzed similarly according to the polytomy size criteria described in Results.</p><p>Bayesian phylodynamics analysis utilized the BDMM-Prime multi-type birth-death model <ref type="bibr" target="#b44">[45,</ref><ref type="bibr" target="#b60">61,</ref><ref type="bibr" target="#b61">62]</ref> and BEAST 2.5 <ref type="bibr" target="#b62">[63]</ref>. We used sequences collected on or before 22-Jan-2020 (N=129; sample YB20200116082 was excluded based upon having 9 substitutions between positions 15101 and 15111). The cutoff date was chosen based upon approximately exponential growth of the cumulative number of sequences through this date, allowing for a relatively simple model of sampling likelihood while including several samples (N&gt;=5) for each of the C18060T, C24034T, and C29095T clades.</p><p>Sequences were aligned using NextClade <ref type="bibr" target="#b53">[54]</ref>. Positions 1-88 and 29709-29903 were masked based upon lacking coverage in at least 5% of samples. A synthetic sample dated 15-Nov-2019, having N at positions 18060, 24034, and 29095, and otherwise identical to lineage A sample WH04 was added, positing that a sample with this ambiguous haplotype could have been collected at this time from the proximal reservoir of the human SARS-CoV-2 outbreak. This was a plausible date of primary human SARS-CoV-2 infections in previous analysis <ref type="bibr" target="#b12">[13]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Model specification is described</head><p>An updated dataset of early SARS-CoV-2 diversity supports a wildlife market origin Z. Hensel &amp; F. Débarre in Supplementary Data 2. To simplify the model given phylogenetic uncertainty, we constrained some parameters e.g. using a fixed death rate as in previous work <ref type="bibr" target="#b63">[64]</ref>.</p><p>Posterior tree distributions were analyzed to construct HIPSTR reconstructions using TreeAnnotator <ref type="bibr" target="#b64">[65]</ref>. Topologies and host-type likelihoods for parent nodes of clades of interest were quantified with Python scripts utilizing TreeSwift <ref type="bibr" target="#b65">[66]</ref>. We discarded 10% of 100,001 trees (sampled every 1,000 of 100 million iterations) and confirmed effective sample sizes above 200 for all inferred parameters using Tracer <ref type="bibr" target="#b66">[67]</ref>. Figures <ref type="figure" target="#fig_0">S1</ref> and <ref type="figure" target="#fig_1">S2</ref> show results of a sensitivity analysis in which we repeated analysis using different prior distributions for the spillover rate, parameterized as an effective reproduction number. Absolute likelihoods that clades were founded by spillovers or had ancestral topologies varied somewhat, but relative likelihoods consistently found C24034T much more likely than C18060T or C29095T to be ancestral.</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_0"><head>Figure 1 .</head><label>1</label><figDesc>Figure 1. Temporal distribution of initial dataset and newly added sequences (samples collected between 24-Dec-2019 and 15-Feb-2020). Sequences in the initial dataset from market-linked samples and within a single familial cluster are colored green and purple, respectively. Other sequences in the initial dataset are colored orange. Newly added sequences are colored blue. The arrow indicates the earliest sequence in added data, which is market-linked.</figDesc><graphic coords="3,44.70,291.70,509.25,348.75" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_1"><head>Figure 2 .</head><label>2</label><figDesc>Figure 2. The cumulative fraction of sequences in lineage A for the initial dataset (orange), sequences added in this study (light blue), and added sequences excluding those collected in Anyang (dark blue). The blue lines are identical prior to the earliest sample from Anyang on 25-Jan-2020. With the exception of sequences from Anyang, newly added sequences have a similar composition of lineage A and B to sequences considered in previous analyses. Time series of binomial proportions with Wilson score intervals (z = 1.4) are shown as confidence bands.</figDesc><graphic coords="4,117.64,149.63,360.00,249.00" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_2"><figDesc>An updated dataset of early SARS-CoV-2 diversity supports a wildlife market origin Z. Hensel &amp; F. Débarre</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_3"><head>Figure 3 .</head><label>3</label><figDesc>Figure 3. Phylogenetic trees of 1,009 samples collected between 24-Dec-2019 and 15-Feb-2020 are shown, rooted on lineage A. Lineage B samples, at the top of the trees, are vertically compressed to focus on lineage A. A given sequence is shown at the same vertical location on both panels. (A) Samples are plotted by time and colored by category: market-linked (green), linked to an early familial cluster (purple), other samples in the initial dataset (orange), and newly added samples (blue). (B) Samples are plotted by divergence from lineage A (a few lineage B samples over 10 mutations divergent from lineage A are not shown). Samples harboring mutations towards either the inferred ancestral sequence RecCA or C24034T are colored yellow. Three clades directly descending from lineage A are annotated; they contain 49 (C29095T), 24 (C18060T), and 32 (C24034T, T26729C, G28077C) descendant samples.</figDesc><graphic coords="5,44.70,73.60,509.25,410.25" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_4"><head>Figure 4 .</head><label>4</label><figDesc>Figure 4. The cumulative fraction of the initial and updated sequencing datasets are shown for three clades with mutations from lineage A (C29095T, top, C18060T, middle, or all three of C24034T, T26729C, and G28077C, bottom). This shows how the overall relative composition of the dataset changes as samples are added over time. Time series of binomial proportions with Wilson score intervals (z = 1.4) are shown as confidence bands.</figDesc><graphic coords="6,117.64,258.68,360.00,245.25" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_5"><head>Figure 5 .</head><label>5</label><figDesc>Figure5. Phylogenetic tree of the region containing SARS-CoV-2 position 24034 and shown to have the most robust phylogenetic signal in previous analysis, and adding the sequence RmBtSY2, also called CX1<ref type="bibr" target="#b42">[43]</ref>, to 111 human and animal sarbecovirus sequences considered in that study<ref type="bibr" target="#b43">[44]</ref>. This corresponds to positions 23734-24183 in the SARS-CoV-2 reference genome. Only the sequences most similar to SARS-CoV-2 are shown. Samples are colored by nucleotide identity at the equivalent position to SARS-CoV-2 24034 in the multiple sequence alignment.</figDesc><graphic coords="8,117.64,433.77,360.00,267.00" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_6"><figDesc>An updated dataset of early SARS-CoV-2 diversity supports a wildlife market origin Z. Hensel &amp; F. Débarre important to evaluate the robustness of this result in future work. Our analysis identified potential multiple spillovers by inferring the host type of internal nodes. The highest independent posterior subtree reconstruction (HIPSTR) supported two introductions, one for lineage A and one for lineage B (Figure6A).</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_7"><head>Figure 6 .</head><label>6</label><figDesc>Figure 6. Consensus tree topology and alternative tree topologies in phylodynamics analysis. (A) Topology and internal node types of the highest independent posterior subtree reconstruction (HIPSTR) tree. Nodes colored by most likely type (proximal host, green, or human, orange). In this topology, the C24034T clade is sister to other lineage A sequences, all lineage A sequences are monophyletic, and the synthetic proximal host sample (terminal green node) is an outgroup to human samples. (B) Three topologies in which the C24034T clade is evolutionarily distinct from the MRCA of other lineage A sequences and lineage B (i.e. an outgroup to the clade containing lineage B and remaining lineage A sequences). Nodes are colored by possible types; the second two topologies have at least two spillovers. Some nodes are shown with split coloring to illustrate how the same topology can have different numbers of spillovers. Percentages indicate the frequencies of these topologies in the posterior tree distribution (N=90,001 sampled trees after discarding burn-in).</figDesc><graphic coords="10,44.70,133.13,509.25,133.50" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_8"><head>An updated dataset of</head><figDesc>early SARS-CoV-2 diversity supports a wildlife market origin Z. Hensel &amp; F. Débarre A24325G (IME-WH02), and only 1 of 2 sequences for S12 has A24325G (IME-WH03). We concluded that swapping IME-WH02 and IME-WH03 was most parsimonious with available sequence data and metadata; a similar change was</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_9"><graphic coords="18,117.64,103.61,360.00,270.00" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_10"><graphic coords="18,117.64,431.46,360.00,270.00" type="bitmap" /></figure>
		</body>
		<back>

			<div type="acknowledgement">
<div><head>Acknowledgments</head><p>We thank <rs type="person">Hui Zeng</rs> and colleagues at <rs type="institution">Beijing Ditan Hospital, Capital Medical University</rs> for making available sequence data in GSA for Human repositories <rs type="grantNumber">HRA000181</rs> and <rs type="grantNumber">HRA000349</rs>. They are gratefully acknowledged along with all other data contributors, i.e., the authors and their originating laboratories responsible for obtaining the specimens, and their submitting laboratories for generating the genetic sequence and metadata and sharing via GISAID and other databases, on which this research is based. Supplementary Data 1 includes names of submitting laboratories and individuals. To view the contributors of each individual sequence for sequences used from GISAID, visit <ref type="url" target="https://doi.org/10.55876/gis8.250331rs">https://doi.org/10.55876/gis8.250331rs</ref>. We thank <rs type="person">Niema Moshiri</rs>, <rs type="person">Jonathan E. Pekar</rs>, and <rs type="person">Alexander Crits-Christoph</rs> for advice on analysis and critical comments on this manuscript.</p></div>
			</div>
			<listOrg type="funding">
				<org type="funding" xml:id="_846bB26">
					<idno type="grant-number">HRA000181</idno>
				</org>
				<org type="funding" xml:id="_fmRG7ea">
					<idno type="grant-number">HRA000349</idno>
				</org>
			</listOrg>

			<div type="availability">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Data Availability</head><p>Our sequencing dataset cannot be published, since it contains sequences exclusively available from GISAID. We provide as</p></div>
			</div>

			<div type="annex">
<div xmlns="http://www.tei-c.org/ns/1.0"><p>Supplementary Data 1 a spreadsheet with information required to reproduce our sequencing dataset. Likewise, phylodynamics configuration files contain sequence data from GISAID; modeling parameters required to reproduce our results are available in Supplementary Data 2.</p><p>An updated dataset of early SARS-CoV-2 diversity supports a wildlife market origin Z. Hensel &amp; F. Débarre</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Supplemental Figures</head><p>Figure <ref type="figure">S1</ref>. Histograms of the number of spillovers in the posterior tree distribution, colored by the prior distribution used for the effective reproduction number for spillover from the proximal host to humans. Trees are likely to have multiple spillovers with all prior distributions that were tested.</p><p>Figure <ref type="figure">S2</ref>. Prior distributions (dashed) and kernel density estimates of posterior distributions (solid) for the effective reproduction number for spillover from the proximal host to humans. Colored by the prior distribution used. Small differences between prior and posterior distributions indicate sensitivity to prior specification.</p></div>			</div>
			<div type="references">

				<listBibl>

<biblStruct status="extracted" xml:id="b0">
	<monogr>
		<title level="m" type="main">How the COVID-19 Outbreak in China Spiraled Out of Control</title>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">L</forename><surname>Yang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Wuhan</forename></persName>
		</author>
		<imprint>
			<date type="published" when="2024">2024</date>
			<publisher>Oxford University Press</publisher>
			<pubPlace>Oxford, New York</pubPlace>
		</imprint>
	</monogr>
	<note type="raw_reference">D. L. Yang, Wuhan: How the COVID-19 Outbreak in China Spiraled Out of Control. Oxford, New York: Oxford University Press, 2024.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b1">
	<analytic>
		<title level="a" type="main">Dissecting the early COVID-19 cases in Wuhan</title>
		<author>
			<persName><forename type="first">M</forename><surname>Worobey</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Science</title>
		<imprint>
			<biblScope unit="volume">374</biblScope>
			<biblScope unit="issue">6572</biblScope>
			<biblScope unit="page" from="1202" to="1204" />
			<date type="published" when="2021-12">Dec. 2021</date>
		</imprint>
	</monogr>
	<note type="raw_reference">M. Worobey, &quot;Dissecting the early COVID-19 cases in Wuhan,&quot; Science, vol. 374, no. 6572, pp. 1202-1204, Dec. 2021, doi: 10/gnqx5k.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b2">
	<monogr>
		<author>
			<persName><surname>万谦宠爱_Yuki</surname></persName>
		</author>
		<ptr target="https://news.ltn.com.tw/news/world/breakingnews/3026710" />
		<title level="m">#Wuhan SARS# That wretched place, the Huanan Seafood Market-I went to investigate it once</title>
		<imprint>
			<date type="published" when="2019-12-30">Dec. 30, 2019</date>
		</imprint>
	</monogr>
	<note>translated Weibo post</note>
	<note type="raw_reference">万谦宠爱_Yuki, &quot;Weibo post (translated): #Wuhan SARS# That wretched place, the Huanan Seafood Market-I went to investigate it once.,&quot; Dec. 30, 2019. [Online]. Available: https://news.ltn.com.tw/news/world/breakingnews/3026710</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b3">
	<analytic>
		<title level="a" type="main">All game shops at Wuhan Huanan Seafood Market have been closed, and disease control personnel are checking the scene with a list</title>
		<author>
			<persName><forename type="first">L</forename><surname>Ping</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Z</forename><surname>Wang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Yang</surname></persName>
		</author>
		<ptr target="https://static.cdsb.com/micropub/Articles/" />
	</analytic>
	<monogr>
		<title level="j">Red Star News</title>
		<imprint>
			<date type="published" when="2025-03-20">Mar. 20, 2025. 201912/5c399cf417e73a1c6fb8fe996d0a6fbb.html</date>
		</imprint>
	</monogr>
	<note type="raw_reference">L. Ping, Z. Wang, and Y. Yang, &quot;All game shops at Wuhan Huanan Seafood Market have been closed, and disease control personnel are checking the scene with a list,&quot; Red Star News. Accessed: Mar. 20, 2025. [Online]. Available: https://static.cdsb.com/micropub/Articles/201912/5c399cf417e73a1c6fb8fe996d0a6fbb.html</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b4">
	<analytic>
		<title level="a" type="main">The Huanan Seafood Wholesale Market in Wuhan was the early epicenter of the COVID-19 pandemic</title>
		<author>
			<persName><forename type="first">M</forename><surname>Worobey</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Science</title>
		<imprint>
			<biblScope unit="volume">377</biblScope>
			<biblScope unit="issue">6609</biblScope>
			<biblScope unit="page" from="951" to="959" />
			<date type="published" when="2022-08">Aug. 2022</date>
		</imprint>
	</monogr>
	<note type="raw_reference">M. Worobey et al., &quot;The Huanan Seafood Wholesale Market in Wuhan was the early epicenter of the COVID-19 pandemic,&quot; Science, vol. 377, no. 6609, pp. 951-959, Aug. 2022, doi: 10/gqj6km.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b5">
	<analytic>
		<title level="a" type="main">Spatiotemporal characteristics and factor analysis of SARS-CoV-2 infections among healthcare workers in Wuhan, China</title>
		<author>
			<persName><forename type="first">P</forename><surname>Wang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Ren</surname></persName>
		</author>
		<author>
			<persName><forename type="first">X</forename><surname>Zhu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">X</forename><surname>Fu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">H</forename><surname>Liu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Hu</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">J. Hosp. Infect</title>
		<imprint>
			<biblScope unit="volume">110</biblScope>
			<biblScope unit="page" from="172" to="177" />
			<date type="published" when="2021-04">Apr. 2021</date>
		</imprint>
	</monogr>
	<note type="raw_reference">P. Wang, H. Ren, X. Zhu, X. Fu, H. Liu, and T. Hu, &quot;Spatiotemporal characteristics and factor analysis of SARS-CoV-2 infections among healthcare workers in Wuhan, China,&quot; J. Hosp. Infect., vol. 110, pp. 172-177, Apr. 2021, doi: 10/g8789n.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b6">
	<analytic>
		<title level="a" type="main">The Emergence and Evolution of SARS-CoV-2</title>
		<author>
			<persName><forename type="first">E</forename><forename type="middle">C</forename><surname>Holmes</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Annu. Rev. Virol</title>
		<imprint>
			<biblScope unit="volume">11</biblScope>
			<biblScope unit="page">8789</biblScope>
			<date type="published" when="2024-09">2024. Sep. 2024</date>
		</imprint>
	</monogr>
	<note type="raw_reference">E. C. Holmes, &quot;The Emergence and Evolution of SARS-CoV-2,&quot; Annu. Rev. Virol., vol. 11, no. Volume 11, 2024, pp. 21-42, Sep. 2024, doi: 10/g8789p.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b7">
	<monogr>
		<title level="m" type="main">Confirmation of the centrality of the Huanan market among early COVID-19 cases</title>
		<author>
			<persName><forename type="first">F</forename><surname>Débarre</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Worobey</surname></persName>
		</author>
		<idno type="DOI">10.48550/arXiv.2403.05859</idno>
		<idno type="arXiv">arXiv:arXiv:2403.05859</idno>
		<imprint>
			<date type="published" when="2024-09">Mar. 09, 2024</date>
		</imprint>
	</monogr>
	<note type="raw_reference">F. Débarre and M. Worobey, &quot;Confirmation of the centrality of the Huanan market among early COVID-19 cases,&quot; Mar. 09, 2024, arXiv: arXiv:2403.05859. doi: 10.48550/arXiv.2403.05859.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b8">
	<analytic>
		<title level="a" type="main">Surveillance of SARS-CoV-2 at the Huanan Seafood Market</title>
		<author>
			<persName><forename type="first">W</forename><forename type="middle">J</forename><surname>Liu</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Nature</title>
		<imprint>
			<biblScope unit="volume">631</biblScope>
			<biblScope unit="issue">8020</biblScope>
			<biblScope unit="page" from="402" to="408" />
			<date type="published" when="2024-07">Jul. 2024</date>
		</imprint>
	</monogr>
	<note type="raw_reference">W. J. Liu et al., &quot;Surveillance of SARS-CoV-2 at the Huanan Seafood Market,&quot; Nature, vol. 631, no. 8020, pp. 402-408, Jul. 2024, doi: 10/j46n.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b9">
	<analytic>
		<title level="a" type="main">Genetic tracing of market wildlife and viruses at the epicenter of the COVID-19 pandemic</title>
		<author>
			<persName><forename type="first">A</forename><surname>Crits-Christoph</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Cell</title>
		<imprint>
			<biblScope unit="volume">187</biblScope>
			<biblScope unit="issue">19</biblScope>
			<biblScope unit="page" from="5468" to="5482" />
			<date type="published" when="2024-09">Sep. 2024</date>
		</imprint>
	</monogr>
	<note type="raw_reference">A. Crits-Christoph et al., &quot;Genetic tracing of market wildlife and viruses at the epicenter of the COVID-19 pandemic,&quot; Cell, vol. 187, no. 19, pp. 5468-5482.e11, Sep. 2024, doi: 10/nh9f.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b10">
	<analytic>
		<title level="a" type="main">Susceptibility of Raccoon Dogs for Experimental SARS-CoV-2 Infection</title>
		<author>
			<persName><forename type="first">C</forename><forename type="middle">M</forename><surname>Freuling</surname></persName>
		</author>
		<idno>doi: 10/ffzf</idno>
	</analytic>
	<monogr>
		<title level="j">Emerg. Infect. Dis</title>
		<imprint>
			<biblScope unit="volume">26</biblScope>
			<biblScope unit="issue">12</biblScope>
			<biblScope unit="page" from="2982" to="2985" />
			<date type="published" when="2020-12">Dec. 2020</date>
		</imprint>
	</monogr>
	<note type="raw_reference">C. M. Freuling et al., &quot;Susceptibility of Raccoon Dogs for Experimental SARS-CoV-2 Infection,&quot; Emerg. Infect. Dis., vol. 26, no. 12, pp. 2982-2985, Dec. 2020, doi: 10/ffzf.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b11">
	<analytic>
		<title level="a" type="main">What we can and cannot learn from SARS-CoV-2 and animals in metagenomic samples from the Huanan market</title>
		<author>
			<persName><forename type="first">F</forename><surname>Débarre</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Virus Evol</title>
		<imprint>
			<biblScope unit="volume">10</biblScope>
			<biblScope unit="issue">1</biblScope>
			<date type="published" when="2024-10">Oct. 2024</date>
		</imprint>
	</monogr>
	<note type="raw_reference">F. Débarre, &quot;What we can and cannot learn from SARS-CoV-2 and animals in metagenomic samples from the Huanan market,&quot; Virus Evol., vol. 10, no. 1, p. vead077, Oct. 2024, doi: 10/g89g3n.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b12">
	<analytic>
		<title level="a" type="main">The molecular epidemiology of multiple zoonotic origins of SARS-CoV-2</title>
		<author>
			<persName><forename type="first">J</forename><forename type="middle">E</forename><surname>Pekar</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Science</title>
		<imprint>
			<biblScope unit="volume">377</biblScope>
			<biblScope unit="issue">6609</biblScope>
			<biblScope unit="page">3</biblScope>
			<date type="published" when="2022-08">Aug. 2022</date>
		</imprint>
	</monogr>
	<note type="raw_reference">J. E. Pekar et al., &quot;The molecular epidemiology of multiple zoonotic origins of SARS-CoV-2,&quot; Science, vol. 377, no. 6609, pp. 960-966, Aug. 2022, doi: 10/gqj4f3.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b13">
	<analytic>
		<title level="a" type="main">Recently reported SARS-CoV-2 genomes suggested to be intermediate between the two early main lineages are instead likely derived</title>
		<author>
			<persName><forename type="first">J</forename><forename type="middle">E</forename><surname>Pekar</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Virus Evol</title>
		<imprint>
			<biblScope unit="volume">11</biblScope>
			<biblScope unit="issue">1</biblScope>
			<date type="published" when="2025-01">Jan. 2025</date>
		</imprint>
	</monogr>
	<note type="raw_reference">J. E. Pekar et al., &quot;Recently reported SARS-CoV-2 genomes suggested to be intermediate between the two early main lineages are instead likely derived,&quot; Virus Evol., vol. 11, no. 1, Jan. 2025, doi: 10/n8np.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b14">
	<analytic>
		<title level="a" type="main">A familial cluster of pneumonia associated with the 2019 novel coronavirus indicating person-to-person transmission: a study of a family cluster</title>
		<author>
			<persName><forename type="first">J</forename><forename type="middle">F W</forename><surname>Chan</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">The Lancet</title>
		<imprint>
			<biblScope unit="volume">395</biblScope>
			<biblScope unit="issue">10223</biblScope>
			<biblScope unit="page" from="514" to="523" />
			<date type="published" when="2020-02">Feb. 2020</date>
		</imprint>
	</monogr>
	<note type="raw_reference">J. F.-W. Chan et al., &quot;A familial cluster of pneumonia associated with the 2019 novel coronavirus indicating person-to-person transmission: a study of a family cluster,&quot; The Lancet, vol. 395, no. 10223, pp. 514-523, Feb. 2020, doi: 10/ggjs7j.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b15">
	<analytic>
		<title level="a" type="main">Recovery of Deleted Deep Sequencing Data Sheds More Light on the Early Wuhan SARS-CoV-2 Epidemic</title>
		<author>
			<persName><forename type="first">J</forename><forename type="middle">D</forename><surname>Bloom</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Mol. Biol. Evol</title>
		<imprint>
			<biblScope unit="volume">38</biblScope>
			<biblScope unit="issue">12</biblScope>
			<biblScope unit="page" from="5211" to="5224" />
			<date type="published" when="2021-12">Dec. 2021</date>
		</imprint>
	</monogr>
	<note type="raw_reference">J. D. Bloom, &quot;Recovery of Deleted Deep Sequencing Data Sheds More Light on the Early Wuhan SARS-CoV-2 Epidemic,&quot; Mol. Biol. Evol., vol. 38, no. 12, pp. 5211-5224, Dec. 2021, doi: 10/gpg89c.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b16">
	<analytic>
		<title level="a" type="main">An Evolutionary Portrait of the Progenitor SARS-CoV-2 and Its Dominant Offshoots in COVID-19 Pandemic</title>
		<author>
			<persName><forename type="first">S</forename><surname>Kumar</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Mol. Biol. Evol</title>
		<imprint>
			<biblScope unit="volume">38</biblScope>
			<biblScope unit="issue">8</biblScope>
			<biblScope unit="page" from="3046" to="3059" />
			<date type="published" when="2021-07">Jul. 2021</date>
		</imprint>
	</monogr>
	<note type="raw_reference">S. Kumar et al., &quot;An Evolutionary Portrait of the Progenitor SARS-CoV-2 and Its Dominant Offshoots in COVID-19 Pandemic,&quot; Mol. Biol. Evol., vol. 38, no. 8, pp. 3046-3059, Jul. 2021, doi: 10/f99c.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b17">
	<monogr>
		<title level="m" type="main">WHO-convened global study of origins of SARS-CoV-2: China Part</title>
		<author>
			<persName><surname>Who</surname></persName>
		</author>
		<ptr target="https://www.who.int/publications/i/item/who-convened-global-study-of-origins-of-sars-cov-2-china-part" />
		<imprint>
			<date type="published" when="2021-03-12">2021. Mar. 12, 2025</date>
		</imprint>
	</monogr>
	<note type="raw_reference">WHO, &quot;WHO-convened global study of origins of SARS-CoV-2: China Part,&quot; 2021. Accessed: Mar. 12, 2025. [Online]. Available: https://www.who.int/publications/i/item/who-convened-global-study-of-origins-of-sars-cov-2-china-part</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b18">
	<analytic>
		<title level="a" type="main">Evolutionary trajectory of diverse SARS-CoV-2 variants at the beginning of COVID-19 outbreak</title>
		<author>
			<persName><forename type="first">J.-X</forename><surname>Lv</surname></persName>
		</author>
		<idno type="DOI">10.1093/ve/veae020</idno>
	</analytic>
	<monogr>
		<title level="j">Virus Evol</title>
		<imprint>
			<biblScope unit="volume">10</biblScope>
			<biblScope unit="issue">1</biblScope>
			<date type="published" when="2024-05">May 2024</date>
		</imprint>
	</monogr>
	<note type="raw_reference">J.-X. Lv et al., &quot;Evolutionary trajectory of diverse SARS-CoV-2 variants at the beginning of COVID-19 outbreak,&quot; Virus Evol., vol. 10, no. 1, p. veae020, May 2024, doi: 10.1093/ve/veae020.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b19">
	<analytic>
		<title level="a" type="main">Characteristics of SARS-CoV-2 transmission in a medium-sized city with traditional communities during the early COVID-19 epidemic in China</title>
		<author>
			<persName><forename type="first">Y</forename><surname>Li</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Virol. Sin</title>
		<imprint>
			<biblScope unit="volume">37</biblScope>
			<biblScope unit="issue">2</biblScope>
			<biblScope unit="page">92</biblScope>
			<date type="published" when="2022-04">Apr. 2022</date>
		</imprint>
	</monogr>
	<note type="raw_reference">Y. Li et al., &quot;Characteristics of SARS-CoV-2 transmission in a medium-sized city with traditional communities during the early COVID-19 epidemic in China,&quot; Virol. Sin., vol. 37, no. 2, pp. 187-197, Apr. 2022, doi: 10/g8z92w.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b20">
	<analytic>
		<title level="a" type="main">Genomic monitoring of SARS-CoV-2 uncovers an Nsp1 deletion variant that modulates type I interferon response</title>
		<author>
			<persName><forename type="first">J</forename><surname>Lin</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Cell Host Microbe</title>
		<imprint>
			<biblScope unit="volume">29</biblScope>
			<biblScope unit="issue">3</biblScope>
			<biblScope unit="page" from="489" to="502" />
			<date type="published" when="2021-03">Mar. 2021</date>
		</imprint>
	</monogr>
	<note type="raw_reference">J. Lin et al., &quot;Genomic monitoring of SARS-CoV-2 uncovers an Nsp1 deletion variant that modulates type I interferon response,&quot; Cell Host Microbe, vol. 29, no. 3, pp. 489-502.e8, Mar. 2021, doi: 10/g88q6d.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b21">
	<analytic>
		<title level="a" type="main">Assessment of two-pool multiplex long-amplicon nanopore sequencing of SARS-CoV-2</title>
		<author>
			<persName><forename type="first">H</forename><surname>Liu</surname></persName>
		</author>
		<idno>doi: 10/gsfr53</idno>
	</analytic>
	<monogr>
		<title level="j">J. Med. Virol</title>
		<imprint>
			<biblScope unit="volume">94</biblScope>
			<biblScope unit="issue">1</biblScope>
			<biblScope unit="page" from="327" to="334" />
			<date type="published" when="2022">2022</date>
		</imprint>
	</monogr>
	<note type="raw_reference">H. Liu et al., &quot;Assessment of two-pool multiplex long-amplicon nanopore sequencing of SARS-CoV-2,&quot; J. Med. Virol., vol. 94, no. 1, pp. 327-334, 2022, doi: 10/gsfr53.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b22">
	<analytic>
		<title level="a" type="main">Rapid genomic characterization of SARS-CoV-2 viruses from clinical specimens using nanopore sequencing</title>
		<author>
			<persName><forename type="first">J</forename><surname>Li</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Sci. Rep</title>
		<imprint>
			<biblScope unit="volume">10</biblScope>
			<biblScope unit="page">17492</biblScope>
			<date type="published" when="2020-10">Oct. 2020</date>
		</imprint>
	</monogr>
	<note type="raw_reference">J. Li et al., &quot;Rapid genomic characterization of SARS-CoV-2 viruses from clinical specimens using nanopore sequencing,&quot; Sci. Rep., vol. 10, p. 17492, Oct. 2020, doi: 10/gk7pgt.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b23">
	<analytic>
		<title level="a" type="main">A critical reexamination of recovered SARS-CoV-2 sequencing data</title>
		<author>
			<persName><forename type="first">F</forename><surname>Débarre</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Z</forename><surname>Hensel</surname></persName>
		</author>
		<idno type="DOI">10.1101/2024.02.15.580500</idno>
	</analytic>
	<monogr>
		<title level="j">bioRxiv</title>
		<imprint>
			<date type="published" when="2024-08-22">Aug. 22, 2024</date>
		</imprint>
	</monogr>
	<note type="raw_reference">F. Débarre and Z. Hensel, &quot;A critical reexamination of recovered SARS-CoV-2 sequencing data,&quot; Aug. 22, 2024, bioRxiv. doi: 10.1101/2024.02.15.580500.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b24">
	<analytic>
		<title level="a" type="main">Multiple clades of SARS-CoV-2 were introduced to Thailand during the first quarter of 2020</title>
		<author>
			<persName><forename type="first">R</forename><surname>Buathong</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Microbiol. Immunol</title>
		<imprint>
			<biblScope unit="volume">65</biblScope>
			<biblScope unit="issue">10</biblScope>
			<biblScope unit="page" from="405" to="409" />
			<date type="published" when="2021">2021</date>
		</imprint>
	</monogr>
	<note type="raw_reference">R. Buathong et al., &quot;Multiple clades of SARS-CoV-2 were introduced to Thailand during the first quarter of 2020,&quot; Microbiol. Immunol., vol. 65, no. 10, pp. 405-409, 2021, doi: 10/g88q6f.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b25">
	<analytic>
		<title level="a" type="main">Genomic characterisation and epidemiology of 2019 novel coronavirus: implications for virus origins and receptor binding</title>
		<author>
			<persName><forename type="first">R</forename><surname>Lu</surname></persName>
		</author>
		<idno>doi: 10/ggjr43</idno>
	</analytic>
	<monogr>
		<title level="j">The Lancet</title>
		<imprint>
			<biblScope unit="volume">395</biblScope>
			<biblScope unit="issue">10224</biblScope>
			<biblScope unit="page" from="565" to="574" />
			<date type="published" when="2020-02">Feb. 2020</date>
		</imprint>
	</monogr>
	<note type="raw_reference">R. Lu et al., &quot;Genomic characterisation and epidemiology of 2019 novel coronavirus: implications for virus origins and receptor binding,&quot; The Lancet, vol. 395, no. 10224, pp. 565-574, Feb. 2020, doi: 10/ggjr43.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b26">
	<analytic>
		<title level="a" type="main">Genomic epidemiology reveals early transmission of SARS-CoV-2 and mutational dynamics in Nanning, China</title>
		<author>
			<persName><forename type="first">D</forename><surname>Bi</surname></persName>
		</author>
		<idno>doi: 10/gs963</idno>
	</analytic>
	<monogr>
		<title level="j">Heliyon</title>
		<imprint>
			<biblScope unit="volume">9</biblScope>
			<biblScope unit="issue">12</biblScope>
			<date type="published" when="2023-12">Dec. 2023</date>
		</imprint>
	</monogr>
	<note type="raw_reference">D. Bi et al., &quot;Genomic epidemiology reveals early transmission of SARS-CoV-2 and mutational dynamics in Nanning, China,&quot; Heliyon, vol. 9, no. 12, Dec. 2023, doi: 10/gs963d.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b27">
	<analytic>
		<title level="a" type="main">Identification of SARS-CoV-2 Variants and Their Clinical Significance in Hefei, China</title>
		<author>
			<persName><forename type="first">X</forename><surname>Cheng</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Front. Med</title>
		<imprint>
			<biblScope unit="volume">8</biblScope>
			<date type="published" when="2022-01">Jan. 2022</date>
		</imprint>
	</monogr>
	<note type="raw_reference">X. Cheng et al., &quot;Identification of SARS-CoV-2 Variants and Their Clinical Significance in Hefei, China,&quot; Front. Med., vol. 8, Jan. 2022, doi: 10/g88q6h.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b28">
	<analytic>
		<title level="a" type="main">Humoral immunity and transcriptome differences of COVID-19 inactivated vacciane and protein subunit vaccine as third booster dose in human</title>
		<author>
			<persName><forename type="first">Y</forename><surname>Zhang</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Front. Immunol</title>
		<imprint>
			<biblScope unit="volume">13</biblScope>
			<date type="published" when="2022-10">Oct. 2022</date>
		</imprint>
	</monogr>
	<note type="raw_reference">Y. Zhang et al., &quot;Humoral immunity and transcriptome differences of COVID-19 inactivated vacciane and protein subunit vaccine as third booster dose in human,&quot; Front. Immunol., vol. 13, Oct. 2022, doi: 10/g88q6j.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b29">
	<analytic>
		<title level="a" type="main">Neutralizing Antibodies against the SARS-CoV-2 Delta and Omicron BA.1 following Homologous CoronaVac Booster Vaccination</title>
		<author>
			<persName><forename type="first">J</forename><surname>Li</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Vaccines</title>
		<imprint>
			<biblScope unit="volume">10</biblScope>
			<biblScope unit="issue">12</biblScope>
			<date type="published" when="2022-12">Dec. 2022</date>
			<pubPlace>Art</pubPlace>
		</imprint>
	</monogr>
	<note type="raw_reference">J. Li et al., &quot;Neutralizing Antibodies against the SARS-CoV-2 Delta and Omicron BA.1 following Homologous CoronaVac Booster Vaccination,&quot; Vaccines, vol. 10, no. 12, Art. no. 12, Dec. 2022, doi: 10/g88q6k.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b30">
	<analytic>
		<title level="a" type="main">A compromised specific humoral immune response against the SARS-CoV-2 receptor-binding domain is related to viral persistence and periodic shedding in the gastrointestinal tract</title>
		<author>
			<persName><forename type="first">F</forename><surname>Hu</surname></persName>
		</author>
		<idno>doi: 10/gk4ppt</idno>
	</analytic>
	<monogr>
		<title level="j">Cell. Mol. Immunol</title>
		<imprint>
			<biblScope unit="volume">17</biblScope>
			<biblScope unit="issue">11</biblScope>
			<biblScope unit="page" from="1119" to="1125" />
			<date type="published" when="2020-11">Nov. 2020</date>
		</imprint>
	</monogr>
	<note type="raw_reference">F. Hu et al., &quot;A compromised specific humoral immune response against the SARS-CoV-2 receptor-binding domain is related to viral persistence and periodic shedding in the gastrointestinal tract,&quot; Cell. Mol. Immunol., vol. 17, no. 11, pp. 1119-1125, Nov. 2020, doi: 10/gk4ppt.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b31">
	<analytic>
		<title level="a" type="main">MINERVA: A Facile Strategy for SARS-CoV-2 Whole-Genome Deep Sequencing of Clinical Samples</title>
		<author>
			<persName><forename type="first">C</forename><surname>Chen</surname></persName>
		</author>
		<idno type="DOI">10.1016/j.molcel.2020.11.030</idno>
	</analytic>
	<monogr>
		<title level="j">Mol. Cell</title>
		<imprint>
			<biblScope unit="volume">80</biblScope>
			<biblScope unit="issue">6</biblScope>
			<biblScope unit="page" from="1123" to="1134" />
			<date type="published" when="2020-12">Dec. 2020</date>
		</imprint>
	</monogr>
	<note type="raw_reference">C. Chen et al., &quot;MINERVA: A Facile Strategy for SARS-CoV-2 Whole-Genome Deep Sequencing of Clinical Samples,&quot; Mol. Cell, vol. 80, no. 6, pp. 1123-1134.e4, Dec. 2020, doi: 10.1016/j.molcel.2020.11.030.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b32">
	<analytic>
		<title level="a" type="main">Phylogenomic tracing of asymptomatic transmission in a COVID-19 outbreak</title>
		<author>
			<persName><forename type="first">J</forename><surname>Zhang</surname></persName>
		</author>
		<idno type="DOI">10.1016/j.xinn.2021.100099</idno>
	</analytic>
	<monogr>
		<title level="j">The Innovation</title>
		<imprint>
			<biblScope unit="volume">2</biblScope>
			<biblScope unit="issue">2</biblScope>
			<biblScope unit="page">100099</biblScope>
			<date type="published" when="2021-05">May 2021</date>
		</imprint>
	</monogr>
	<note type="raw_reference">J. Zhang et al., &quot;Phylogenomic tracing of asymptomatic transmission in a COVID-19 outbreak,&quot; The Innovation, vol. 2, no. 2, p. 100099, May 2021, doi: 10.1016/j.xinn.2021.100099.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b33">
	<analytic>
		<title level="a" type="main">Genomic surveillance of COVID-19 cases in Beijing</title>
		<author>
			<persName><forename type="first">P</forename><surname>Du</surname></persName>
		</author>
		<idno type="DOI">10.1038/s41467-020-19345-0</idno>
	</analytic>
	<monogr>
		<title level="j">Nat. Commun</title>
		<imprint>
			<biblScope unit="volume">11</biblScope>
			<biblScope unit="issue">1</biblScope>
			<biblScope unit="page">5503</biblScope>
			<date type="published" when="2020-10">Oct. 2020</date>
		</imprint>
	</monogr>
	<note type="raw_reference">P. Du et al., &quot;Genomic surveillance of COVID-19 cases in Beijing,&quot; Nat. Commun., vol. 11, no. 1, p. 5503, Oct. 2020, doi: 10.1038/s41467-020-19345-0.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b34">
	<analytic>
		<title level="a" type="main">Two-step fitness selection for intra-host variations in SARS-CoV-2</title>
		<author>
			<persName><forename type="first">J</forename><surname>Li</surname></persName>
		</author>
		<idno type="DOI">10.1016/j.celrep.2021.110205</idno>
	</analytic>
	<monogr>
		<title level="j">Cell Rep</title>
		<imprint>
			<biblScope unit="volume">38</biblScope>
			<biblScope unit="issue">2</biblScope>
			<biblScope unit="page">110205</biblScope>
			<date type="published" when="2022-01">Jan. 2022</date>
		</imprint>
	</monogr>
	<note type="raw_reference">J. Li et al., &quot;Two-step fitness selection for intra-host variations in SARS-CoV-2,&quot; Cell Rep., vol. 38, no. 2, p. 110205, Jan. 2022, doi: 10.1016/j.celrep.2021.110205.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b35">
	<analytic>
		<title level="a" type="main">Assessment of microbiota in the gut and upper respiratory tract associated with SARS-CoV-2 infection</title>
		<author>
			<persName><forename type="first">J</forename><surname>Li</surname></persName>
		</author>
		<idno type="DOI">10.1186/s40168-022-01447-0</idno>
	</analytic>
	<monogr>
		<title level="j">Microbiome</title>
		<imprint>
			<biblScope unit="volume">11</biblScope>
			<biblScope unit="issue">1</biblScope>
			<biblScope unit="page">38</biblScope>
			<date type="published" when="2023-03">Mar. 2023</date>
		</imprint>
	</monogr>
	<note type="raw_reference">J. Li et al., &quot;Assessment of microbiota in the gut and upper respiratory tract associated with SARS-CoV-2 infection,&quot; Microbiome, vol. 11, no. 1, p. 38, Mar. 2023, doi: 10.1186/s40168-022-01447-0.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b36">
	<analytic>
		<title level="a" type="main">A pneumonia outbreak associated with a new coronavirus of probable bat origin</title>
		<author>
			<persName><forename type="first">P</forename><surname>Zhou</surname></persName>
		</author>
		<idno>doi: 10/ggj5cg</idno>
	</analytic>
	<monogr>
		<title level="j">Nature</title>
		<imprint>
			<biblScope unit="volume">579</biblScope>
			<biblScope unit="issue">7798</biblScope>
			<biblScope unit="page" from="270" to="273" />
			<date type="published" when="2020-03">Mar. 2020</date>
		</imprint>
	</monogr>
	<note type="raw_reference">P. Zhou et al., &quot;A pneumonia outbreak associated with a new coronavirus of probable bat origin,&quot; Nature, vol. 579, no. 7798, pp. 270-273, Mar. 2020, doi: 10/ggj5cg.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b37">
	<analytic>
		<title level="a" type="main">Genomic Diversity of Severe Acute Respiratory Syndrome-Coronavirus 2 in Patients With Coronavirus Disease 2019</title>
		<author>
			<persName><forename type="first">Z</forename><surname>Shen</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Clin. Infect. Dis</title>
		<imprint>
			<biblScope unit="volume">71</biblScope>
			<biblScope unit="issue">15</biblScope>
			<biblScope unit="page" from="713" to="720" />
			<date type="published" when="2020-07">Jul. 2020</date>
		</imprint>
	</monogr>
	<note type="raw_reference">Z. Shen et al., &quot;Genomic Diversity of Severe Acute Respiratory Syndrome-Coronavirus 2 in Patients With Coronavirus Disease 2019,&quot; Clin. Infect. Dis., vol. 71, no. 15, pp. 713-720, Jul. 2020, doi: 10/ggq35t.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b38">
	<analytic>
		<title level="a" type="main">Nanopore Targeted Sequencing for the Accurate and Comprehensive Detection of SARS-CoV-2 and Other Respiratory Viruses</title>
		<author>
			<persName><forename type="first">M</forename><surname>Wang</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Small</title>
		<imprint>
			<biblScope unit="volume">16</biblScope>
			<biblScope unit="issue">32</biblScope>
			<date type="published" when="2020">2002169. 2020</date>
		</imprint>
	</monogr>
	<note type="raw_reference">M. Wang et al., &quot;Nanopore Targeted Sequencing for the Accurate and Comprehensive Detection of SARS-CoV-2 and Other Respiratory Viruses,&quot; Small, vol. 16, no. 32, p. 2002169, 2020, doi: 10/ghc6tc.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b39">
	<monogr>
		<title level="m" type="main">He once worked at the Huanan Seafood Market and is now the first recovered patient in Jingzhou. He considers himself lucky</title>
		<author>
			<persName><forename type="first">Y</forename><surname>Zhang</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Y</forename><surname>Huai</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Shan</forename><forename type="middle">S</forename></persName>
		</author>
		<ptr target="https://news.qq.com/rain/a/20200129A0GMVD00" />
		<imprint>
			<date type="published" when="2025-03-14">Mar. 14, 2025</date>
		</imprint>
	</monogr>
	<note>腾讯网 (Tencent)</note>
	<note type="raw_reference">Zhang Y., Huai Y., and Shan S., &quot;He once worked at the Huanan Seafood Market and is now the first recovered patient in Jingzhou. He considers himself lucky.,&quot; 腾讯网 (Tencent). Accessed: Mar. 14, 2025. [Online]. Available: https://news.qq.com/rain/a/20200129A0GMVD00</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b40">
	<monogr>
		<title level="m" type="main">Genome-wide data inferring the evolution and population demography of the novel pneumonia coronavirus (SARS-CoV-2)</title>
		<author>
			<persName><forename type="first">B</forename><surname>Fang</surname></persName>
		</author>
		<idno type="DOI">10.1101/2020.03.04.976662</idno>
		<imprint>
			<date type="published" when="2020-05-11">May 11, 2020</date>
		</imprint>
	</monogr>
	<note type="raw_reference">B. Fang et al., &quot;Genome-wide data inferring the evolution and population demography of the novel pneumonia coronavirus (SARS-CoV-2),&quot; May 11, 2020, bioRxiv. doi: 10.1101/2020.03.04.976662.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b41">
	<analytic>
		<title level="a" type="main">A novel SARS-CoV-2 related coronavirus in bats from Cambodia</title>
		<author>
			<persName><forename type="first">D</forename><surname>Delaune</surname></persName>
		</author>
		<idno>doi: 10/gnfwv4</idno>
	</analytic>
	<monogr>
		<title level="j">Nat. Commun</title>
		<imprint>
			<biblScope unit="volume">12</biblScope>
			<biblScope unit="issue">1</biblScope>
			<biblScope unit="page">6563</biblScope>
			<date type="published" when="2021-11">Nov. 2021</date>
		</imprint>
	</monogr>
	<note type="raw_reference">D. Delaune et al., &quot;A novel SARS-CoV-2 related coronavirus in bats from Cambodia,&quot; Nat. Commun., vol. 12, no. 1, p. 6563, Nov. 2021, doi: 10/gnfwv4.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b42">
	<analytic>
		<title level="a" type="main">Individual bat virome analysis reveals co-infection and spillover among bats and virus zoonotic potential</title>
		<author>
			<persName><forename type="first">J</forename><surname>Wang</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Nat. Commun</title>
		<imprint>
			<biblScope unit="volume">14</biblScope>
			<biblScope unit="issue">1</biblScope>
			<biblScope unit="page">4079</biblScope>
			<date type="published" when="2023-07">Jul. 2023</date>
		</imprint>
	</monogr>
	<note type="raw_reference">J. Wang et al., &quot;Individual bat virome analysis reveals co-infection and spillover among bats and virus zoonotic potential,&quot; Nat. Commun., vol. 14, no. 1, p. 4079, Jul. 2023, doi: 10/g89j5v.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b43">
	<analytic>
		<title level="a" type="main">Phylogeography of horseshoe bat sarbecoviruses in Vietnam and neighbouring countries. Implications for the origins of SARS-CoV and SARS-CoV-2</title>
		<author>
			<persName><forename type="first">A</forename><surname>Hassanin</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Mol. Ecol</title>
		<imprint>
			<biblScope unit="volume">33</biblScope>
			<biblScope unit="issue">18</biblScope>
			<biblScope unit="page">2024</biblScope>
		</imprint>
	</monogr>
	<note type="raw_reference">A. Hassanin et al., &quot;Phylogeography of horseshoe bat sarbecoviruses in Vietnam and neighbouring countries. Implications for the origins of SARS-CoV and SARS-CoV-2,&quot; Mol. Ecol., vol. 33, no. 18, p. e17486, 2024, doi: 10/g82t4f.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b44">
	<monogr>
		<title level="m" type="main">Bayesian phylodynamic inference of multi-type population trajectories using genomic data</title>
		<author>
			<persName><forename type="first">T</forename><forename type="middle">G</forename><surname>Vaughan</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Stadler</surname></persName>
		</author>
		<idno type="DOI">10.1101/2024.11.26.625381</idno>
		<imprint>
			<date type="published" when="2024-01">Dec. 01, 2024</date>
		</imprint>
	</monogr>
	<note type="raw_reference">T. G. Vaughan and T. Stadler, &quot;Bayesian phylodynamic inference of multi-type population trajectories using genomic data,&quot; Dec. 01, 2024, bioRxiv. doi: 10.1101/2024.11.26.625381.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b45">
	<monogr>
		<title level="m" type="main">The mutation rate of SARS-CoV-2 is highly variable between sites and is influenced by sequence context, genomic region, and RNA structure</title>
		<author>
			<persName><forename type="first">H</forename><forename type="middle">K</forename><surname>Haddox</surname></persName>
		</author>
		<idno type="DOI">10.1101/2025.01.07.631013</idno>
		<imprint>
			<date type="published" when="2025-08">Jan. 08, 2025</date>
		</imprint>
	</monogr>
	<note type="raw_reference">H. K. Haddox et al., &quot;The mutation rate of SARS-CoV-2 is highly variable between sites and is influenced by sequence context, genomic region, and RNA structure,&quot; Jan. 08, 2025, bioRxiv. doi: 10.1101/2025.01.07.631013.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b46">
	<analytic>
		<title level="a" type="main">Viral and host factors related to the clinical outcome of COVID-19</title>
		<author>
			<persName><forename type="first">X</forename><surname>Zhang</surname></persName>
		</author>
		<idno type="DOI">10.1038/s41586-020-2355-0</idno>
	</analytic>
	<monogr>
		<title level="j">Nature</title>
		<imprint>
			<biblScope unit="volume">583</biblScope>
			<biblScope unit="issue">7816</biblScope>
			<biblScope unit="page" from="437" to="440" />
			<date type="published" when="2020-07">Jul. 2020</date>
		</imprint>
	</monogr>
	<note type="raw_reference">X. Zhang et al., &quot;Viral and host factors related to the clinical outcome of COVID-19,&quot; Nature, vol. 583, no. 7816, pp. 437-440, Jul. 2020, doi: 10.1038/s41586-020-2355-0.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b47">
	<analytic>
		<title level="a" type="main">The prevalence of antibodies to SARS-CoV-2 among blood donors in China</title>
		<author>
			<persName><forename type="first">L</forename><surname>Chang</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Nat. Commun</title>
		<imprint>
			<biblScope unit="volume">12</biblScope>
			<biblScope unit="issue">1</biblScope>
			<biblScope unit="page">1383</biblScope>
			<date type="published" when="2021-03">Mar. 2021</date>
		</imprint>
	</monogr>
	<note type="raw_reference">L. Chang et al., &quot;The prevalence of antibodies to SARS-CoV-2 among blood donors in China,&quot; Nat. Commun., vol. 12, no. 1, p. 1383, Mar. 2021, doi: 10/g9bpff.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b48">
	<analytic>
		<title level="a" type="main">Serosurvey for SARS-CoV-2 among blood donors in Wuhan</title>
		<author>
			<persName><forename type="first">L</forename><surname>Chang</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Protein Cell</title>
		<imprint>
			<biblScope unit="volume">14</biblScope>
			<biblScope unit="issue">1</biblScope>
			<biblScope unit="page" from="28" to="36" />
			<date type="published" when="2019-12">September to December 2019. Jan. 2023</date>
		</imprint>
	</monogr>
	<note type="raw_reference">L. Chang et al., &quot;Serosurvey for SARS-CoV-2 among blood donors in Wuhan, China from September to December 2019,&quot; Protein Cell, vol. 14, no. 1, pp. 28-36, Jan. 2023, doi: 10/g9bpfg.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b49">
	<analytic>
		<title level="a" type="main">Minimap2: pairwise alignment for nucleotide sequences</title>
		<author>
			<persName><forename type="first">H</forename><surname>Li</surname></persName>
		</author>
		<idno>doi: 10/gdhbqt</idno>
	</analytic>
	<monogr>
		<title level="j">Bioinformatics</title>
		<imprint>
			<biblScope unit="volume">34</biblScope>
			<biblScope unit="issue">18</biblScope>
			<biblScope unit="page" from="3094" to="3100" />
			<date type="published" when="2018-09">Sep. 2018</date>
		</imprint>
	</monogr>
	<note type="raw_reference">H. Li, &quot;Minimap2: pairwise alignment for nucleotide sequences,&quot; Bioinformatics, vol. 34, no. 18, pp. 3094-3100, Sep. 2018, doi: 10/gdhbqt.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b50">
	<analytic>
		<title level="a" type="main">ViralConsensus: a fast and memory-efficient tool for calling viral consensus genome sequences directly from read alignment data</title>
		<author>
			<persName><forename type="first">N</forename><surname>Moshiri</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Bioinformatics</title>
		<imprint>
			<biblScope unit="volume">39</biblScope>
			<biblScope unit="issue">5</biblScope>
			<date type="published" when="2023-05">May 2023</date>
		</imprint>
	</monogr>
	<note type="raw_reference">N. Moshiri, &quot;ViralConsensus: a fast and memory-efficient tool for calling viral consensus genome sequences directly from read alignment data,&quot; Bioinformatics, vol. 39, no. 5, p. btad317, May 2023, doi: 10/g9bpfs.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b51">
	<analytic>
		<title level="a" type="main">WHO clarifies details of early covid patients in Wuhan after errors in virus report</title>
		<author>
			<persName><forename type="first">E</forename><surname>Dou</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><surname>Rauhala</surname></persName>
		</author>
		<ptr target="https://www.washingtonpost.com/world/asia_pacific/covid-wuhan-outbreak-who/2021/07/15/" />
	</analytic>
	<monogr>
		<title level="s">The Washington Post</title>
		<imprint>
			<biblScope unit="page">47</biblScope>
			<date type="published" when="2021-07-15">Jul. 15, 2021. Mar. 14, 2025. 51e7e8a6-e2c6-11</date>
		</imprint>
	</monogr>
	<note type="raw_reference">E. Dou and E. Rauhala, &quot;WHO clarifies details of early covid patients in Wuhan after errors in virus report,&quot; The Washington Post, Jul. 15, 2021. Accessed: Mar. 14, 2025. [Online]. Available: https://www.washingtonpost.com/world/asia_pacific/covid-wuhan-outbreak-who/2021/07/15/51e7e8a6-e2c6-11 eb-88c5-4fd6382c47cb_story.html</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b52">
	<analytic>
		<title level="a" type="main">Nextstrain: real-time tracking of pathogen evolution</title>
		<author>
			<persName><forename type="first">J</forename><surname>Hadfield</surname></persName>
		</author>
		<idno>doi: 10/gdkbqx</idno>
	</analytic>
	<monogr>
		<title level="j">Bioinformatics</title>
		<imprint>
			<biblScope unit="volume">34</biblScope>
			<biblScope unit="issue">23</biblScope>
			<biblScope unit="page" from="4121" to="4123" />
			<date type="published" when="2018-12">Dec. 2018</date>
		</imprint>
	</monogr>
	<note type="raw_reference">J. Hadfield et al., &quot;Nextstrain: real-time tracking of pathogen evolution,&quot; Bioinformatics, vol. 34, no. 23, pp. 4121-4123, Dec. 2018, doi: 10/gdkbqx.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b53">
	<analytic>
		<title level="a" type="main">Nextclade: clade assignment, mutation calling and quality control for viral genomes</title>
		<author>
			<persName><forename type="first">I</forename><surname>Aksamentov</surname></persName>
		</author>
		<author>
			<persName><forename type="first">C</forename><surname>Roemer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">E</forename><forename type="middle">B</forename><surname>Hodcroft</surname></persName>
		</author>
		<author>
			<persName><forename type="first">R</forename><forename type="middle">A</forename><surname>Neher</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">J. Open Source Softw</title>
		<imprint>
			<biblScope unit="volume">6</biblScope>
			<biblScope unit="issue">67</biblScope>
			<biblScope unit="page">3773</biblScope>
			<date type="published" when="2021-11">Nov. 2021</date>
		</imprint>
	</monogr>
	<note type="raw_reference">I. Aksamentov, C. Roemer, E. B. Hodcroft, and R. A. Neher, &quot;Nextclade: clade assignment, mutation calling and quality control for viral genomes,&quot; J. Open Source Softw., vol. 6, no. 67, p. 3773, Nov. 2021, doi: 10/gpgs4p.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b54">
	<analytic>
		<title level="a" type="main">MAFFT Multiple Sequence Alignment Software Version 7: Improvements in Performance and Usability</title>
		<author>
			<persName><forename type="first">K</forename><surname>Katoh</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">M</forename><surname>Standley</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Mol. Biol. Evol</title>
		<imprint>
			<biblScope unit="volume">30</biblScope>
			<biblScope unit="issue">4</biblScope>
			<biblScope unit="page" from="4" to="9" />
			<date type="published" when="2013-04">Apr. 2013</date>
		</imprint>
	</monogr>
	<note type="raw_reference">K. Katoh and D. M. Standley, &quot;MAFFT Multiple Sequence Alignment Software Version 7: Improvements in Performance and Usability,&quot; Mol. Biol. Evol., vol. 30, no. 4, pp. 772-780, Apr. 2013, doi: 10/f4qwp9.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b55">
	<analytic>
		<title level="a" type="main">IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era</title>
		<author>
			<persName><forename type="first">B</forename><forename type="middle">Q</forename><surname>Minh</surname></persName>
		</author>
		<idno>doi: 10/ggkxzj</idno>
	</analytic>
	<monogr>
		<title level="j">Mol. Biol. Evol</title>
		<imprint>
			<biblScope unit="volume">37</biblScope>
			<biblScope unit="issue">5</biblScope>
			<biblScope unit="page" from="1530" to="1534" />
			<date type="published" when="2020-05">May 2020</date>
		</imprint>
	</monogr>
	<note type="raw_reference">B. Q. Minh et al., &quot;IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era,&quot; Mol. Biol. Evol., vol. 37, no. 5, pp. 1530-1534, May 2020, doi: 10/ggkxzj.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b56">
	<analytic>
		<title level="a" type="main">ETE 3: Reconstruction, Analysis, and Visualization of Phylogenomic Data</title>
		<author>
			<persName><forename type="first">J</forename><surname>Huerta-Cepas</surname></persName>
		</author>
		<author>
			<persName><forename type="first">F</forename><surname>Serra</surname></persName>
		</author>
		<author>
			<persName><forename type="first">P</forename><surname>Bork</surname></persName>
		</author>
		<idno>doi: 10/gfzpph</idno>
	</analytic>
	<monogr>
		<title level="j">Mol. Biol. Evol</title>
		<imprint>
			<biblScope unit="volume">33</biblScope>
			<biblScope unit="issue">6</biblScope>
			<biblScope unit="page" from="1635" to="1638" />
			<date type="published" when="2016-06">Jun. 2016</date>
		</imprint>
	</monogr>
	<note type="raw_reference">J. Huerta-Cepas, F. Serra, and P. Bork, &quot;ETE 3: Reconstruction, Analysis, and Visualization of Phylogenomic Data,&quot; Mol. Biol. Evol., vol. 33, no. 6, pp. 1635-1638, Jun. 2016, doi: 10/gfzpph.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b57">
	<monogr>
		<title level="m" type="main">A global phylogeny of SARS-CoV-2 sequences from GISAID</title>
		<author>
			<persName><forename type="first">R</forename><surname>Lanfear</surname></persName>
		</author>
		<idno type="DOI">10.5281/zenodo.4289383</idno>
		<imprint>
			<date type="published" when="2020-11-24">Nov. 24, 2020</date>
		</imprint>
	</monogr>
	<note type="raw_reference">R. Lanfear, A global phylogeny of SARS-CoV-2 sequences from GISAID. (Nov. 24, 2020). Zenodo. doi: 10.5281/zenodo.4289383.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b58">
	<analytic>
		<title level="a" type="main">Ultrafast Sample placement on Existing tRees (UShER) enables real-time phylogenetics for the SARS-CoV-2 pandemic</title>
		<author>
			<persName><forename type="first">Y</forename><surname>Turakhia</surname></persName>
		</author>
		<idno>doi: 10/gjx423</idno>
	</analytic>
	<monogr>
		<title level="j">Nat. Genet</title>
		<imprint>
			<biblScope unit="volume">53</biblScope>
			<biblScope unit="issue">6</biblScope>
			<biblScope unit="page" from="809" to="816" />
			<date type="published" when="2021-06">Jun. 2021</date>
		</imprint>
	</monogr>
	<note type="raw_reference">Y. Turakhia et al., &quot;Ultrafast Sample placement on Existing tRees (UShER) enables real-time phylogenetics for the SARS-CoV-2 pandemic,&quot; Nat. Genet., vol. 53, no. 6, pp. 809-816, Jun. 2021, doi: 10/gjx423.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b59">
	<analytic>
		<title level="a" type="main">Pandemic-scale phylogenomics reveals the SARS-CoV-2 recombination landscape</title>
		<author>
			<persName><forename type="first">Y</forename><surname>Turakhia</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Nature</title>
		<imprint>
			<biblScope unit="volume">609</biblScope>
			<biblScope unit="issue">7929</biblScope>
			<biblScope unit="page" from="994" to="997" />
			<date type="published" when="2022-09">Sep. 2022</date>
		</imprint>
	</monogr>
	<note type="raw_reference">Y. Turakhia et al., &quot;Pandemic-scale phylogenomics reveals the SARS-CoV-2 recombination landscape,&quot; Nature, vol. 609, no. 7929, pp. 994-997, Sep. 2022, doi: 10/gvm7xb.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b60">
	<analytic>
		<title level="a" type="main">Phylodynamics with Migration: A Computational Framework to Quantify Population Structure from Genomic Data</title>
		<author>
			<persName><forename type="first">D</forename><surname>Kühnert</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Stadler</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><forename type="middle">G</forename><surname>Vaughan</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><forename type="middle">J</forename><surname>Drummond</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Mol. Biol. Evol</title>
		<imprint>
			<biblScope unit="volume">33</biblScope>
			<biblScope unit="issue">8</biblScope>
			<biblScope unit="page" from="2102" to="2116" />
			<date type="published" when="2016-08">Aug. 2016</date>
		</imprint>
	</monogr>
	<note type="raw_reference">D. Kühnert, T. Stadler, T. G. Vaughan, and A. J. Drummond, &quot;Phylodynamics with Migration: A Computational Framework to Quantify Population Structure from Genomic Data,&quot; Mol. Biol. Evol., vol. 33, no. 8, pp. 2102-2116, Aug. 2016, doi: 10/f8v3gf.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b61">
	<analytic>
		<title level="a" type="main">Robust Phylodynamic Analysis of Genetic Sequencing Data from Structured Populations</title>
		<author>
			<persName><forename type="first">J</forename><surname>Scire</surname></persName>
		</author>
		<author>
			<persName><forename type="first">J</forename><surname>Barido-Sottani</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Kühnert</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><forename type="middle">G</forename><surname>Vaughan</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Stadler</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Viruses</title>
		<imprint>
			<biblScope unit="volume">14</biblScope>
			<biblScope unit="issue">8</biblScope>
			<date type="published" when="2022-08">Aug. 2022</date>
			<pubPlace>Art</pubPlace>
		</imprint>
	</monogr>
	<note type="raw_reference">J. Scire, J. Barido-Sottani, D. Kühnert, T. G. Vaughan, and T. Stadler, &quot;Robust Phylodynamic Analysis of Genetic Sequencing Data from Structured Populations,&quot; Viruses, vol. 14, no. 8, Art. no. 8, Aug. 2022, doi: 10/g7vrxr.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b62">
	<analytic>
		<title level="a" type="main">BEAST 2.5: An advanced software platform for Bayesian evolutionary analysis</title>
		<author>
			<persName><forename type="first">R</forename><surname>Bouckaert</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">PLOS Comput. Biol</title>
		<imprint>
			<biblScope unit="volume">15</biblScope>
			<biblScope unit="issue">4</biblScope>
			<biblScope unit="page" from="24" to="28" />
			<date type="published" when="2019">1006650, Apr. 2019</date>
		</imprint>
	</monogr>
	<note type="raw_reference">R. Bouckaert et al., &quot;BEAST 2.5: An advanced software platform for Bayesian evolutionary analysis,&quot; PLOS Comput. Biol., vol. 15, no. 4, p. e1006650, Apr. 2019, doi: 10/gg24p8.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b63">
	<analytic>
		<title level="a" type="main">Decomposing the sources of SARS-CoV-2 fitness variation in the United States</title>
		<author>
			<persName><forename type="first">L</forename><surname>Kepler</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Hamins-Puertolas</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><forename type="middle">A</forename><surname>Rasmussen</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Virus Evol</title>
		<imprint>
			<biblScope unit="volume">7</biblScope>
			<biblScope unit="issue">2</biblScope>
			<biblScope unit="page">4</biblScope>
			<date type="published" when="2021-12">Dec. 2021</date>
		</imprint>
	</monogr>
	<note type="raw_reference">L. Kepler, M. Hamins-Puertolas, and D. A. Rasmussen, &quot;Decomposing the sources of SARS-CoV-2 fitness variation in the United States,&quot; Virus Evol., vol. 7, no. 2, p. veab073, Dec. 2021, doi: 10/gsw8v4.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b64">
	<monogr>
		<title level="m" type="main">HIPSTR: highest independent posterior subtree reconstruction in TreeAnnotator X</title>
		<author>
			<persName><forename type="first">G</forename><surname>Baele</surname></persName>
		</author>
		<idno type="DOI">10.1101/2024.12.08.627395</idno>
		<imprint>
			<date type="published" when="2024-12-10">Dec. 10, 2024</date>
		</imprint>
	</monogr>
	<note type="raw_reference">G. Baele et al., &quot;HIPSTR: highest independent posterior subtree reconstruction in TreeAnnotator X,&quot; Dec. 10, 2024, bioRxiv. doi: 10.1101/2024.12.08.627395.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b65">
	<analytic>
		<title level="a" type="main">TreeSwift: A massively scalable Python tree package</title>
		<author>
			<persName><forename type="first">N</forename><surname>Moshiri</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">SoftwareX</title>
		<imprint>
			<biblScope unit="volume">11</biblScope>
			<biblScope unit="page">85</biblScope>
			<date type="published" when="2020-01">Jan. 2020</date>
		</imprint>
	</monogr>
	<note type="raw_reference">N. Moshiri, &quot;TreeSwift: A massively scalable Python tree package,&quot; SoftwareX, vol. 11, p. 100436, Jan. 2020, doi: 10/g6w85m.</note>
</biblStruct>

<biblStruct status="extracted" xml:id="b66">
	<analytic>
		<title level="a" type="main">Posterior Summarization in Bayesian Phylogenetics Using Tracer 1.7</title>
		<author>
			<persName><forename type="first">A</forename><surname>Rambaut</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><forename type="middle">J</forename><surname>Drummond</surname></persName>
		</author>
		<author>
			<persName><forename type="first">D</forename><surname>Xie</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Baele</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><forename type="middle">A</forename><surname>Suchard</surname></persName>
		</author>
		<idno>doi: 10</idno>
	</analytic>
	<monogr>
		<title level="j">Syst. Biol</title>
		<imprint>
			<biblScope unit="volume">67</biblScope>
			<biblScope unit="issue">5</biblScope>
			<biblScope unit="page" from="901" to="904" />
			<date type="published" when="2018-09">Sep. 2018</date>
		</imprint>
	</monogr>
	<note type="raw_reference">A. Rambaut, A. J. Drummond, D. Xie, G. Baele, and M. A. Suchard, &quot;Posterior Summarization in Bayesian Phylogenetics Using Tracer 1.7,&quot; Syst. Biol., vol. 67, no. 5, pp. 901-904, Sep. 2018, doi: 10/gdf52n.</note>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
