TY - JOUR
T1 - The curse of the uncultured fungus
AU - Abarenkov, Kessy
AU - Kristiansson, Erik
AU - Ryberg, Martin
AU - Nogal-Prata, Sandra
AU - Gomez-Martinez, Daniela
AU - Stueer-Patowsky, Katrin
AU - Jansson, Tobias
AU - Polme, Sergei
AU - Ghobad-Nejhad, Masoomeh
AU - Corcoll, Natalia
AU - Scharn, Ruud
AU - Sanchez-Garcia, Marisol
AU - Khomich, Maryia
AU - Wurzbacher, Christian
AU - Nilsson, R. Henrik
PY - 2022
Y1 - 2022
N2 - The international DNA sequence databases abound in fungal sequences not annotated beyond the kingdom level, typically bearing names such as "uncultured fungus". These sequences beget lowresolution mycological results and invite further deposition of similarly poorly annotated entries. What do these sequences represent? This study uses a 767,918-sequence corpus of public full-length that represent truly unidentifiable fungal taxa - and what proportion of them that would have deposition. Our results suggest that more than 70% of these sequences would have been trivial to identify to at least the order/family level at the time of sequence deposition, hinting that factors other than poor availability of relevant reference sequences explain the low-resolution names. We speculate that researchers' perceived lack of time and lack of insight into the ramifications of this problem are the main explanations for the low-resolution names. We were surprised to find that more than a fifth of these sequences seem to have been deposited by mycologists rather than researchers unfamiliar with the consequences of poorly annotated fungal sequences in molecular repositories. The proportion of these needlessly poorly annotated sequences does not decline over time, suggesting that this problem must not be left unchecked.
AB - The international DNA sequence databases abound in fungal sequences not annotated beyond the kingdom level, typically bearing names such as "uncultured fungus". These sequences beget lowresolution mycological results and invite further deposition of similarly poorly annotated entries. What do these sequences represent? This study uses a 767,918-sequence corpus of public full-length that represent truly unidentifiable fungal taxa - and what proportion of them that would have deposition. Our results suggest that more than 70% of these sequences would have been trivial to identify to at least the order/family level at the time of sequence deposition, hinting that factors other than poor availability of relevant reference sequences explain the low-resolution names. We speculate that researchers' perceived lack of time and lack of insight into the ramifications of this problem are the main explanations for the low-resolution names. We were surprised to find that more than a fifth of these sequences seem to have been deposited by mycologists rather than researchers unfamiliar with the consequences of poorly annotated fungal sequences in molecular repositories. The proportion of these needlessly poorly annotated sequences does not decline over time, suggesting that this problem must not be left unchecked.
KW - Data interoperability
KW - data mining
KW - DNA barcoding
KW - scientific practice
KW - species identification
KW - taxonomic
KW - annotation
KW - Data interoperability
KW - data mining
KW - DNA barcoding
KW - scientific practice
KW - species identification
KW - taxonomic
KW - annotation
UR - https://res.slu.se/id/publ/115967
U2 - 10.3897/mycokeys.86.76053
DO - 10.3897/mycokeys.86.76053
M3 - Journal article
SN - 1314-4057
SP - 177
EP - 194
JO - MycoKeys
JF - MycoKeys
IS - 86
ER -