Making floras computable: Audited translation of flora of China as reusable taxonomic knowledge
Bin-Bin Liua,b, Fang-Zhe Aic, Chen-Yang Liaod, Chao Xua,b, Hai-Rui Luoe, Jing Xuane, Yi-Gang Songf,*, Min Lie,*     
a. Key Laboratory of Systematic and Evolutionary Botany/State Key Laboratory of Plant Diversity and Specialty Crops, Institute of Botany, Chinese Academy of Sciences, Beijing 100093, China;
b. China National Botanical Garden, Beijing 100093, China;
c. Beijing Baidu Netcom Science Technology Co., Ltd., No. 10 Shangdi 10th Street, Beijing 100085, China;
d. College of Architecture and Environment, Sichuan University, Chengdu 610065, China;
e. Big Data and AI Biodiversity Conservation Research Center, Institute of Botany, Chinese Academy of Sciences, Beijing 100093, China;
f. Key Laboratory of National Forestry and Grassland Administration on East China Plant Conservation and Utilization, Shanghai Chenshan Botanical Garden, Shanghai 201602, China
Keywords: Floras    Taxonomic descriptions    Large language models    Terminology alignment    Audit    Versioning    

Floras and monographs are among the most information-dense syntheses in taxonomy, but they were written primarily for expert reading rather than computational reuse. In Large Language Model (LLM)-assisted translation, the most consequential failures are often small: a measurement range may be simplified, a unit or special symbol may be changed, a negation cue may be lost, or a name string may no longer be stable enough for linking. Fluent Chinese output therefore does not, by itself, guarantee fidelity to the elements that matter most in downstream reuse. Here, using Rosaceae in Flora of China (FoC) as a worked example, we argue that flora translation should be treated as a controlled release process rather than a one-off language conversion. In this paper, "computable" has a deliberately limited meaning: translated units should remain searchable across languages, checkable for key structure and integrity-critical elements, linkable through normalized scientific names and author strings, and reusable for tasks such as trait-oriented extraction or structured knowledge representation. We outline a minimum blueprint for attaching terminology control, entity normalization, Quality Assurance (QA) flags, correction records, and versioned updates to translated flora units. An overview of the release-oriented workflow and outputs is shown in Fig. 1.

Fig. 1 From flora text to taxonomic knowledge infrastructure. Authoritative flora content (keys and descriptions) is segmented into stable units with persistent identifiers and translated using an LLM guided by domain knowledge resources (a bilingual morphology glossary and a curated person-name/author database). Outputs then undergo terminology alignment and entity normalization (scientific names; author/name strings), followed by automated, reproducible QA that flags reuse-critical failures (numeric ranges and units, negation cues, special symbols and formatting, and key-structure fidelity). The resulting bilingual corpus is released as a versioned, citable artifact with QA logs, changelogs, and documented limitations, enabling downstream reuse (cross-lingual search, trait extraction from descriptive units, and machine-assisted linking to names, literature, and knowledge-graph representations) and iterative community correction.
1. Why translation now requires an infrastructure perspective

Floristic treatments and monographs remain among the most authoritative syntheses in plant systematics. Their value, however, is still mostly carried by prose: descriptions, dichotomous keys, distribution statements, names, and author citations were written for expert use, not as structured records. This becomes a practical limitation when floristic information needs to be searched across languages, compared among taxa, linked to names, traits, occurrences, images, or literature, or reused in downstream data systems. Biodiversity informatics has shown that such reuse depends on explicit identifiers, shared semantics, and interoperable data practices, not on text availability alone (Wieczorek et al., 2012; Wilkinson et al., 2016).

This bottleneck becomes more evident in modern systematics, where reticulation, introgression, incomplete lineage sorting (ILS), and polyploidy often complicate species boundaries and produce conflicts among evidence streams. The consequence is a stronger need for evidence chains that can be inspected, revised, and compared across iterations. Our phylogenomic studies in Rosaceae and Campanulaceae illustrate this point: taxonomic interpretations may depend on sampling, data choices, and model assumptions, and therefore benefit from records that preserve provenance and support structured comparison (Liu et al., 2022; Jin et al., 2023, 2025; Lin et al., 2025; Xie et al., 2025; Xu et al., 2025). In this setting, translation is useful not only for improving access but also for making floristic knowledge traceable across languages.

The appropriate goal is therefore not a single final translated flora. A more useful target is a corpus in which failures are detectable and locatable, changes are tracked and attributable, and repeated terms and names remain linkable through terminology and entity control. This follows the same general logic that has made community standards important in biodiversity informatics: interoperability depends on explicit semantics, identifiers and governance, not on text availability alone (Wieczorek et al., 2012; Wilkinson et al., 2016). Our proposal is not to import an external Natural Language Processing (NLP) agenda into floristics. It is to formalize a need already present in taxonomy itself: authoritative descriptive knowledge should be easier to trace, inspect, correct, and reuse across languages, revisions, and downstream data systems.

2. What changes when large language models enter the pipeline

Large language models change the scale at which floristic text can be translated and checked. They can process large corpora rapidly, but speed is useful only if the output can be treated as a stable and inspectable record. For a translated flora, this means that each unit should remain searchable, linkable to names and traits, and correctable without losing its connection to the source text. These requirements are practical rather than abstract. Tasks such as parsing key structure, extracting diagnostic traits from descriptions, or linking translated names and citation strings to external biodiversity resources often fail because a small meaning-bearing element has been lost, altered, or reformatted, not because the sentence is unreadable.

LLMs therefore change more than translation speed. They also make it possible to apply explicit constraints to large numbers of floristic units (Fig. 1). Keys require structural fidelity, including numbering, couplet pairing, and branching logic. Descriptions depend on ranges, units, negation cues, special symbols, and consistent terminology. Citations and author strings must remain stable enough to support linking to authority files. These requirements do not come from the model; they come from the structure of floristic knowledge itself. The practical implication is that translated output should be evaluated and then improved against these reuse-critical constraints rather than judged solely by surface fluency.

This shifts quality control from a one-time reading judgment to part of the release process. A useful release should report the model and version used, the main prompting rules, terminology, and name resources, machine-checkable QA outputs, a version tag, and a changelog. Such practices are consistent with broader calls for transparent documentation of model-mediated datasets (Bender et al., 2021; Bommasani et al., 2021; Xu et al., 2025). Here, however, the reason is specifically taxonomic: floristic knowledge should remain linkable across languages and updateable across time, without making each correction an isolated rewrite.

3. A workflow that treats floras as structured, versioned records

We propose a minimal and extensible workflow in which a translated flora is handled as a set of versioned records, not as a single block of text (Fig. 1). Each record retains a stable identifier and includes: (ⅰ) the source English text, (ⅱ) the Chinese translation, (ⅲ) terminology mapping annotations or normalization outputs, (ⅳ) entity-normalization fields for scientific names and people or author names, and (ⅴ) machine-readable QA flags. This design is intentionally pragmatic. It does not assume that full semantic parsing is immediately available. Instead, it provides a stable record structure to which additional annotation, correction, and semantic information can be added without losing provenance.

Three resource layers are especially important. The first is a bilingual morphology glossary, which anchors key terms and reduces terminological drift across taxa and translation runs. The second is a curated person-name and author resource, used to stabilize human names and citation strings. This error class is easy to overlook, but it can affect attribution, traceability, and linking. The third is a normalization layer for scientific names, so that names can function as reliable join keys across biodiversity resources (Patterson et al., 2016; Pyle, 2016; Mozzherin et al., 2017).

This workflow builds on earlier work showing that taxonomic and biodiversity texts become more reusable when terminology, entities, and character statements are made explicit. Semantic annotation and parsing studies have shown that morphological descriptions can be segmented into structured statements and reused in phenotypic analyses (Cui, 2012; Thessen et al., 2012; Endara et al., 2018). More recent work has extended this direction through deep learning (DL) and LLM-assisted methods for taxonomic entity recognition, text mining, and trait extraction from ecological and botanical corpora (Le Guillarme and Thuiller, 2022; Farrell et al., 2024; Scheepens et al., 2024; Marcos et al., 2025; Domazetoski et al., 2025; Schmidt-Lebuhn and Knerr, 2025). Together, these studies show that scientific text can be processed and linked more effectively when explicit computational rules and resources are provided.

Our purpose is not to claim novelty for semantic extraction itself, or to replace existing biodiversity NLP efforts. We focus on a narrower problem in floristics and taxonomy: how translated flora content can be released with stable units, QA annotations, controlled terminology, entity normalization, and documented updates. In this sense, the contribution is not just text parsing, but the treatment of LLM-assisted flora translation as a release process, with integrity checks and change logs as part of the output. To support reuse and inspection, we distribute the resulting bilingual, QA-annotated units as a browsable public resource through the iPlant Flora of China portal (www.iplant.cn/foc), consistent with the release model in Fig. 1.

4. Minimum sufficient QA and the role of Rosaceae as a stress test

A Correspondence is not a full translation-evaluation study, and the Rosaceae results reported here are not intended as a comprehensive benchmark across plant lineages, language pairs, or model settings. We use Rosaceae as a worked example and stress test for the release workflow. The aim is minimum sufficient QA: a compact set of automated checks that can run without extensive expert labor, target high-impact and machine-detectable failures, and return diagnostics that can guide correction (Fig. 1). These checks should be read as integrity indicators for selected reuse-critical elements, not as evidence of full semantic equivalence. They do not systematically capture all meaning-bearing constructions in taxonomic descriptions, especially comparative, relative, and vague locative expressions, nor can they exclude all fluent paraphrases that shift meaning. This interpretation is consistent with broader Machine Translation (MT) evaluation practice: aggregate metrics can be informative, but domain-critical constraints often require targeted checks and explicit error taxonomies, especially when meaning-bearing elements are sparse and brittle (Mathur et al., 2020; Rei et al., 2020; Papineni et al., 2002).

Keys and descriptions pose different QA problems. A key encodes decision logic, so small structural errors can be serious even when the translated language is fluent. Numbering, couplet pairing, and branching logic must remain intact. Descriptions are more heterogeneous and contain many numeric and symbolic patterns where minor normalization can affect reuse. For example, a range written as "(2–)3–5(–7) mm" may be flattened to "3–7 mm", parentheses may be dropped, or "×" may be replaced by a word; each change can remove information that downstream parsing expects. A missing negation cue, such as "not", can be even more consequential because it may invert meaning. These errors are small on the page, but they are large failures for structured reuse.

For the Rosaceae slice translated with qwen-max under knowledge-base constraints, we evaluated 1724 key units and 3548 in-scope description units at the line/segment level. In keys, critical-integrity checks, including numbers, ranges, units, negation cues, special symbols, and formatting constraints, passed for 93.97% of units; entity/name integrity passed for 96.00%; and glossary-term presence was detected in 94.49%. In descriptions, critical integrity passed at 75.87%; entity integrity at 80.50%; and glossary-term presence at 46.42%. Here, "pass" means that the relevant integrity conditions were satisfied under rule-based checks designed to detect reuse-breaking changes. It does not mean that the translation is semantically identical to the source text, or that all diagnostically important expressions have been captured. Failures were not evenly distributed. In descriptions, they focused mainly on numeric-range handling and symbol normalization, whereas negation-cue failures were rarer but prioritized because they can change meaning.

These results are best interpreted as a diagnostic profile. They show where correction effort is likely to be most useful. Keys benefit from strict structure-preserving rules, terminology control, and entity normalization. Descriptions require particular attention to numeric ranges, symbols, and domain-aware post-processing. The same profile should not be assumed for other families, translation directions, or model configurations without further testing. The practical lesson is simple: make failures visible, correct them, and carry those corrections into the next release.

5. Beyond readability: reuse, linking, and cumulative correction

Once a flora corpus is translated under explicit constraints and released with machine-readable QA, it can be reused in ways that ordinary parallel text cannot support (Fig. 1). Terminology-controlled text supports cross-taxon search and cross-lingual retrieval. Normalized scientific names and author strings make it easier to link translated content to databases, literature, and other biodiversity resources. QA flags help users distinguish units that passed selected integrity checks from those that require caution or correction. For end users, the value is practical: a translated flora becomes easier to search, extract from, link to, and revise, rather than serving only as readable prose.

This perspective also serves systematics directly. In clades shaped by reticulation and polyploidy, hypotheses about relationships and circumscription may change as sampling, models, and evidence integration improve. Our phylogenomic studies reflect this dynamic (Liu et al., 2022; Jin et al., 2023, 2025; Lin et al., 2025; Xie et al., 2025; Xu et al., 2025). A versioned, QA-annotated flora corpus provides a complementary record of descriptive and key knowledge. It makes updates traceable and supports cumulative corrections, rather than leaving revisions opaque or silent.

Automated pipelines can also amplify biases and propagate errors when documentation is weak. The Machine Learning (ML) community has argued that documentation tools such as dataset "datasheets" and model "cards" provide basic scaffolding for responsible reuse (Gebru et al., 2018; Mitchell et al., 2019). For flora translation, a similar discipline is needed: a concise release checklist should state the scope, constraints, QA design, and known limitations of the corpus. Automated QA should therefore be treated as a first filtering layer for reuse-critical failures. Finer-grained semantic fidelity in descriptive text, including deeper meaning shifts or fluent but inaccurate paraphrases, remains a target for expert audit or future semantic-checking modules.

6. A practical call: community standards for living, versioned floras

LLM-assisted flora translation should be treated as a living release process, not as a one-off conversion (Fig. 1). Each release should be citable, version-tagged, accompanied by a concise QA report and changelog, and anchored to explicit knowledge-base snapshots, including terminology resources and authority files. This approach also gives expert review a clearer role. Expert time can be directed to flagged failures, glossary curation, and correction policies, rather than to undifferentiated proofreading of entire corpora.

Based on the failure modes above, we suggest adopting a minimum reporting standard for LLM-translated floras and monographs (Box 1), analogous in spirit to FAIR-aligned stewardship and biodiversity data standards (Wieczorek et al., 2012; Wilkinson et al., 2016). The aim is practical rather than rhetorical: translated floras should be released with enough information for users to know what was translated, how it was checked, what resources were used, what limitations remain, and how later corrections are recorded.

Box 1 Minimum reporting items for LLM-assisted translation of floras and monographs as versioned releases
To make LLM-translated floras reproducible, comparable across projects, and fit for computational reuse, each release should report the items below. When both are included, report keys and descriptions separately.
1) Source and scope (Required). State the source work (title, edition/version, year), translation direction, and any licensing constraints. Specify what is included and excluded (e.g., keys, descriptions, synonymy/citations, distribution). Define the translation/evaluation unit (e.g., key couplet line; description sentence/segment) and provide stable IDs linking each source unit to its translation and QA record.
2) Model and run settings (Required). Report model provider/name and the exact version (or access date). Report generation settings (e.g., temperature, top-p, max length) and whether runs are deterministic. Summarize how constraints were applied (in-generation vs post-processing) and whether retrieval/grounding was used.
3) Knowledge resources (Required). Document the morphology glossary (provenance, version/date, enforcement/mapping rule). Document the person/author-name resource (provenance, version/date, matching rule, fallback when unmatched). State how scientific name strings were handled, including rules/tools for hybrids, ranks, authorship abbreviations, and uncertain/partial names.
4) Automated QA and metric definitions (Required). List the QA checks and any thresholds, reported separately for keys vs descriptions. At minimum, include checks for: (i) meaning-critical elements (numbers/ranges, units, negation cues, special symbols), (ii) key structure (numbering, couplet pairing, branching consistency), and (iii) entity integrity (scientific names; people/author names). Define what counts as "in scope" and what counts as a "pass," and provide a compact error taxonomy (≈5–10 categories) highlighting dominant high-impact failures. Projects should also state explicitly whether non-terminological but meaning-bearing constructions (e.g., comparative, relative, or vague locative expressions) are included in or excluded from automated checks. Tip: explicitly preserve parse-critical formats (e.g., "(2)35(7) mm", "×", "±") or flag deviations.
5) Corrections and human audit (If applicable; recommended). State how corrections are made (rule-based normalization, targeted re-translation, and/or human edits) and what triggers each pathway. If human audit is performed, report sampling, criteria (terminology accuracy; semantic fidelity including ranges/negation and detection of meaning shifts not captured by automated checks; consistency), and how disagreements were resolved. Indicate whether corrected units are rechecked by the QA suite before release.
6) Release, versioning, and intended use (Required). Provide a version tag and changelog (what changed: model, resources, rules, scope). Preserve provenance needed for reproduction (source IDs, timestamps, configuration). State where the corpus is released and where QA logs/code are archived. Clarify intended use (reading vs computational reuse) and how QA/uncertainty flags should be interpreted by downstream users.

In this sense, living floras should be defined as versioned, provenance-rich, and linkable releases, not as webpages that can be edited without a visible record. A simple community norm would make this operational: every LLM-translated flora or monograph release, whether a paper, dataset, or project deliverable, should include a minimum reporting checklist (Box 1) and a citable version tag with a changelog. To implement this principle in practice, we have released the translated outputs online via the iPlant Flora of China portal, where users can browse volumes and families and access corresponding source materials in a stable, citable setting. Future improvements, including model updates, terminology expansion and correction rules, can then be tied to explicit version tags and visible diffs, rather than silently changing "the translation" as a moving target. Future work should also test the portability of this workflow across additional families, language pairs, and model configurations, so that corpus-specific failure modes can be assessed more systematically.

Acknowledgements

We thank the Qwen (Tongyi Qianwen) platform developed by Alibaba Cloud for providing AI-assisted translation support and efficient multilingual processing (Qwen-MT/Qwen3 series; accessed December 2025). We thank Na He-Ya (Sichuan University) for her valuable contributions to preparing the figure. This work was supported by the National Natural Science Foundation of China (grant numbers 32570240 to B.B.L., 32270216 to B.B.L., and 32000163 to B.B.L.), the Youth Innovation Promotion Association CAS (2023086 to B.B.L.), and the Biological Resources Programme, Chinese Academy of Sciences (CAS-TAX-24-013 to B.B.L.).

CRediT authorship contribution statement

Bin–Bin Liu: Conceptualization, Methodology, Project administration, Supervision, Writing – original draft, Writing – review & editing. Min Li: Conceptualization, Methodology, Project administration, Supervision, Writing – review & editing. Yi-Gang Song: Conceptualization, Supervision, Writing – review & editing. Fang-Zhe Ai: Methodology, Investigation, Resources, Writing – review & editing. Chen-Yang Liao: Visualization and Data curation – review & editing. Chao Xu: Data curation, Methodology, Investigation, Writing – review & editing. Hai-Rui Luo: Data curation. Jing Xuan: Data curation.

Data availability

Aggregate QA outputs, including summary statistics, per-unit pass/fail flags, and a small set of illustrative error examples, are publicly available through the project repository and associated release materials. The translated Flora of China content is publicly accessible via the iPlant FoC portal (www.iplant.cn/foc). A practical Rosaceae-based step-by-step guide for reproducing the QA workflow is also provided through the public project documentation associated with this study. Redistribution of full text outside the portal will follow applicable licensing constraints.

Code availability

All scripts and configuration files required to reproduce the reported QA metrics are available in a public GitHub repository (https://github.com/PhyloAI/foc-translation-infrastructure).

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Appendix A. Supplementary data

Supplementary data to this article can be found online at https://doi.org/10.1016/j.pld.2026.05.015.

References
Bender, E.M., Gebru, T., McMillan-Major, A., et al., 2021. On the dangers of stochastic parrots: can language models be too big?. In: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. Association for Computing Machinery. Virtual Event, Canada, pp. 610–623. https://doi.org/10.1145/3442188.3445922.
Bommasani, R., Hudson, D.A., Adeli, E., et al., 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258v3. https://doi.org/10.48550/arXiv.2108.07258.
Cui H., 2012. CharaParser for fine-grained semantic annotation of organism morphological descriptions. J. Am. Soc. Inf. Sci. Technol., 63: 738-754. DOI:10.1002/asi.22618
Domazetoski V., Kreft H., Bestova H., et al, 2025. Using large language models to extract plant functional traits from unstructured text. Appl. Plant Sci., 13: e70011. DOI:10.1002/aps3.70011
Endara L., Cui H., Burleigh J.G., 2018. Extraction of phenotypic traits from taxonomic descriptions for the tree of life using natural language processing. Appl. Plant Sci., 6: e1035. DOI:10.1002/aps3.1035
Farrell M.J., Le Guillarme N., Brierley L., et al, 2024. The changing landscape of text mining: a review of approaches for ecology and evolution. Proc. Biol. Sci. B-Biol. Sci., 291: 20240423. DOI:10.1098/rspb.2024.0423
Gebru, T., Morgenstern, J., Vecchione, B., et al., 2018. Datasheets for datasets. arXiv preprint arXiv:1803.09010. https://doi.org/10.48550/arXiv.1803.09010.
Jin Z.T., Hodel R.G.J., Ma D.K., et al, 2023. Nightmare or delight: taxonomic circumscription meets reticulate evolution in the phylogenomic era. Mol. Phy-logenet. Evol., 189: 107914. DOI:10.1016/j.ympev.2023.107914
Jin Z.T., Lin X.H., Ma D.K., et al, 2025. Unravelling the web of life: incomplete lineage sorting and hybridisation as primary mechanisms over polyploidisation in the evolutionary dynamics of pear species. Mol. Ecol. Resour., 25: e70029. DOI:10.1111/1755-0998.70029
Le Guillarme N., Thuiller W., 2022. TaxoNERD: deep neural models for the recognition of taxonomic entities in the ecological and evolutionary literature. Methods Ecol. Evol., 13: 625-641. DOI:10.1111/2041-210X.13778
Lin X.H., Xie S.Y., Ma D.K., et al, 2025. Phylogenomic insights into Adenophora and its allies (Campanulaceae): revisiting generic delimitation and hybridization dynamics. Plant Divers., 47: 576-592. DOI:10.1016/j.pld.2025.05.010
Liu B.B., Ren C., Kwak M., et al, 2022. Phylogenomic conflict analyses in the apple genus Malus s.l. reveal widespread hybridization and allopolyploidy driving diversification, with insights into the complex biogeographic history in the Northern Hemisphere. J. Integr. Plant Biol., 64: 1020-1043. DOI:10.1111/jipb.13246
Marcos D., van de Vlasakker R., Athanasiadis I.N., et al, 2025. Fully automatic extraction of morphological traits from the web: utopia or reality?. Appl. Plant Sci., 13: e70005. DOI:10.1002/aps3.70005
Mathur, N., Wei, J., Freitag, M., et al., 2020. Results of the WMT20 metrics shared task. In: Proceedings of the Fifth Conference on Machine Translation. Association for Computational Linguistics, pp. 688–725. Online.
Mitchell, M., Wu, S., Zaldivar, A., et al., 2019. Model cards for model reporting. In: Proceedings of the Conference on Fairness, Accountability, and Transparency. Association for Computing Machinery, Atlanta, GA, USA, pp. 220–229.
Mozzherin D.Y., Myltsev A.A., Patterson D.J., 2017. "gnparser": a powerful parser for scientific names based on parsing Expression Grammar. BMC Bioinformat-ics, 18: 279. DOI:10.1186/s12859-017-1663-3
Papineni, K., Roukos, S., Ward, T., et al., 2002. Bleu: a method for automatic evaluation of machine translation. In: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Philadelphia, Pennsylvania, USA, pp. 311–318.
Patterson D., Mozzherin D., Shorthouse D.P., et al, 2016. Challenges with using names to link digital biodiversity information. Biodivers. Data J., 4: e8080. DOI:10.3897/BDJ.4.e8080
Pyle R.L., 2016. Towards a global names architecture: the future of indexing scientific names. ZooKeys, 550: 261-281. DOI:10.3897/zookeys.550.10009
Rei, R., Stewart, C., Farinha, A.C., et al., 2020. COMET: a neural framework for MT evaluation. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, pp. 2685–2702. Online.
Scheepens D., Millard J., Farrell M., et al, 2024. Large language models help facilitate the automated synthesis of information on potential pest controllers. Methods Ecol. Evol., 15: 1261-1273. DOI:10.1111/2041-210X.14341
Schmidt-Lebuhn A.N., Knerr N., 2025. Large Language Models can extract morphological data from taxonomic descriptions, but their stochastic nature makes automation challenging: a test on Australian Asteraceae. PhytoKeys, 261: 189-210. DOI:10.3897/phytokeys.261.158396
Thessen A.E., Cui H., Mozzherin D., 2012. Applications of natural language processing in biodiversity science. Adv. Bioinf., 2012: 391574. DOI:10.1155/2012/391574
Wieczorek J., Bloom D., Guralnick R., et al, 2012. Darwin Core: an evolving community-developed biodiversity data standard. PLoS One, 7: e29715. DOI:10.1371/journal.pone.0029715
Wilkinson M.D., Dumontier M., Aalbersberg I.J., et al, 2016. The FAIR Guiding Principles for scientific data management and stewardship. Sci. Data, 3: 160018. DOI:10.1038/sdata.2016.18
Xie S.Y., Lin X.H., Wang J.R., et al, 2025. Unraveling evolutionary pathways: allopolyploidization and introgression in polyploid Prunus (Rosaceae). Plant J., 123: e70320. DOI:10.1111/tpj.70320
Xu C., Jin Z.T., Xie S.Y., et al, 2025. Ortho2Web: a workflow for disentangling the roles of hybridization and allopolyploidization in reticulation within Campanulaceae. BMC Biology, 23: 322. DOI:10.1186/s12915-025-02426-1