<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://wiki.biomarkerkb.org/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=MariaKim</id>
	<title>BiomarkerKB Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://wiki.biomarkerkb.org/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=MariaKim"/>
	<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/Special:Contributions/MariaKim"/>
	<updated>2026-09-21T10:04:41Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.42.3</generator>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Frequently_Asked_Questions&amp;diff=310</id>
		<title>Frequently Asked Questions</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Frequently_Asked_Questions&amp;diff=310"/>
		<updated>2026-09-19T17:18:38Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;The frequently asked questions are a collection of user questions related to the BiomarkerKB frontend, backend, and data. The answers to these questions contain definition and explanations of terms, such as Single Biomarker or Multicomponent Biomarker. The list of questions is subdivided into questions related to [[#Biomarker FAQs|biomarker]] and [[#General FAQs|general questions]]. You can also use the BiomarkerKB [https://biomarkerkb.org/contact-us/ contact] page to reach out to us with any additional questions or queries.&lt;br /&gt;
&lt;br /&gt;
== General FAQs ==&lt;br /&gt;
=== Where can I find more information on the project? ===&lt;br /&gt;
* The project webpage can be found [https://biomarkerkb.org/about/ here].&lt;br /&gt;
&lt;br /&gt;
=== Is the project repository publicly available? ===&lt;br /&gt;
* You can view all the project repositories [https://github.com/clinical-biomarkers here].&lt;br /&gt;
&lt;br /&gt;
=== What are the biomarker scores and how are scores assigned for the biomarkers? ===&lt;br /&gt;
* The biomarker-score-calculator and default scoring algorithm can be found [https://github.com/clinical-biomarkers/biomarker-score-calculator here]. The biomarker scores can be seen on the full JSON data model responses from the API.&lt;br /&gt;
&lt;br /&gt;
=== Why are some biomarkers assigned a score of 0? ===&lt;br /&gt;
* Biomarkers with a default score of 0 are manually assigned a 0 score and are pending a manual review. The review of the biomarker can include a spot check, full manual quality checking, NLP based methods, and discussions with the submitter/resource. Until the review is complete the biomarker will keep a score of 0 and after the review is complete the biomarker will be scored using the biomarker score calculator tool.&lt;br /&gt;
&lt;br /&gt;
=== How to download all the current dataset files using CLI? ===&lt;br /&gt;
* The BiomarkerKB dataset can be downloaded using the command &amp;lt;code&amp;gt;wget -r -l1 -np -nd -R &amp;quot;index.html*&amp;quot; &amp;lt;nowiki&amp;gt;https://data.biomarkerkb.org/ln2data/releases/data/current/reviewed/&amp;lt;/nowiki&amp;gt;&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Biomarker FAQs ==&lt;br /&gt;
&lt;br /&gt;
=== What is the difference biomarker types and biomarker roles? ===&lt;br /&gt;
&#039;&#039;&#039;Biomarker type&#039;&#039;&#039; refers to the methodology or modality used to measure a biomarker. According to the FDA-NIH BEST glossary, biomarker types include molecular, histologic, radiographic, and physiologic characteristics.&amp;lt;ref&amp;gt;[https://www.fda.gov/drugs/biomarker-qualification-program/about-biomarkers-and-qualification#BEST_Glossary About Biomarkers and Qualification – FDA BEST Glossary]&amp;lt;/ref&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Biomarker role&#039;&#039;&#039; describes the clinical or scientific purpose a biomarker serves. Recognized biomarker roles are:&lt;br /&gt;
&lt;br /&gt;
* Susceptibility/risk&lt;br /&gt;
* Diagnostic&lt;br /&gt;
* Monitoring&lt;br /&gt;
* Prognostic&lt;br /&gt;
* Predictive&lt;br /&gt;
* Pharmacodynamic/response&lt;br /&gt;
* Safety&lt;br /&gt;
&lt;br /&gt;
For further detail on biomarker roles, see the BEST Resource.&amp;lt;ref&amp;gt;[https://www.ncbi.nlm.nih.gov/books/NBK338448/ BEST (Biomarkers, EndpointS, and other Tools) Resource – NCBI Bookshelf]&amp;lt;/ref&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== What is the difference between single-component, multi-component, and composite biomarkers? ===&lt;br /&gt;
Biomarkers are classified into three categories based on how many components they comprise and how those components relate to one another.&lt;br /&gt;
==== Key definitions ====&lt;br /&gt;
; Entity&lt;br /&gt;
: The biological object or concept being assessed. See [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary entity examples in the Biomarker Controlled Vocabulary].&lt;br /&gt;
; Component&lt;br /&gt;
: A single instance of an assessed entity — that is, a single measured analyte.&lt;br /&gt;
==== Single-component biomarker ====&lt;br /&gt;
A single-component biomarker consists of exactly one component (a single measured analyte).&lt;br /&gt;
&lt;br /&gt;
For further detail, see [[Single biomarker]].&lt;br /&gt;
==== Multi-component biomarker ====&lt;br /&gt;
A multi-component biomarker (MCB) is a defined combination or defined set of two or more individual biomarkers whose values, when considered together in a specified way, yield a meaningful result (e.g., a lipid panel of total cholesterol, LDL, HDL, and triglycerides). MCBs have two subtypes:&lt;br /&gt;
* Integrative biomarker&lt;br /&gt;
: An integrative biomarker entity is composed of multiple component measurements that are generated separately, often using different measurement methods or data sources (e.g., imaging measurements, EEG features, or laboratory values) rather than as part of a single omics experiment. All components are known. A defined mathematical or computational algorithm combines these separate measurements into a single quantitative value, score, or index that is interpreted to have a specific biological, clinical, or diagnostic meaning.&lt;br /&gt;
* Pattern biomarker&lt;br /&gt;
: A pattern biomarker entity is composed of a defined set of component measurements that are generated or analyzed together to identify a pattern or signature result. All components may not be known. This may include omics-derived patterns — such as proteomic, metabolomic, glycomic, transcriptomic, or multi-analyte signatures — as well as other multiplex measurement panels. The key feature is that component measurements are interpreted collectively, rather than individually, to produce a biomarker result with a specific meaning.&lt;br /&gt;
&lt;br /&gt;
For further detail, see [[Multi-component biomarker]].&lt;br /&gt;
==== Composite biomarker ====&lt;br /&gt;
A composite biomarker entity is one in which the entity being measured is itself composed of multiple molecular components that together form a single structural or functional unit. Components may be covalently linked (e.g., a glycan attached to a protein) or associated through non-covalent interactions (e.g., protein complexes or RNA–protein complexes). In some cases, the measurement may reflect the composite entity as a whole without establishing which individual component is responsible for the measured signal. The defining feature is that the molecular components are physically associated and are being considered together as one biomarker entity.&lt;br /&gt;
&lt;br /&gt;
For further detail, see [[Multi-entity biomarker]].&lt;br /&gt;
&lt;br /&gt;
==== Notes on Terminology ====&lt;br /&gt;
These definitions are informed by the FDA-NIH Biomarker Working Group terminology&amp;lt;ref&amp;gt;[https://www.ncbi.nlm.nih.gov/books/NBK326791/ BEST Resource: Biomarker Terminology – NCBI Bookshelf]&amp;lt;/ref&amp;gt;&amp;lt;ref&amp;gt;[https://www.ncbi.nlm.nih.gov/books/NBK610679/ BEST Resource: Additional Terminology – NCBI Bookshelf]&amp;lt;/ref&amp;gt; but have been adapted in places to better describe the specific biomarker concepts represented in BiomarkerKB.&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
Daniall Masood, Mariia Kim, Jeet Vora, Robel Kahsay, Patrick McNeely, Sean Kim, Cyrus Chun Hong Au Yeung, Sujeet Kulkarni, Darren A. Natale, Srinivasan Ramachandran, Shakti Gupta, Mano Maurya, Cristian G. Bologa, Thomas S. DeNapoli, Vincent T. Metzger, Praveen Kumar, Nasheath Ahmed, John Erol Evangelista, Nia Lingam, Sean C. Kelly, Jorge L. Sepulveda, Avi Ma’ayan, Jonathan Silverstein, Deanne M. Taylor, Daniel J. Crichton, Ashish Mahabal, Jeremy J. Yang, Christophe G. Lambert, Shankar Subramaniam, Michael Tiemeyer, Rene Ranzinger, Raja Mazumder (2026). &#039;&#039;&#039;&amp;quot;BiomarkerKB: An Integrated Knowledgebase Supporting Biomarker-Centric Exploration of Biomedical Data&amp;quot;&#039;&#039;&#039;. &#039;&#039;Patterns&#039;&#039;, DOI 10.1016/j.patter.2026.101636.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;references /&amp;gt;&lt;br /&gt;
== External links ==&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;BiomarkerKB&#039;&#039;&#039;: https://biomarkerkb.org/&lt;br /&gt;
&amp;lt;div style=&amp;quot;float: right;&amp;quot;&amp;gt;  [[#top|[top]]]&amp;lt;/div&amp;gt;&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=309</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=309"/>
		<updated>2026-09-18T16:53:35Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* assessed_entity_type */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Instructions to submit Biomarker Data ==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted are provided [[Data Submission/Data Upload#Headers|below]].&lt;br /&gt;
&lt;br /&gt;
# Create a TSV file with the agreed upon fields which correspond to the biomarker data model.&lt;br /&gt;
# Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please provide metadata and description on how biomarker data was collected. This is important for adding submitted data to the Biomarker Data page as each dataset needs a BioCompute Object (BCO). Examples of BCOs are available on the [https://data.biomarkerkb.org/BMK_000001 biomarker data page].&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB data model fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
&lt;br /&gt;
To make reporting easier, please use these templates with color-coded fields: orange for mandatory and green for optional.&lt;br /&gt;
* Disease biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/disease_biomarker_template.xlsx Download the Disease Biomarker Template]&lt;br /&gt;
** Example of a disease biomarker: [https://biomarkerkb.org/biomarker/AN5370-8 AN5370-8]&lt;br /&gt;
* Exposure agent biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/exposure_agent_biomarker_template.xlsx Download the Exposure Agent Biomarker Template]&lt;br /&gt;
** Example of an exposure agent biomarker: [https://biomarkerkb.org/biomarker/BMKB151582-1 BMKB151582-1]&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...).&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): sub-index within &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt;. Component counter within each biomarker (integer). To learn the difference between single and multicomponent biomarkers, see [[Single biomarker]] and [[Multicomponent biomarker]].&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): sub-index within &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt;. Entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein. See [[Multi-entity biomarker]].&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt; (optional): taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]] or [[Data Submission/Data Upload#exposure_agent|exposure_agent]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]] or [[Data Submission/Data Upload#exposure_agent_id|exposure_agent_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
Follow the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
&amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow.&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
Only the following entity types are allowed:&lt;br /&gt;
* protein&lt;br /&gt;
* glycan&lt;br /&gt;
* DNA&lt;br /&gt;
* RNA&lt;br /&gt;
* cell&lt;br /&gt;
* lipid&lt;br /&gt;
* image&lt;br /&gt;
* metabolite&lt;br /&gt;
* element&lt;br /&gt;
* gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; should be reported in all lowercase.&lt;br /&gt;
* Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; (from Disease Ontology, MONDO, or SNOMED or NCIt). Format as &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
* Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
If reporting more than one, use separate rows. Example:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt; !! ... !! &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || diagnostic&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || monitoring&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
* Example: feces&lt;br /&gt;
&lt;br /&gt;
If reporting more than one, use separate rows. Example:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt; !! ... !! &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || blood || UBERON:0000178&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || urine || UBERON:0001088&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; in the following column should be from UBERON. Format as &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon. Leave blank if not applicable.&lt;br /&gt;
* Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Format as &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality). Leave blank if not applicable.&lt;br /&gt;
&lt;br /&gt;
* Example: 77354-9&lt;br /&gt;
&lt;br /&gt;
=== evidence ===&lt;br /&gt;
One or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
&lt;br /&gt;
=== evidence_source ===&lt;br /&gt;
Report in the format &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
If reporting more than one, use separate rows. Example:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt; !! ... !! &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || Insert a quote from the paper || PubMed:26243686&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || Insert a quote from the paper || PubMed:25096510&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== annotations ===&lt;br /&gt;
Provide extra annotations from your DCC with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field. For example, relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Controlled_Vocabulary_and_Keywords&amp;diff=297</id>
		<title>Controlled Vocabulary and Keywords</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Controlled_Vocabulary_and_Keywords&amp;diff=297"/>
		<updated>2026-09-17T16:40:18Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;nowiki&amp;gt;---------------------------------------------------------------------------&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB Controlled Vocabulary ==&lt;br /&gt;
Biomarker Knowledgebase Project&lt;br /&gt;
&lt;br /&gt;
This page is no longer maintained. Please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary Biomarker Controlled Vocabulary GitHub page] for the latest information.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;---------------------------------------------------------------------------&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Biomarker Entity Types Controlled Vocabulary File ===&lt;br /&gt;
Description: Controlled vocabulary of biomarker entity types used in BiomarkerKB.&lt;br /&gt;
&lt;br /&gt;
Name:        biomarker_entity_types.txt&lt;br /&gt;
&lt;br /&gt;
Release:     2025_09_02&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;---------------------------------------------------------------------------&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This document lists the biomarker entity types used in the BiomarkerKB curation system. When a biomarker entity is paired with measurable.txt terms (increase, decrease, presence, absence etc.), they form the biomarker. The biomarker entity types are:&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;   * gene&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;   * protein&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;   * metabolite&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;   * glycan&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;   * DNA&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;   * RNA&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;   * cell&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;   * lipid&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;   * image&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Each entry consists of the following line codes:&lt;br /&gt;
&lt;br /&gt;
 ----  ----------  ---------------------------  ---------------------------------&lt;br /&gt;
&lt;br /&gt;
 Code  Name        Content                      Occurrences&lt;br /&gt;
&lt;br /&gt;
 ----  ----------  ---------------------------  ---------------------------------&lt;br /&gt;
&lt;br /&gt;
 ID    Identifier  Biomarker entity name        Once; starts an entry&lt;br /&gt;
&lt;br /&gt;
 AC    Accession   Unique identifier (BM-xxxx)  Once&lt;br /&gt;
&lt;br /&gt;
 DE    Definition  Definition/description       Once; line wrapping possible&lt;br /&gt;
&lt;br /&gt;
 HI    Hierarchy                                           Optional; zero or more&lt;br /&gt;
&lt;br /&gt;
 SY    Synonym                                          Optional; zero or more&lt;br /&gt;
&lt;br /&gt;
 EQ    Equivalent  External identifier with   Optional; zero or more&lt;br /&gt;
&lt;br /&gt;
                             same meaning&lt;br /&gt;
&lt;br /&gt;
 EX    Example     A biomarker that uses    Optional; zero or once; line&lt;br /&gt;
&lt;br /&gt;
                            the indicated entity         wrapping possible&lt;br /&gt;
&lt;br /&gt;
 NT    Notes                                                Optional; zero or once; line&lt;br /&gt;
&lt;br /&gt;
                                                                    wrapping possible&lt;br /&gt;
&lt;br /&gt;
 //    Terminator  End of entry                      Once; ends an entry&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;---------------------------------------------------------------------------&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
ID   sequence&lt;br /&gt;
&lt;br /&gt;
AC   BM-0001&lt;br /&gt;
&lt;br /&gt;
DE   Nucleic acid sequences that function as units of heredity and which code for&lt;br /&gt;
&lt;br /&gt;
DE   the basic instructions for the development, reproduction, and maintenance&lt;br /&gt;
&lt;br /&gt;
DE   of an organism.&lt;br /&gt;
&lt;br /&gt;
HI   sequence&lt;br /&gt;
&lt;br /&gt;
EQ   MESH:D008969; Molecular Sequence Data&lt;br /&gt;
&lt;br /&gt;
EQ   SO:0000704; gene&lt;br /&gt;
&lt;br /&gt;
EX   BRCA1; &amp;lt;nowiki&amp;gt;https://biomarkerkb.org/canonical/AN4824&amp;lt;/nowiki&amp;gt;   &lt;br /&gt;
&lt;br /&gt;
//&lt;br /&gt;
&lt;br /&gt;
ID   protein&lt;br /&gt;
&lt;br /&gt;
AC   BM-0002&lt;br /&gt;
&lt;br /&gt;
DE   Amino acid chain formed by ribosome-mediated translation of a genetically-&lt;br /&gt;
&lt;br /&gt;
DE   encoded mRNA, and any post-translationally modified derivatives.&lt;br /&gt;
&lt;br /&gt;
HI   protein&lt;br /&gt;
&lt;br /&gt;
EQ   MESH:D011506; Proteins&lt;br /&gt;
&lt;br /&gt;
EQ   PR:000000001; protein&lt;br /&gt;
&lt;br /&gt;
EX   Interleukin-6; &amp;lt;nowiki&amp;gt;https://biomarkerkb.org/canonical/AN6278&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
//&lt;br /&gt;
&lt;br /&gt;
ID   metabolite&lt;br /&gt;
&lt;br /&gt;
AC   BM-0003&lt;br /&gt;
&lt;br /&gt;
DE   A small molecule involved in metabolism.&lt;br /&gt;
&lt;br /&gt;
HI   metabolite&lt;br /&gt;
&lt;br /&gt;
EQ   CHEBI:25212; metabolite&lt;br /&gt;
&lt;br /&gt;
EX   UREA; &amp;lt;nowiki&amp;gt;https://biomarkerkb.org/canonical/AN6341&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
NT &#039;&#039;  MESH term for metabolite being requested.&#039;&#039;  &lt;br /&gt;
&lt;br /&gt;
//&lt;br /&gt;
&lt;br /&gt;
ID   glycan&lt;br /&gt;
&lt;br /&gt;
AC   BM-0004&lt;br /&gt;
&lt;br /&gt;
DE   Any oligosaccharide, polysaccharide or their derivatives consisting of &lt;br /&gt;
&lt;br /&gt;
DE   monosaccharides or monosaccharide derivatives linked by glycosidic bonds.&lt;br /&gt;
&lt;br /&gt;
HI   glycan&lt;br /&gt;
&lt;br /&gt;
EQ   CHEBI:167559&lt;br /&gt;
&lt;br /&gt;
EX   N-glycan; &amp;lt;nowiki&amp;gt;https://biomarkerkb.org/canonical/AN6729&amp;lt;/nowiki&amp;gt; &lt;br /&gt;
&lt;br /&gt;
//&lt;br /&gt;
&lt;br /&gt;
ID   DNA&lt;br /&gt;
&lt;br /&gt;
AC   BM-0005&lt;br /&gt;
&lt;br /&gt;
DE   A polymer of deoxyribose-containing nucleotides linked by phosphodiester bonds.&lt;br /&gt;
&lt;br /&gt;
HI   DNA&lt;br /&gt;
&lt;br /&gt;
EQ   MESH:D004247; DNA&lt;br /&gt;
&lt;br /&gt;
EX   cfDNA; &amp;lt;nowiki&amp;gt;https://biomarkerkb.org/canonical/AN6380&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
//&lt;br /&gt;
&lt;br /&gt;
ID   RNA&lt;br /&gt;
&lt;br /&gt;
AC   BM-00096&lt;br /&gt;
&lt;br /&gt;
DE   A polymer of ribose-containing nucleotides linked by phosphodiester bonds.&lt;br /&gt;
&lt;br /&gt;
HI   RNA&lt;br /&gt;
&lt;br /&gt;
EQ   MESH:D012313; RNA&lt;br /&gt;
&lt;br /&gt;
EX   miRNA-21; &amp;lt;nowiki&amp;gt;https://biomarkerkb.org/canonical/AN6498&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
//&lt;br /&gt;
&lt;br /&gt;
ID   cell&lt;br /&gt;
&lt;br /&gt;
AC   BM-0007&lt;br /&gt;
&lt;br /&gt;
DE   An organism (or part thereof) that is a maximally connected compartment  &lt;br /&gt;
&lt;br /&gt;
DE   surrounded by a plasma membrane.&lt;br /&gt;
&lt;br /&gt;
HI   cell&lt;br /&gt;
&lt;br /&gt;
EQ   MESH:D002477; Cells&lt;br /&gt;
&lt;br /&gt;
EQ   CL:0000000; cell&lt;br /&gt;
&lt;br /&gt;
EX   WBC; &amp;lt;nowiki&amp;gt;https://biomarkerkb.org/canonical/AN6280&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
//&lt;br /&gt;
&lt;br /&gt;
ID   lipid&lt;br /&gt;
&lt;br /&gt;
AC   BM-0008&lt;br /&gt;
&lt;br /&gt;
DE   A class of biomolecules including fats, oils, and certain hormones.&lt;br /&gt;
&lt;br /&gt;
HI   lipid&lt;br /&gt;
&lt;br /&gt;
EQ   MESH:D008055; Lipids&lt;br /&gt;
&lt;br /&gt;
EQ   CHEBI:18059; lipid&lt;br /&gt;
&lt;br /&gt;
EX   very long chain fatty acid; &amp;lt;nowiki&amp;gt;https://biomarkerkb.org/canonical/AN6187&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
//&lt;br /&gt;
&lt;br /&gt;
ID   image  &lt;br /&gt;
&lt;br /&gt;
AC   BM-0018  &lt;br /&gt;
&lt;br /&gt;
DE   Any visual display of structural or functional patterns of organs or tissues for &lt;br /&gt;
&lt;br /&gt;
DE   clinical evaluation.  &lt;br /&gt;
&lt;br /&gt;
HI   image  &lt;br /&gt;
&lt;br /&gt;
EQ   NCIT:C48179; Image&lt;br /&gt;
&lt;br /&gt;
EX   presence of ground-glass opacity; &lt;br /&gt;
&lt;br /&gt;
EX   &amp;lt;nowiki&amp;gt;https://data.oncomx.org/allbiomarkers/biomarker/A0076&amp;lt;/nowiki&amp;gt; &lt;br /&gt;
&lt;br /&gt;
//&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Biomarker Reporting Terms Controlled Vocabulary File ===&lt;br /&gt;
         Biomarker Knowledgebase Project&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;---------------------------------------------------------------------------&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Description: Controlled vocabulary of standardized terms used to describe&lt;br /&gt;
&lt;br /&gt;
biomarker status, behavior, or detection in curated data entries.&lt;br /&gt;
&lt;br /&gt;
Name:        measured.txt&lt;br /&gt;
&lt;br /&gt;
Release:     2025_09_09&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;---------------------------------------------------------------------------&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
 This document lists the controlled reporting terms used to describe&lt;br /&gt;
&lt;br /&gt;
 biomarkers in the BiomarkerKB curation system. These nouns are used&lt;br /&gt;
&lt;br /&gt;
 in fields like `biomarker` to standardize natural language descriptions and&lt;br /&gt;
&lt;br /&gt;
 can be modified using adjectives (“significant”, “gradual”, etc).&lt;br /&gt;
&lt;br /&gt;
 These terms are classified into categories such as:&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;      * increased&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
 Each entry consists of the following line codes:&lt;br /&gt;
&lt;br /&gt;
   ----  ----------  ---------------------------  ---------------------------------&lt;br /&gt;
&lt;br /&gt;
 Code  Name        Content                      Occurrences&lt;br /&gt;
&lt;br /&gt;
 ----  ----------  ---------------------------  ---------------------------------&lt;br /&gt;
&lt;br /&gt;
 ID    Identifier  Reporting term               Once; starts an entry&lt;br /&gt;
&lt;br /&gt;
 AC    Accession   Unique identifier (RT-xxxx)  Once&lt;br /&gt;
&lt;br /&gt;
 DE    Definition  Definition/description       Once; line wrapping possible&lt;br /&gt;
&lt;br /&gt;
 HI    Hierarchy                                Optional; zero or more&lt;br /&gt;
&lt;br /&gt;
 SY    Synonym                                  Optional; zero or more&lt;br /&gt;
&lt;br /&gt;
 EQ    Equivalent  External identifier with     Optional; zero or more&lt;br /&gt;
&lt;br /&gt;
                     same meaning&lt;br /&gt;
&lt;br /&gt;
 EX    Example                                  Optional; zero or once&lt;br /&gt;
&lt;br /&gt;
 NT    Notes                                    Optional; zero or once; line&lt;br /&gt;
&lt;br /&gt;
                                                  wrapping possible&lt;br /&gt;
&lt;br /&gt;
 //    Terminator  End of entry                 Once; ends an entry&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;---------------------------------------------------------------------------&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
ID   increased&lt;br /&gt;
&lt;br /&gt;
AC   RT-0001&lt;br /&gt;
&lt;br /&gt;
DE   Indicates an assessed biomarker entity level is higher than normal by&lt;br /&gt;
&lt;br /&gt;
DE   a clinically relevant degree.&lt;br /&gt;
&lt;br /&gt;
HI   abundance&lt;br /&gt;
&lt;br /&gt;
EQ   PATO:0002300; increased quality&lt;br /&gt;
&lt;br /&gt;
EX   increased IL6 level; &amp;lt;nowiki&amp;gt;https://biomarkerkb.org/canonical/AN6278&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
//&lt;br /&gt;
&lt;br /&gt;
ID   decreased&lt;br /&gt;
&lt;br /&gt;
AC   RT-0002&lt;br /&gt;
&lt;br /&gt;
DE   Indicates an assessed biomarker entity level is lower than normal by&lt;br /&gt;
&lt;br /&gt;
DE   a clinically relevant degree.&lt;br /&gt;
&lt;br /&gt;
HI   abundance&lt;br /&gt;
&lt;br /&gt;
EQ   PATO:0002301; decreased quality&lt;br /&gt;
&lt;br /&gt;
EX   decreased albumin level; &amp;lt;nowiki&amp;gt;https://biomarkerkb.org/canonical/AN6351&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
//&lt;br /&gt;
&lt;br /&gt;
ID   presence of&lt;br /&gt;
&lt;br /&gt;
AC   RT-0006&lt;br /&gt;
&lt;br /&gt;
DE   Indicates that the assessed entity is present.&lt;br /&gt;
&lt;br /&gt;
HI   presence&lt;br /&gt;
&lt;br /&gt;
EX   presence of rs180177132 mutation in PALB2; &lt;br /&gt;
&lt;br /&gt;
EX   &amp;lt;nowiki&amp;gt;https://biomarkerkb.org/canonical/AV9568&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
//&lt;br /&gt;
&lt;br /&gt;
ID   difference&lt;br /&gt;
&lt;br /&gt;
AC   RT-0007&lt;br /&gt;
&lt;br /&gt;
DE   A statistic that is a subtraction of one quantity from another.&lt;br /&gt;
&lt;br /&gt;
HI   expression&lt;br /&gt;
&lt;br /&gt;
EQ   STATO:0000613; difference&lt;br /&gt;
&lt;br /&gt;
EX   differential expression of TNMD; &amp;lt;nowiki&amp;gt;https://biomarkerkb.org/canonical/BB1486&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
//&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=296</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=296"/>
		<updated>2026-08-27T14:09:18Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Instructions to submit Biomarker Data ==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted are provided [[Data Submission/Data Upload#Headers|below]].&lt;br /&gt;
&lt;br /&gt;
# Create a TSV file with the agreed upon fields which correspond to the biomarker data model.&lt;br /&gt;
# Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please provide metadata and description on how biomarker data was collected. This is important for adding submitted data to the Biomarker Data page as each dataset needs a BioCompute Object (BCO). Examples of BCOs are available on the [https://data.biomarkerkb.org/BMK_000001 biomarker data page].&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB data model fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
&lt;br /&gt;
To make reporting easier, please use these templates with color-coded fields: orange for mandatory and green for optional.&lt;br /&gt;
* Disease biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/disease_biomarker_template.xlsx Download the Disease Biomarker Template]&lt;br /&gt;
** Example of a disease biomarker: [https://biomarkerkb.org/biomarker/AN5370-8 AN5370-8]&lt;br /&gt;
* Exposure agent biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/exposure_agent_biomarker_template.xlsx Download the Exposure Agent Biomarker Template]&lt;br /&gt;
** Example of an exposure agent biomarker: [https://biomarkerkb.org/biomarker/BMKB151582-1 BMKB151582-1]&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...).&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): sub-index within &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt;. Component counter within each biomarker (integer). To learn the difference between single and multicomponent biomarkers, see [[Single biomarker]] and [[Multicomponent biomarker]].&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): sub-index within &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt;. Entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein. See [[Multi-entity biomarker]].&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt; (optional): taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]] or [[Data Submission/Data Upload#exposure_agent|exposure_agent]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]] or [[Data Submission/Data Upload#exposure_agent_id|exposure_agent_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
Follow the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
&amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow.&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; should be reported in all lowercase.&lt;br /&gt;
* Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; (from Disease Ontology, MONDO, or SNOMED or NCIt). Format as &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
* Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
If reporting more than one, use separate rows. Example:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt; !! ... !! &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || diagnostic&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || monitoring&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
* Example: feces&lt;br /&gt;
&lt;br /&gt;
If reporting more than one, use separate rows. Example:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt; !! ... !! &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || blood || UBERON:0000178&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || urine || UBERON:0001088&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; in the following column should be from UBERON. Format as &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon. Leave blank if not applicable.&lt;br /&gt;
* Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Format as &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality). Leave blank if not applicable.&lt;br /&gt;
&lt;br /&gt;
* Example: 77354-9&lt;br /&gt;
&lt;br /&gt;
=== evidence ===&lt;br /&gt;
One or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
&lt;br /&gt;
=== evidence_source ===&lt;br /&gt;
Report in the format &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
If reporting more than one, use separate rows. Example:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt; !! ... !! &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || Insert a quote from the paper || PubMed:26243686&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || Insert a quote from the paper || PubMed:25096510&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== annotations ===&lt;br /&gt;
Provide extra annotations from your DCC with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field. For example, relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Multi-entity_biomarker&amp;diff=295</id>
		<title>Multi-entity biomarker</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Multi-entity_biomarker&amp;diff=295"/>
		<updated>2026-08-27T13:35:57Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Also known as complex biomarker.&lt;br /&gt;
&lt;br /&gt;
== Single-component multi-entity biomarker ==&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
== Multicomponent multi-entity biomarker ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
See also [https://biomarkerkb.org/biomarker/BMKB203684-81 BMKB203684-81].&lt;br /&gt;
&lt;br /&gt;
This article is a stub. You can help us by adding missing information.&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Multicomponent_biomarker&amp;diff=294</id>
		<title>Multicomponent biomarker</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Multicomponent_biomarker&amp;diff=294"/>
		<updated>2026-08-27T13:33:38Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4.&lt;br /&gt;
&lt;br /&gt;
== Multicomponent single-entity biomarker ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
See also [https://biomarkerkb.org/biomarker/AN6165-1#Biomarker-Components AN6165-1].&lt;br /&gt;
&lt;br /&gt;
== Multicomponent multi-entity biomarker ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Single_biomarker&amp;diff=293</id>
		<title>Single biomarker</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Single_biomarker&amp;diff=293"/>
		<updated>2026-08-27T13:32:29Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;A single-component, single-entity biomarker.&lt;br /&gt;
&lt;br /&gt;
== Examples ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt; !! ...&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || Increased 3-hydroxy-3-methylglutaric acid || ...&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
See also [https://biomarkerkb.org/biomarker/AN8635-2#Biomarker-Components AN8635-2].&lt;br /&gt;
&lt;br /&gt;
This article is a stub. You can help us by adding missing information.&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Frequently_Asked_Questions&amp;diff=292</id>
		<title>Frequently Asked Questions</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Frequently_Asked_Questions&amp;diff=292"/>
		<updated>2026-08-27T13:28:52Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* How to download all the &amp;#039;current&amp;#039; dataset files using CLI? */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;The frequently asked questions are a collection of user questions related to the BiomarkerKB frontend, backend, and data. The answers to these questions contain definition and explanations of terms, such as Single Biomarker or Multicomponent Biomarker. The list of questions is subdivided into questions related to [[#Biomarker FAQs|biomarker]] and [[#General FAQs|general questions]]. You can also use the BiomarkerKB [https://biomarkerkb.org/contact-us/ contact] page to reach out to us with any additional questions or queries.&lt;br /&gt;
&lt;br /&gt;
== General FAQs ==&lt;br /&gt;
=== Where can I find more information on the project? ===&lt;br /&gt;
* The project webpage can be found [https://biomarkerkb.org/about/ here].&lt;br /&gt;
&lt;br /&gt;
=== Is the project repository publicly available? ===&lt;br /&gt;
* You can view all the project repositories [https://github.com/clinical-biomarkers here].&lt;br /&gt;
&lt;br /&gt;
=== What are the biomarker scores and how are scores assigned for the biomarkers? ===&lt;br /&gt;
* The biomarker-score-calculator and default scoring algorithm can be found [https://github.com/clinical-biomarkers/biomarker-score-calculator here]. The biomarker scores can be seen on the full JSON data model responses from the API.&lt;br /&gt;
&lt;br /&gt;
=== Why are some biomarkers assigned a score of 0? ===&lt;br /&gt;
* Biomarkers with a default score of 0 are manually assigned a 0 score and are pending a manual review. The review of the biomarker can include a spot check, full manual quality checking, NLP based methods, and discussions with the submitter/resource. Until the review is complete the biomarker will keep a score of 0 and after the review is complete the biomarker will be scored using the biomarker score calculator tool.&lt;br /&gt;
&lt;br /&gt;
=== How to download all the current dataset files using CLI? ===&lt;br /&gt;
* The BiomarkerKB dataset can be downloaded using the command &amp;lt;code&amp;gt;wget -r -l1 -np -nd -R &amp;quot;index.html*&amp;quot; &amp;lt;nowiki&amp;gt;https://data.biomarkerkb.org/ln2data/releases/data/current/reviewed/&amp;lt;/nowiki&amp;gt;&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Biomarker FAQs ==&lt;br /&gt;
=== What is the difference between single, multicomponent, and multi-entity biomarkers? ===&lt;br /&gt;
* A single biomarker consists of exactly one component and one entity (a single measured analyte).&lt;br /&gt;
* A multicomponent biomarker combines two or more independently measured components assessed together, each having a distinct &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (e.g., a lipid panel of total cholesterol, LDL, HDL, and triglycerides).&lt;br /&gt;
* A multi-entity biomarker occurs when at least one of those components is itself made up of multiple molecular entities sharing the same &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; but distinct &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; values (e.g., a glycan bound to a protein)&lt;br /&gt;
* A biomarker can be multicomponent, multi-entity, or both at once.&lt;br /&gt;
&lt;br /&gt;
For more details, see [[Single biomarker]], [[Multicomponent biomarker]], and [[Multi-entity biomarker]].&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
Daniall Masood, Mariia Kim, Jeet Vora, Robel Kahsay, Patrick McNeely, Sean Kim, Cyrus Chun Hong Au Yeung, Sujeet Kulkarni, Darren A. Natale, Srinivasan Ramachandran, Shakti Gupta, Mano Maurya, Cristian G. Bologa, Thomas S. DeNapoli, Vincent T. Metzger, Praveen Kumar, Nasheath Ahmed, John Erol Evangelista, Nia Lingam, Sean C. Kelly, Jorge L. Sepulveda, Avi Ma’ayan, Jonathan Silverstein, Deanne M. Taylor, Daniel J. Crichton, Ashish Mahabal, Jeremy J. Yang, Christophe G. Lambert, Shankar Subramaniam, Michael Tiemeyer, Rene Ranzinger, Raja Mazumder (2026). &#039;&#039;&#039;&amp;quot;BiomarkerKB: An Integrated Knowledgebase Supporting Biomarker-Centric Exploration of Biomedical Data&amp;quot;&#039;&#039;&#039;. &#039;&#039;Patterns&#039;&#039;, DOI 10.1016/j.patter.2026.101636.&lt;br /&gt;
&lt;br /&gt;
== External links ==&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;BiomarkerKB&#039;&#039;&#039;: https://biomarkerkb.org/&lt;br /&gt;
&amp;lt;div style=&amp;quot;float: right;&amp;quot;&amp;gt;  [[#top|[top]]]&amp;lt;/div&amp;gt;&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Frequently_Asked_Questions&amp;diff=291</id>
		<title>Frequently Asked Questions</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Frequently_Asked_Questions&amp;diff=291"/>
		<updated>2026-08-27T13:27:25Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;The frequently asked questions are a collection of user questions related to the BiomarkerKB frontend, backend, and data. The answers to these questions contain definition and explanations of terms, such as Single Biomarker or Multicomponent Biomarker. The list of questions is subdivided into questions related to [[#Biomarker FAQs|biomarker]] and [[#General FAQs|general questions]]. You can also use the BiomarkerKB [https://biomarkerkb.org/contact-us/ contact] page to reach out to us with any additional questions or queries.&lt;br /&gt;
&lt;br /&gt;
== General FAQs ==&lt;br /&gt;
=== Where can I find more information on the project? ===&lt;br /&gt;
* The project webpage can be found [https://biomarkerkb.org/about/ here].&lt;br /&gt;
&lt;br /&gt;
=== Is the project repository publicly available? ===&lt;br /&gt;
* You can view all the project repositories [https://github.com/clinical-biomarkers here].&lt;br /&gt;
&lt;br /&gt;
=== What are the biomarker scores and how are scores assigned for the biomarkers? ===&lt;br /&gt;
* The biomarker-score-calculator and default scoring algorithm can be found [https://github.com/clinical-biomarkers/biomarker-score-calculator here]. The biomarker scores can be seen on the full JSON data model responses from the API.&lt;br /&gt;
&lt;br /&gt;
=== Why are some biomarkers assigned a score of 0? ===&lt;br /&gt;
* Biomarkers with a default score of 0 are manually assigned a 0 score and are pending a manual review. The review of the biomarker can include a spot check, full manual quality checking, NLP based methods, and discussions with the submitter/resource. Until the review is complete the biomarker will keep a score of 0 and after the review is complete the biomarker will be scored using the biomarker score calculator tool.&lt;br /&gt;
&lt;br /&gt;
=== How to download all the &#039;current&#039; dataset files using CLI? ===&lt;br /&gt;
* The BiomarkerKB dataset can be downloaded using the command &amp;lt;code&amp;gt;wget -r -l1 -np -nd -R &amp;quot;index.html*&amp;quot; &amp;lt;nowiki&amp;gt;https://data.biomarkerkb.org/ln2data/releases/data/current/reviewed/&amp;lt;/nowiki&amp;gt;&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Biomarker FAQs ==&lt;br /&gt;
=== What is the difference between single, multicomponent, and multi-entity biomarkers? ===&lt;br /&gt;
* A single biomarker consists of exactly one component and one entity (a single measured analyte).&lt;br /&gt;
* A multicomponent biomarker combines two or more independently measured components assessed together, each having a distinct &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (e.g., a lipid panel of total cholesterol, LDL, HDL, and triglycerides).&lt;br /&gt;
* A multi-entity biomarker occurs when at least one of those components is itself made up of multiple molecular entities sharing the same &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; but distinct &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; values (e.g., a glycan bound to a protein)&lt;br /&gt;
* A biomarker can be multicomponent, multi-entity, or both at once.&lt;br /&gt;
&lt;br /&gt;
For more details, see [[Single biomarker]], [[Multicomponent biomarker]], and [[Multi-entity biomarker]].&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
Daniall Masood, Mariia Kim, Jeet Vora, Robel Kahsay, Patrick McNeely, Sean Kim, Cyrus Chun Hong Au Yeung, Sujeet Kulkarni, Darren A. Natale, Srinivasan Ramachandran, Shakti Gupta, Mano Maurya, Cristian G. Bologa, Thomas S. DeNapoli, Vincent T. Metzger, Praveen Kumar, Nasheath Ahmed, John Erol Evangelista, Nia Lingam, Sean C. Kelly, Jorge L. Sepulveda, Avi Ma’ayan, Jonathan Silverstein, Deanne M. Taylor, Daniel J. Crichton, Ashish Mahabal, Jeremy J. Yang, Christophe G. Lambert, Shankar Subramaniam, Michael Tiemeyer, Rene Ranzinger, Raja Mazumder (2026). &#039;&#039;&#039;&amp;quot;BiomarkerKB: An Integrated Knowledgebase Supporting Biomarker-Centric Exploration of Biomedical Data&amp;quot;&#039;&#039;&#039;. &#039;&#039;Patterns&#039;&#039;, DOI 10.1016/j.patter.2026.101636.&lt;br /&gt;
&lt;br /&gt;
== External links ==&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;BiomarkerKB&#039;&#039;&#039;: https://biomarkerkb.org/&lt;br /&gt;
&amp;lt;div style=&amp;quot;float: right;&amp;quot;&amp;gt;  [[#top|[top]]]&amp;lt;/div&amp;gt;&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Frequently_Asked_Questions&amp;diff=290</id>
		<title>Frequently Asked Questions</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Frequently_Asked_Questions&amp;diff=290"/>
		<updated>2026-08-27T12:08:08Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;The frequently asked questions are a collection of user questions related to the BiomarkerKB frontend, backend, and data. The answers to these questions contain definition and explanations of terms, such as Single Biomarker or Multicomponent Biomarker. The list of questions is subdivided into questions related to [[#Biomarker FAQs|biomarker]] and [[#General FAQs|general questions]]. You can also use the BiomarkerKB [https://biomarkerkb.org/contact-us/ contact] page to reach out to us with any additional questions or queries.&lt;br /&gt;
&lt;br /&gt;
== General FAQs ==&lt;br /&gt;
=== Where can I find more information on the project? ===&lt;br /&gt;
* The project webpage can be found [https://biomarkerkb.org/about/ here].&lt;br /&gt;
&lt;br /&gt;
=== Is the project repository publicly available? ===&lt;br /&gt;
* You can view all the project repositories [https://github.com/clinical-biomarkers here].&lt;br /&gt;
&lt;br /&gt;
=== What are the biomarker scores and how are scores assigned for the biomarkers? ===&lt;br /&gt;
* The biomarker-score-calculator and default scoring algorithm can be found [https://github.com/clinical-biomarkers/biomarker-score-calculator here]. The biomarker scores can be seen on the full JSON data model responses from the API.&lt;br /&gt;
&lt;br /&gt;
=== Why are some biomarkers assigned a score of 0? ===&lt;br /&gt;
* Biomarkers with a default score of 0 are manually assigned a 0 score and are pending a manual review. The review of the biomarker can include a spot check, full manual quality checking, NLP based methods, and discussions with the submitter/resource. Until the review is complete the biomarker will keep a score of 0 and after the review is complete the biomarker will be scored using the biomarker score calculator tool.&lt;br /&gt;
&lt;br /&gt;
=== How to download all the &#039;current&#039; dataset files using CLI? ===&lt;br /&gt;
* The BiomarkerKB dataset can be downloaded using the command &amp;lt;code&amp;gt;wget -r -l1 -np -nd -R &amp;quot;index.html*&amp;quot; &amp;lt;nowiki&amp;gt;https://data.biomarkerkb.org/ln2data/releases/data/current/reviewed/&amp;lt;/nowiki&amp;gt;&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Biomarker FAQs ==&lt;br /&gt;
=== What is a single biomarker? ===&lt;br /&gt;
See [[Single biomarker]].&lt;br /&gt;
&lt;br /&gt;
=== What is a multicomponent biomarker? ===&lt;br /&gt;
See [[Multicomponent biomarker]].&lt;br /&gt;
&lt;br /&gt;
=== What is a multi-entity biomarker? ===&lt;br /&gt;
See [[Multi-entity biomarker]].&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
Daniall Masood, Mariia Kim, Jeet Vora, Robel Kahsay, Patrick McNeely, Sean Kim, Cyrus Chun Hong Au Yeung, Sujeet Kulkarni, Darren A. Natale, Srinivasan Ramachandran, Shakti Gupta, Mano Maurya, Cristian G. Bologa, Thomas S. DeNapoli, Vincent T. Metzger, Praveen Kumar, Nasheath Ahmed, John Erol Evangelista, Nia Lingam, Sean C. Kelly, Jorge L. Sepulveda, Avi Ma’ayan, Jonathan Silverstein, Deanne M. Taylor, Daniel J. Crichton, Ashish Mahabal, Jeremy J. Yang, Christophe G. Lambert, Shankar Subramaniam, Michael Tiemeyer, Rene Ranzinger, Raja Mazumder (2026). &#039;&#039;&#039;&amp;quot;BiomarkerKB: An Integrated Knowledgebase Supporting Biomarker-Centric Exploration of Biomedical Data&amp;quot;&#039;&#039;&#039;. &#039;&#039;Patterns&#039;&#039;, DOI 10.1016/j.patter.2026.101636.&lt;br /&gt;
&lt;br /&gt;
== External links ==&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;BiomarkerKB&#039;&#039;&#039;: https://biomarkerkb.org/&lt;br /&gt;
&amp;lt;div style=&amp;quot;float: right;&amp;quot;&amp;gt;  [[#top|[top]]]&amp;lt;/div&amp;gt;&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=289</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=289"/>
		<updated>2026-08-27T12:05:34Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Instructions to submit Biomarker Data ==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted are provided [[Data Submission/Data Upload#Headers|below]].&lt;br /&gt;
&lt;br /&gt;
# Create a TSV file with the agreed upon fields which correspond to the biomarker data model.&lt;br /&gt;
# Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please provide metadata and description on how biomarker data was collected. This is important for adding submitted data to the Biomarker Data page as each dataset needs a BioCompute Object (BCO). Examples of BCOs are available on the [https://data.biomarkerkb.org/BMK_000001 biomarker data page].&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB data model fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
&lt;br /&gt;
To make reporting easier, please use these templates with color-coded fields: orange for mandatory and green for optional.&lt;br /&gt;
* Disease Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/disease_biomarker_template.xlsx Download the Disease Biomarker Template]&lt;br /&gt;
* Exposure Agent Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/exposure_agent_biomarker_template.xlsx Download the Exposure Agent Biomarker Template]&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...).&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): sub-index within &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt;. Component counter within each biomarker (integer). To learn the difference between single and multicomponent biomarkers, see [[Single biomarker]] and [[Multicomponent biomarker]].&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): sub-index within &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt;. Entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein. See [[Multi-entity biomarker]].&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]] or [[Data Submission/Data Upload#exposure_agent|exposure_agent]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]] or [[Data Submission/Data Upload#exposure_agent_id|exposure_agent_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
Follow the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
&amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow.&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; should be reported in all lowercase.&lt;br /&gt;
* Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; (from Disease Ontology, MONDO, or SNOMED or NCIt). Format as &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
* Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
If reporting more than one, use separate rows. Example:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt; !! ... !! &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || diagnostic&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || monitoring&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
* Example: feces&lt;br /&gt;
&lt;br /&gt;
If reporting more than one, use separate rows. Example:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt; !! ... !! &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || blood || UBERON:0000178&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || urine || UBERON:0001088&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; in the following column should be from UBERON. Format as &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon. Leave blank if not applicable.&lt;br /&gt;
* Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Format as &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality). Leave blank if not applicable.&lt;br /&gt;
&lt;br /&gt;
* Example: 77354-9&lt;br /&gt;
&lt;br /&gt;
=== evidence ===&lt;br /&gt;
One or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
&lt;br /&gt;
=== evidence_source ===&lt;br /&gt;
Report in the format &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
If reporting more than one, use separate rows. Example:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt; !! ... !! &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || Insert a quote from the paper || PubMed:26243686&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || Insert a quote from the paper || PubMed:25096510&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== annotations ===&lt;br /&gt;
Provide extra annotations from your DCC with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field. For example, relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Multicomponent_biomarker&amp;diff=288</id>
		<title>Multicomponent biomarker</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Multicomponent_biomarker&amp;diff=288"/>
		<updated>2026-08-27T12:04:57Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4.&lt;br /&gt;
&lt;br /&gt;
== Multicomponent single-entity biomarker ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Multicomponent multi-entity biomarker ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Multi-entity_biomarker&amp;diff=287</id>
		<title>Multi-entity biomarker</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Multi-entity_biomarker&amp;diff=287"/>
		<updated>2026-08-27T12:04:24Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Also known as complex biomarker.&lt;br /&gt;
&lt;br /&gt;
== Single-component multi-entity biomarker ==&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
== Multicomponent multi-entity biomarker ==&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
This article is a stub. You can help us by adding missing information.&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Multicomponent_biomarker&amp;diff=286</id>
		<title>Multicomponent biomarker</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Multicomponent_biomarker&amp;diff=286"/>
		<updated>2026-08-27T11:57:25Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4.&lt;br /&gt;
&lt;br /&gt;
Example of a multicomponent single-entity biomarker:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Single_biomarker&amp;diff=285</id>
		<title>Single biomarker</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Single_biomarker&amp;diff=285"/>
		<updated>2026-08-27T11:54:53Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;A single-component, single-entity biomarker.&lt;br /&gt;
&lt;br /&gt;
Example:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt; !! ...&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || Increased 3-hydroxy-3-methylglutaric acid || ...&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
This article is a stub. You can help us by adding missing information.&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=284</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=284"/>
		<updated>2026-08-27T11:52:41Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Instructions to submit Biomarker Data ==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted are provided [[Data Submission/Data Upload#Headers|below]].&lt;br /&gt;
&lt;br /&gt;
# Create a TSV file with the agreed upon fields which correspond to the biomarker data model.&lt;br /&gt;
# Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please provide metadata and description on how biomarker data was collected. This is important for adding submitted data to the Biomarker Data page as each dataset needs a BioCompute Object (BCO). Examples of BCOs are available on the [https://data.biomarkerkb.org/BMK_000001 biomarker data page].&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB data model fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
&lt;br /&gt;
To make reporting easier, please use these templates with color-coded fields: orange for mandatory and green for optional.&lt;br /&gt;
* Disease Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/disease_biomarker_template.xlsx Download the Disease Biomarker Template]&lt;br /&gt;
* Exposure Agent Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/exposure_agent_biomarker_template.xlsx Download the Exposure Agent Biomarker Template]&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]] or [[Data Submission/Data Upload#exposure_agent|exposure_agent]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]] or [[Data Submission/Data Upload#exposure_agent_id|exposure_agent_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
Follow the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
&amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow.&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; should be reported in all lowercase.&lt;br /&gt;
* Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; (from Disease Ontology, MONDO, or SNOMED or NCIt). Format as &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
* Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
If reporting more than one, use separate rows. Example:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt; !! ... !! &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || diagnostic&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || monitoring&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
* Example: feces&lt;br /&gt;
&lt;br /&gt;
If reporting more than one, use separate rows. Example:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt; !! ... !! &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || blood || UBERON:0000178&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || urine || UBERON:0001088&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; in the following column should be from UBERON. Format as &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon. Leave blank if not applicable.&lt;br /&gt;
* Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Format as &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality). Leave blank if not applicable.&lt;br /&gt;
&lt;br /&gt;
* Example: 77354-9&lt;br /&gt;
&lt;br /&gt;
=== evidence ===&lt;br /&gt;
One or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
&lt;br /&gt;
=== evidence_source ===&lt;br /&gt;
Report in the format &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
If reporting more than one, use separate rows. Example:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt; !! ... !! &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || Insert a quote from the paper || PubMed:26243686&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || Insert a quote from the paper || PubMed:25096510&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== annotations ===&lt;br /&gt;
Provide extra annotations from your DCC with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field. For example, relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=283</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=283"/>
		<updated>2026-08-27T11:16:26Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Instructions to submit Biomarker Data ==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted are provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
&lt;br /&gt;
To make reporting easier, please use these templates with color-coded fields: orange for mandatory and green for optional.&lt;br /&gt;
* Disease Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/disease_biomarker_template.xlsx Download the Disease Biomarker Template]&lt;br /&gt;
* Exposure Agent Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/exposure_agent_biomarker_template.xlsx Download the Exposure Agent Biomarker Template]&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]] or [[Data Submission/Data Upload#exposure_agent|exposure_agent]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]] or [[Data Submission/Data Upload#exposure_agent_id|exposure_agent_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
Follow the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
&amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
* Example: feces&lt;br /&gt;
&lt;br /&gt;
If reporting more than one, use separate rows. Example:&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt; !! ... !! &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || blood || UBERON:0000178&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || increased protein X || ... || urine || UBERON:0001088&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=282</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=282"/>
		<updated>2026-08-27T11:03:14Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* BiomarkerKB dataset datamodel fields */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Instructions to submit Biomarker Data ==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted are provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
&lt;br /&gt;
To make reporting easier, please use these templates with color-coded fields: orange for mandatory and green for optional.&lt;br /&gt;
* Disease Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/disease_biomarker_template.xlsx Download the Disease Biomarker Template]&lt;br /&gt;
* Exposure Agent Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/exposure_agent_biomarker_template.xlsx Download the Exposure Agent Biomarker Template]&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]] or [[Data Submission/Data Upload#exposure_agent|exposure_agent]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]] or [[Data Submission/Data Upload#exposure_agent_id|exposure_agent_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
&amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=281</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=281"/>
		<updated>2026-08-27T11:00:31Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* assessed_biomarker_entity */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Instructions to submit Biomarker Data ==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted are provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
&lt;br /&gt;
Biomarker data must be reported using dedicated templates with color-coded mandatory (orange) and optional (green) columns.&lt;br /&gt;
* Disease Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/disease_biomarker_template.xlsx Download the Disease Biomarker Template]&lt;br /&gt;
* Exposure Agent Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/exposure_agent_biomarker_template.xlsx Download the Exposure Agent Biomarker Template]&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]] or [[Data Submission/Data Upload#exposure_agent|exposure_agent]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]] or [[Data Submission/Data Upload#exposure_agent_id|exposure_agent_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
&amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=280</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=280"/>
		<updated>2026-08-27T11:00:16Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* condition_id */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Instructions to submit Biomarker Data ==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted are provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
&lt;br /&gt;
Biomarker data must be reported using dedicated templates with color-coded mandatory (orange) and optional (green) columns.&lt;br /&gt;
* Disease Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/disease_biomarker_template.xlsx Download the Disease Biomarker Template]&lt;br /&gt;
* Exposure Agent Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/exposure_agent_biomarker_template.xlsx Download the Exposure Agent Biomarker Template]&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]] or [[Data Submission/Data Upload#exposure_agent|exposure_agent]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]] or [[Data Submission/Data Upload#exposure_agent_id|exposure_agent_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=279</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=279"/>
		<updated>2026-08-27T11:00:02Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* condition */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Instructions to submit Biomarker Data ==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted are provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
&lt;br /&gt;
Biomarker data must be reported using dedicated templates with color-coded mandatory (orange) and optional (green) columns.&lt;br /&gt;
* Disease Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/disease_biomarker_template.xlsx Download the Disease Biomarker Template]&lt;br /&gt;
* Exposure Agent Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/exposure_agent_biomarker_template.xlsx Download the Exposure Agent Biomarker Template]&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]] or [[Data Submission/Data Upload#exposure_agent|exposure_agent]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]] or [[Data Submission/Data Upload#exposure_agent_id|exposure_agent_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
&amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=278</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=278"/>
		<updated>2026-08-27T10:59:44Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* specimen_id */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Instructions to submit Biomarker Data ==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted are provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
&lt;br /&gt;
Biomarker data must be reported using dedicated templates with color-coded mandatory (orange) and optional (green) columns.&lt;br /&gt;
* Disease Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/disease_biomarker_template.xlsx Download the Disease Biomarker Template]&lt;br /&gt;
* Exposure Agent Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/exposure_agent_biomarker_template.xlsx Download the Exposure Agent Biomarker Template]&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]] or [[Data Submission/Data Upload#exposure_agent|exposure_agent]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]] or [[Data Submission/Data Upload#exposure_agent_id|exposure_agent_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
&amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=277</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=277"/>
		<updated>2026-08-27T10:59:25Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Instructions to submit Biomarker Data ==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted are provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
&lt;br /&gt;
Biomarker data must be reported using dedicated templates with color-coded mandatory (orange) and optional (green) columns.&lt;br /&gt;
* Disease Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/disease_biomarker_template.xlsx Download the Disease Biomarker Template]&lt;br /&gt;
* Exposure Agent Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/exposure_agent_biomarker_template.xlsx Download the Exposure Agent Biomarker Template]&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]] or [[Data Submission/Data Upload#exposure_agent|exposure_agent]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]] or [[Data Submission/Data Upload#exposure_agent_id|exposure_agent_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Workflow_%26_Data_Model&amp;diff=276</id>
		<title>Data Workflow &amp; Data Model</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Workflow_%26_Data_Model&amp;diff=276"/>
		<updated>2026-08-27T10:59:14Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Data Workflow &amp;amp; Data Model =&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Multi-entity_biomarker&amp;diff=275</id>
		<title>Multi-entity biomarker</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Multi-entity_biomarker&amp;diff=275"/>
		<updated>2026-08-26T21:26:32Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: Created page with &amp;quot;Also known as complex biomarker. This article is a stub.&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Also known as complex biomarker. This article is a stub.&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Multicomponent_biomarker&amp;diff=274</id>
		<title>Multicomponent biomarker</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Multicomponent_biomarker&amp;diff=274"/>
		<updated>2026-08-26T21:25:54Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: Created page with &amp;quot;This article is a stub.&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;This article is a stub.&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Single_biomarker&amp;diff=273</id>
		<title>Single biomarker</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Single_biomarker&amp;diff=273"/>
		<updated>2026-08-26T21:25:15Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: Created page with &amp;quot;This article is a stub.&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;This article is a stub.&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Frequently_Asked_Questions&amp;diff=272</id>
		<title>Frequently Asked Questions</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Frequently_Asked_Questions&amp;diff=272"/>
		<updated>2026-08-26T02:22:58Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;The frequently asked questions are a collection of user questions related to the BiomarkerKB frontend, backend, and data. The answers to these questions contain definition and explanations of terms, such as Single Biomarker or Multicomponent Biomarker. The list of questions is subdivided into questions related to [[#Biomarker FAQs|biomarker]] and [[#General FAQs|general questions]]. You can also use the BiomarkerKB [https://biomarkerkb.org/contact-us/ contact] page to reach out to us with any additional questions or queries.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== General FAQs ==&lt;br /&gt;
&lt;br /&gt;
=== Where can I find more information on the project? ===&lt;br /&gt;
* The project webpage can be found [https://biomarkerkb.org/about/ here].&lt;br /&gt;
&lt;br /&gt;
=== Is the project repository publicly available? ===&lt;br /&gt;
* You can view all the project repositories [https://github.com/clinical-biomarkers here].&lt;br /&gt;
&lt;br /&gt;
=== What are the biomarker scores and how are scores assigned for the biomarkers?===&lt;br /&gt;
* The biomarker-score-calculator and default scoring algorithm can be found [https://github.com/clinical-biomarkers/biomarker-score-calculator here]. The biomarker scores can be seen on the full JSON data model responses from the API.&lt;br /&gt;
&lt;br /&gt;
=== Why are some biomarkers assigned a score of 0?===&lt;br /&gt;
* Biomarkers with a default score of 0 are manually assigned a 0 score and are pending a manual review. The review of the biomarker can include a spot check, full manual quality checking, NLP based methods, and discussions with the submitter/resource. Until the review is complete the biomarker will keep a score of 0 and after the review is complete the biomarker will be scored using the biomarker score calculator tool.&lt;br /&gt;
&lt;br /&gt;
=== How to download all the &#039;current&#039; dataset files using CLI?===&lt;br /&gt;
* The BiomarkerKB dataset can be downloaded using the command &amp;lt;code&amp;gt;wget -r -l1 -np -nd -R &amp;quot;index.html*&amp;quot; &amp;lt;nowiki&amp;gt;https://data.biomarkerkb.org/ln2data/releases/data/current/reviewed/&amp;lt;/nowiki&amp;gt;&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Biomarker FAQs ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
Daniall Masood, Mariia Kim, Jeet Vora, Robel Kahsay, Patrick McNeely, Sean Kim, Cyrus Chun Hong Au Yeung, Sujeet Kulkarni, Darren A. Natale, Srinivasan Ramachandran, Shakti Gupta, Mano Maurya, Cristian G. Bologa, Thomas S. DeNapoli, Vincent T. Metzger, Praveen Kumar, Nasheath Ahmed, John Erol Evangelista, Nia Lingam, Sean C. Kelly, Jorge L. Sepulveda, Avi Ma’ayan, Jonathan Silverstein, Deanne M. Taylor, Daniel J. Crichton, Ashish Mahabal, Jeremy J. Yang, Christophe G. Lambert, Shankar Subramaniam, Michael Tiemeyer, Rene Ranzinger, Raja Mazumder (2026). &#039;&#039;&#039;&amp;quot;BiomarkerKB: An Integrated Knowledgebase Supporting Biomarker-Centric Exploration of Biomedical Data&amp;quot;&#039;&#039;&#039;. &#039;&#039;Patterns&#039;&#039;, DOI 10.1016/j.patter.2026.101636.&lt;br /&gt;
&lt;br /&gt;
== External links ==&lt;br /&gt;
&lt;br /&gt;
*&#039;&#039;&#039;BiomarkerKB&#039;&#039;&#039;: https://biomarkerkb.org/&lt;br /&gt;
&amp;lt;div style=&amp;quot;float: right;&amp;quot;&amp;gt;  [[#top|[top]]]&amp;lt;/div&amp;gt;&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=271</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=271"/>
		<updated>2026-08-26T01:52:48Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* BiomarkerKB dataset datamodel fields */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Instructions to submit Biomarker Data ==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted are provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
&lt;br /&gt;
Biomarker data must be reported using dedicated templates with color-coded mandatory (orange) and optional (green) columns.&lt;br /&gt;
* Disease Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/disease_biomarker_template.xlsx Download the Disease Biomarker Template]&lt;br /&gt;
* Exposure Agent Biomarkers: [https://data.biomarkerkb.org/ln2downloads/templates/current/exposure_agent_biomarker_template.xlsx Download the Exposure Agent Biomarker Template]&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]] or [[Data Submission/Data Upload#exposure_agent|exposure_agent]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]] or [[Data Submission/Data Upload#exposure_agent_id|exposure_agent_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=270</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=270"/>
		<updated>2026-08-26T01:43:41Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Instructions to submit Biomarker Data ==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted are provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out. A template is available at (link TBA with color-coded mandatory and optional columns). There will be two templates: one for disease biomarkers (only condition id required), one for exposure agent biomarkers (either exposure agent or exposure agent id required).&lt;br /&gt;
&lt;br /&gt;
[https://data.biomarkerkb.org/ln2downloads/curator_agent/2026_08_11/biomarkers-formatted.tsv Disease biomarker template] - this link will be replaced, the template will have its dedicated directory.&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]] or [[Data Submission/Data Upload#exposure_agent|exposure_agent]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]] or [[Data Submission/Data Upload#exposure_agent_id|exposure_agent_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=269</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=269"/>
		<updated>2026-08-25T17:09:40Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* Headers */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out. A template is available at (link TBA with color-coded mandatory and optional columns). There will be two templates: one for disease biomarkers (only condition id required), one for exposure agent biomarkers (either exposure agent or exposure agent id required).&lt;br /&gt;
&lt;br /&gt;
[https://data.biomarkerkb.org/ln2downloads/curator_agent/2026_08_11/biomarkers-formatted.tsv Disease biomarker template] - this link will be replaced, the template will have its dedicated directory.&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]] or [[Data Submission/Data Upload#exposure_agent|exposure_agent]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]] or [[Data Submission/Data Upload#exposure_agent_id|exposure_agent_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=268</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=268"/>
		<updated>2026-08-25T17:08:38Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* Headers */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out. A template is available at (link TBA with color-coded mandatory and optional columns). There will be two templates: one for disease biomarkers (only condition id required), one for exposure agent biomarkers (either exposure agent or exposure agent id required).&lt;br /&gt;
&lt;br /&gt;
[https://data.biomarkerkb.org/ln2downloads/curator_agent/2026_08_11/biomarkers-formatted.tsv Disease biomarker template] - this link will be replaced, the template will have its dedicated directory.&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Release_Notes&amp;diff=267</id>
		<title>Data Release Notes</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Release_Notes&amp;diff=267"/>
		<updated>2026-08-20T21:12:17Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* Version 3.5.1 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Versioning Format ==&lt;br /&gt;
The versioning format follows a three-digit structure: X.Y.Z.&lt;br /&gt;
* The first digit (X) changes when a major update is introduced, such as changes in the data model.&lt;br /&gt;
* The second digit (Y) increments when new data is added.&lt;br /&gt;
* The third digit (Z) is updated for bug fixes or minor changes.&lt;br /&gt;
&lt;br /&gt;
== Version 3.5.1 ==&lt;br /&gt;
Planned: Aug 20th, 2026&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* The [https://data.biomarkerkb.org/BMK_000003 BMK_000003] dataset (&amp;lt;code&amp;gt;biomarkers_llm_glycan.tsv&amp;lt;/code&amp;gt;) has been retired and replaced by a new LLM-mined glycan dataset (BioCompute Object TBA).&lt;br /&gt;
&lt;br /&gt;
== Version 3.4.2 ==&lt;br /&gt;
Date: Aug 6th, 2026&lt;br /&gt;
=== Biomarker Knowledge Graph ===&lt;br /&gt;
* Fixed an issue where biomarker labels in the KG did not reflect standardized biomarker terminology.&lt;br /&gt;
* The JSON-to-NT conversion script now uses controlled vocabulary terms for biomarkers instead of original biomarker names.&lt;br /&gt;
* Expanded and corrected edge predicate mappings (&amp;lt;code&amp;gt;biomarker_change&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;specimen_sampled_from&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_role_indicator&amp;lt;/code&amp;gt;) to use standardized 9-digit OBCI URIs; previously missing predicates have been added, resolving cases where not all biomarker types appeared in &amp;quot;Select Relation.&amp;quot;&lt;br /&gt;
=== Biomarker Ontology ===&lt;br /&gt;
* Updated Metadata &amp;gt; Details page: &amp;quot;Biomarker Ontology&amp;quot; now displays as &amp;quot;Ontology for Biomarkers of Clinical Importance (OBCI).&amp;quot;&lt;br /&gt;
* Ontology nodes are now sorted alphanumerically.&lt;br /&gt;
* Updated hierarchy view instructional text to &amp;quot;Please click on a term on left side to explore more.&amp;quot;&lt;br /&gt;
* Published a new hierarchy-display-specific ontology version that resolves the multiple-parent issue.&lt;br /&gt;
&lt;br /&gt;
== Version 3.4.1 ==&lt;br /&gt;
Date: July 23rd, 2026&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* Added the caDSR ([https://cadsr.cancer.gov/onedata/Home.jsp NCI Cancer Data Standards Registry and Repository]) dataset: a pilot set of ~30 manually curated, standardized biomarker entries derived from caDSR permissible values, mapped to the BiomarkerKB schema. Conditions are broadly mapped to general cancer (DOID:162) pending future organ/tissue-specific enrichment. Extraction code and mapping scripts are available on [https://github.com/clinical-biomarkers/biomarker-extraction GitHub].&lt;br /&gt;
=== Backend Updates ===&lt;br /&gt;
* Added a new [https://api.biomarkerkb.org/ Global Search API] (&amp;lt;code&amp;gt;/biomarker/search_global/&amp;lt;/code&amp;gt;) enabling search results to be broken out by section (Biomarker Entity, Biomarker Term, Condition/Disease, Specimen, Cross-References, Biomarker ID).&lt;br /&gt;
* Each result section now includes an &amp;lt;code&amp;gt;order&amp;lt;/code&amp;gt; attribute to control display order on the frontend, standardized as:&lt;br /&gt;
  1. Biomarker Entity&lt;br /&gt;
  2. Biomarker ID&lt;br /&gt;
  3. Biomarker Term&lt;br /&gt;
  4. Condition/Disease&lt;br /&gt;
  5. Specimen&lt;br /&gt;
  6. Cross-References&lt;br /&gt;
* Relabeled &amp;quot;crossref&amp;quot; section to &amp;quot;Cross-References&amp;quot; in the UI.&lt;br /&gt;
&lt;br /&gt;
== Version 3.3.1 ==&lt;br /&gt;
Date: July 13th, 2026&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* Split merged datasets back into their original four separate outputs to match the legacy pipeline structure and restore BCO consistency. Affected datasets: [https://data.biomarkerkb.org/BMK_000001 BMK_000001], [https://data.biomarkerkb.org/BMK_000005 BMK_000005], [https://data.biomarkerkb.org/BMK_000014 BMK_000014], [https://data.biomarkerkb.org/BMK_000015 BMK_000015].&lt;br /&gt;
* Reverted Biomarker ID generation back to the original format (&amp;lt;code&amp;gt;XX1234&amp;lt;/code&amp;gt;, two-letter prefix + four-digit numeric ID, e.g., &amp;lt;code&amp;gt;AN6256&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;AN6256-1&amp;lt;/code&amp;gt;), correcting a regression that had introduced a non-standard format (e.g., &amp;lt;code&amp;gt;BMKB151581-1&amp;lt;/code&amp;gt;).&lt;br /&gt;
=== Backend Updates ===&lt;br /&gt;
* Automated Docker image cleanup to reduce disk usage on the build server.&lt;br /&gt;
* Increased local disk space allocation for the VM.&lt;br /&gt;
=== Bug Fixes ===&lt;br /&gt;
* Fixed column misalignment in old &amp;lt;code&amp;gt;masterlist&amp;lt;/code&amp;gt; datasets caused by the missing &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; field. New datasets (e.g., &amp;lt;code&amp;gt;biomarkers_exposure_agents.tsv&amp;lt;/code&amp;gt;) include this field; old datasets did not, causing a shift across all rows. Old datasets have been updated to include &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
== Version 3.2.2 ==&lt;br /&gt;
Date: June 22nd, 2026&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* The Ontology page has been updated to the latest version.&lt;br /&gt;
* The Knowledge Graph files, &amp;lt;all-biomarkers-nt.tar.gz&amp;gt; and &amp;lt;owlnets.tar.gz&amp;gt;, are now up-to-date with the current data.&lt;br /&gt;
* Added Dataset Badges to the BiomarkerKB Component Section, linking to data.biomarkerkb.org for all datasets integrated in BiomarkerKB, displayed alongside existing PMID and data source badges.&lt;br /&gt;
=== Backend Updates ===&lt;br /&gt;
* Implemented logic so that search results can be dissected by section (e.g., biomarker component, publication, evidence) and provided this information to the front end for the new intermediate/global search results page. Created a new Global Search API.&lt;br /&gt;
* Increased the upload size limit from 2MB to 100MB on the BiomarkerKB Wiki.&lt;br /&gt;
=== Bug Fixes ===&lt;br /&gt;
* All data sources are now displayed correctly under &amp;quot;Data Source&amp;quot; in the Advanced Search.&lt;br /&gt;
&lt;br /&gt;
== Version 3.2.1 ==&lt;br /&gt;
Date: April 23rd, 2026&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* Added a new dataset containing user-submitted biomarkers.&lt;br /&gt;
* Added a new manually curated dataset of biomarkers with exposure agents.&lt;br /&gt;
=== Bug Fixes ===&lt;br /&gt;
* Publications and exposure agents are now displayed correctly on the biomarker details pages.&lt;br /&gt;
&lt;br /&gt;
== Version 3.1.1 ==&lt;br /&gt;
Date: April 9th, 2026&lt;br /&gt;
=== Backend Updates ===&lt;br /&gt;
Complete overhaul of the backend data pipeline architecture:&lt;br /&gt;
* Improved ETL processes for greater reliability and scalability&lt;br /&gt;
* Enhanced data validation and error handling across pipeline stages&lt;br /&gt;
* Optimized performance for faster data processing and reduced runtime&lt;br /&gt;
* Refactored codebase for maintainability and extensibility&lt;br /&gt;
&lt;br /&gt;
== Version 2.4.3 ==&lt;br /&gt;
Date: February 26th, 2026&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* Added a new image-based biomarker from the OncoMX dataset.&lt;br /&gt;
* Fixed UniProtKB biomarkers that incorrectly included exposure agents.&lt;br /&gt;
=== Backend Updates ===&lt;br /&gt;
* Updated ChEBI API integration to properly parse JSON responses.&lt;br /&gt;
&lt;br /&gt;
== Version 2.4.2 ==&lt;br /&gt;
Date: February 19th, 2026&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* An archive file of all biomarker NTriples is now available for download at [https://data.biomarkerkb.org/BMK_000019 data.biomarkerkb.org/BMK_000019].&lt;br /&gt;
&lt;br /&gt;
== Version 2.4.1 ==&lt;br /&gt;
Date: February 12th, 2026&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* A master list of all biomarkers present in BiomarkerKB is now available for download at [https://data.biomarkerkb.org/BMK_000007 data.biomarkerkb.org/BMK_000007].&lt;br /&gt;
&lt;br /&gt;
== Version 2.4.0 ==&lt;br /&gt;
Date: February 5, 2026&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* MarkerDB data has been removed due to its license being free for academic use only.&lt;br /&gt;
=== Bug Fixes ===&lt;br /&gt;
* Fixed the issue where glycan biomarkers were being assigned incorrect GlyTouCan IDs in the controlled vocabulary field.&lt;br /&gt;
* An advanced search by some data sources, e.g., ClinVar, now yields biomarkers from the data source in question instead of showing all biomarkers.&lt;br /&gt;
* Duplicate entity normal range rows have been removed where applicable.&lt;br /&gt;
* Entity type casing in searches and search filters has been corrected.&lt;br /&gt;
* GWAS and SenNet biomarkers have their controlled vocabulary terms displayed consistently.&lt;br /&gt;
&lt;br /&gt;
== Version 2.3.0 ==&lt;br /&gt;
Date: January 12, 2026&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* New dataset: Top 50 Clinically Relevant Disease Biomarkers created and manually curated by Sparsh Gupta.&lt;br /&gt;
=== Backend Updates ===&lt;br /&gt;
* New [https://biomarkerkb.org/biomarker-search/ Advanced Search] type: users can now search biomarkers by Data Source. The following data sources are currently available:&lt;br /&gt;
** &amp;lt;code&amp;gt;cgi&amp;lt;/code&amp;gt; (Cancer Genome Interpreter)&lt;br /&gt;
** &amp;lt;code&amp;gt;civic&amp;lt;/code&amp;gt; (CIViC)&lt;br /&gt;
** &amp;lt;code&amp;gt;clinvar&amp;lt;/code&amp;gt; (ClinVar)&lt;br /&gt;
** &amp;lt;code&amp;gt;edrn&amp;lt;/code&amp;gt; (Early Detection Research Network)&lt;br /&gt;
** &amp;lt;code&amp;gt;gwas&amp;lt;/code&amp;gt; (Genome-Wide Association Studies)&lt;br /&gt;
** &amp;lt;code&amp;gt;llm_glycan&amp;lt;/code&amp;gt; (LLM-extracted glycan biomarkers)&lt;br /&gt;
** &amp;lt;code&amp;gt;markerdb&amp;lt;/code&amp;gt; (MarkerDB)&lt;br /&gt;
** &amp;lt;code&amp;gt;mw&amp;lt;/code&amp;gt; (Metabolomics Workbench)&lt;br /&gt;
** &amp;lt;code&amp;gt;oncomx&amp;lt;/code&amp;gt; (OncoMX)&lt;br /&gt;
** &amp;lt;code&amp;gt;opentargets&amp;lt;/code&amp;gt; (OpenTargets)&lt;br /&gt;
** &amp;lt;code&amp;gt;PMC_biomarker_sets&amp;lt;/code&amp;gt; (PubMed Central)&lt;br /&gt;
** &amp;lt;code&amp;gt;sennet&amp;lt;/code&amp;gt; (SenNet Consortium)&lt;br /&gt;
** &amp;lt;code&amp;gt;top_50&amp;lt;/code&amp;gt; (Top-50 clinically relevant biomarkers)&lt;br /&gt;
** &amp;lt;code&amp;gt;upkb_reviewed_v2&amp;lt;/code&amp;gt; (UniProtKB)&lt;br /&gt;
=== Bug Fixes ===&lt;br /&gt;
* The &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt; field in TSV files is now constructed based on the &amp;lt;code&amp;gt;biomarker_id&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;biomarker_orig&amp;lt;/code&amp;gt; tuple. Previously it only used &amp;lt;code&amp;gt;biomarker_id&amp;lt;/code&amp;gt; as key, introducing inconsistencies in biomarkers that had multiple &amp;lt;code&amp;gt;biomarker_component&amp;lt;/code&amp;gt; objects.&lt;br /&gt;
&lt;br /&gt;
== Version 2.2.0 ==&lt;br /&gt;
Date: December 22, 2025&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* Electronic Health Records data has been added to creatinine biomarkers.&lt;br /&gt;
* New dataset: senescence biomarkers from [https://docs.sennetconsortium.org/biomarkers/ SenNet Consortium].&lt;br /&gt;
=== Backend Updates ===&lt;br /&gt;
* On the API level, each biomarker now contains a new field: &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt; which shows the standardized biomarker name. Original biomarker names are now shown in the &amp;lt;code&amp;gt;biomarker_orig&amp;lt;/code&amp;gt; field.&lt;br /&gt;
&lt;br /&gt;
== Version 2.1.0 ==&lt;br /&gt;
Date: December 11, 2025&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* Added the LLM-extracted glycan biomarker dataset provided by Cyrus Chun Hong Au Yeung.&lt;br /&gt;
=== Backend Updates ===&lt;br /&gt;
* The incorrect download links on the [https://data.biomarkerkb.org Data Portal] have been fixed.&lt;br /&gt;
* LOINC codes are no longer tied to specimen IDs.&lt;br /&gt;
&lt;br /&gt;
== Version 2.0.2 ==&lt;br /&gt;
Date: December 4, 2025&lt;br /&gt;
=== Bug Fixes ===&lt;br /&gt;
* LOINC codes are no longer tied to specimen (UBERON) IDs.&lt;br /&gt;
* For biomarkers that could not be mapped to [[Controlled Vocabulary and Keywords|Controlled Vocabulary]] the original biomarker name is displayed, followed by &amp;quot;in review&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== Version 2.0.1 ==&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* Added cross-references to the Common Fund Data Ecosystem ([https://commonfund.nih.gov/dataecosystem CFDE]) Data Coordinating Centers and other resources:&lt;br /&gt;
** [https://www.gtexportal.org/home/ GTEx]&lt;br /&gt;
** [https://pharos.nih.gov/ Pharos]&lt;br /&gt;
** [https://reactome.org/ Reactome]&lt;br /&gt;
** [https://undiagnosed.hms.harvard.edu/ Undiagnosed Diseases Network]&lt;br /&gt;
** [https://idg.reactome.org/ Illuminating the Druggable Genome (IDG) Reactome Portal]&lt;br /&gt;
** [https://www.metabolomicsworkbench.org/ Metabolomics Workbench]&lt;br /&gt;
** [https://maayanlab.cloud/sigcom-lincs SigCom LINCS]&lt;br /&gt;
&lt;br /&gt;
== Version 2.0.0 ==&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* The biomarker field is now standardized using controlled vocabulary terms.&lt;br /&gt;
* Added metabolite as an &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt; to &amp;lt;code&amp;gt;mw_loinc_biomarkers.tsv&amp;lt;/code&amp;gt;.&lt;br /&gt;
* Added [https://rnacentral.org/ RNAcentral] cross-reference support.&lt;br /&gt;
* Added Electronic Health Records Normal ranges data from Oracle Health for Troponin I as an example.&lt;br /&gt;
&lt;br /&gt;
== Version 1.0.6 ==&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* Added a new dataset: MW LOINC biomarkers (&amp;lt;code&amp;gt;mw_loinc_biomarkers.tsv&amp;lt;/code&amp;gt;).&lt;br /&gt;
* Added [https://ncithesaurus.nci.nih.gov/ National Cancer Institute Thesaurus] and [https://www.rcsb.org/ Protein Data Bank] cross-references.&lt;br /&gt;
=== Backend Updates ===&lt;br /&gt;
* Added the &amp;lt;code&amp;gt;display_name&amp;lt;/code&amp;gt; field to the &amp;lt;code&amp;gt;format-converter&amp;lt;/code&amp;gt; so data source names appear with correct casing.&lt;br /&gt;
&lt;br /&gt;
== Version 1.0.5 ==&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* Updated the Troponin biomarker value &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; for consistency.&lt;br /&gt;
* Added normal ranges from Electronic Health Records provided by the University of New Mexico for Troponin biomarkers.&lt;br /&gt;
* Added Cell Ontology and Protein Ontology cross-references.&lt;br /&gt;
=== Backend Updates ===&lt;br /&gt;
* Updated all script paths to use &amp;lt;code&amp;gt;data_source.conf&amp;lt;/code&amp;gt; and validated data source names.&lt;br /&gt;
&lt;br /&gt;
== Version 1.0.4 ==&lt;br /&gt;
This release introduces new datasets, cross-references, and bug fixes.&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* Added Cancer Genome Interpreter data on cancer biomarkers from MetaKB.&lt;br /&gt;
* Added Metabolomics Workbench LOINC data on metabolite biomarkers.&lt;br /&gt;
* Added Cell Ontology and Protein Ontology cross-references.&lt;br /&gt;
=== Bug Fixes ===&lt;br /&gt;
* Fixed issue where cookie preferences weren&#039;t being saved when selecting &amp;quot;Allow&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
== Version 1.0.3 ==&lt;br /&gt;
This release introduces new cross-references and updates to ensure compatibility with external resources.&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* NCBI cross-references added across gene biomarker entries.&lt;br /&gt;
* ChEBI cross-references integrated for small molecules and metabolites.&lt;br /&gt;
=== Backend Updates ===&lt;br /&gt;
* ChEBI API migration: Updated all programmatic links from the legacy SOAP services to the new REST API endpoints, following ChEBI’s platform migration.&lt;br /&gt;
** Old services retired 1 September 2025.&lt;br /&gt;
** New stable API: [https://www.ebi.ac.uk/chebi/backend/api/docs ChEBI REST API docs]&lt;br /&gt;
** New data products and beta interface available at [https://www.ebi.ac.uk/chebi/beta/ ChEBI 2.0].&lt;br /&gt;
== Version 1.0.2 ==&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* Published updated [https://www.metabolomicsworkbench.org/ Metabolomics Workbench] data.&lt;br /&gt;
* Published sample data from the [https://edrn.nci.nih.gov/ Early Detection Research Network].&lt;br /&gt;
=== Backend Updates ===&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; database names now retain their original casing for accuracy and consistency.&lt;br /&gt;
* EDRN identifiers were added to the [https://github.com/clinical-biomarkers/format-converter/blob/main/mapping_data/namespace_map.json namespace map].&lt;br /&gt;
* [https://www.genenames.org/ HUGO Gene Nomenclature Committee] (HGNC) was added to the cross-reference JSON file.&lt;br /&gt;
* Fixed an issue where &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; values without tags were previously dropped; these are now preserved.&lt;br /&gt;
* Added a user-guided spelling correction function to improve data entry quality.&lt;br /&gt;
* The TSV-to-JSON converter now automatically checks for header spelling errors.&lt;br /&gt;
* Introduced &amp;lt;code&amp;gt;_suggest_header_corrections&amp;lt;/code&amp;gt; to flag and propose fixes for misspelled headers.&lt;br /&gt;
* Enhanced &amp;lt;code&amp;gt;_stream_tsv&amp;lt;/code&amp;gt; with a call to &amp;lt;code&amp;gt;_check_header_spelling&amp;lt;/code&amp;gt; to prevent invalid headers from being processed.&lt;br /&gt;
&lt;br /&gt;
== Version 1.0.1 ==&lt;br /&gt;
=== Data Updates ===&lt;br /&gt;
* Added &amp;lt;code&amp;gt;xrefs.tsv&amp;lt;/code&amp;gt; to the list of datasets.&lt;br /&gt;
=== Backend Updates ===&lt;br /&gt;
* Fixed ID formatting issues in NCBI and UniProt references within &amp;lt;code&amp;gt; oncomx.tsv&amp;lt;/code&amp;gt;, removing erroneous spaces (e.g., &amp;lt;code&amp;gt; NCBI: 3288&amp;lt;/code&amp;gt; → &amp;lt;code&amp;gt; NCBI:3288&amp;lt;/code&amp;gt;) and extraneous text (e.g., &amp;lt;code&amp;gt;&amp;quot;(composition)&amp;quot;&amp;lt;/code&amp;gt;). Affected biomarkers included AN6295-1, AN6756-1, AN6728-1, and others.&lt;br /&gt;
* Merged assessed entity type synonyms.&lt;br /&gt;
&lt;br /&gt;
== Version 1.0.0 ==&lt;br /&gt;
* BiomarkerKB data portal available with OncoMX, OpenTargets, MarkerDB, ClinVar, PubMed Central Biomarker Gene Set Curation, MW, UniProtKB, GWAS, CIViC biomarker data.&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=266</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=266"/>
		<updated>2026-08-18T18:05:58Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* BiomarkerKB dataset datamodel fields */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out. A template is available at (link TBA with color-coded mandatory and optional columns). There will be two templates: one for disease biomarkers (only condition id required), one for exposure agent biomarkers (either exposure agent or exposure agent id required).&lt;br /&gt;
&lt;br /&gt;
[https://data.biomarkerkb.org/ln2downloads/curator_agent/2026_08_11/biomarkers-formatted.tsv Disease biomarker template] - this link will be replaced, the template will have its dedicated directory.&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=265</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=265"/>
		<updated>2026-08-18T18:02:38Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* BiomarkerKB dataset datamodel fields */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out. A template is available at (link TBA with color-coded mandatory and optional columns). There will be two templates: one for disease biomarkers (only condition id required), one for exposure agent biomarkers (either exposure agent or exposure agent id required).&lt;br /&gt;
&lt;br /&gt;
[https://data.biomarkerkb.org/ln2downloads/curator_agent/2026_08_11/biomarkers-formatted.tsv Disease biomarker template]&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=264</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=264"/>
		<updated>2026-08-18T17:57:49Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* BiomarkerKB dataset datamodel fields */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out. A template is available at (link TBA with color-coded mandatory and optional columns). There will be two templates: one for disease biomarkers (only condition id required), one for exposure agent biomarkers (either exposure agent or exposure agent id required).&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected. This field is going to be broken down into three items: change, aspect, and entity.&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=263</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=263"/>
		<updated>2026-08-18T17:56:58Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* BiomarkerKB dataset datamodel fields */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out. A template is available at (link TBA with color-coded mandatory and optional columns). There will be two templates: one for disease biomarkers (only condition id required), one for exposure agent biomarkers (either exposure agent or exposure agent id required).&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=262</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=262"/>
		<updated>2026-08-18T17:55:17Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* BiomarkerKB dataset datamodel fields */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out. A template is available at (link TBA with color-coded mandatory and optional columns).&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected&lt;br /&gt;
* &amp;lt;code&amp;gt;taxonomy_id&amp;lt;/code&amp;gt;: taxonomy ID of the organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; (optional): see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt; (optional): see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=261</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=261"/>
		<updated>2026-08-18T17:49:16Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* BiomarkerKB dataset datamodel fields */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out. A template is available at (link TBA with color-coded mandatory and optional columns).&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;: entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected&lt;br /&gt;
* &amp;lt;code&amp;gt;organism_name&amp;lt;/code&amp;gt;: organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;: see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=260</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=260"/>
		<updated>2026-08-18T17:48:20Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* BiomarkerKB dataset datamodel fields */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
== BiomarkerKB dataset datamodel fields ==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out. A template is available at (insert link here).&lt;br /&gt;
&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;: entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected&lt;br /&gt;
* &amp;lt;code&amp;gt;organism_name&amp;lt;/code&amp;gt;: organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;: see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=259</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=259"/>
		<updated>2026-08-18T17:47:21Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* BiomarkerKB dataset datamodel fields */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
==BiomarkerKB dataset datamodel fields==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (&#039;&#039;&#039;required&#039;&#039;&#039;): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;: entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected&lt;br /&gt;
* &amp;lt;code&amp;gt;organism_name&amp;lt;/code&amp;gt;: organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;: see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=258</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=258"/>
		<updated>2026-08-18T17:46:32Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* BiomarkerKB dataset datamodel fields */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
==BiomarkerKB dataset datamodel fields==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (required): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (required): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (required): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;: entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected&lt;br /&gt;
* &amp;lt;code&amp;gt;organism_name&amp;lt;/code&amp;gt;: organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;: see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=257</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=257"/>
		<updated>2026-08-18T17:45:57Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* BiomarkerKB dataset datamodel fields */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
==BiomarkerKB dataset datamodel fields==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; (mandatory): biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; (mandatory): component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; (): entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;: entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected&lt;br /&gt;
* &amp;lt;code&amp;gt;organism_name&amp;lt;/code&amp;gt;: organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;: see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent single-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== Single-component multi-entity (&amp;quot;complex&amp;quot;) biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent multi-entity biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=256</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=256"/>
		<updated>2026-08-18T17:36:34Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* &amp;quot;Complex&amp;quot; biomarker */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
==BiomarkerKB dataset datamodel fields==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt;: biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt;: component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt;: entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;: entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected&lt;br /&gt;
* &amp;lt;code&amp;gt;organism_name&amp;lt;/code&amp;gt;: organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;: see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== &amp;quot;Complex&amp;quot; biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|660px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent and &amp;quot;complex&amp;quot; biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=255</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=255"/>
		<updated>2026-08-18T17:36:23Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* &amp;quot;Complex&amp;quot; biomarker */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
==BiomarkerKB dataset datamodel fields==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt;: biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt;: component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt;: entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;: entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected&lt;br /&gt;
* &amp;lt;code&amp;gt;organism_name&amp;lt;/code&amp;gt;: organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;: see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== &amp;quot;Complex&amp;quot; biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|650px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent and &amp;quot;complex&amp;quot; biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=254</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=254"/>
		<updated>2026-08-18T17:36:14Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* &amp;quot;Complex&amp;quot; biomarker */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
==BiomarkerKB dataset datamodel fields==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt;: biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt;: component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt;: entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;: entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected&lt;br /&gt;
* &amp;lt;code&amp;gt;organism_name&amp;lt;/code&amp;gt;: organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;: see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== &amp;quot;Complex&amp;quot; biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|600px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent and &amp;quot;complex&amp;quot; biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=253</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=253"/>
		<updated>2026-08-18T17:36:04Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* &amp;quot;Complex&amp;quot; biomarker */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
==BiomarkerKB dataset datamodel fields==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt;: biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt;: component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt;: entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;: entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected&lt;br /&gt;
* &amp;lt;code&amp;gt;organism_name&amp;lt;/code&amp;gt;: organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;: see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== &amp;quot;Complex&amp;quot; biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|550px]]&lt;br /&gt;
&lt;br /&gt;
===== Multicomponent and &amp;quot;complex&amp;quot; biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=252</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=252"/>
		<updated>2026-08-18T17:32:27Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* Multicomponent and complex biomarker */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
==BiomarkerKB dataset datamodel fields==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt;: biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt;: component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt;: entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;: entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected&lt;br /&gt;
* &amp;lt;code&amp;gt;organism_name&amp;lt;/code&amp;gt;: organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;: see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== &amp;quot;Complex&amp;quot; biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|530px]]&lt;br /&gt;
===== Multicomponent and &amp;quot;complex&amp;quot; biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=251</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=251"/>
		<updated>2026-08-18T17:30:59Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* BiomarkerKB dataset datamodel fields */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
==BiomarkerKB dataset datamodel fields==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt;: biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt;: component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt;: entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;: entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected&lt;br /&gt;
* &amp;lt;code&amp;gt;organism_name&amp;lt;/code&amp;gt;: organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;: see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== &amp;quot;Complex&amp;quot; biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|530px]]&lt;br /&gt;
===== Multicomponent and complex biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 3 || 1 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
In this example, component 1 is a glycoprotein, component 2 is a DNA molecule, and component 3 is an RNA molecule.&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
	<entry>
		<id>https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=250</id>
		<title>Data Submission/Data Upload</title>
		<link rel="alternate" type="text/html" href="https://wiki.biomarkerkb.org/index.php?title=Data_Submission/Data_Upload&amp;diff=250"/>
		<updated>2026-08-18T17:25:50Z</updated>

		<summary type="html">&lt;p&gt;MariaKim: /* Examples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Instructions to submit Biomarker Data==&lt;br /&gt;
To submit data for the BiomarkerKB Portal, the biomarker data model must be followed. Instructions on how to format the data for submission, where to send it, and creating a BCO for the data submitted will be provided below.&lt;br /&gt;
&lt;br /&gt;
# Biomarker data collected should follow the biomarker data model.&lt;br /&gt;
# &amp;quot;Core&amp;quot; fields should be filled in from the data source where biomarker data is collected. Core fields:&lt;br /&gt;
## &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; OR &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;&lt;br /&gt;
## &amp;lt;code&amp;gt;component_group&amp;lt;/code&amp;gt; containing integers (1, 2, 3...) from 1 to N where N is the number of components. Normally N would simply be equal to the number of rows, unless your data contains multicomponent biomarkers. A multicomponent biomarker must have the same integer in all rows related to that biomarker.&lt;br /&gt;
# Other fields and annotations may also be collected from the data source, however if data is missing it can also be inferred or mapped from other sources.&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt; is one or more exact citations from the evidence source (in most cases, it will be the PubMed publication).&lt;br /&gt;
# Apply the following standards to the data when possible:&lt;br /&gt;
## &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;DOID:0080600&amp;lt;/code&amp;gt;. Refer to https://disease-ontology.org/do/.&lt;br /&gt;
## &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;UBERON:0000178&amp;lt;/code&amp;gt;. Refer to https://www.ebi.ac.uk/ols4/ontologies/uberon.&lt;br /&gt;
## &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;LOINC:100153-6&amp;lt;/code&amp;gt;. Refer to https://loinc.org/ (you may need to create an account to access the search functionality).&lt;br /&gt;
## &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt; = &amp;lt;code&amp;gt;SOURCE:ID&amp;lt;/code&amp;gt;, for example &amp;lt;code&amp;gt;PubMed:32677844&amp;lt;/code&amp;gt;&lt;br /&gt;
## For &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt; please refer to the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary GitHub documentation] for which standards to follow&lt;br /&gt;
# Provide extra annotations from your DCC/data with the agreed upon standards from the Biomarker Annotation RFC. This data does not have to follow the data model and can be submitted in a separate file or can be added in the &amp;lt;code&amp;gt;comment&amp;lt;/code&amp;gt; field .&lt;br /&gt;
## For example: Relevant EHR data/LOINC data for biomarkers/biomarker entities can be included in a separate sheet.&lt;br /&gt;
# Create a tsv/json file with the agreed upon fields which correspond to the biomarker data model. The data dictionary provides details on what the different fields represent.&lt;br /&gt;
## The preferred method for data submission is a json file as it will help ingest the data into the existing data efficiently. However, tsv file submissions are ok as well. In the GitHub, &amp;lt;code&amp;gt;data_conversion.py&amp;lt;/code&amp;gt; script exists in the Data Conversion Folder and it will handle tsv to json file conversion and json to tsv file conversion as well.&lt;br /&gt;
## The [BiomarkerKB data page] has examples of tsv data submissions and how the data should be formatted with the appropriate biomarker fields. Example&lt;br /&gt;
# For panel biomarkers, if the biomarkers are part of the same panel, the biomarker_id value for each biomarker should be any string value that can uniquely identify which rows are part of the same biomarker panel. Documentation&lt;br /&gt;
# If curating data in tsv format: If biomarker rows are part of the same biomarker entry but differ on specimen, evidence, or role, then the biomarker_id for each row should be any string value that can uniquely identify which rows are part of the same biomarker.&lt;br /&gt;
&lt;br /&gt;
=== Submission === &lt;br /&gt;
Once your data is formatted and cleaned, please send it to mazumder_lab@gwu.edu.&lt;br /&gt;
# Concurrently with submitting data please fill out the BCO Information: Biomarker Data Google Form.&lt;br /&gt;
## This will give metadata and description on how biomarker data was collected and is important for adding submitted data to the Biomarker Data page. An example of a previous BCO is provided in the sheet and available on the biomarker data page as well. [https://hivelab.biochemistry.gwu.edu/biomarker-partnership/data/BCO_000435 Example]&lt;br /&gt;
# If there are any further questions please consult the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for contributing data or reach out to Daniall using the email above.&lt;br /&gt;
&lt;br /&gt;
==BiomarkerKB dataset datamodel fields==&lt;br /&gt;
&lt;br /&gt;
This is the standard way to report biomarker data. This section covers how biomarkers should be reported and how other fields should be filled out.&lt;br /&gt;
=== Headers ===&lt;br /&gt;
Your file must contain the following headers:&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt;: biomarker counter - integer (1, 2, 3...)&lt;br /&gt;
** &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt;: component counter within each biomarker (integer). In most cases, a biomarker would have only one component, unless it is a multicomponent biomarker. A multicomponent biomarker is a group of biological markers tested together to provide a comprehensive assessment of health, e.g., a lipid panel (blood test) that measures total cholesterol, LDL, HDL, and triglycerides. These four entities would occupy four rows with the &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; from 1 to 4 (see example table in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt;: entity counter within each component (integer). This index is used to represent complex entities such as glycoforms or protein complexes. E.g., in a biomarker where a glycoprotein is being measured, within one component entity 1 is a glycan, while entity 2 is a protein (see example in the &amp;quot;Examples&amp;quot; section).&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;: entity ID for each entity_index&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_biomarker_entity|assessed_biomarker_entity]]&lt;br /&gt;
*** &amp;lt;code&amp;gt;assessed_entity_type&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#assessed_entity_type|assessed_entity_type]]&lt;br /&gt;
* &amp;lt;code&amp;gt;biomarker_controlled_vocab&amp;lt;/code&amp;gt;: this replaces the field &amp;quot;biomarker&amp;quot; (see below) to emphasize that controlled vocabulary is expected&lt;br /&gt;
* &amp;lt;code&amp;gt;organism_name&amp;lt;/code&amp;gt;: organism in which the biomarker has been measured (human, mouse, zebrafish...)&lt;br /&gt;
* &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition|condition]]&lt;br /&gt;
* &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#condition_id|condition_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#best_biomarker_role|best_biomarker_role]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;: see [[Data_Submission/Data_Upload#specimen|specimen]]&lt;br /&gt;
* &amp;lt;code&amp;gt;specimen_id&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#specimen_id|specimen_id]]&lt;br /&gt;
* &amp;lt;code&amp;gt;loinc_code&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#loinc_code|loinc_code]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence_source&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence_source|evidence_source]]&lt;br /&gt;
* &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;: see [[Data Submission/Data Upload#evidence|evidence]]&lt;br /&gt;
&lt;br /&gt;
==== Examples ====&lt;br /&gt;
===== Multicomponent biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 1 || 1 || total cholesterol&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 2 || 1 || LDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 3 || 1 || HDL&lt;br /&gt;
|-&lt;br /&gt;
| 1 || 4 || 1 || triglycerides&lt;br /&gt;
|}&lt;br /&gt;
===== &amp;quot;Complex&amp;quot; biomarker =====&lt;br /&gt;
[[File:Data upload example table.png|530px]]&lt;br /&gt;
===== Multicomponent and complex biomarker =====&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! &amp;lt;code&amp;gt;biomarker_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;component_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;entity_index&amp;lt;/code&amp;gt; !! &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 1 || glycan A&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 1 || 2 || protein X&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 1 || some DNA&lt;br /&gt;
|-&lt;br /&gt;
| 2 || 2 || 2 || some RNA&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Biomarker representation framework ===&lt;br /&gt;
A biomarker is not simply a gene, protein, metabolite, or other biological entity. A biomarker must include a defined measurement or change concept — such as presence, absence, increase, or decrease — describing what is observed. For example, EGFR alone is not a biomarker, but a specific EGFR mutation used for diagnostic, prognostic, or treatment-selection purposes is. Likewise, &amp;quot;IL6&amp;quot; alone is not a biomarker, but &amp;quot;increased IL6 expression&amp;quot; in a defined clinical context may be.&lt;br /&gt;
&lt;br /&gt;
The fields below fall into two groups. Core fields directly align with the biomarker definition: &amp;lt;code&amp;gt;biomarker&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;assessed_biomarker_entity_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;condition_id&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;exposure_agent&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;exposure_agent_id&amp;lt;/code&amp;gt;. Contextual fields enrich the representation: &amp;lt;code&amp;gt;specimen&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;best_biomarker_role&amp;lt;/code&amp;gt;, and &amp;lt;code&amp;gt;evidence&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
In the BiomarkerKB accession model, the canonical biomarker concept represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;), and disease- or condition-specific records are represented as child records linked to that canonical biomarker.&lt;br /&gt;
&lt;br /&gt;
=== biomarker_id ===&lt;br /&gt;
A unique identifier assigned to each canonical biomarker concept. The canonical biomarker represents the measured change or observation (e.g. &amp;quot;increased IL6 expression&amp;quot;); disease- or condition-specific records are child records that share the same biomarker_id while differing in condition, specimen, or evidence. biomarker_id is assigned by the biomarkerKB data processing scripts automatically so the field can be left blank.&lt;br /&gt;
&lt;br /&gt;
=== biomarker ===&lt;br /&gt;
The biomarker field is the most important as follows the [https://github.com/clinical-biomarkers/biomarker-controlled-vocabulary BiomarkerKB Controlled Vocabulary] for standardized reporting. There are several distinctions here and changes are made based on the entity being reported. The text should be in lowercase except when a gene name appears then it should remain all uppercase.&lt;br /&gt;
Examples&lt;br /&gt;
* Increased level of protein SPP1/UPKB:P10451&lt;br /&gt;
* Increased expression of RNA PCA3/HGNC:8637&lt;br /&gt;
* Increased expression of gene B2M PCA3/NCBI:567&lt;br /&gt;
* Increased methylation in gene VIM/NCBI:7431&amp;lt;br /&amp;gt;&lt;br /&gt;
For more examples please refer to the [https://data.biomarkerkb.org/ BiomarkerKB Data Page]&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity ===&lt;br /&gt;
assessed_biomarker_entity is the entity in which the change is assessed.&lt;br /&gt;
Should start off with a capital letter but if it is just a gene then it should remain in all capitals (e.g Myosin-binding protein H-like or IL6).&lt;br /&gt;
If the entity type is anything but a gene the whole name should be typed out.&lt;br /&gt;
&lt;br /&gt;
=== assessed_biomarker_entity_id ===&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot;&lt;br /&gt;
!Assessed Entity Type&lt;br /&gt;
!Resource (in order of preference/availability)&lt;br /&gt;
|-&lt;br /&gt;
|Carbohydrate&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Cell&lt;br /&gt;
|Cell Ontology (CO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Chemical Element&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|DNA&lt;br /&gt;
|National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Gene&lt;br /&gt;
|NCBI&lt;br /&gt;
|-&lt;br /&gt;
|Gene (mutation)&lt;br /&gt;
|NCBI dbSNP&lt;br /&gt;
|-&lt;br /&gt;
|Glycan&lt;br /&gt;
|GlyTouCan Accession (GTC) -&amp;gt; PubChem (PCCID)&lt;br /&gt;
|-&lt;br /&gt;
|Lipoprotein&lt;br /&gt;
|Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Metabolite&lt;br /&gt;
|PubChem (PCCID) -&amp;gt; Chemical Entities of Biological Interest (ChEBI)&lt;br /&gt;
|-&lt;br /&gt;
|Peptide&lt;br /&gt;
|Protein Ontology (PRO)&lt;br /&gt;
|-&lt;br /&gt;
|Protein&lt;br /&gt;
|Uniprot (UPKB) -&amp;gt; Protein Data Bank (PDB) -&amp;gt; Protein Ontology (PRO) -&amp;gt; National Cancer Institute Thesaurus (NCIt)&lt;br /&gt;
|-&lt;br /&gt;
|Protein Complex&lt;br /&gt;
|Protein Ontology (PRO) -&amp;gt; Gene Ontology (GO)&lt;br /&gt;
|-&lt;br /&gt;
|RNA&lt;br /&gt;
|HUGO Gene Nomenclature Committee (HGNC) -&amp;gt; RNA Central (RNAC)&lt;br /&gt;
|-&lt;br /&gt;
|miRNA&lt;br /&gt;
|miRBase (MRB)&lt;br /&gt;
|}&lt;br /&gt;
Refer to the [https://github.com/clinical-biomarkers/biomarker-partnership/blob/main/supplementary_files/documentation/contributing_data.md GitHub Documentation] for the correct resource.&lt;br /&gt;
&lt;br /&gt;
=== assessed_entity_type ===&lt;br /&gt;
Report in all lowercase.&lt;br /&gt;
Example: gene&lt;br /&gt;
&lt;br /&gt;
=== condition ===&lt;br /&gt;
`condition` should be reported in all lowercase.&lt;br /&gt;
Example: colon cancer&lt;br /&gt;
&lt;br /&gt;
=== condition_id ===&lt;br /&gt;
`condition_id` (from Disease Ontology, MONDO, or SNOMED or NCIt) should be provided in the following column.&lt;br /&gt;
Example: DOID:219&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent ===&lt;br /&gt;
Report in all lowercase. The exposure_agent documents any external stimulus, treatment, environmental factor, or intervention relevant to the biomarker&#039;s expression or activity. It provides context for biomarkers that respond to specific exposures rather than intrinsic disease processes (for example, response biomarkers). Leave blank if not applicable.&lt;br /&gt;
Example: cisplatin&lt;br /&gt;
&lt;br /&gt;
=== exposure_agent_id ===&lt;br /&gt;
The ontology identifier for the exposure_agent, provided in the following column. Leave blank if not applicable.&lt;br /&gt;
Example: CHEBI:27899&lt;br /&gt;
&lt;br /&gt;
=== best_biomarker_role ===&lt;br /&gt;
Report in all lowercase. Refer to the [BEST Resource](https://www.ncbi.nlm.nih.gov/books/NBK326791/) to infer the correct biomarker role. Accepted role terms are:&lt;br /&gt;
&lt;br /&gt;
* diagnostic: Detects or confirms the presence of a disease or condition, or identifies individuals with a specific disease subtype.&lt;br /&gt;
* monitoring: Assesses the status of a disease, medical condition, or exposure to a medical product over time.&lt;br /&gt;
* predictive: Identifies which patients are more or less likely to respond favorably or unfavorably to a specific treatment or exposure.&lt;br /&gt;
* prognostic: Identifies the likelihood of a clinical event, disease recurrence, or progression in patients with an already established disease or condition.&lt;br /&gt;
* response: Shows that a biological response has occurred in a patient after being exposed to a medical product or environmental agent.&lt;br /&gt;
* risk: Indicates the potential for an individual to develop a disease or condition in the future.&lt;br /&gt;
* safety: Measures or indicates the likelihood, nature, or severity of adverse effects, toxicity, or organ injury.&lt;br /&gt;
&lt;br /&gt;
Example: diagnostic&lt;br /&gt;
&lt;br /&gt;
=== specimen ===&lt;br /&gt;
Report in all lowercase. Leave blank if not applicable.&lt;br /&gt;
Example: feces&lt;br /&gt;
&lt;br /&gt;
=== specimen_id ===&lt;br /&gt;
`specimen_id` in the following column should be from UBERON. Leave blank if not applicable.&lt;br /&gt;
Example: UBERON:0001988&lt;br /&gt;
&lt;br /&gt;
=== loinc_code ===&lt;br /&gt;
Report the Logical Observation Identifiers Names and Codes (LOINC) code corresponding to the test or measurement (e.g. 77354-9). Leave blank if not applicable.&lt;br /&gt;
Example: 77354-9&lt;/div&gt;</summary>
		<author><name>MariaKim</name></author>
	</entry>
</feed>