Skip to content

Mapping a CIViC oncogenicity assertion to GKM

Notebook source

Rendered from notebooks/civic/vignettes/oncogenicity/civic-oncogenicity-gkm.ipynb.

This notebook follows CIViC Oncogenic Assertion 202 into the GA4GH Genomic Knowledge Model (GKM). It accompanies the CIViC oncogenicity vignette.

It starts with CIViC's assertion page, then follows the claim, assessment, variant context, evidence, and approval into connected GKM records.

Before you begin

This notebook assumes familiarity with the GKM Toolkit, the GA4GH Reference Implementations, and the CIViC bundle.

New to how GKM data is represented or explored? Start with Explore GKM Toolkit features with CIViC bundle examples. It walks through loading a bundle example, inspecting its schema and collections, resolving linked records, and preparing data for another tool.

CIViC oncogenicity assertions

CIViC is a community-curated knowledgebase of cancer variant interpretations.

A Molecular Profile supplies the variant context and can group one or more CIViC Variants. An Evidence Item records a curator-reviewed interpretation of an observation from a source publication.

An Oncogenic Assertion connects a Molecular Profile and disease, then records CIViC's classification.

Oncogenicity assertions also record oncogenicity codes from the ClinGen/CGC/VICC guidelines (Horak et al. 2022). Each code identifies a guideline criterion met by the assertion.

CIViC's information model describes these records and their relationships in more detail.

CIViC Assertion 202

CIViC Assertion 202 classifies RET M918T as likely oncogenic in medullary thyroid carcinoma.

Its summary shows the claim, classification, oncogenicity codes, approval status, and attached Evidence Items.

CIViC Assertion 202

GKM represents the claim as a proposition and CIViC's assessment as a Statement. The following sections inspect those records.

Mapping the CIViC assertion to GKM

Load the published GKM bundle

CIViCpy creates GKM-compatible records and bundles from CIViC data. Load the published CIViC bundle.

Note

The examples read the published CIViC bundle. On first use, the Toolkit may download and cache it locally.

In [1]
import json

from ga4gh.gkm.bundles import BundleRepository, load_repository_bundle

repository = BundleRepository(refresh=False)
civic_bundle = load_repository_bundle(repository, "civic", refresh=False)
civic_bundle

Output:

Bundle(name='civic', collections=17)

Retrieve the oncogenic assertion from the bundle

Retrieve CIViC AID 202. The output identifies its GKM record type.

In [2]
assertion_id = "civic.aid:202"
assertion = civic_bundle.assertion[assertion_id]
type(assertion)

Output:

ga4gh.va_spec.ccv_2022.models.VariantOncogenicityStatement

Assessment structure

CIViC records an assessment for the interpretation. Direction records that CIViC supports it. Significance records the classification: likely oncogenic.

CIViC assertion assessment

GKM represents the assessment as a Statement. The next output shows the Statement and its assessment fields.

In [3]
print(f"Assertion type: {type(assertion).__name__}")
print(f"Assessment direction: {assertion.direction} the proposition")
classification = assertion.classification.primaryCoding.code.root
print(f"Classification: {classification}")

proposition = assertion.proposition
variant = proposition.subject
disease = proposition.object
gene = proposition.geneContextQualifier
origin = proposition.alleleOriginQualifier
evidence_lines = assertion.hasEvidenceLines
framework = evidence_lines[0].specifiedBy.name

Output:

Assertion type: VariantOncogenicityStatement
Assessment direction: supports the proposition
Classification: likely oncogenic

CIViC identifies this as an oncogenicity assertion. CIViCpy maps it to VariantOncogenicityStatement, a specific VA-Spec Statement profile that records CIViC's assessment and classification.

Claim structure

The Statement links to a VariantOncogenicityProposition, which holds the variant, gene context, allele origin, and disease.

civic assertion claim

The output shows the fields that make up the claim.

In [4]
print(f"Subject variant: {variant.name}")
print(f"Gene context: {gene.name}")
print(f"Allele origin: {origin.name}")
print(f"Disease: {disease.name}")
print(f"Claim: {origin.name} {variant.name} is oncogenic for {disease.name}")

Output:

Subject variant: RET M918T
Gene context: RET
Allele origin: somatic
Disease: Medullary Thyroid Carcinoma
Claim: somatic RET M918T is oncogenic for Medullary Thyroid Carcinoma

The VariantOncogenicityProposition states: somatic RET M918T is oncogenic for medullary thyroid carcinoma. The separate Statement records CIViC's assessment of that claim.

Mapping the variant and disease

Storing the molecular profile and its variant context

The assertion links to a CIViC Molecular Profile, which provides its variant context.

CIViC Assertion Molecular Profile

CIViCpy represents the Molecular Profile as a Cat-VRS CategoricalVariant, which becomes the proposition's subject.

In [5]
members = [member.root for member in variant.members]
civic_variant = next(
    mapping.coding
    for mapping in variant.mappings or []
    if mapping.coding.system == "https://civicdb.org/links/variant/"
)

print(f"Categorical variant: {variant.name} ({variant.id})")

Output:

Categorical variant: RET M918T (civic.mpid:113)

The output identifies the CategoricalVariant. The CIViC Molecular Profile page shows that M918T is attached to it:

CIViC Molecular Profile

The linked CIViC Variant supplies the sequence context. CIViC stores its genomic, coding, and protein representations under one Variant ID:

CIViC Variant

GKM represents those sequence contexts as VRS Alleles. The output lists the corresponding identifiers.

In [6]
print(f"CIViC Variant ID: {civic_variant.id}")
print("VRS Allele IDs by sequence context:")
context_labels = {
    "hgvs.c": "Coding",
    "hgvs.g": "Genomic",
    "hgvs.p": "Protein",
}
protein_allele = next(
    member
    for member in members
    if any(expression.syntax == "hgvs.p" for expression in member.expressions)
)
for member in members:
    expressions = [expr.value for expr in member.expressions]
    context = context_labels[member.expressions[0].syntax]
    print(f"  {context}:")
    print(f"    VRS ID: {member.id}")
    print(f"    HGVS Description(s): {', '.join(expressions)}\n")

Output:

CIViC Variant ID: civic.vid:113
VRS Allele IDs by sequence context:
  Coding:
    VRS ID: ga4gh:VA.TZBjEPHhLRYxssQopcOQLWEBQrwzhH3T
    HGVS Description(s): NM_020975.4:c.2753T>C

  Genomic:
    VRS ID: ga4gh:VA.ON-Q17mJBYx3unmQ8GiqllzEphxR-Fie
    HGVS Description(s): NC_000010.10:g.43617416T>C, NC_000010.11:g.43121968T>C

  Protein:
    VRS ID: ga4gh:VA.hEybNB_CeKflfFhT5AKOU5i1lgZPP-aS
    HGVS Description(s): NP_065681.1:p.Met918Thr, ENSP00000347942.3:p.Met918Thr

Full VRS Allele representation

Each VRS Allele has a precise, computable identifier. The protein-level member is shown next.

In [7]
print(f"\nExample VRS Protein Allele: {protein_allele.id}")
print(f"Name: {protein_allele.name}")
for expression in protein_allele.expressions:
    print(f"  {expression.syntax}: {expression.value}")

Output:

Example VRS Protein Allele: ga4gh:VA.hEybNB_CeKflfFhT5AKOU5i1lgZPP-aS
Name: RET M918T
  hgvs.p: NP_065681.1:p.Met918Thr
  hgvs.p: ENSP00000347942.3:p.Met918Thr

This JSON is the structured representation behind the protein-level identifier.

In [8]
print(json.dumps(protein_allele.model_dump(exclude_none=True), indent=2))

Output:

{
  "id": "ga4gh:VA.hEybNB_CeKflfFhT5AKOU5i1lgZPP-aS",
  "type": "Allele",
  "name": "RET M918T",
  "digest": "hEybNB_CeKflfFhT5AKOU5i1lgZPP-aS",
  "expressions": [
    {
      "syntax": "hgvs.p",
      "value": "NP_065681.1:p.Met918Thr"
    },
    {
      "syntax": "hgvs.p",
      "value": "ENSP00000347942.3:p.Met918Thr"
    }
  ],
  "location": {
    "id": "ga4gh:SL.oIeqSfOEuqO7KNOPt8YUIa9vo1f6yMao",
    "type": "SequenceLocation",
    "digest": "oIeqSfOEuqO7KNOPt8YUIa9vo1f6yMao",
    "sequenceReference": {
      "type": "SequenceReference",
      "refgetAccession": "SQ.jMu9-ItXSycQsm4hyABeW_UfSNRXRVnl"
    },
    "start": 917,
    "end": 918,
    "sequence": "M"
  },
  "state": {
    "type": "LiteralSequenceExpression",
    "sequence": "T"
  }
}

The coding and genomic Alleles use a similar representation.

CIViC disease as a mapped concept

CIViC links each assertion to a Disease record. Here, medullary thyroid carcinoma is the disease in the oncogenicity claim.

CIViC disease as shown in assertion

The linked disease record provides the CIViC term that GKM maps to a normalized disease concept:

CIViC Disease Medullary Thyroid Carcinoma

The CIViC Disease record supplies the display name. GKM retains it and adds a computable disease identifier, helping other systems recognize the same disease concept:

In [9]
print("Disease:", disease.name)
for mapping in disease.mappings or []:
    coding = mapping.coding
    print(f"  {mapping.relation}: {coding.system}{coding.code}")

Output:

Disease: Medullary Thyroid Carcinoma
  exactMatch: https://disease-ontology.org/?id=root='DOID:3973'

Disease concept JSON

The next output shows the full mapped disease concept:

In [10]
print(json.dumps(disease.model_dump(exclude_none=True), indent=2))

Output:

{
  "id": "civic.did:15",
  "type": "MappableConcept",
  "name": "Medullary Thyroid Carcinoma",
  "conceptType": "Disease",
  "mappings": [
    {
      "coding": {
        "system": "https://disease-ontology.org/?id=",
        "code": "DOID:3973"
      },
      "relation": "exactMatch"
    }
  ]
}

Mapping the evidence behind the classification

Oncogenicity criteria as Evidence Lines

CIViC stores oncogenicity codes on the assertion to document the guideline criteria used for its classification. Assertion 202 meets OM1, OS2, OP4, OP1, and OP3.

The screenshot shows the oncogenicity codes in the assertion summary.

CIViC Assertion 202 Summary

GKM represents each criterion as a VA-Spec Evidence Line. The output lists each line's method, strength, and score.

In [11]
print(f"{'Code':<6}{'Method':<32}{'Strength':<14}{'Score'}")
print("-" * 62)
evidence_codes = [
    line.evidenceOutcome.primaryCoding.code.root for line in evidence_lines
]
for line in evidence_lines:
    code = line.evidenceOutcome.primaryCoding.code.root
    method = line.specifiedBy.methodType
    strength = line.strengthOfEvidenceProvided.primaryCoding.code.root
    score = line.scoreOfEvidenceProvided
    print(f"{code:<6}{method:<32}{strength:<14}{score}")

total_score = sum(line.scoreOfEvidenceProvided for line in evidence_lines)
print(f"\nTotal classification score: {total_score}")
evidence_summary = ", ".join(
    f"{line.specifiedBy.methodType.replace('_', ' ')} "
    f"({line.evidenceOutcome.primaryCoding.code.root})"
    for line in evidence_lines
)

Output:

Code  Method                          Strength      Score
--------------------------------------------------------------
OM1   functional_domain_assessment    moderate      2
OS2   functional_data_assessment      strong        4
OP4   population_data_assessment      supporting    1
OP1   in_silico_impact_assessment     supporting    1
OP3   somatic_hotspot_assessment      supporting    1

Total classification score: 9

CIViC Evidence Items and scored criteria

The Evidence Lines contribute scores of 2 (OM1), 4 (OS2), and 1 each (OP4, OP1, and OP3), for a total of 9.

CIViC associates Evidence Items with the assertion but does not link an individual item to an oncogenicity code. CIViCpy maps the assertion-level links to Statement.hasEvidence and each scored criterion to an EvidenceLine. EvidenceLine.hasEvidenceItems remains empty because CIViC does not identify supporting Evidence Lines for individual codes.

In [12]
for line in evidence_lines:
    criterion = line.evidenceOutcome.primaryCoding.code.root
    print(criterion, "Evidence Line hasEvidenceItems:", line.hasEvidenceItems)

Output:

OM1 Evidence Line hasEvidenceItems: None
OS2 Evidence Line hasEvidenceItems: None
OP4 Evidence Line hasEvidenceItems: None
OP1 Evidence Line hasEvidenceItems: None
OP3 Evidence Line hasEvidenceItems: None

Linking CIViC Evidence Items to the assertion

CIViC lists Evidence Items separately from the oncogenicity codes:

CIViC Assertion 202 EIDs

CIViCpy stores these URLs in assertion.hasEvidence. The output lists the Evidence Items linked to the assertion.

In [13]
evidence_item_urls = assertion.hasEvidence or []
evidence_item_links = []
print("Evidence Item URLs")
for evidence_item in evidence_item_urls:
    civic_eid_url = evidence_item.root
    evidence_item_links.append(f"[{civic_eid_url}]({civic_eid_url})")
    print(f"- {civic_eid_url}")
evidence_item_bullets = "\n".join(f"- {item}" for item in evidence_item_links)

Output:

Evidence Item URLs
- https://civicdb.org/links/evidence/74
- https://civicdb.org/links/evidence/12800
- https://civicdb.org/links/evidence/78
- https://civicdb.org/links/evidence/12711
- https://civicdb.org/links/evidence/12805
- https://civicdb.org/links/evidence/11723
- https://civicdb.org/links/evidence/12709

The output lists each CIViC Evidence Item associated with the assertion.

Each Evidence Line's specifiedBy method identifies the ClinGen/CGC/VICC guideline that defines how its criterion is weighed. The method's reportedIn field points to the following guideline document.

In [14]
guideline_document = evidence_lines[0].specifiedBy.reportedIn
print(guideline_document.name)
print(guideline_document.title)
print("PMID:", guideline_document.pmid)

Output:

Horak et al., 2022, Genet Med.
Standards for the classification of pathogenicity of somatic variants in cancer (oncogenicity): Joint recommendations of Clinical Genome Resource (ClinGen), Cancer Genomics Consortium (CGC), and Variant Interpretation for Cancer Consortium (VICC)
PMID: 35101336

Mapping CIViC approval

CIViC records its curation and review history with the assertion. In GKM, a VA-Spec Contribution connects the approval activity, agent, and date.

CIViC Assertion 202 Approvals

In [15]
approval = next(
    contribution
    for contribution in assertion.contributions or []
    if contribution.activityType.startswith("approval")
)
approval_date = str(approval.date).split(" ")[0]

for contribution in assertion.contributions or []:
    contributor = contribution.contributor
    print(f"{contribution.activityType} by {contributor.name} on {approval_date}")

Output:

approval.last_reviewed by CIViC on 2026-04-16

The approval output records the activity, agent, and date.

Reusing the connected interpretation

CIViC stores the Molecular Profile, Disease, Evidence Items, and assertion as related records. The bundle uses pointers for internal records and hasEvidence URLs for the assertion's Evidence Items.

In [16]
connected = civic_bundle.normalize(assertion)
print(f"Connected record: {connected['id']}")
print(f"Classification: {connected['classification']['primaryCoding']['code']}")
print(f"Variant: {connected['proposition']['subject']['name']}")
print(f"Disease: {connected['proposition']['object']['name']}")
print("Evidence lines:", ", ".join(evidence_codes))
print(f"Approving agent: {connected['contributions'][0]['contributor']['name']}")

Output:

Connected record: civic.aid:202
Classification: likely oncogenic
Variant: RET M918T
Disease: Medullary Thyroid Carcinoma
Evidence lines: OM1, OS2, OP4, OP1, OP3
Approving agent: CIViC

Construct the complete interpretation

The final cell assembles the connected records for display or exchange. It lists the assertion's Evidence Items and scored Evidence Lines separately because CIViC does not link an item to an individual oncogenicity code.

In [17]
from IPython.display import Markdown
from IPython.display import display as ipython_display

interpretation = (
    f"**{assertion.id.upper()}:**\n\n"
    f"**{origin.name.capitalize()}** **{variant.name}** is "
    f"**{classification}** for **{disease.name}**, evaluated under the "
    f"**{framework}** framework.\n\n"
    "The classification has a score of "
    f"**{total_score}** and is supported by {evidence_summary} evidence. CIViC associates "
    f"the following Evidence Items with the assertion:\n\n"
    f"{evidence_item_bullets}\n\n"
    "The classification was "
    f"{approval.activityType} by **{approval.contributor.name}** on "
    f"**{approval_date}**."
)
ipython_display(Markdown(interpretation))

Output:

CIVIC.AID:202:

Somatic RET M918T is likely oncogenic for Medullary Thyroid Carcinoma, evaluated under the ClinGen/CGC/VICC Guidelines for Oncogenicity, 2022 framework.

The classification has a score of 9 and is supported by functional domain assessment (OM1), functional data assessment (OS2), population data assessment (OP4), in silico impact assessment (OP1), somatic hotspot assessment (OP3) evidence. CIViC associates the following Evidence Items with the assertion:

The classification was approval.last_reviewed by CIViC on 2026-04-16.

Why GKM matters for exchange

The connected record can be serialized as JSON. An application that supports GKM can use CIViC's claim, classification, evidence, and provenance.

For example, ClinVar This can use GKM-formatted JSON to prepare and track ClinVar submissions.