Optimizing Real-World Data Collection: Genomics Electronic Data Capture Instrument

Author

David M. Miller, Sophia Z. Shalhout

Published

April 30, 2021

Abstract
We publish the data dictionary for our genomics data collection instrument

Overview

Advances in clinical oncology require a deep understanding of cancer biology. Clinico-Genomic Data (CGD) obtained from routine clinical practice can greatly increase our comprehension of tumor biology. However, there are a number barriers that impede capitalization of these critical real-world data (RWD). Paramount amongst these obstacles include prolonged time-to-analyses secondary to the difficulties of capturing data from heterogeneous sources, as well as the challenges of processing vast amounts of genomic information. These hurdles increase time-to-insight from RWD and threaten our ability to fully maximize on advances in molecular and information technologies.

In the real-world setting, genomic information resides in a variety of formats. The most common is a report from an institutional molecular pathology department or a commercial vendor. This information is often presented to the clinician or clinical research team in a semi-structured form. Collecting CGD of a patient cohort in a structured electronic data capture (EDC) system can facilitate analysis and maximize time-to-analytics and time-to-action.

We previously published an overview of a methodology and design of a REDCap-based system to facilitate capture of real-world data. Here, we provide an updated version of that instrument as well as the data dictionary for that capture tool in order to optimize the time-to-discovery for other real-world evidence projects.

Real-World Data Collection Pipeline for Clinico-Genomic Data:

The overall schema of the pipeline for the capture of real-world clinico-genomic information can be seen in the following figure:

Challenges to Clinico-Genomic Data Capture

Heterogeneous Platforms

As depicted above, CGD may be stored directly in a Electronic Health Record (EHR) if it is generated by the institution’s “in-house” molecular pathology department. More commonly, however, it may contained in a PDF from a commercial vendor. For example, here we provide a sample report from a PDF provided by Guardant Health. Of note, this is a sample report and does not represent actual patient data.

Sample Report #1

Sample Report #2

Another report, however, may have a substantively different presentation and format. For example, this a synthetic, but representative version of an institutional molecular pathology report. Again, this is a sample report and does not represent actual patient data.

Structured Data Capture to Optimize Real-World Clinico-Genomic Data

The heterogeneity in presentation and content of the various sources of CGD presents obstacles for analysis and interpretation. Therefore, utilizing tools to homogenize the data structures for which CGD is stored is essential. To address this problem, we developed a Genomics Instrument that can be used with the Research Electronic Data Capture system REDCap1,2.

Our Genomics Instrument captures the following CGD:

In addition we incorporate our in-instrument Quality Control procedure as previously described


Example of Genomics Instrument in REDCap

Here we provide an example of the capture of the above mentioned Guardant Report in REDCap



Structured Data Systems

Numeric Suffix Linker System

To develop a cohesive data structure we created a Numeric Suffix Linker System (NSLS). The NSLS links related elements of a genomic report with an underscore and a numeric (e.g “_1”). A given gene with a specific nucleotide and amino acid variant will be created with a tripartite group linked with the same underscore and numeric. For example, the ATM mutation found above in Sample Report #2.
ATM ENSP00000278616.4:p.Pro1166Ser (ENST00000278616.4:c.3496C>T)
will be structured in the Genomics Instrument as follows:

Data scientists can then link these structured data with the strings “variant” and “_1”, as these characters provide unique pairing with the data dictionary of the Genomics Instrument

Facilitated Analysis

Data collected in the Genomics Instrument can be easily exported into the Integrated Development Environment (IDE) of your choice, as depicted below. There, analyses of these CGD can yield novel insights into tumor biology.

Data Dictionary

Here we provide the data dictionary used to develop the Genomics Instrument. It can be downloaded here and used to recreate this data form in total.

Summary

Clinico-Genomic Data obtained in routine clinical practice can help better understand the genetic determinants of illness. Here we provide the rationale of a structured-data approach to the capture of real-world CGD as well as the data dictionary to lower the start-up efforts of other clinical investigators.

License

Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.