Genomic Variant Normalization for Document Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current approaches to normalizing genomic variants in databases are limited in extracting complex or structural variations and are simplistic, leading to incomplete information retrieval during searches, as there is no universal standard for annotation, resulting in missed data and relationships.

Innovation Solution

A rules-based approach for normalizing annotations and annotating documents within a content repository, using unique identifiers to standardize gene variants and mapping their relationships, allowing for comprehensive search and retrieval of genomic information across various forms such as DNA, RNA, and amino acid sequences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If current approaches to normalizing variants are used, then simple variant annotations can be processed, but complex or structural variations cannot be extracted

Engineering Contradiction:
Improveability to extract different types of variationsVSAvoidinformation about complex or structural variations
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent applies universality by creating a normalization system that handles multiple types of genomic variations (SNVs, indels, copy number variations, structural variations) through a single unified framework. The system uses a comprehensive dictionary that maps various representation formats of gene variants to standardized forms, enabling the extraction and normalization of diverse variation types that previous simplistic approaches could not handle.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Quantity of substance

If a single variant is represented in many different manners within a corpus of documents, then comprehensive information is stored, but search technologies may identify only a subset of available information

Engineering Contradiction:
Improveamount of information about gene variantsVSAvoidinformation retrieved during search
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent introduces a normalization dictionary as an intermediary layer between the diverse representations of gene variants in documents and the search queries. This dictionary maps multiple representation formats (different nomenclatures, synonyms, formats) to standardized variant identifiers, enabling search technologies to retrieve all information about a gene variant regardless of how it was represented in the original documents.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If no universal standard for annotation is used, then flexibility in representing variants is maintained, but information completeness and retrieval accuracy deteriorate

Engineering Contradiction:
Improveflexibility in variant representationVSAvoidaccuracy of variant identification
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies parameter changes by transforming the representation parameters of gene variants from diverse, inconsistent formats into a standardized format. The normalization process changes the parameters of variant representation (nomenclature, formatting, identification) while preserving the underlying biological meaning, thereby improving identification accuracy without losing representation flexibility.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10949607B2Automated document filtration with normalized annotation for document searching and access
Publication Date: 2021.03.16 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10949607B2 patent drawing
  • US10949607B2 patent drawing
  • US10949607B2 patent drawing

AI summary

Computer-based methods, systems, and computer readable media for managing documents within a content repository or documents within the document subsets are provided. Variant annotations within the documents may be normalized to a standard nomenclature. A request is processed for the documents including one or more search terms, where the search terms pertain to one or more from a group of genes/gene variants, drugs, and cancer terms. Documents are identified that satisfy the request by comparing the one or more search terms to the normalized annotations and specific sections of the documents, and determining a relevance of a document based on the comparison and a frequency of the one or more search terms in each of the specific sections. The identified documents are ranked in accordance with a priority based on the determined relevance.