Virtual Landscape for Ranking Biological Sequences
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems are unable to effectively identify new chemical entities or missing structures in low-dimensional visualizations and generate additional data related to chemical compounds, particularly for biologic identifiers like nucleotide or protein sequences, which require additional processing and analysis for evaluation in lower dimensional visualizations.
Innovation Solution
A computer-implemented method that generates a virtual n-dimensional manifold using a manifold-generator module to convert coded representations of chemical or biologic sequences into a virtual space, where unsupervised learning algorithms place these representations based on similarity, allowing for the filtering and ranking of new chemical or biologic entities not initially present in the dataset.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If dimensionality reduction techniques are used to visualize chemical structures in low-dimensional space, then visualization complexity is reduced and relationships become easier to see, but the system cannot generate or identify new chemical entities that are absent from the original dataset
Solution Approach 1:
The patent embeds high-dimensional chemical structure data into low-dimensional space using dimensionality reduction techniques (such as t-SNE or autoencoders) to create visual representations that are easier to interpret. Simultaneously, it introduces a generative model that operates in this learned low-dimensional space to synthesize new chemical structures, effectively adding a generative dimension to the visualization system without losing the ability to create novel entities.
2Productivity
If conventional analysis systems process large variable datasets and present lower dimensional visualizations, then data processing efficiency is improved, but the systems are entirely unable to identify absent chemical structures that conform to the reduced dimensional space
Solution Approach 1:
The system implements a feedback loop where the generative model continuously proposes new chemical structures, these proposals are evaluated against the learned representation of known chemicals, and the results are used to refine both the visualization and the generation process. This feedback mechanism ensures that newly generated structures are chemically valid and conform to the patterns learned from the original dataset, preventing information loss while maintaining processing efficiency.
Solution Approach 2:
The patent performs preliminary dimensionality reduction and pattern learning on the training dataset before attempting to identify or generate new structures. This preliminary action creates a robust low-dimensional representation that captures essential chemical relationships, enabling both efficient processing and accurate identification of novel structures that fit within the learned chemical space.
3Manufacturing precision
If biologic identifiers such as nucleotide or protein sequences are converted into coded forms for analysis, then data standardization is improved, but additional processing and analysis are required to evaluate sequences in lower dimensional visualizations
Solution Approach 1:
The patent employs a universal coding scheme and dimensionality reduction approach that handles multiple types of biologic identifiers (nucleotide sequences, protein sequences, and other molecular data) through the same processing pipeline. The generative model is designed to work with this universal representation, reducing the need for separate processing paths and thereby lowering overall system complexity while maintaining high standardization across different data types.
Data Source
AI summary
The present invention is directed to generating an n-dimensional map using the results of a query for compounds enumerated within a collection of documents describing a particular biological target of interest and a curated set of sequences, such as but not limited to, protein or nucleotide sequences not enumerated in the collection of documents. Both sets of sequences (document coded and curated coded) are converted into coded forms and placed in the n-dimensional map. One or more processors are configured to evaluate the distance between the curated coded forms and the closest cluster of document coded forms. Based on the distance between a coded form and the document coded forms, the curated coded forms can be ranked regarding the likelihood of interacting with the particular biological target.


