Neural Network Biomolecule Embedding for Plasma Protein Noise Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for early disease detection, particularly in cancer, face challenges in characterizing plasma proteins due to their wide concentration range and complex biochemical workflows, which are not practical for discovery studies, limiting the validation and replication of biomarkers and clinical performance.
Innovation Solution
A neural network is trained to generate embeddings of biomolecule descriptors by reducing dimensionality and filtering noise, using a latent space with a denoised embedding, allowing for the identification of novel biomarkers and improved biomolecule-based biomarker discovery through multi-surface panel platforms for plasma biomolecule profiling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If complex biochemical workflows are used to characterize plasma proteins, then measurement precision may be improved, but device complexity and ease of operation deteriorate
Solution Approach 1:
The patent replaces complex mechanical/biochemical workflows with a computational machine learning system. The neural network processes raw plasma protein data directly, substituting multiple biochemical steps with an information-processing system that achieves comparable or superior measurement precision without the operational complexity
Solution Approach 2:
The patent introduces machine learning models as an intermediary between raw plasma protein data and biomarker identification. This intermediary layer filters noise and extracts meaningful patterns, improving measurement precision while simplifying the overall process by consolidating multiple biochemical interpretation steps into a computational framework
2Measurement precision
If complex biochemical workflows are used, then measurement precision may be improved, but ease of operation worsens
Solution Approach 1:
The computational system replaces manual biochemical workflows with automated machine learning processing, significantly improving ease of operation. The neural network can process large datasets without manual intervention, making discovery studies more practical and scalable while maintaining measurement precision
3Manufacturing precision
If dimensionality reduction is applied to biomolecule descriptors, then loss of information may occur, but manufacturing precision (embedding quality) is improved
Solution Approach 1:
The patent extracts only the most relevant features from high-dimensional biomolecule descriptors through neural network processing. The model identifies and retains meaningful patterns while discarding noise and redundant information, achieving high embedding quality without significant loss of useful information
Solution Approach 2:
The neural network applies different processing weights to different dimensions of the biomolecule descriptors, treating informative features differently from noisy features. This local quality approach preserves critical information while filtering out noise, improving embedding quality selectively
4Reliability
If plasma proteins are analyzed for biomarker discovery, then reliability of biomarker identification is improved, but object-affected harmful factors (noise from wide concentration range) increase
Solution Approach 1:
The patent transforms the concentration data through neural network processing, converting raw intensity values across wide concentration ranges into normalized embedding spaces. This parameter transformation reduces the harmful effect of concentration variability while preserving the reliability of biomarker identification patterns
Data Source
AI summary
In some aspects, the present disclosure describes a method for determining a biological state associated with a polyamino acid descriptor. In some cases, the method comprises receiving the polyamino acid descriptor comprising at least one dimension representing a polyamino acid association with a given assay method. In some cases, the method comprises generating, in a latent space, a latent descriptor based at least in part on the polyamino acid descriptor, and wherein the latent descriptor comprises sufficiently fewer dimensions than the polyamino acid descriptor such that at least a portion of information in the polyamino acid descriptor is lost in the latent descriptor. In some cases, the method comprises determining, based at least in part on the latent descriptor, the biological state associated with the polyamino acid descriptor.


