Synthetic Biological Data Generation via Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for assessing biological aging rely heavily on physical biological samples, which may not always be feasible, and there is a need for technologies that can predict biological age without direct sampling, especially for interventions targeting individual-level aging processes.

Innovation Solution

A method involving machine learning platforms that receive real biological data signatures, generate input vectors, and produce synthetic biological data signatures, allowing for predictions based on attributes like age, sex, and tissue types, enabling simulations for both aging acceleration and rejuvenation without requiring physical samples.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If physical biological samples are used for assessing biological aging, then measurement precision is improved, but ease of operation deteriorates due to the need for invasive sampling procedures

Engineering Contradiction:
Improvebiological age assessment accuracyVSAvoidsampling feasibility
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent creates synthetic biological data signatures that copy the essential information from real biological samples without requiring physical sampling. The generative model learns the underlying distribution of biological aging markers and generates synthetic data that preserves the statistical properties and relationships found in real samples, enabling non-invasive biological age assessment

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces a machine learning model as an intermediary between the need for biological sample analysis and the desire for non-invasive assessment. The model is trained on real biological data and then generates synthetic signatures that mediate the assessment process, allowing predictions without direct sample contact

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If real biological data signatures are collected from multiple subjects, then manufacturing precision of the aging model is improved, but loss of time increases due to extensive data collection requirements

Engineering Contradiction:
Improveaging model accuracyVSAvoiddata collection time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-training the generative model on a large dataset of real biological signatures before actual use. This upfront data collection and model training creates a robust foundation that can then generate accurate synthetic data without requiring time-consuming real-time sample collection for each assessment

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The generative model becomes self-sufficient after training, generating its own synthetic biological data signatures without requiring continuous access to real biological samples. The model serves itself by internally generating the data needed for assessments, eliminating the need for ongoing external data collection

Inventive Principle:
Principle #25Self-service

3Ease of operation

If synthetic biological data is generated using machine learning platforms, then ease of operation is improved by eliminating physical sampling, but reliability may deteriorate without direct biological validation

Engineering Contradiction:
Improvenon-invasive assessmentVSAvoidbiological data accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent implements feedback mechanisms where the generative model is trained on real biological data and its outputs are continuously refined by comparing generated synthetic signatures against actual sample data. This feedback loop ensures that the synthetic data maintains high fidelity to real biological patterns while enabling non-invasive operation

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent creates a dynamic system where the generative model can adapt and refine its synthetic data generation based on the specific characteristics of the input query. The model dynamically adjusts the generated signatures to match the biological plausibility and statistical properties learned from training data, ensuring reliability while maintaining operational ease

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20220310196A1Synthetic biological characteristic generator based on real biological data signatures
Publication Date: 2022.09.29 INSILICO MEDICINE IP LTD
  • US20220310196A1 patent drawing
  • US20220310196A1 patent drawing
  • US20220310196A1 patent drawing

AI summary

Creating synthetic biological data for a subject can include: (a) receiving a real biological data signature derived from a biological sample of the subject; (b) creating input vectors based on the real biological data signature; (c) inputting the input vectors into a machine learning platform; (d) generating a predicted biological data signature of the subject based on the input vectors, wherein the predicted biological data signature includes synthetic biological data specific to the subject; and (e) preparing a report that includes the synthetic biological data of the subject. Biological pathway activation signatures can be genomics, transcriptomics, proteomics, metabolomics, lipidomics, glycomics, methylomics, or secretomics. Conditioning latent codes of the input vectors in a latent space of the machine learning platform with at least one constraint of an attribute of the subject is performed so the predicted biological data signature is based on the at least one constraint.