Nucleotide Sequence Encoding With Cube-Based Symmetry Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The lack of a rigorous mathematical analysis of genetic information due to the absence of a 'grammar' relating symbolic representations of biological sequences hampers the exploitation of symmetries in nucleotide sequences.

Innovation Solution

A method of representing nucleotide sequences as a cube or tetrahedron, with each vertex assigned a nucleotide base, and encoding them in a matrix form to exploit symmetries, allowing for the detection of modifications and defining symmetry-invariant markers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If symbolic representations (strings of characters) are used to describe biological sequences, then the representation is simple and intuitive, but the ability to perform rigorous mathematical analysis is hampered due to lack of grammar structure

Engineering Contradiction:
Improveease of representationVSAvoidmathematical analysis capability
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent transforms the symbolic representation parameters by mapping nucleotide sequences to geometric structures (cube vertices, tetrahedron vertices, matrix elements) instead of simple character strings. This parameter change enables the application of group theory and symmetry operations, providing rigorous mathematical analysis capability while preserving the essential information of the biological sequences.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces geometric dimensions by representing nucleotide sequences as paths through cube or tetrahedron vertices, or as matrix elements. This dimensional transformation from 1D string to 3D geometric structure or 2D matrix enables exploitation of spatial symmetries and mathematical operations that are not available in traditional symbolic representations.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Device complexity

If traditional symbolic representations are used, then the data structure is simple, but symmetries in nucleotide sequences cannot be exploited for analysis

Engineering Contradiction:
Improvedata structure complexityVSAvoidsymmetry exploitation capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent employs both symmetric and asymmetric representations depending on the analysis needs. The cube and tetrahedron provide symmetric frameworks that capture conservation laws, while the specific sequence paths and matrix assignments can be asymmetric to reflect biological reality. This duality enables both symmetry exploitation and biological accuracy.

Inventive Principle:
Principle #4Asymmetry

Solution Approach 2:

The geometric representation framework serves multiple functions: it provides a visualizable structure, enables group theory analysis, captures symmetry operations, and maintains sequence information. The cube/tetrahedron/matrix representation is universally applicable to any nucleotide sequence while adapting to different analytical requirements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If geometric representations (cube, tetrahedron, matrix) are used to represent nucleotide sequences, then mathematical analysis and symmetry exploitation are enabled, but the representation complexity increases

Engineering Contradiction:
Improvemathematical analysis capabilityVSAvoidrepresentation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a geometric copy or model of the nucleotide sequence rather than replacing the original sequence data. The cube vertices, tetrahedron vertices, or matrix elements serve as geometric copies that encode the sequence information, enabling mathematical analysis without losing the biological meaning. This copying approach adds analytical capability while keeping the original sequence intact.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250292867A1Encoding genetic sequencing information and uses thereof
Publication Date: 2025.09.18 INST PASTEUR DE DAKAR
  • US20250292867A1 patent drawing
  • US20250292867A1 patent drawing
  • US20250292867A1 patent drawing

AI summary

A computer-implemented method of representing a selected sequence of RNA, and/or DNA nucleotides and/or nucleotide analogues, comprising the steps of (a) defining a cube using nucleotide bases, wherein each of the eight vertices of the cube is assigned to a nitrogen-containing base from the subset [A (adenine), C (cytosine), G (guanine), T (thymine)], wherein nitrogen-containing bases T (thymine) and U (uracil) are herein understood as equivalent and thus interchangeable, and wherein each of the vertices of the cube is assigned to a nucleotide base from the subset such that for each vertex, the assigned base is directly connected to every other base type of the subset through an edge, (b) assigning the first nucleotide base of the selected sequence of nucleotides to a vertex of the cube to which the nitrogen-containing base type of the nucleotide has been assigned and sequentially assigning each subsequent nucleotide of the selected sequence of nucleotides to a corresponding vertex in the cube or the tetrahedron, such that the assigned vertex is either directly connected to the vertex of the previous nucleotide base through an edge of the cube; wherein in case the nucleotide base is equal to the previous nucleotide base, the nucleotide base is assigned to the same vertex of the cube as the vertex to which the previous nucleotide base has been assigned to; and wherein the selected sequence of nucleotides and/or nucleotide analogues comprises nitrogen-containing bases selected from A (adenine), C (cytosine), G (guanine), T (thymine), and/or U (uracil).