A bioinformatics processing system and method applied to synthetic biology
By acquiring real-time proteomics data for polymorphic context decoding and cross-domain graph structure deconstruction, multi-scale causal structure maps are generated, overcoming the limitations of traditional methods in processing low-similarity proteins. This enables high-precision protein functional design and expression pathway reconstruction, adapting to the multi-dimensional regulatory needs of synthetic biology.
Patent Information
- Application Number
- CN202511145822.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Traditional protein data processing methods have limitations when dealing with low-similarity proteins, predicting novel structures or unknown functional regions, and are difficult to meet the multi-dimensional requirements of synthetic biology for predictable functions, controllable design, and tunable pathways of protein modules in specific engineering scenarios.
By acquiring real-time proteomics data, we perform polymorphic context decoding and cross-domain graph structure deconstruction to generate multi-scale causal structure maps, conduct protein functional expression function reduction and scheduling simulations, dynamically assemble conformation decision graphs, perform multi-scenario virtual response simulations and path adaptability verifications, and finally perform multi-dimensional clustering and variability screening to form the optimal expression conformation set.
It achieves cross-scale, multi-dimensional, and high-precision modeling of protein expression regulation, improves the coupling expression ability of protein behavioral information and structural semantics, ensures the accuracy, stability and executability of the model, and adapts to the engineering needs of complex regulatory networks.
Smart Images

Figure CN120708691B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of protein-related data processing, and in particular to a biological information processing system and method applied to synthetic biology. BACKGROUND
[0002] Protein function design, path optimization and component combination are important functional modules in synthetic biology. Protein structure and function relationship modeling, sequence design optimization and function prediction have wide application prospects in gene circuit construction, metabolic pathway modification and cell factory regulation. Traditional protein data processing methods are mainly based on sequence alignment, structure template mapping and expert experience rules, including BLAST, PSI-BLAST, HMMER, SWISS-MODEL and other tools. These methods have certain practical value in the initial homology recognition of protein sequences and structure template alignment, but have great limitations in processing low-similarity proteins, predicting novel structures or unknown functional regions. In addition, traditional methods often rely on static databases and manual annotation, lack the ability to model dynamic structural changes, environmental response behavior and cross-scale regulation paths of proteins, and are difficult to meet the current multi-dimensional needs of protein modules in specific engineering scenarios in synthetic biology, such as function predictability, design controllability and path adjustability. SUMMARY
[0003] Therefore, it is necessary to provide a biological information processing system and method applied to synthetic biology to solve at least one of the above technical problems.
[0004] To achieve the above-mentioned purpose, a biological information processing method applied to synthetic biology comprises the following steps:
[0005] Step S1: acquiring real-time proteomics data, and decoding the real-time proteomics data in a polymorphic context to obtain a protein expression context matrix;
[0006] Step S2: deconstructing and reconstructing the behavior characteristics in the protein expression context matrix based on a cross-domain graph structure to obtain a reconstructed protein matrix; analyzing the biological response path-protein domain interaction based on the reconstructed protein matrix to obtain a multi-scale causal structure graph;
[0007] Step S3: reducing the path nodes in the multi-scale causal structure graph based on a protein function expression function to obtain a function expression minimum module combination unit; and performing structure-function transformation rule extraction based on the function expression minimum module combination unit to obtain a function mapping unit set;
[0008] Step S4: Taking the functional mapping unit set as input, performing protein scheduling simulation, and performing unit configuration path validity and structure stability evaluation according to the protein scheduling simulation result, thereby dynamically splicing to form a configuration decision graph;
[0009] Step S5: Performing multi-scenario virtual response simulation on the paths in the configuration decision graph, and performing path adaptability verification and scoring on the multi-scenario virtual simulation results to obtain a path adaptability feedback table;
[0010] Step S6: Multi-dimensional clustering and variability screening are performed on the path adaptability feedback table to obtain an optimal expression configuration set.
[0011] The application relates to a biological information processing method applied to synthetic biology, which forms a cross-scale, multi-dimensional and high-precision protein expression regulation modeling process around key links such as protein function design, expression path modeling and structure-function mapping, and can effectively make up for the deficiencies of traditional methods in low-similarity area identification, new structure-function prediction and expression path reconstruction. By introducing expression semantic node nested extraction and cross-domain feature relationship modeling, the coupling expression capability between protein behavior information and structure semantics can be significantly improved, and the expression integrity of the model to the context semantics of the protein is further enhanced. The construction of the structure compression representation graph can realize the effective compression of the graph scale on the basis of ensuring the integrity of the information structure, and reduce the calculation redundancy and complexity in the subsequent structure projection and graph modeling process. The coordinate coding of the domain space projection combined with the coordinate system of the standard protein domain database not only ensures the biological rationality of the node mapping, but also improves the spatial positioning accuracy of the behavior node in the domain. The context coupling retrieval and neighborhood cross-validation operation strengthens the semantic and spatial consistency between the mapped nodes and the database domains, and improves the accuracy and stability of the structure mapping. The chimeric metric modeling fuses the domain function clustering cluster information and the structure generic similarity, so that the generated chimeric graph has structure-function dual consistency, and the situation of structure conflict or function loss is avoided. The path perception modeling and graph convolution aggregation can fully extract the key response path and its causal response strength in the expression path, and the generated causal response weight matrix not only has good interpretability, but also can provide clear index support for subsequent path pruning and function reconstruction. Setting the maximum path depth as 6 helps to control the complexity of the expression path, avoid the dilution and accumulation of interference caused by too long path, and set the sparse control parameter as 0.03, which can remove redundant expressions under the premise of ensuring function integrity, improve the calculation efficiency and expression accuracy of the model, and set the hash distance threshold as 0.15 and the intra-cluster similarity as 85%, which can accurately control the granularity of expression function aggregation and strengthen the semantic consistency between function expressions. In the expression function pruning stage, the double-index screening mechanism of the causal regulation weight threshold 0.6 and the average gradient weight 0.65 can not only ensure the core position of the reserved sub-function in the path, but also effectively eliminate the redundant part with insufficient expression activity. The setting of the matching tolerance threshold 0.1 improves the selectivity and reliability of the matching of the expression topology stability rule, and helps to ensure that the expression combination unit constructed subsequently has stable structure and high availability. In the function structure graph modeling stage, the structure graph fingerprint is generated through topological rearrangement, and the symbolic function mapping is performed combined with the structure-function mapping relationship, so that the path mapping and structure identification have structural uniqueness and functional accuracy. The generation of the rule expression substructure set combined with the rule reduction operation helps to extract the expression regulation core structure, and provides a controllable input space for subsequent transformation path modeling.The double threshold setting of a score ≥ 0.85 and a structure energy level ≤ 0.25 adopted in the path scoring ensures that the finally screened stable configuration path has both high functional accessibility and low structure coupling energy, so as to realize the dynamic optimal matching of the expression path and the structure configuration. In addition, by constructing the path execution feature set and the fusion model, the execution complexity and conflict characteristics of the path are uniformly modeled, so that the generated scheduling path not only realizes the collaborative combination at the structure level, but also has good stability and scheduling robustness at the execution level, ensuring that the protein modules have engineering executability and path optimizability under the complex regulation network.
[0012] Optionally, the present specification also provides a biological information processing system applied to synthetic biology, for executing the biological information processing method applied to synthetic biology as described above, and the biological information processing system applied to synthetic biology comprises:
[0013] A context decoding module is configured to acquire real-time proteomics data, and perform polymorphic context decoding on the real-time proteomics data to obtain a protein expression context matrix.
[0014] An interaction modeling module is configured to perform cross-domain graph structure deconstruction and reconstruction on behavior characteristics in the protein expression context matrix to obtain a reconstructed protein matrix, and analyze biological response path-protein domain interaction based on the reconstructed protein matrix to obtain a multi-scale causal structure graph.
[0015] A functional transformation analysis module is configured to perform protein function expression function reduction on path nodes in the multi-scale causal structure graph to obtain a functional expression minimum module combination unit, and perform structure-function transformation rule extraction based on the functional expression minimum module combination unit to obtain a functional mapping unit set.
[0016] A protein scheduling simulation module is configured to take the functional mapping unit set as input to perform protein scheduling simulation, and perform unit configuration path effectiveness and structure stability evaluation based on a protein scheduling simulation result, so as to dynamically splice a configuration decision graph.
[0017] A virtual response simulation module is configured to perform multi-scenario virtual response simulation on paths in the configuration decision graph, and perform path adaptability verification and scoring on multi-scenario virtual simulation results to obtain a path adaptability feedback table.
[0018] A variability screening module is configured to perform multi-dimensional clustering and variability screening on the path adaptability feedback table to obtain an optimal expression configuration set.
[0019] The application is a biological information processing system applied to synthetic biology, which can realize any one of the biological information processing methods applied to synthetic biology of the application, and is used as a medium for joint operation and signal transmission between modules to complete the biological information processing method applied to synthetic biology. The modules in the system cooperate with each other, thereby improving the precision, stability and controllability of protein information in the process of structure analysis, function identification and regulation path construction. BRIEF DESCRIPTION OF DRAWINGS
[0020] Other features, objects and advantages of the application will become more apparent from the following detailed description of non-limiting embodiments made with reference to the accompanying drawings:
[0021] Fig. 1 A step flow diagram of the biological information processing method applied to synthetic biology of the application;
[0022] Fig. 2 A detailed step flow diagram of step S1 in the application;
[0023] Fig. 3 A detailed step flow diagram of step S2 in the application;
[0024] The implementation of the object of the application, functional features and advantages will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0025] The technical method of the application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.
[0026] In addition, the accompanying drawings are only schematic diagrams of the application, and are not necessarily drawn to scale. The same reference signs in the drawings represent the same or similar parts, and thus repeated descriptions thereof will be omitted. Some block diagrams shown in the drawings are functional entities, which do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0027] It should be understood that, although the terms "first", "second" or the like can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the example embodiments, a first element can be referred to as a second element, and similarly a second element can be referred to as a first element. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0028] To achieve the above object, please refer to Figs. 1 to 3 The application provides a biological information processing method applied to synthetic biology, which comprises the following steps:
[0029] Step S1: acquiring real-time proteomics data, and performing polymorphic context decoding on the real-time proteomics data to obtain a protein expression context matrix;
[0030] In this embodiment, a Thermo Fisher Orbitrap Exploris 480 mass spectrometer is combined with an Agilent 1290 Infinity II high-throughput liquid chromatography platform to collect multiple batches of protein samples of a HeLa cell strain under different induction states. The resolution set during mass spectrometry collection is 120,000, the full scan range is set to 350-1,800 m / z, and the liquid phase gradient elution time is set to 90 minutes. Public protein expression data related to the selected cell state in the cell expression database (such as Human Protein Atlas) is introduced synchronously. After collecting the data, the data streams from different sources are processed synchronously in time, the time stamps are aligned by using a linear interpolation method, and the inconsistency problem in data collection time is solved. Subsequently, according to a polymorphic context decoding model, the protein expression signal is decomposed into different context states, and the model comprises a three-layer nested structure: the first layer is an environmental factor coding layer, which is used to capture different stimulation conditions inside and outside the cell; the second layer is an expression pattern mapping layer, which represents the expression difference of the protein under different states; and the third layer is a behavior fusion layer, which integrates the expression data to form a three-dimensional protein expression context matrix with a dimension of [N protein x M time point x K context state]. The matrix has a tensor structure, each element corresponds to the expression amount of the protein under a specific time and context, and the numerical range is normalized to 0-1, which is convenient for subsequent processing.
[0031] Step S2: performing cross-domain graph structure deconstruction and reconstruction on the behavior characteristics in the protein expression context matrix to obtain a reconstructed protein matrix; analyzing biological response path-protein domain interaction based on the reconstructed protein matrix to obtain a multi-scale causal structure atlas;
[0032] In this embodiment, the behavior features in the protein expression context matrix are extracted to construct a cross-domain graph structure. The graph consists of nodes and edges, with nodes representing proteins and their expression behavior channels, and edges representing the interaction between proteins in different biological response paths. The construction of the graph is based on the adjacency matrix A and the feature matrix X, which defines the connection between nodes and contains the behavior feature vectors of protein expression. For this graph, graph decomposition techniques are used to remove low correlation edges by edge weight threshold screening, and a sparse adjacency matrix is reconstructed to generate a reconstructed protein matrix. Subsequently, by combining the standard protein domain database, the behavior nodes in the graph are mapped to the corresponding domain coordinate system to realize behavior-structure embedding. Through the layer-by-layer aggregation of the multi-scale graph, the causal relationship between protein domains on the path is extracted to obtain a multi-scale causal structure graph. This graph is represented as a multi-layer directed graph structure, with different scales of causal paths between layers.
[0033] Step S3: Perform protein function expression function reduction on the path nodes in the multi-scale causal structure graph to obtain a function expression minimal module combination unit; and perform structure-function transformation rule extraction based on the function expression minimal module combination unit to obtain a function mapping unit set;
[0034] In this embodiment, in the multi-scale causal structure graph, the protein nodes and their function expression information on each path are extracted, the causal chain is unfolded, the maximum path depth is limited to 6, and a path expression function set is formed. Symbolic simplification is performed on these function sets, with a threshold of 0.03, and low-weight expression functions are removed to generate a function sparse structure tensor. Local sensitive hashing method is used to process the expression functions, with a hash distance threshold of 0.15, and the expression functions are clustered based on a similarity of ≥ 85%. Combining the causal weight graph and the gradient propagation flow graph, core sub-functions are selected and low-weight functions are pruned to obtain a function pruning set. According to the pre-defined stable function patterns in the expression topology stability rule library, the functions are matched and selected to form an expression topology stability graph. Finally, the pruning functions are reorganized and sorted to output the function expression minimal module combination unit, which is a combination of 3-7 core function sub-units, facilitating subsequent function mapping.
[0035] Step S4: Taking the function mapping unit set as input, perform protein scheduling simulation, and evaluate the unit configuration path effectiveness and structure stability according to the protein scheduling simulation results, to dynamically splice a configuration decision graph;
[0036] In this embodiment, the functional mapping unit set is taken as the input of the scheduling simulation, and the simulation platform sets up the execution process of simulating protein scheduling, including scheduling path, functional unit dependency relationship and scheduling sequence. A unit scheduling graph is constructed, in which the nodes represent functional mapping units and the edges represent dependency relationships, and the directed graph structure clearly shows the execution sequence. For the unit scheduling graph, the bottleneck link in the path, the feedback loop and the timing conflict are evaluated, the path conflict vector is extracted, and the parameters include the timing conflict intensity. The path connection tightness and the energy level transition distribution are calculated, and the structure energy distribution map is generated. Combined with the path conflict vector and the energy map, a path performance score table is formed, and the score ≥ 0.85 is an excellent path. Hierarchical clustering is performed on the score table, combined with evolutionary trend analysis, and the path with stable structure and excellent performance is screened out, and the redundant path is cut off, and finally the configuration decision graph is dynamically spliced to form a directed weighted graph, and the weight reflects the path stability and execution efficiency.
[0037] Step S5: Perform multi-scene virtual response simulation on the paths in the configuration decision graph, and verify and score the multi-scene virtual simulation results to obtain a path adaptability feedback table;
[0038] In this embodiment, based on the configuration decision graph, the paths with higher scores are selected, and response simulation is performed in three virtual biological scenes, respectively simulating cell proliferation environment, stress response environment and metabolic regulation environment. 1) Cell proliferation environment: simulate the protein regulation dynamics in the cell cycle, set the cell cycle parameter to 24 hours, and the simulation time step is 5 minutes. 2) Stress response environment: simulate the protein expression adjustment of cells to external environmental stress (such as heat shock, oxidative stress), set the stress intensity gradient to 0~1, and the simulation duration is 1 hour. 3) Metabolic regulation environment: simulate the regulation of proteins in the metabolic network, the concentration change range of metabolic products is 0~100 μM, the simulation time step is 10 minutes, and the duration is 2 hours. In each scene, a multi-parameter biological response model is used to simulate the activation of protein expression nodes in the path, and a response curve is generated. The multi-parameter biological response model used in the simulation process is derived from an improved version of the classical cell signaling model, combined with protein expression dynamics and molecular regulation mechanisms. The model is composed of multiple coupled differential equations, and the core includes activation function, feedback regulation module and regulation coupling term: the activation function adopts Sigmoid function form, which describes the response of protein expression amount to activation signal, and the form is where the parameters control the steepness of the curve, The activation threshold is represented. The feedback regulation module designs a negative feedback loop to simulate the dynamic balance of protein expression levels by adjusting the protein degradation rate and the concentration of transcriptional repressor. The regulatory coupling term is used to describe the interaction effect between proteins and the influence of environmental factors, which adopts a linear weighted superposition form. The weight parameters are preset according to biological experimental data, and the range is adjusted between 0.1 and 0.9. The model input is the initial expression state of each protein in the pathway and the environmental parameter, and the output is the expression intensity in the form of time series. The model structure is verified by biological experiments, which can accurately reflect the time dynamics and intensity changes of protein expression response. The model parameters can be adjusted according to different simulation scenarios to ensure the biological relevance and practical applicability of the simulation results. In the simulation, the following indicators are collected: pathway activation time (unit: minutes), peak expression intensity (0-1 normalization), response duration (time length above threshold 0.7), and scene-specific regulation efficiency index. All data are summarized into a "pathway adaptability feedback table". The feedback table is saved in matrix form, with rows corresponding to pathways and columns corresponding to indicators of each scene. The data uses double-precision floating-point format to ensure simulation accuracy. Subsequently, the comprehensive adaptability score is calculated by the weighted average method, and the weights are set to 0.4 (cell proliferation), 0.35 (stress response), and 0.25 (metabolic regulation) according to the importance of the scene, and normalized to facilitate scoring, sorting, and screening.
[0039] Step S6: Multidimensional clustering and variability screening are performed on the pathway adaptability feedback table to obtain an optimal expression configuration set.
[0040] In this embodiment, multi-dimensional time series statistical analysis is performed on the path adaptability feedback table to construct a path adaptability feature matrix. The rows of the matrix represent path numbers, and the columns represent adaptability index changes at different time steps. The variance, kurtosis and skewness of each path on each index are calculated, and the threshold is set to variance > 0.02 to screen out a subset of paths with significant adaptability fluctuations. The subset is screened for variability, and the similarity between time series is evaluated using the dynamic time warping (DTW) method to further cluster into multiple variability path clusters. For each variability path cluster, analyze the adaptability score trend to identify potential optimization points and deficiencies. Combined with the path simulation model, design targeted adjustment strategies (such as functional module reordering, path node activation threshold adjustment, etc.), simulate and verify the variability path and optimize iteratively. During optimization, record the adaptability index changes before and after optimization in real time to ensure effective improvement. Finally, integrate the optimized path set with the original stable path set to establish a unified path expression configuration set. This set contains two parts: one is the path that has been verified in multiple scenarios and has stable adaptability, and the other is the path whose performance has been significantly improved after optimization. Through a comprehensive scoring mechanism, reorder the optimal expression configuration set to output the optimal expression configuration set for subsequent experimental design and protein function regulation applications. The set is stored in the form of a weighted directed graph, with nodes representing functional modules and edge weights reflecting path reliability and regulation efficiency, facilitating subsequent path selection and dynamic adjustment.
[0041] Especially important is that the variability screening in step S6 is specifically:
[0042] Based on the path adaptability feedback table, extract the adaptability fluctuation data of each path to construct a path adaptability feature matrix;
[0043] In this embodiment, the adaptability performance of each path at different time points is extracted from the multi-scenario index data in the path adaptability feedback table to construct a path adaptability feature matrix. The rows of the matrix represent different paths, and the columns correspond to the time series adaptability index in each simulation scenario, such as peak expression intensity and response duration. The matrix elements are double-precision floating-point numbers, and the matrix dimensions are usually path number x time step x index dimension. Through this matrix, the dynamic adaptability change characteristics of the path can be comprehensively reflected, facilitating subsequent multi-dimensional analysis.
[0044] Statistical analysis is performed on the path adaptability feature matrix to calculate the variance, skewness and kurtosis of each index, and a set of paths with significant fluctuations is identified, resulting in a set of significant fluctuation paths;
[0045] In this embodiment, for the constructed path adaptability feature matrix, statistical quantities are calculated for each index, including variance (measuring the degree of data fluctuation), skewness (describing the asymmetry of data distribution), and kurtosis (reflecting the sharpness of data distribution). For example, when calculating the variance of the peak expression intensity sequence, a time window length of 30 time points is selected, and the skewness and kurtosis are calculated based on the entire time series. By threshold screening of these statistical characteristics (such as variance threshold set to 0.05, skewness threshold ±0.3, and kurtosis threshold 3), a set of paths with significant fluctuations is preliminarily screened out, facilitating the identification of potential abnormal dynamic responses.
[0046] Perform multi-scale time series analysis on the set of significant fluctuation paths, mark abnormal timing fluctuations, and obtain an abnormal fluctuation path marker vector;
[0047] In this embodiment, the time series data of each path in the set of screened significant fluctuation paths is input into a multi-scale time series decomposition framework, and is decomposed into trend, periodic, and random fluctuation components. By setting threshold rules (variance > 0.05 indicates significant fluctuation, indicating the presence of dynamic instability), abnormal detection is performed on the random component, and abnormal timing fluctuation events are marked. The abnormal marking result is stored in the form of a binary vector, with the vector length equal to the time series length, 1 representing an abnormal point and 0 representing normal. This step helps reveal the sudden changes and abnormal performance of path adaptability in the time dimension.
[0048] Perform false positive path screening on the abnormal fluctuation path marker vector to obtain a real variation path subset;
[0049] In this embodiment, for the obtained abnormal fluctuation path marker vector, a false positive exclusion mechanism is designed in combination with common noise and model error characteristics in historical simulation data. By setting a minimum continuous abnormal length threshold (e.g., not less than 5 time points of continuous abnormal points) and an abnormal frequency threshold, sporadic and non-persistent false positive fluctuation paths are eliminated. Finally, a path subset with real biological variation significance is screened out, ensuring that subsequent analysis focuses on reliable dynamic abnormal paths.
[0050] Perform clustering analysis on the real variation path subset, group according to variation type and amplitude, and form variation path clusters;
[0051] In this embodiment, the real variation path subset is grouped into path clusters based on its variation type (such as peak surge, sustained decline, periodic fluctuation, etc.) and variation amplitude index using feature similarity measurement. Each cluster contains paths with similar performance, facilitating the analysis of their common dynamic behavior patterns and potential mechanisms. The clustering result is represented by constructing a cluster center vector and a cluster internal variance matrix, and the cluster center vector represents the typical adaptability change trend of the group of paths.
[0052] The evolution trend of the variant path cluster is combined with the adaptability score to screen paths with significant variation and optimization feasibility, and a set of paths that can be optimized is generated.
[0053] In this embodiment, the evolution trend analysis (such as trend persistence and fluctuation stability) of the variant path cluster is combined with the adaptability score of the original path to perform a multi-dimensional screening strategy, and paths that exhibit significant variation and optimization potential in simulation are retained. The screening conditions include that the duration of the evolution trend is not less than 50 time steps, and the adaptability score is higher than 0.8, which ensures that the path has biological significance and engineering feasibility in terms of functional expression and regulation. The set of paths that can be optimized formed finally will be used for subsequent path integration and construction of the optimal expression configuration set, promoting the optimization design of protein regulation strategy.
[0054] Optionally, step S1 is specifically:
[0055] Step S11: Collect real-time proteomics data through a mass spectrometer, a high-throughput liquid chromatography platform, and a cell expression database, and perform time series alignment and data decoupling on the real-time proteomics data to obtain a protein expression original data set;
[0056] In this embodiment, a Thermo Fisher Orbitrap Exploris 480 mass spectrometer combined with an Agilent 1290 Infinity II high-throughput liquid chromatography platform is used to collect multiple batches of protein samples of HeLa cell lines under different induction states. The resolution set during mass spectrometry collection is 120,000, the full scan range is set to 350–1,800 m / z, and the liquid chromatography gradient elution time is set to 90 minutes. The public protein expression data related to the selected cell state in the cell expression database (such as Human Protein Atlas) is simultaneously introduced. The obtained data is divided into time axis according to 5 seconds as time interval, and the experimental data and the database static data are time-mapped and registered through a synchronous alignment module. After peak recognition, isotope rejection, retention time recalibration, and other processing are performed on the collected data, a multi-source mapping decoupling method is used to decouple and output protein ID, expression abundance, mass spectrometry signal-to-noise ratio, and other fields, thereby generating a protein expression original data set.
[0057] It is noted that the Thermo Fisher Orbitrap Exploris 480 is a high-resolution mass spectrometer used to accurately measure the mass and relative abundance of molecular ions, commonly used in proteomics analysis, as it can provide very high mass accuracy and sensitivity to help identify and quantify proteins in complex biological samples. The Agilent 1290 Infinity II high-throughput liquid chromatography platform is an ultra-high performance liquid chromatography (UHPLC) system used to separate samples before entering the mass spectrometer, separating components in complex mixtures by chemical properties (such as hydrophobicity, polarity) to allow the mass spectrometer to more accurately detect and analyze. When used together, the liquid chromatograph is responsible for separation, and the mass spectrometer is responsible for detection and analysis, allowing high-throughput, accurate molecular identification and quantification of complex biological samples.
[0058] Step S12: Perform high-dimensional semantic feature mapping conversion on the protein expression raw data set to construct a protein expression semantic tensor;
[0059] In this embodiment, based on the obtained protein expression raw data set, by setting five-dimensional semantic space projection (protein function, cell location, regulation state, biological process, expression abundance), each protein expression event is encoded as a five-tuple . With the definition in GeneOntology (Gene Ontology Database) as a reference, a nested semantic mapping dictionary is constructed, and each row in the original expression matrix is replaced and mapped using the dictionary. A three-order tensor is used to represent the semantic structure, and the tensor dimension is (protein_index x time_point x semantic_dimension). For example, for a sample containing 500 proteins, 60 time points and 5-dimensional semantic information, the following semantic tensor is constructed: ; Each element represents the projection of the protein's semantic vector at the specified time point, such as [nucleus, activated state, high expression, involved in transcription regulation, transcription factor] mapped to [0.89, 0.74, 0.92, 0.81, 0.93]. Finally, a semantic tensor is formed that can be used for subsequent context modeling.
[0060] Step S13: Perform multi-state context nested modeling according to the protein expression semantic tensor and extract the nested relationship of environmental regulation factors to construct a context nested representation model;
[0061] In this embodiment, dynamic context nesting modeling is performed based on the semantic tensor structure and the semantic similarity between proteins in their regulatory states. The nesting relationship is established by constructing a three-layer nested structure of "state—factor—expression" to build a multi-granularity expression environment model. This model is represented in tensor expansion form as: Context_Nesting_Model={context_layer_1: state vector} ,context_layer_2: Nested set of environmental factors F_i={f_ij|j=1...k},context_layer_3: Nested matrix representing response }; where f_ij represents the state The j-th environmental factor. For example, under TGF-β induction, Corresponding induced state, Including pH value, temperature, Concentration and other environmental factors, This involves understanding the specific regulatory responses of these factors to protein expression. Through tensor analysis and conditional screening, a structured set of nested context models is ultimately obtained for the dynamic interpretation of expression states.
[0062] Step S14: Analyze the expression behavior vector sequence of proteins in various states based on the context nested representation model, thereby constructing the dynamic matrix of expression states;
[0063] In this embodiment, the context-nested model is input into the expression behavior extraction module. By identifying changes in the protein's expression vector within specific nested contexts, a dynamic sequence of expression states is generated. For example, under different stress conditions, the expression vector of protein P12345...
[0064] The expressive behavior exhibits significant fluctuations, and its expressive behavior sequence is represented in the form of a state vector as follows: The expression behaviors of all proteins were summarized, and the structure was constructed as follows: ;in For protein quantity, The duration is specified. Status labels (such as high expression, stress response, and inhibitory expression) are also added to each expression behavior vector, giving it a traceable contextual meaning for subsequent modeling.
[0065] Step S15: Fuse the expression state dynamic matrix and perform low-rank semantic compression to construct the protein expression context matrix.
[0066] In this embodiment, the dynamic expression state matrix described above undergoes low-rank representation compression processing to retain the main trends of change and remove redundant features. The first k=20 principal components are retained through singular value filtering to construct the low-rank expression state matrix: ; On the basis of low-rank semantics, the expression matrix under different time periods and environments is fused to construct a protein expression context matrix The dimension of the matrix is: ; wherein is the compressed semantic dimension, for example, set as d = 32, and each row represents the context expression profile of a protein. In terms of structure, each element is the comprehensive expression after fusing the expression state, nested environmental factor and semantic attribute, which constitutes a unified input basis for subsequent structural analysis and functional reduction.
[0067] Optionally, step S13 specifically comprises:
[0068] Step S131: Extracting the state identification factor and experimental metadata in the protein expression semantic tensor to obtain a factor candidate set;
[0069] In this embodiment, the content of a specific dimension is extracted from the constructed protein expression semantic tensor as the input source of the state identification factor and experimental metadata. The protein expression semantic tensor is set as , wherein is the number of proteins, is the time point, is the semantic dimension (such as function, location, regulation state, biological process, expression level). From the regulation state dimension of the tensor and the experimental metadata field (induction method, culture condition, stress type, etc.), the identification factor is extracted, and the fields such as the collection batch, culture temperature, treatment duration and cell type in the experimental metadata are associated to form a joint candidate factor description set, which is used for subsequent path construction.
[0070] Step S132: Constructing an initial nested path graph based on the factor candidate set and performing context correlation evaluation to obtain an initial nested path graph;
[0071] In this embodiment, the candidate factor set F candidate is taken as a node to construct an initial nested path graph , wherein V represents the candidate factor, and E represents the context semantic co-occurrence relationship between two factors. The context relationship is constructed through the following two types of sources: 1) co-occurrence frequency analysis: calculating the simultaneous occurrence times of each factor combination under different experimental conditions; 2) semantic similarity analysis: evaluating the regulation similarity between factors based on the semantic embedding space (such as GO definition), and taking the vector angle less than 30° as the threshold condition for establishing connection. The minimum co-occurrence frequency threshold is set to 3, and the minimum semantic similarity is set to 0.7, to obtain the following initial nested path graph connection diagram:
[0072]
[0073]
[0074] The edge weight of the graph represents the coupling strength of the context semantics, and the structure is used to support subsequent subgraph partitioning.
[0075] Step S133: Perform context-coupled subgraph partitioning on the initial nested path graph to construct a factor-coupled subgraph set.
[0076] In this embodiment, the initial nested path graph is partitioned into a set of factor-coupled subgraphs. The entire graph is then segmented by context-coupled strength, and factor pairs with significant semantic coupling are divided into the same subgraph. The continuity of the path is used as the basis for context semantic propagation, and a connected context factor sequence is extracted as a factor-coupled subgraph. For example, the following coupled subgraph set is constructed by path analysis:
[0077]
[0078] Each subgraph represents a potential nested expression path, indicating the expression regulation sequence relationship between multiple regulatory factors in a specific experimental context. The minimum number of nodes for each subgraph is set to 2, and the maximum length is set to 5, to avoid isolated nodes or lengthy paths.
[0079] Step S134: Map each path pair in the factor-coupled subgraph set back to the behavior channel in the protein expression semantic tensor, to obtain a context-nested tensor structure.
[0080] In this embodiment, the paths in each factor-coupled subgraph are mapped back to the original behavior channel in the protein expression semantic tensor T_expr. The mapping logic is based on the index position of the semantic dimension (regulatory state) corresponding to the path factor in the tensor and the channel number, and performs index reverse retrieval. For example, if the subgraph path is , all protein behavior vector sequences in the tensor that match the two labels in the regulatory state dimension are located, and a tensor structure is generated by integration. The obtained set of context-nested tensor structures provides a structural mapping between the regulation path and the protein expression behavior, which is used for expression state modeling.
[0081] Step S135: Construct a context-nested representation model based on the context-nested tensor structure.
[0082] In this embodiment, based on the set of tensor structures constructed by T_nested, a complete context nested representation model C_nested_model is established by fusing the behavior tensor corresponding to each subgraph path, which is used to capture the nested relationship of the multi-path context regulating the protein expression behavior. The form of the context nested representation model can be: C_nested_model={context_unit_i:{subgraph_id:i,factor_sequence:[ ],behavior_profile: ,interaction_strength_matrix: , }};
[0083] Wherein, interaction_strength_matrix represents the joint influence strength between factors in the path, which is obtained by evaluating the consistency of the internal expression response of the tensor, and is represented by a standard deviation matrix or a correlation matrix. The whole model can be used to support the modeling of protein expression under the joint action of multiple state factors, expression prediction or functional linkage analysis, and has good structural interpretability and nested logic integrity.
[0084] Optionally, step S14 is specifically:
[0085] Step S141: parsing the context nested representation model to extract the state master factor sequence and response channel structure, thereby constructing a state channel index table;
[0086] In this embodiment, the context nested representation model C_nested_model is structurally parsed to extract the state master factor sequence contained in each context unit and the response channel information mapped thereby. The specific form of the structure contained in each context unit in the model C_nested_model can be represented as: context_unit_i={factor_sequence:[ ],behavior_profile ,interaction_strength_matrix: }; wherein, factor_sequence is the master factor sequence, and the S' dimension in behavior_profile is one-to-one corresponding to the tensor channel mapping relationship. By traversing all context units in the model, the master factor sequence and the behavior channel index mapped thereby in the protein expression semantic tensor T_expr are extracted, thereby establishing a state-channel corresponding index table. The specific form of the structure of the state-channel corresponding index table can be represented as: The index table is used to guide subsequent behavior feature extraction and expression modeling steps, to ensure that the extracted data and state labels are closely bound, and to avoid behavior channel mismatch.
[0087] Step S142: According to the state channel index table, multi-channel behavior data is extracted from the protein expression semantic tensor, a protein expression behavior feature space is constructed, and state expression vector aggregation modeling is performed to obtain an expression behavior time sequence group;
[0088] In this embodiment, based on the state channel index table Index_table, multiple channel indexes corresponding to each state label are extracted from the protein expression semantic tensor , and are integrated to form a multi-channel behavior sub-tensor group T_sub. For example, for "hypoxia_response", channels 3, 5, and 9 are extracted, and a behavior tensor segment is constructed. All state segments are combined to form a behavior feature space, and the structure of the behavior feature space is as follows: ; then, time dimension aggregation modeling is performed on each state tensor segment, and the behavior change of a protein in a certain state is represented by a state expression vector v_state(t), which is defined as follows: ; and an expression behavior time sequence group is obtained: ; each time sequence is used to depict the evolution behavior of protein expression under a specific regulatory state.
[0089] Step S143: The expression behavior time sequence group is modeled to obtain the expression vector variation pattern under a preset different time resolution, and a regulatory time window is marked to form a nested expression trajectory set;
[0090] In this embodiment, for the expression behavior time sequence group S_behavior, three time resolution windows are set: 2 hours, 6 hours, and 24 hours, and time sequence segments are extracted to observe the change trend of the expression vector under different scales. In each resolution, a sliding window vector comparison is performed to identify the expression change interval, and a change rate exceeding a set threshold (such as ±15%) is used as a marker basis for state transition. For example, in the "oxidative_stress" expression sequence: 1) the expression vector change rate in the 12th to 16th hour segment reaches 20%, which is marked as a regulatory time window W1=[12h, 16h]; 2) the change trend in the 30th to 36th hour segment is stable, and is not marked. The change time window markers of all states are combined to construct a nested expression trajectory set: ; the trajectory set is used for subsequent modeling of the state causal structure.
[0091] Step S144: A state vector graph model is constructed based on the nested expression trajectory set to dynamically deduce the causal regulation relationship between different expression states, and an output state regulation response graph is output.
[0092] In this embodiment, in this step, a state vector graph model G_state = (V, E, W) is constructed based on the nested expression trajectory set. Wherein: V is a set of state factors, such as ; E represents a directed edge between states that exist in a regulatory relationship; W is an edge weight matrix, representing the regulatory weight between different states. The weight deduction is established based on the expression change response relationship within the time window. For example, if the expression change of "hypoxia_response" occurs within 2 hours before "oxidative_stress" and the similarity is higher than a set value (such as cosine similarity > 0.85), an edge is established: ( ); the graph structure is as follows:
[0093]
[0094]
[0095]
[0096] The graph expresses the regulatory causal structure of protein expression driven by multiple state factors.
[0097] Step S145: The time evolution trajectory of the nested expression trajectory set is fused with the causal regulatory weight in the state regulatory response graph, and the time sequence weight is normalized and reconstructed to obtain an expression state dynamic matrix.
[0098] In this embodiment, the time evolution information of the nested expression trajectory set is fused with the edge weight weight in the state causal graph to establish a final expression state dynamic matrix . The fusion method is: for each expression vector sequence of a state v_i, the edge relationship in the state vector graph is weighted and combined. In order to prevent weight imbalance, normalization processing is performed: 1) the regulatory weight vector W[:, i] of each state is normalized to the [0, 1] interval; 2) z-score standardization is performed on the fused value after normalization, so that the dynamic matrix expression of each state has comparability. Finally, the expression state dynamic matrix M_dyn expresses the aggregated response behavior of each protein under the regulation of different state factors over time, and the form is as follows: ; wherein the first row corresponds to the aggregated expression intensity of the "hypoxia_response" state at different time points, the second row corresponds to the "oxidative_stress" state, and so on. The matrix can be used as the core input basis for subsequent protein expression behavior reconstruction, function mapping, and regulatory strategy deduction.
[0099] Optionally, step S2 is specifically:
[0100] Step S21: Perform expression semantic node nested extraction on the protein expression context matrix to construct a cross-domain feature relationship graph;
[0101] In this embodiment, based on the obtained protein expression context matrix (wherein is the number of proteins, is the number of time points, is the number of context semantic channels), semantic nested extraction is performed on the expression nodes in each semantic channel. The nested extraction operation is based on the context semantic similarity, and selects expression node pairs with a semantic cosine similarity greater than 0.85 between context channels to form a semantic nested pair set: ; Subsequently, a feature entity node graph G_feat=(V, E, A) is constructed according to the nested pair set, wherein: 1) V is a set of protein expression semantic nodes; 2) E is a semantic nested edge; 3) is the semantic feature vector of each node, and the dimension d is set to 128. In addition, to enhance the expression behavior interaction relationship between context channels, a cross-channel edge weight correction mechanism is introduced in the graph structure, and edges are added to nodes with consistent expression trend directions, with a threshold of <20° in a sliding window. The final cross-domain feature relationship graph serves as the basis for subsequent structure analysis.
[0102] Step S22: Perform graph sparse reconstruction on the expression behavior cross-domain feature graph to obtain a structure compressed representation graph;
[0103] In this embodiment, the feature relationship graph G_feat is structurally sparse and reconstructed to reduce the problem of redundant edges and excessive connections. The specific processing method is as follows: First, calculate the edge connection degree of each node, and count its local edge weight aggregation distribution; for edges with a connection degree less than the average μ and an edge weight lower than the lower quartile Q1, they are removed; at the same time, for some node pairs that show strong association but have no structural connection, if their semantic cosine similarity is >0.9 and the Euclidean distance of the time behavior curve is (default ), a new edge is added. After the above pruning and edge supplement mechanism, the structure compressed graph is reconstructed, wherein the number of nodes remains unchanged, but the number of edges is reduced by about 35% on average compared to the original graph. This graph structure serves as the input form of the structure domain mapping, facilitating subsequent spatial alignment operations.
[0104] Step S23: Connect the standard protein structure domain database, map the behavior nodes in the structure compressed representation graph to the domain coordinate system in the standard protein structure domain database, and perform hybrid modeling to generate a behavior-structure hybrid graph;
[0105] In this embodiment, the behavior nodes in the compressed graph G_sparse are mapped to a standard protein domain database (such as the Pfam database or the SCOPe database). The standard domain model provided by the database includes protein families, functional regions, and three-dimensional coordinate fragments. The specific form of the standard domain model of the database can be: ; the protein position index corresponding to the behavior node in the graph is mapped to the corresponding domain coordinates by analyzing its belonging protein, time point, and semantic label. Chimeric modeling uses structure alignment evaluation indicators such as structure center angle and residue overlap to screen the optimal chimeric pair. Form a behavior-structure chimeric graph , wherein: is a set of behavior nodes; is a set of domain nodes; is a behavior node and domain mapping edge, and the edge weight is the chimeric fitting degree (0-1).
[0106] Step S24: Perform path-aware extended modeling on the behavior-structure chimeric graph to generate a response path graph set;
[0107] In this embodiment, based on the behavior-structure chimeric graph , the path-aware extension operation is performed. First, define the path extension window (the default length is 5 steps), and from each chimeric behavior node, combine its behavior trend direction and the physical contact topology of the domain to construct a local response path: 1) If the expression rate of the continuous 3-step node is >20%, and its domain belongs to the same functional module (such as the kinase structure module), the path is extended to a potential response chain; 2) If the behavior trend is reversed (such as the expression decreases and then recovers), a branch path is constructed and marked as a turning node. All path combinations form a response path graph set , each graph corresponds to a protein regulation response process.
[0108] Step S25: Perform multi-scale graph convolution feature aggregation on the response path graph set to extract a causal response weight matrix;
[0109] In this embodiment, for each path subgraph in the response path graph set , multi-scale graph convolution feature aggregation modeling is performed. The context behavior vector of each node in the graph is used as the initial feature, the scale range is set to {1, 2, 4} (representing 1-order, 2-order and 4-order neighborhoods), and the feature propagation and weighted integration guided by the graph structure are performed. After aggregation, the global expression vector of each path graph is extracted, the behavior response correlation between the graph and the state label is calculated, and a causal response weight matrix , where: row and column indexes are response path numbers; values are response coupling degrees (e.g., cross-regulation response strength, with a value range of [0, 1]) between paths.
[0110] Step S26: According to the causal significance in the causal response weight matrix and the path density threshold, each path in the response path graph set is weighted, reconstructed and screened, thereby constructing a multi-scale causal structure graph.
[0111] In this embodiment, each path in the response path graph set is weighted, reconstructed and screened, based on two core indicators: 1) a causal significance threshold : a weight value from M_causal, and a threshold value is set as , and paths below this value are removed; and 2) a path density threshold : defined as the ratio of the number of edges in the path graph to the maximum possible number of edges, and set as = 0.45, and below this value is considered as weak behavioral connection and not included in the final graph. The retained path set is fused to construct a multi-scale causal structure graph G_causal_final=(V,E,W), and the edge weight in the graph is inherited from the weight in the causal response matrix after screening, and the node attribute remains the context behavior vector inherited from the original graph. The graph provides a structural basis for subsequent protein regulation mechanism inference, expression prediction or target identification.
[0112] Optionally, step S23 specifically comprises:
[0113] Step S231: Connect a standard protein domain database, and perform semantic alignment on the semantic features of each behavioral node in the structure compressed representation graph and the domain description vector of the standard protein domain database to obtain a semantic index mapping table;
[0114] In this embodiment, a standard protein domain database (e.g., Pfam-A or SCOPe) is connected, and the built-in domain description information is extracted to construct a domain semantic description vector set D_struct={ , ,..., }, where each represents a semantic description vector of a domain, with a dimension of 256. At the same time, the behavioral node semantic vector (inherited from the previous graph) in the structure compressed representation graph is normalized. On this basis, cosine similarity matching is performed on all behavioral nodes and domain description vectors, and a matching threshold of 0.7 is set to construct a semantic index mapping table , where each behavioral node can be mapped to at most 3 domains, sorted by similarity. The mapping table serves as a basis for subsequent spatial projection and positioning, ensuring the accuracy of domain semantic matching.
[0115] Step S232: Perform domain space projection on each behavior node in the structure compressed representation graph based on the semantic index mapping table, and perform coordinate coding in combination with the domain coordinate system of the standard protein domain database to obtain a domain positioning tensor;
[0116] In this embodiment, according to the generated semantic index mapping table T_map, each behavior node v_i in the graph is mapped to the three-dimensional coordinate system corresponding to the domain d_j. The standard domain database provides the domain coordinate origin, the orientation matrix, and the functional block coordinates. Taking the Pfam domain PF00001 as an example, the description form of the domain PF00001 includes: Centroid: (x=12.4, y=8.3, z=5.1), Orientation: [0.6, 0.2, 0.7], FunctionalBlocks: { [15-25], [55-75]}; in combination with the time index t_i and the channel coordinate c_i of the behavior node in the expression context matrix, the behavior node is mapped to the relative offset vector in the domain coordinate space, and then the standard domain orientation transformation matrix is used for coordinate coding to obtain a domain positioning tensor . wherein n is the number of behavior nodes, 3 is the dimension of the three-dimensional space, and K is the number of candidate domains. The tensor is used for spatial neighborhood structure matching and context coupling verification. Step S233: Perform context coupling retrieval and neighborhood cross verification on the domain positioning tensor to obtain a structure matching enhancement matrix;
[0117] In this embodiment, the domain positioning tensor is subjected to context coupling retrieval. Specifically, for any two behavior nodes
[0118] , if they are mapped to the same domain and the distance between their positioning vectors in the Euclidean space is less than 5, they constitute a spatial coupling candidate pair. At the same time, cross verification is performed in combination with the expression semantic similarity (the threshold is set to 0.8). All node pairs that meet the conditions are constructed into an enhancement matrix , wherein each element in the matrix represents the structure matching enhancement score of the node pair, and the value range is [0, 1]. The matrix provides a weight basis for subsequent structure mapping relationship construction. Step S234: Construct an initial behavior-structure mapping graph based on the structure matching enhancement matrix;
[0119] In this embodiment, based on the enhancement matrix M_enhance, an initial behavior-structure mapping graph
[0120] is constructed. Where: 1) V is the set of behavior nodes; 2) E is the number of nodes passed through. 3) W is the set of edges defined by a median score greater than 0.7; 4) W is the set of edge weights, where each edge... The weight is The graph does not contain structural domain nodes; it only preserves structural correlation connections between behavioral nodes. To ensure graph connectivity, virtual edges are added to nodes with low edge density, connecting them to their most similar structural neighborhood centers. This initial graph is used for subsequent structural functional clustering and chimeric modeling.
[0121] Step S235: Perform chimerism measurement modeling using functional cluster information and structural class similarity of domains in the standard protein domain database, and output the chimerism association tensor;
[0122] In this embodiment, the domain functional clusters (such as "kinase class" and "transporter subclass") and their structural category labels (such as...) recorded in the standard domain database are combined. fold, (Using bucket structures, etc.), the similarity score is assigned to the structural domain mapped to each behavioral node. For example, if two behavioral nodes are mapped to PF00069 and PF07714 respectively, their functional clustering is consistent (both are Ser / Thr kinases), and the similarity score can reach 0.9. A chimeric association tensor is constructed. The dimensions in this tensor have the following meanings: the first and second dimensions are the behavior node indices; the first channel of the third dimension is the functional clustering consistency score, and the second channel is the class similarity score. The values in this tensor serve as a quantitative indicator of the degree of chimerism and compatibility between behavior node pairs.
[0123] Step S236: Use the chimeric correlation tensor to assign chimeric strength labels to the edges in the initial behavior-structure mapping graph, and perform graph edge recalibration and connection optimization to construct the behavior-structure chimeric graph.
[0124] In this embodiment, the score value in the chimeric correlation tensor T_embed is projected onto the edge weight attribute in the initial behavior-structure mapping graph G_init, and each edge e_{ij} is reassigned a comprehensive chimeric strength score, specifically using the following weighted combination: ;in , , This is used to adjust the weights of the three embedding influence factors. After recalibrating all edges, a connection optimization operation is performed: edges with embedding strength below 0.5 are deleted; if a node's connectivity is below 2, it is automatically connected to the node with the highest embedding score. The final optimized behavior-structure embedding graph is obtained. This graph possesses both semantic and structural alignment capabilities, and can serve as a foundational graph structure for regulatory path modeling and causal analysis.
[0125] Optionally, the protein function expression function in step S3 is specifically reduced as follows:
[0126] The causal chain expansion is performed on each expression path in the multi-scale causal structure graph, and the maximum path depth is set to 6, the function expression sequence in each expression path is extracted, and a path expression function set is constructed;
[0127] In this embodiment, when processing the expression path in the multi-scale causal structure graph G_causal=(V,E,W), for each path, the causal weight direction is expanded from the starting node, and at most 6 layers of connection depth are recursively traversed. Each node contains an expression function label f_i, such as f_i=tanh(Wx+b), which corresponds to the activation function expression form. For any path , the function expression sequence on it is extracted, and the nesting level and input-output variable flow between functions are recorded. The expression function sequences of all paths are uniformly packaged into a set , which is used as the basic function set for subsequent structure modeling. To control the extraction process of redundant paths, only paths with a weight sum greater than 2.5 are retained.
[0128] The sparse control parameter is set to 0.03, and the path expression function set is symbolically and sparsely modeled to generate a function sparse structure tensor;
[0129] In this embodiment, for the constructed expression function set F_all, the sparse control parameter is set to 0.03 to limit the density of symbols in each expression function sequence. In the modeling process, first, each function expression is symbolically converted into a standard infix expression form (such as sigmoid(Wx+b) is converted into sigmoid,+,W,x,b). The sparse modeling aims to control the structure complexity, and after converting the function representation into a symbol chain form, operators, variables or sub-expressions with a frequency lower than the threshold are removed, where is the total number of all symbol instances in F_all. The final output sparse structure tensor is obtained, where is the number of paths, l is the maximum function length (taking 16), is the symbol dimension (encoded as a 128-dimensional vector). This tensor reflects the function symbol structure pattern of each path, which is used for subsequent similarity clustering.
[0130] The function expression in the function sparse structure tensor is processed by local sensitive hashing, the hash distance threshold is set to 0.15, and the intra-cluster function similarity is greater than or equal to 85% as the clustering condition, and an expression function aggregation tensor is output;
[0131] In this embodiment, Each function expression in T_agg is regarded as a high-dimensional sparse vector, which is mapped to a low-dimensional hash space by Locality-Sensitive Hashing (LSH). The hash distance threshold is set to 0.15, and 4 groups of parallel hash function families are used for processing (20 hash functions in each group), to aggregate function expressions with similar symbolic structures. The aggregation criterion is that if the average Hamming distance of two function expressions in all hash function groups does not exceed 0.15, and the matching degree at the symbolic structure level exceeds 85%, they are considered as cluster members. Finally, the function aggregation tensor T_agg is output where K is the number of function clustering clusters, is the number of functions in each cluster (dynamic change), is the length of the symbolic vector (128 dimensions), which is used to identify function clusters with homogeneous expression functions.
[0132] Based on the gradient propagation graph and the causal regulation weight graph of each expression function cluster in the expression function aggregation tensor, the sub-functions with causal regulation weight ≥ 0.6 and average gradient weight ≥ 0.65 are set as core sub-functions, and pruning is performed on the low-weight functions in the expression function cluster, to output the function pruning set;
[0133] In this embodiment, based on the function clustering clusters in T_agg, the gradient propagation graph of each cluster is extracted where the edge weight is the local gradient value, and the node is the sub-function identifier. At the same time, combined with the causal regulation graph the causal edge weight is set as follows: if the average gradient propagation weight of a sub-function in is ≥ 0.65; and its causal regulation weight in is ≥ 0.6; it is marked as a core sub-function. Low-weight sub-functions that do not meet the above conditions are removed from , their positions and structural expressions are recorded, and a function pruning set F_pruned = {f_i} is constructed, which is used for subsequent topological stability modeling. This pruning process can effectively control the complexity of function expressions in each cluster, with an average pruning ratio of about 30%.
[0134] The structure transformation deduction is performed on the sub-functions in the function pruning set, and matched with the stable function patterns in the preset expression topological stability rule library, with a matching tolerance threshold of 0.1, to obtain the expression topological stability graph;
[0135] In this embodiment, the transformation deduction process is performed on each sub-function structural expression in F_pruned, including variable renaming, hierarchical unfolding, and nested reconstruction. For example, is converted to . Then match the function template in the expression topology stability rule base in the form of structure diagram (node is operator, edge is data flow), the library contains about 200 typical stable function patterns, the matching tolerance is set to 0.1 (i.e. structural similarity ≥ 0.9). The function expression form matched successfully is added to the expression topology stability atlas , which is used to identify function nodes and flow directions with stable effects under multi-path.
[0136] Based on the expression topology stability atlas and the function pruning set, the sub-function is recombined, connected and sorted to obtain the minimum functional expression unit.
[0137] In this embodiment, based on the obtained G_stable and F_pruned, the minimum functional expression unit recombination process is performed. First, the functional level of all core sub-functions is divided (such as activation class, gating class, normalization class), the node connection combination is performed according to the topology connection direction, and the expression sub-graph is constructed. The nodes in each sub-graph are rearranged according to the topology sorting to ensure the continuity of the data dependency chain, and the combined expression complexity (such as total symbol number, function depth, etc.) is calculated. After removing the redundant repeated paths, the minimum functional expression combination unit set is output, each unit represents a function combination sub-graph with stable structure and clear expression function, which can be used for bottom-level explanation of the regulation expression mechanism and construction of the modular expression system.
[0138] Optionally, the structure-function transformation rule extraction in step S3 is specifically:
[0139] Taking the minimum functional expression combination unit set as input, the symbolic function topology, path link structure and causal response factor are extracted, and the expression-function structure atlas is constructed;
[0140] In this embodiment, taking the generated minimum functional expression combination unit set as the basis, the symbolic function sequence in each unit u_i is extracted, and the topological nesting relationship in the graph structure is labeled. Each expression unit is represented as a directed structure graph , wherein the node V_ui represents the symbolic function (such as ReLU, Sigmoid, multiplication, etc.), and the edge E_ui represents the input-output relationship between functions. All G_ui are uniformly coded to construct the whole graph , and the causal response factor tensor (indicating the response intensity of each expression unit to m functional indicators) is fused, and the response factor is labeled on the edge of the structure graph as a weight vector. Finally, the expression-function structure atlas is formed, which is used for subsequent topology feature induction and function explanation.
[0141] Topological rearrangement modeling is performed on the expression-function structure map, and the structure map fingerprint of each expression unit in the topological rearrangement modeling result is extracted to construct a functional structure fingerprint set;
[0142] In this embodiment, topological rearrangement modeling is performed on G_expr-func. The graph is reordered using the principle of minimizing structural entropy, prioritizing representation units with higher connectivity centrality and response factor values. During the rearrangement process, nodes in each representation subgraph G_ui are renumbered, and its normalized topological arrangement vector is output. This is to uniformly represent the logical execution order. Subsequently, based on the topological adjacency matrix A_ui and the node function label sequence... Constructing a structural graph fingerprint , where d is the function embedding dimension (set to 64), and each function label is embedded into the function semantic vector through a lookup table. The structural fingerprints of all representation units constitute the functional structure fingerprint set. .
[0143] The structure graph fingerprints in the functional structure fingerprint set are compared with the structure-function mapping correspondence in the topological stability rule base by vector similarity comparison, and symbol function mapping and local subgraph recognition are performed to obtain the structure-function symbol correspondence matrix.
[0144] In this embodiment, the fingerprint F_ui in the above-mentioned functional structure fingerprint set is compared with the standard structure template in the rule base for expressing topological stability. Perform vector similarity comparison. Each standard template is jointly constructed from a sequence of function labels and a standard topological graph, represented as a triple ( Cosine similarity was used to compare the similarity between F_ui and each S_i fingerprint vector, with a similarity threshold of 0.85. Successful matches were mapped to a structure-function symbol correspondence matrix. ,in This indicates a symbolic correspondence between the representation unit u_i and the template structure S_j. This matrix can serve as a semantic guide for subsequent local graph recognition and representation reconstruction.
[0145] Based on the symbolic relationships in the structure-function symbolic correspondence matrix, rule reduction and expression-driven recoding are performed to obtain a set of rule-expressed substructures.
[0146] In this embodiment, a set of expression units with similar expressive functions but redundant topological structures is extracted based on the matching relationships in M_sym. For the expression units in this set, their common sub-expression structures (such as shared activation function sequences, input-output links, etc.) are extracted, and redundant paths, redundant connections, and duplicate sub-functions are removed to generate a set of regular expression sub-structures. Each substructure r_i is represented as a directed subgraph, accompanied by a function expressing the label and response value range. For example, substructure It can be composed of three function nodes [Sigmoid → Multiply → Add], and the response weight in the range of [0.4, 0.8] is attached to the edge weight. Through structure normalization and symbol relabeling, the driving recoding of the expression unit is completed, which ensures to reduce the graph structure complexity while retaining the function key path.
[0147] According to the rule expression substructure set, the directed graph modeling and path scoring are performed to obtain the structure-function transformation path graph;
[0148] In this embodiment, all rule expression substructures in R_expr are used to construct the structure-function transformation path graph , wherein each node represents an expression substructure, and the edge represents a feasible path from one structure to another structure. The edge weight is calculated by using a comprehensive scoring method: according to the response weight improvement ratio (ΔC); the expression complexity reduction amplitude (ΔN); and the topology similarity improvement (ΔS) with the structure template; and the three are weighted as the path score value S . The path traversal is performed on G_path, and all reachable paths and their scores are recorded, and finally the structure-function transformation path graph and the path score table are output , which is used to identify the preferred function structure combination path.
[0149] The high robustness structure conversion path is selected from the structure-function transformation path graph, and thus the function mapping unit set is induced.
[0150] In this embodiment, based on the score table P_scores of G_path, the paths with a score value not less than 0.85 are selected, and the structural robustness change in the perturbation experiment (such as the function perturbation amount ±10%) is further evaluated, and the robustness score threshold is set to 0.8. For the paths meeting the conditions, the minimum expression structure link between the start and end nodes is extracted to construct the function mapping unit. Finally, the function mapping unit set is output, wherein each M_i represents a structure-function bridging unit with high stability and high matching degree between topology change and function expression, which is used to guide the subsequent expression reconstruction and function modularization integration.
[0151] Optionally, the unit configuration path effectiveness and structure stability evaluation in step S4 is specifically:
[0152] According to the function mapping unit set and the preliminary scheduling scheme in the protein scheduling simulation result, a directed graph structure representing the dependency relationship and execution order between the function mapping units in the scheduling path is constructed, and a unit scheduling graph is generated;
[0153] In this embodiment, the function mapping unit set is obtained, and the preliminary scheduling scheme in the protein scheduling simulation result On the basis of the above, the dependence of each M_i on the upstream and downstream relationship, functional trigger boundary and time sequence are extracted, and a directed graph G_sched=(V_sched, E_sched) is constructed, wherein the node V_sched represents the functional mapping unit in the scheduling, and the edge E_sched represents the sequence and causal dependence chain between functions. The execution time window [t_start, t_end] and the required resource identifier of each node M_i are allocated, and the conditional trigger label (such as expression gate signal, energy input, etc.) is attached to the edge. The generated unit scheduling graph supports the modeling of parallelism and dependence of multi-path scheduling structure, and provides structural support for subsequent conflict detection.
[0154] The bottleneck and feedback closed loop modeling of the unit scheduling graph is performed, and the path timing conflict feature is extracted according to the bottleneck and feedback closed loop modeling result, to obtain a path conflict vector group;
[0155] In this embodiment, the structure backtracking analysis is performed on the scheduling graph G_sched, and all ring subgraphs with feedback paths are identified The cumulative execution time and state reaction interval of the nodes in each ring are aggregated and modeled. If there is an execution delay difference greater than a set threshold ΔT_max=0.12s in a certain closed loop path, it is marked as a timing bottleneck feedback ring. At the same time, all reentrant points and shared resource nodes on all paths are collected, and a path overlap matrix is constructed, and the number of timing conflicts between each pair of paths is counted. Based on R_conflict and the bottleneck closed loop link, a path conflict vector group is extracted, each vector representing the conflict density, feedback strength, timing conflict strength, conflict frequency and conflict duration of the path, and other key resource competition indexes.
[0156] The connection compactness, energy level transition distribution and synergistic configuration pressure of the graph nodes in the unit scheduling graph are modeled, and a structure energy distribution map is generated;
[0157] In this embodiment, the structure characteristic analysis is performed on each node in G_sched, and its topological connection compactness (mainly based on the connection degree and local average clustering coefficient of the node), energy level transition frequency (reflecting the number of intermediate energy state changes required from the input state to the excited state) and local synergistic pressure (based on the tensor distribution of the adjacency unit interaction force matrix) are extracted. The above three parameters are fused to form a node energy vector , and a node energy tensor is constructed on the whole graph, and finally a structure energy distribution map is drawn, reflecting the load hotspots and activation centers of each region in the unit scheduling graph during execution, for execution efficiency evaluation.
[0158] The scheduling path execution efficiency is evaluated based on the unit scheduling graph and the path conflict vector group.
[0159] In this embodiment, the actual scheduling efficiency of each complete path P_i from the starting point to the end point is evaluated based on the combination of the scheduling graph G_sched and the path conflict vector group C_vec. The efficiency score uses the following formula group: the baseline execution time ; the actual delay time ; the path span length L_i; the unit execution efficiency . Δ_conflict (path timing conflict delay) represents the cumulative amount of execution delay time caused by scheduling conflicts in the key nodes in the path in terms of resource occupation, input-output dependency, etc. Its calculation method is as follows: for each path P_i in the scheduling graph G_sched, there are m conflict points (including resource sharing conflicts, data dependency conflicts, etc.), and each conflict point corresponds to a delay δ_j; the conflict points are obtained through the path conflict vector group C_vec; ; where δ_j( ) can be estimated according to the resource queuing waiting time, the idle time after the trigger conflict condition is triggered, or the system specified conflict resolution scheduling delay. Δ_feedback( ) represents the logical delay and state synchronization lag introduced to meet the feedback control consistency in the case of scheduling graph with feedback structure paths (such as function A depends on B, and B indirectly depends on A). Its estimation is based on the following structure: closed loop path ; the longest logic waiting time τ_max( ) in the path; the difference τ_sync( ) between the synchronization trigger signals; the state refresh frequency f_update( ) in the same cycle; then it can be calculated: ; if there are multiple feedback loops, the maximum or average value involved in the path P_i is evaluated, or simulation measurement is performed according to the modeling accuracy requirement. Additional evaluation parameter records are attached to the output of each path, including the delay source node number, the conflict vector dimension, the feedback influence factor, etc. The path efficiency report table is formed, providing quantitative reference for path optimization and configuration selection.
[0160] The execution efficiency of the scheduling path is evaluated, and the structure energy distribution map is executed to obtain the effectiveness-stability dual-index quantitative score, and the path performance score table is obtained;
[0161] In this embodiment, the execution efficiency Eff_i and the average energy level value E_avg of the corresponding path in the structure energy tensor T_energy are input into the dual-index scoring function . The full score of the scoring is set to 1.0, and the performance score of each path is calculated to form the scoring table The higher the score, the higher the execution efficiency and the lower the structure energy consumption, and the higher the priority reservation value. In addition, the path with a score lower than 0.6 is recorded in the report to record the combination of the structure bottleneck position and the low efficiency factor, and to provide the optimization direction of the refined path.
[0162] Based on the path performance score table, the scheduling paths in the unit scheduling graph are hierarchically clustered and associated mapped, and the evolution trend simulation is performed to obtain a path evolution decision graph;
[0163] In this embodiment, the path performance score table Score_path is clustered by multiple indexes between paths, and the path similarity matrix is constructed using the structure similarity measure Sim(P_i, P_j) combined with the score difference The clustering threshold θ=0.75 is used to divide the paths into multiple functional behavior clusters. The score trend modeling is performed within each cluster to construct a path score evolution trajectory graph Each edge represents the natural transition feasibility between paths in the score improvement direction (such as local reconstruction from inefficient paths to efficient paths). Finally, the path evolution decision graph G_decision is output, where each subgraph represents a possible path evolution strategy, which can be used to support the upgrade and improvement or redundancy reduction of the scheduling path.
[0164] The paths with a score≥0.85 in the path evolution decision graph and the paths with an average structure energy level≤0.25 of the path nodes are executed for redundancy pruning and structure verification to screen out a stable configuration path set;
[0165] The stable configuration path set is dynamically spliced to form a configuration decision graph.
[0166] In this embodiment, all path sets P_stable with a path score s_i≥0.85 and an average structure energy level E_avg of the path nodes≤0.25 are screened out from G_decision. Node redundancy detection is performed on each path, including: repeated calculation function; redundant nested loop; empty operation node. The above redundancies are pruned, and the structure matching interface and the stability rule library are called to verify the logical integrity of the pruned path. Under the premise of ensuring the functional continuity and stability, the screened path set is dynamically spliced based on the input-output compatibility, structure continuity and average response time difference not exceeding 0.03s. The splicing result forms a configuration decision graph G_config=(V_conf, E_conf), which provides an efficient and stable graph-level execution model for functional deployment.
[0167] Especially important is that the execution efficiency of the scheduling path is evaluated as follows:
[0168] Extract the path length, dependency depth, node connection method, and number of feedback loops from each scheduling path in the unit scheduling graph, and associate the temporal conflict information in the path conflict vector group to generate a path execution feature set.
[0169] In this embodiment, based on the unit scheduling graph G_sched=(V,E), all sets of valid scheduling paths are identified from it. For each path P_i, extract the following four types of structural features in sequence: 1) Path length: record the total number of nodes in the path, denoted as . 1) Unit: hop count; 2) Dependency depth: Extract the length of the longest directed dependency chain based on the dependency edge hierarchy in the path. 3) Node connection method: Statistically analyze the distribution of connection types between nodes in the path (e.g., "sequential", "parallel", "backflow"), and label the connection dimension vector of each node; 4) Number of feedback loops: Extract the number of feedback substructures based on whether there are self-loops or cross-layer loops on the path. Subsequently, a mapping relationship is established between path P_i and the temporal conflict record C_vec(i) in the path conflict vector group, and the conflict type corresponding to the path is extracted from it. (Initiation conflicts, resource contention conflicts, etc.), conflict density (Number of conflict events per unit time) and duration of conflict The final path execution feature set is structured as follows: .
[0170] Calculate the path execution complexity based on the dependency hierarchy, conflict density, and feedback structure distribution in the path execution feature set;
[0171] In this embodiment, the path execution feature set from the previous step is used as input to construct an execution complexity index for each path. The specific implementation method is as follows: based on the dependency hierarchy. Set dependency coefficient weights =0.35, based on the number of feedback closed loops Assign closed-loop factor weights =0.25, and based on the connection method, the tension tensor Structural perturbation values (e.g., maximum connectivity difference) set connection complexity weights. =0.4. The overall execution complexity score for this path is calculated as follows: Among them, tension tensor This represents the distribution matrix of connection weights for each node in the path. The more dispersed the values, the stronger the connection instability during the scheduling process.
[0172] Based on the path execution complexity, a multi-index fusion model is constructed in combination with the timing conflict strength, conflict frequency and conflict duration characteristics in the path conflict vector group;
[0173] In this embodiment, the timing conflict characteristic values in the path conflict vector group, such as the conflict strength , the conflict frequency and the conflict average duration , are combined with the execution complexity index C_exec(i) to construct a multi-index fusion model. The input weight proportion of the model is set as: the complexity term =0.4, the conflict strength =0.3, the conflict frequency =0.2, and the conflict duration =0.1. Then, the model fusion score is: wherein, the unit is normalized conflict tension, the number of conflict events per unit time, and the average duration (seconds).
[0174] The multi-index fusion model is used for weighted scoring, and the weighted scoring result is nonlinearly mapped and converted to generate the scheduling path execution efficiency.
[0175] In this embodiment, the fusion score value is taken as input, and is converted into path execution efficiency by constructing a set of nonlinear mapping formulas based on empirical functions. The nonlinear mapping adopts a logarithmic decay type scoring function: ; wherein is the current highest fusion value; the conversion result is limited in the range of [0, 1], and the value closer to 1 indicates higher scheduling efficiency. The mapping function can effectively compress the weight of the high complexity task path, so that its weight in the scoring system is reduced, which is beneficial to the scheduling priority sorting.
[0176] It is noted that the English phrases contained in the parameters in the embodiments of the present application are named according to the data type or data meaning, and the naming of these parameters is not limited in the present application.
[0177] Optionally, the present specification also provides a biological information processing system applied to synthetic biology, for executing the biological information processing method applied to synthetic biology as described above, and the biological information processing system applied to synthetic biology comprises:
[0178] A context decoding module is configured to acquire real-time proteomics data, and perform polymorphic context decoding on the real-time proteomics data to obtain a protein expression context matrix.
[0179] The interaction modeling module is configured to perform cross-domain graph structure decomposition and reconstruction on the behavior features in the protein expression context matrix, to obtain a reconstructed protein matrix; and analyze biological response path-protein domain interaction based on the reconstructed protein matrix, to obtain a multi-scale causal structure graph;
[0180] The function transformation analysis module is configured to perform protein function expression function reduction on the path nodes in the multi-scale causal structure graph, to obtain a function expression minimum module combination unit; and perform structure-function transformation rule extraction according to the function expression minimum module combination unit, to obtain a function mapping unit set.
[0181] The protein scheduling simulation module is configured to take the function mapping unit set as input, to perform protein scheduling simulation, and perform unit configuration path effectiveness and structure stability evaluation according to the protein scheduling simulation result, so as to dynamically splice to form a configuration decision graph.
[0182] The virtual response simulation module is configured to perform multi-scenario virtual response simulation on the paths in the configuration decision graph, and perform path adaptability verification and scoring on the multi-scenario virtual simulation result, to obtain a path adaptability feedback table.
[0183] The variability screening module is configured to perform multi-dimensional clustering and variability screening on the path adaptability feedback table, to obtain an optimal expression configuration set.
[0184] Therefore, from any viewpoint, the embodiments should be considered as exemplary and non-limiting, the scope of the present application being defined by the appended claims and not by the above description, and it is intended to embrace all variations falling within the meaning and range of equivalents of the elements of the claims.
[0185] The above description is merely one specific implementation of the application, which enables a person skilled in the art to understand or implement the application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the application. Therefore, the application shall not be limited to these embodiments shown herein, but shall conform to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A bioinformatics processing method applied to synthetic biology, characterized by, The method comprises the following steps: Step S1: acquiring real-time proteomics data, and performing polymorphic context decoding on the real-time proteomics data to obtain a protein expression context matrix; Step S2: performing cross-domain graph structure decomposition and reconstruction on the behavior characteristics in the protein expression context matrix to obtain a reconstructed protein matrix; analyzing the biological response path-protein domain interaction based on the reconstructed protein matrix to obtain a multi-scale causal structure graph; Step S3: performing protein function expression function reduction on the path nodes in the multi-scale causal structure graph to obtain a function expression minimum module combination unit; performing structure-function transformation rule extraction according to the function expression minimum module combination unit to obtain a function mapping unit set; Step S4: taking the function mapping unit set as input, performing protein scheduling simulation, and performing unit configuration path effectiveness and structure stability evaluation according to the protein scheduling simulation result, so as to dynamically splice to form a configuration decision graph; Step S5: performing multi-scenario virtual response simulation on the paths in the configuration decision graph, and performing path adaptability verification and scoring on the multi-scenario virtual simulation result to obtain a path adaptability feedback table; Step S6: performing multi-dimensional clustering and variability screening on the path adaptability feedback table to obtain an optimal expression configuration set.
2. The bioinformatics processing method for synthetic biology according to claim 1, wherein, Step S1 is specifically: Step S11: acquiring real-time proteomics data through a mass spectrometer, a high-throughput liquid phase platform and a cell expression database, and performing time sequence alignment and data decoupling on the real-time proteomics data to obtain a protein expression original data set; Step S12: performing high-dimensional semantic feature mapping conversion on the protein expression original data set, so as to construct a protein expression semantic tensor; Step S13: performing multi-state context nested modeling according to the protein expression semantic tensor and extracting the nested relationship of environmental regulation factors, so as to construct a context nested representation model; Step S14: analyzing the expression behavior vector sequence of the protein under various states based on the context nested representation model, so as to constitute an expression state dynamic matrix; Step S15: fusing the expression state dynamic matrix and performing low-rank semantic compression, so as to construct a protein expression context matrix.
3. The bioinformatics processing method for synthetic biology according to claim 2, wherein, Step S13 is specifically: Step S131: extracting state identifier factors and experimental metadata in the protein expression semantic tensor to obtain a factor candidate set; Step S132: constructing an initial nested path graph based on the factor candidate set, and performing context correlation evaluation to obtain an initial nested path graph; Step S133: performing context coupling subgraph segmentation on the initial nested path graph to construct a factor coupling subgraph set; Step S134: mapping each path pair in the factor coupling subgraph set back to the behavior channel in the protein expression semantic tensor, so as to obtain a context nested tensor structure; Step S135: constructing a context nested representation model based on the context nested tensor structure.
4. The bioinformatics processing method for synthetic biology according to claim 2, wherein, Step S14 is specifically: Step S141: analyzing the context nested representation model to extract a state master factor sequence and a response channel structure, so as to construct a state channel index table; Step S142: According to the state channel index table, the multi-channel behavior data is extracted from the protein expression semantic tensor, the protein expression behavior feature space is constructed, the state expression vector aggregation modeling is performed, and the expression behavior time sequence group is obtained; Step S143: Under the preset different time resolution, the expression behavior time sequence group is respectively modeled to express the vector variation mode, and the regulation time window is marked to form a nested expression trajectory set; Step S144: Based on the nested expression trajectory set, a state vector graph model is constructed to dynamically deduce the causal regulation relationship between different expression states, and output a state regulation response graph; Step S145: The time evolution trajectory of the nested expression trajectory set is fused with the causal regulation weight in the state regulation response graph, and the time sequence weight is normalized and reconstructed to obtain an expression state dynamic matrix.
5. The bioinformatics processing method for synthetic biology according to claim 1, wherein, Step S2 is specifically: Step S21: Perform expression semantic node nested extraction on the protein expression context matrix to construct a cross-domain feature relationship graph; Step S22: Perform graph sparse reconstruction on the expression behavior cross-domain feature graph to obtain a structure compressed representation graph; Step S23: Connect the standard protein domain database, map the behavior nodes in the structure compressed representation graph to the domain coordinate system in the standard protein domain database, and perform hybrid modeling to generate a behavior-structure hybrid graph; Step S24: Perform path perception expansion modeling on the behavior-structure hybrid graph to generate a response path graph group; Step S25: Perform multi-scale graph convolution feature aggregation on the response path graph group to extract a causal response weight matrix; Step S26: According to the causal significance and path density threshold in the causal response weight matrix, each path in the response path graph group is weighted and reconstructed and screened, thereby constructing a multi-scale causal structure graph.
6. The bioinformatics processing method for synthetic biology according to claim 5, wherein, Step S23 is specifically: Step S231: Connect the standard protein domain database, and perform semantic alignment of the semantic features of each behavior node in the structure compressed representation graph with the domain description vector of the standard protein domain database to obtain a semantic index mapping table; Step S232: Perform domain space projection on each behavior node in the structure compressed representation graph based on the semantic index mapping table, and perform coordinate coding in combination with the domain coordinate system of the standard protein domain database to obtain a domain positioning tensor; Step S233: Perform context coupling retrieval and neighborhood cross-validation on the domain positioning tensor to obtain a structure matching enhancement matrix; Step S234: Construct an initial behavior-structure mapping graph based on the structure matching enhancement matrix; Step S235: Perform hybrid measurement modeling using the function cluster information and structure generic similarity of the domains in the standard protein domain database to output a hybrid correlation tensor; Step S236: Use the hybrid correlation tensor to give the hybrid strength label to the edges in the initial behavior-structure mapping graph, and perform graph edge re-labeling and connection optimization to construct a behavior-structure hybrid graph.
7. The bioinformatics processing method for synthetic biology according to claim 1, wherein, The protein function expression function reduction in step S3 is specifically: The causal chain expansion is performed on each expression path in the multi-scale causal structure map, the maximum path depth is set to 6, the function expression sequence in each expression path is extracted, and a path expression function set is constructed; The sparse control parameter is set to 0.03, the path expression function set is subjected to symbolic sparse modeling, and a function sparse structure tensor is generated; The function expression in the function sparse structure tensor is subjected to local sensitive hashing processing, the hashing distance threshold is set to 0.15, and the aggregation condition is that the similarity of the functions in the cluster is greater than or equal to 85%, and an expression function aggregation tensor is output; Based on the gradient propagation flow graph and the causal regulation weight graph of each expression function cluster in the expression function aggregation tensor, the sub-function with a causal regulation weight greater than or equal to 0.6 and an average gradient weight greater than or equal to 0.65 is set as a core sub-function, the low-weight function in the expression function cluster is pruned, and a function pruning set is output; The sub-function in the function pruning set is subjected to structure transformation deduction, and is matched with the stable function mode in the preset expression topology stability rule library, the matching tolerance threshold is set to 0.1, and an expression topology stability map is obtained; Based on the expression topology stability map and the function pruning set, the sub-function is recombined, connected and sorted to obtain a functional expression minimum combination unit.
8. The bioinformatics processing method for synthetic biology according to claim 1, wherein, The structure-function transformation rule extraction in step S3 is specifically: Taking the functional expression minimum combination unit set as input, the symbolic function topology, path link structure and causal response factor are extracted, and an expression-function structure map is constructed; The expression-function structure map is subjected to topological rearrangement modeling, and the structure graph fingerprint of each expression unit in the topological rearrangement modeling result is extracted, thereby constructing a functional structure fingerprint set; The structure graph fingerprint in the functional structure fingerprint set is subjected to vector similarity comparison with the structure-function mapping corresponding relationship in the expression topology stability rule library, symbolic function mapping and local subgraph identification are performed, and a structure-function symbolic correspondence matrix is obtained; Based on the symbolic relationship in the structure-function symbolic correspondence matrix, rule reduction and expression driven recoding are performed to obtain a rule expression substructure set; According to the rule expression substructure set, a directed graph is modeled and a path score is obtained, and a structure-function transformation path graph is obtained; From the structure-function transformation path graph, a high-robustness structure conversion path is selected, and a functional mapping unit set is induced.
9. The bioinformatics processing method for synthetic biology according to claim 1, wherein, The unit configuration path effectiveness and structure stability evaluation in step S4 is specifically: According to the functional mapping unit set and the preliminary scheduling scheme in the protein scheduling simulation result, a directed graph structure representing the dependence relationship and execution order between the functional mapping units in the scheduling path is constructed, and a unit scheduling graph is generated; The bottleneck and feedback closed loop modeling is performed on the unit scheduling graph, and the path timing conflict features are extracted according to the bottleneck and feedback closed loop modeling results, and a path conflict vector group is obtained; The connection compactness, energy level transition distribution and synergistic configuration pressure of the graph nodes in the unit scheduling graph are modeled, and a structure energy distribution map is generated; The scheduling path execution efficiency is evaluated based on the unit scheduling graph and the path conflict vector group; The effectiveness-stability dual-index quantitative score is performed by using the evaluation scheduling path execution efficiency and the structure energy distribution map, and a path performance score table is obtained; Based on the path performance score table, the scheduling paths in the unit scheduling graph are hierarchically clustered and associated mapped, and the evolutionary trend simulation is performed to obtain a path evolution decision graph; The paths with scores greater than or equal to 0.85 in the path evolution decision graph and the paths with average structure energy levels less than or equal to 0.25 are subjected to redundancy pruning and structure verification to screen a stable configuration path set; The stable configuration path set is dynamically spliced to form a configuration decision graph.
10. A bioinformatics processing system applied to synthetic biology, characterized by, The biological information processing system for synthetic biology comprises: A context decoding module is configured to acquire real-time proteomics data and perform polymorphic context decoding on the real-time proteomics data to obtain a protein expression context matrix; An interaction modeling module is configured to perform cross-domain graph structure deconstruction and reconstruction on the behavior characteristics in the protein expression context matrix to obtain a reconstructed protein matrix, analyze biological response path-protein domain interactions based on the reconstructed protein matrix, and obtain a multi-scale causal structure graph; A functional transformation analysis module is configured to perform protein function expression function reduction on the path nodes in the multi-scale causal structure graph to obtain a functional expression minimum module combination unit, perform structure-function transformation rule extraction based on the functional expression minimum module combination unit, and obtain a functional mapping unit set; A protein scheduling simulation module is configured to perform protein scheduling simulation based on the functional mapping unit set, and perform unit configuration path effectiveness and structure stability evaluation based on the protein scheduling simulation result to dynamically splice a configuration decision graph; A virtual response simulation module is configured to perform multi-scenario virtual response simulation on the paths in the configuration decision graph, and perform path adaptability verification and scoring on the multi-scenario virtual simulation result to obtain a path adaptability feedback table; A variability screening module is configured to perform multi-dimensional clustering and variability screening on the path adaptability feedback table to obtain an optimal expression configuration set.
Citation Information
Patent Citations
Gene fusions and gene variants associated with cancer
CN118910253A
Mining method and system for synthetic biological functional element and storage medium
CN119673283A