Biological information processing system and method applied to synthetic biology

Through a cross-scale, multi-dimensional protein expression regulation modeling process, the limitations of traditional methods in low-similarity protein identification and function prediction are overcome, high-precision protein function design and pathway regulation are achieved, and engineering applications in synthetic biology are supported.

CN120708691AActive Publication Date: 2025-09-26ZHEJIANG HUIJIA BIOTECH CO LTD

Patent Information

Application Number
CN202511145822.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-09-26
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Traditional protein data processing methods have limitations in dealing with low-similarity proteins, predicting novel structures or unknown functional regions, and are unable to meet the multi-dimensional requirements of synthetic biology for predictable functions, controllable designs, and adjustable pathways of protein modules in specific engineering scenarios.

Method used

By adopting steps such as polymorphic context decoding, cross-domain graph structure deconstruction and reconstruction, functional expression function reduction, protein scheduling simulation and multi-scenario virtual response simulation, a cross-scale, multi-dimensional and high-precision protein expression regulation modeling process is formed. By introducing nested extraction of expression semantic nodes and cross-domain feature relationship modeling, the coupling expression capability of protein behavior information and structural semantics is improved.

Benefits of technology

It achieves high-precision regulation of protein function design, expression pathway modeling and structure-function mapping, enhances the model's ability to recognize low-similarity regions, ensures the stability and controllability of the pathway, and supports the engineering application of protein modules in complex regulatory networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708691A_ABST
    Figure CN120708691A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of protein related data processing, in particular to a biological information processing system and method applied to synthetic biology. The method comprises the following steps: acquiring real-time proteomics data, and performing polymorphic context decoding to obtain a protein expression context matrix; analyzing interaction of behavior characteristics in the protein expression context matrix to obtain a multi-scale causal structure map; executing structure-function transformation rule extraction in the multi-scale causal structure atlas to obtain a function mapping unit set; performing evaluation based on the function mapping unit set so as to form a configuration decision diagram; simulating paths in the configuration decision diagram, and performing path adaptability scoring to obtain a path adaptability feedback table; and optimizing the path adaptability feedback table to obtain an optimal expression configuration set. According to the method, the precision, the stability and the controllability of protein information in the process of structural analysis, function recognition and regulation path construction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of protein-related data processing, and in particular to a biological information processing system and method applied to synthetic biology. Background Art

[0002] As important functional modules in synthetic biology, protein function design, pathway optimization, and component assembly are key areas of research. Modeling the structure-function relationship, sequence design optimization, and function prediction of proteins hold broad application prospects in areas such as gene circuit construction, metabolic pathway modification, and cell factory regulation. Traditional protein data processing methods are often based on sequence alignment, structure-template mapping, and expert rule-based construction, primarily including tools such as BLAST, PSI-BLAST, HMMER, and SWISS-MODEL. These methods have some practical value in identifying homology and aligning structure templates in the early stages of protein sequences, but they have significant limitations when processing low-similarity proteins and predicting novel structures or regions of unknown function. Furthermore, traditional methods often rely on static databases and manual annotation, lacking the ability to model dynamic structural changes, environmental response behaviors, and cross-scale regulatory pathways. This makes it difficult to meet the multi-dimensional demands of synthetic biology for predictable functions, controllable designs, and adjustable pathways for protein modules in specific engineering scenarios. Summary of the Invention

[0003] Based on this, it is necessary for the present invention to provide a biological information processing system and method applied to synthetic biology to solve at least one of the above technical problems.

[0004] To achieve the above objectives, a biological information processing method applied to synthetic biology comprises the following steps: Step S1: Acquire real-time proteomics data and perform polymorphic context decoding on the real-time proteomics data to obtain a protein expression context matrix; Step S2: Deconstruct and reconstruct the behavioral features in the protein expression context matrix across domains to obtain a reconstructed protein matrix; analyze the biological response pathway-protein domain interactions based on the reconstructed protein matrix to obtain a multi-scale causal structure map; Step S3: performing protein function expression function reduction on the path nodes in the multi-scale causal structure map to obtain the minimum module combination unit of function expression; performing structure-function transformation rule extraction based on the minimum module combination unit of function expression to obtain a function mapping unit set; Step S4: Using the functional mapping unit set as input, perform protein scheduling simulation, and evaluate the effectiveness of the unit configuration path and the structural stability based on the protein scheduling simulation results, thereby dynamically assembling to form a configuration decision diagram; Step S5: Perform multi-scenario virtual response simulation on the paths in the configuration decision diagram, and perform path adaptability verification and scoring on the multi-scenario virtual simulation results to obtain a path adaptability feedback table; Step S6: Perform multidimensional clustering and variability screening on the path adaptability feedback table to obtain the optimal expression configuration set.

[0005] The present invention describes a biological information processing method for synthetic biology, which focuses on key links such as protein function design, expression pathway modeling and structure-function mapping, and forms a set of cross-scale, multi-dimensional and high-precision protein expression regulation modeling processes, which can effectively make up for the shortcomings of traditional methods in low-similarity region identification, new structure function prediction and expression pathway reconstruction. By introducing nested extraction of expression semantic nodes and cross-domain feature relationship modeling, the coupling expression ability between protein behavior information and structural semantics can be significantly improved, thereby enhancing the model's expression integrity of protein contextual semantics. The construction of a structural compression representation graph can achieve effective compression of the map scale on the basis of ensuring the integrity of the information structure, reducing the computational redundancy and complexity in the subsequent structure projection and map modeling process. The domain space projection is combined with the coordinate system of the standard protein domain database for coordinate encoding, which not only ensures the biological rationality of the node mapping, but also improves the spatial positioning accuracy of the behavior node in the domain. The context-coupled retrieval and neighborhood cross-validation operations strengthen the semantic and spatial consistency between the mapping node and the database domain, and improve the accuracy and stability of the structure mapping. Chimera metric modeling integrates domain functional clustering information with structural category similarity, resulting in a chimera graph with dual structure-function consistency, avoiding structural conflicts or functional loss. Pathway-aware modeling and graph convolutional aggregation fully extract key response pathways and their causal response strengths within the expression pathway. The resulting causal response weight matrix is ​​not only highly interpretable but also provides clear metrics for subsequent pathway pruning and functional reorganization. Setting a maximum path depth of 6 helps control the complexity of expression pathways, avoiding the dilution of propagated information and the accumulation of interference caused by excessively long pathways. Setting the sparsity control parameter to 0.03 eliminates redundant expressions while ensuring function integrity, improving model computational efficiency and expression accuracy. Setting a hash distance threshold of 0.15 and an intra-cluster similarity of 85% precisely controls the granularity of expression function aggregation and strengthens semantic consistency between function expressions. During the expression function pruning phase, a dual-metric screening mechanism of a causal regulation weight threshold of 0.6 and an average gradient weight of 0.65 ensures that retained sub-functions retain their regulatory core role within the pathway while effectively eliminating redundant components with insufficient expression activity. Setting a matching tolerance threshold of 0.1 improves the selectivity and reliability of expression topology stability rule matching, helping to ensure that subsequently constructed expression combination units possess stable structures and high availability. During the functional structure graph modeling phase, topological rearrangement is used to generate a structural graph fingerprint, and symbolic function mapping is performed in conjunction with the structure-function mapping relationship to ensure structural uniqueness and functional accuracy in path mapping and structure identification. The generation of a set of rule-expressed substructures, combined with rule reduction operations, helps to extract the core structure of expression regulation, providing a controllable input space for subsequent transformation path modeling.The dual thresholds of a score ≥ 0.85 and a structural energy level ≤ 0.25 used in pathway scoring ensure that the ultimately selected stable configuration pathways possess both high functional accessibility and low structural coupling energy, thereby achieving dynamic optimal matching between expression pathways and structural configurations. Furthermore, by constructing a pathway execution feature set and a fusion model to uniformly model pathway execution complexity and conflict characteristics, the generated scheduling pathways not only achieve synergistic combinations at the structural level but also exhibit good stability and scheduling robustness at the execution level, ensuring that protein modules are both engineerable and optimizable within complex regulatory networks.

[0006] Optionally, this specification further provides a biological information processing system applied to synthetic biology, for executing the biological information processing method applied to synthetic biology as described above, the biological information processing system applied to synthetic biology comprising: A context decoding module is used to acquire real-time proteomics data and perform polymorphic context decoding on the real-time proteomics data to obtain a protein expression context matrix; The interaction modeling module is used to deconstruct and reconstruct the cross-domain graph structure of the behavioral features in the protein expression context matrix to obtain a reconstructed protein matrix; based on the reconstructed protein matrix, the biological response pathway-protein domain interaction is analyzed to obtain a multi-scale causal structure map; Functional transformation analysis module, used to reduce protein functional expression functions of path nodes in multi-scale causal structure maps to obtain the minimum functional expression module combination unit; based on the minimum functional expression module combination unit, structure-function transformation rules are extracted to obtain a functional mapping unit set; The protein scheduling simulation module is used to perform protein scheduling simulation using the functional mapping unit set as input, and evaluate the effectiveness of the unit configuration path and the structural stability based on the protein scheduling simulation results, thereby dynamically assembling and forming a configuration decision diagram; The virtual response simulation module is used to perform multi-scenario virtual response simulation on the path in the configuration decision diagram, and to verify and score the path adaptability of the multi-scenario virtual simulation results to obtain a path adaptability feedback table; The variability screening module is used to perform multidimensional clustering and variability screening on the path adaptability feedback table to obtain the optimal expression configuration set.

[0007] The bioinformation processing system applied to synthetic biology of the present invention can implement any of the bioinformation processing methods applied to synthetic biology of the present invention, and is used to combine the operations and signal transmission media between various modules to complete the bioinformation processing method applied to synthetic biology. The modules within the system cooperate with each other, thereby improving the accuracy, stability and controllability of protein information in the process of structural analysis, function identification and regulatory pathway construction. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments thereof made with reference to the following drawings: Figure 1 Schematic diagram of the steps of the biological information processing method applied to synthetic biology of the present invention; Figure 2 Detailed step flow diagram of step S1 in the present invention; Figure 3 Detailed step flow diagram of step S2 in the present invention; The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0009] The following is a clear and complete description of the technical method of the present invention in conjunction with the accompanying drawings. It is obvious that the embodiments described are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of the present invention.

[0010] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor and / or microcontroller approaches.

[0011] It should be understood that although the terms "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. The term "and / or" as used herein includes any and all combinations of one or more of the listed associated items.

[0012] To achieve this, please refer to Figures 1 to 3 The present invention provides a biological information processing method applied to synthetic biology, the method comprising the following steps: Step S1: Acquire real-time proteomics data and perform polymorphic context decoding on the real-time proteomics data to obtain a protein expression context matrix; In this example, a Thermo Fisher Orbitrap Exploris 480 mass spectrometer coupled with an Agilent 1290 Infinity II high-throughput liquid chromatography platform was used to acquire multiple batches of protein samples from HeLa cells under different induced states. During mass spectrometry acquisition, a resolution of 120,000 was set, the full scan range was set from 350–1,800 m / z, and the liquid phase gradient elution time was set to 90 minutes. Publicly available protein expression data related to the selected cell states from cell expression databases (such as the Human Protein Atlas) were also simultaneously imported. After data acquisition, the data streams from different sources were time-synchronized, and timestamps were aligned using linear interpolation to resolve inconsistencies in data acquisition time. Subsequently, the protein expression signal is decomposed into distinct contextual states using a multi-state context decoding model. This model comprises three nested layers: the first is an environmental factor encoding layer, capturing different intracellular and extracellular stimulus conditions; the second is an expression pattern mapping layer, representing the differences in protein expression under different states; and the third is a behavioral fusion layer, which integrates the expression data into a three-dimensional protein expression context matrix with dimensions [N proteins × M time points × K contextual states]. This matrix is ​​a tensor structure, with each element corresponding to the protein expression level at a specific time and context. The values ​​are normalized to a range of 0–1 to facilitate subsequent processing.

[0013] Step S2: Deconstruct and reconstruct the behavioral features in the protein expression context matrix across domains to obtain a reconstructed protein matrix; analyze the biological response pathway-protein domain interactions based on the reconstructed protein matrix to obtain a multi-scale causal structure map; In this embodiment, the behavioral features in the protein expression context matrix are extracted to construct a cross-domain graph structure. The graph consists of nodes and edges. The nodes represent proteins and their expression behavior channels, and the edges represent the interactions between proteins in different biological response paths. The construction of the graph is based on the adjacency matrix A and the feature matrix X. The adjacency matrix defines the connection relationship between nodes, and the feature matrix X contains the behavioral feature vectors of protein expression. For this graph, the graph deconstruction technology is used to filter out low-correlation edges through edge weight thresholds, and reconstruct a sparse adjacency matrix. , and then generate a reconstructed protein matrix. Subsequently, combined with a standard protein domain database, the behavioral nodes in the graph are mapped to the corresponding domain coordinate system, achieving behavior-structure chimera. Through layer-by-layer aggregation of the multi-scale graph, the causal relationships between protein domains along the path are extracted, resulting in a multi-scale causal structure graph. This graph is represented as a multi-layer directed graph structure, with causal paths at different scales between layers.

[0014] Step S3: performing protein function expression function reduction on the path nodes in the multi-scale causal structure map to obtain the minimum module combination unit of function expression; performing structure-function transformation rule extraction based on the minimum module combination unit of function expression to obtain a function mapping unit set; In this embodiment, in the multi-scale causal structure map, the protein nodes and their functional expression information on each path are extracted, the causal chain is expanded, the maximum path depth is limited to 6, and a path expression function set is formed. Symbolic simplification is performed on these function sets, the threshold is set to 0.03, low-weight expression functions are eliminated, and a function sparse structure tensor is generated. The expression function is processed using the local sensitive hashing method, the hash distance threshold is set to 0.15, and it is clustered based on the expression function similarity ≥ 85%. Combining the causal weight graph with the gradient propagation flow graph, the core sub-functions are screened out, the low-weight functions are pruned, and the function clipping set is obtained. According to the predefined stable function pattern in the expression topology stability rule library, the functions are matched and screened to form an expression topology stability map. Finally, the clipped functions are reorganized and sorted, and the minimum module combination unit of functional expression is output, specifically a combination containing 3 to 7 core functional sub-units, which is convenient for subsequent functional mapping.

[0015] Step S4: Using the functional mapping unit set as input, perform protein scheduling simulation, and evaluate the effectiveness of the unit configuration path and the structural stability based on the protein scheduling simulation results, thereby dynamically assembling to form a configuration decision diagram; In this embodiment, the function mapping unit set is used as the input of the scheduling simulation, and the simulation platform is set to simulate the execution process of protein scheduling, including the scheduling path, functional unit dependency and scheduling order. A unit scheduling graph is constructed, in which the nodes represent the function mapping units, the edges represent the dependency, and the directed graph structure clearly defines the execution order. For the unit scheduling graph, the bottleneck links, feedback loops and timing conflicts in the path are evaluated, and the path conflict vector is extracted. The parameters include the timing conflict intensity. The path connection density and energy level transition distribution are calculated to generate a structural energy distribution map. Combining the path conflict vector with the energy map, a path performance score table is formed, and a score ≥0.85 is an excellent path. The score table is hierarchically clustered and combined with evolutionary trend analysis to screen paths with stable structure and excellent performance, cut redundant paths, and finally dynamically splice to form a configuration decision diagram, which is expressed as a directed weighted graph, and the weight reflects the path stability and execution efficiency.

[0016] Step S5: Perform multi-scenario virtual response simulation on the paths in the configuration decision diagram, and perform path adaptability verification and scoring on the multi-scenario virtual simulation results to obtain a path adaptability feedback table; In this example, based on the configuration decision graph, pathways with higher scores were selected and response simulations were performed in three virtual biological scenarios: simulating the cell proliferation environment, the stress response environment, and the metabolic regulation environment. 1) Cell proliferation environment: This simulates the dynamics of protein regulation during the cell cycle, with a cell cycle parameter set to 24 hours and a simulation time step of 5 minutes. 2) Stress response environment: This simulates the regulation of protein expression in cells in response to external environmental stresses (such as heat shock and oxidative stress), with a stress intensity gradient set to 0–1 and a simulation duration of 1 hour. 3) Metabolic regulation environment: This simulates protein regulation in the metabolic network, with metabolite concentrations ranging from 0 to 100 μM and a simulation time step of 10 minutes for 2 hours. In each scenario, a multi-parameter biological response model was used to simulate the activation of protein expression nodes in the pathway and generate a response curve. The multi-parameter biological response model used in the simulation is derived from a modified version of the classic cell signaling model, combining protein expression dynamics and molecular regulatory mechanisms. The model consists of multiple coupled differential equations, the core of which includes activation function, feedback regulation module and regulatory coupling term: the activation function adopts the Sigmoid function form to describe the response of protein expression to activation signal, which is in the form of , where the parameters Control the steepness of the curve, represents the activation threshold. The feedback regulation module designs a negative feedback loop to simulate the dynamic equilibrium of protein expression levels by regulating the protein degradation rate and the concentration of transcriptional repressors. The regulatory coupling term describes the interaction effects between proteins and the influence of environmental factors. It adopts a linear weighted superposition approach, with weight parameters preset based on biological experimental data and adjustable between 0.1 and 0.9. The model inputs are the initial expression states of each protein in the pathway and environmental parameters, and the output is a time series of expression intensities. This model structure has been validated through biological experiments and can accurately reflect the temporal dynamics and intensity changes of protein expression responses. Model parameters can be adjusted according to different simulation scenarios to ensure the biological relevance and real-world applicability of the simulation results. During the simulation, the following key metrics were collected: pathway activation time (in minutes), peak expression intensity (normalized from 0 to 1), response duration (the duration exceeding the threshold of 0.7), and a scenario-specific regulatory efficiency index. All data are summarized in a "pathway adaptability feedback table" stored in matrix format, with rows corresponding to pathways and columns corresponding to indicators for each scenario. Data is formatted in double-precision floating-point format to ensure simulation accuracy. Subsequently, the comprehensive adaptability score was calculated using the weighted average method, with the weights set to 0.4 (cell proliferation), 0.35 (stress response), and 0.25 (metabolic regulation) according to the importance of the scenario, and normalized to facilitate score sorting and screening.

[0017] Step S6: Perform multidimensional clustering and variability screening on the path adaptability feedback table to obtain the optimal expression configuration set.

[0018] In this embodiment, a multidimensional time series statistical analysis is performed on the path adaptability feedback table to construct a path adaptability feature matrix. The rows of the matrix represent the path number, and the columns represent the changes in adaptability indicators at different time steps. The variance, kurtosis, and skewness of each indicator are calculated for each path, with a threshold set to variance > 0.02. A subset of paths with significant adaptability fluctuations is identified. This subset is then screened for variability, and the dynamic time warping (DTW) method is used to assess the similarity between time series. These clusters are then clustered into multiple variant path clusters. For each variant path cluster, the adaptability score trends are analyzed to identify potential optimization points and deficiencies. In conjunction with a path simulation model, targeted adjustment strategies (such as functional module reordering and path node activation threshold adjustment) are designed. The variant paths are then simulated and iteratively optimized. During the optimization process, changes in adaptability indicators before and after optimization are recorded in real time to ensure effective improvements. Finally, the optimized path set is integrated with the original stable path set to establish a unified path representation configuration set. This set consists of two components: pathways that have been validated across multiple scenarios and demonstrate stable adaptability, and pathways that have been optimized to significantly improve performance. A comprehensive scoring mechanism is used to re-rank the pathways and output the optimal expression configuration set for subsequent experimental design and protein function regulation. This set is stored as a weighted directed graph, with nodes representing functional modules and edge weights reflecting pathway reliability and regulatory efficiency, facilitating subsequent pathway selection and dynamic adjustment.

[0019] It is particularly important that the variability screening in step S6 is specifically: Based on the path adaptability feedback table, the adaptability fluctuation data of each path is extracted to construct the path adaptability feature matrix; In this embodiment, the adaptability performance of each path at different time points is extracted from the multi-scenario indicator data in the path adaptability feedback table to construct a path adaptability feature matrix. The rows of this matrix represent different paths, and the columns correspond to the time series adaptability indicators for each simulation scenario, such as peak expression intensity and response duration. The matrix elements are double-precision floating-point numbers, and the matrix dimensions are typically calculated as the number of paths × the number of time steps × the indicator dimension. This matrix comprehensively reflects the dynamic adaptability characteristics of the paths, facilitating subsequent multidimensional analysis.

[0020] Perform statistical analysis on the path adaptability characteristic matrix, calculate the variance, skewness and kurtosis of each indicator, preliminarily identify the path set with significant fluctuation amplitude, and obtain the significant fluctuation path set; In this embodiment, for the constructed path adaptability feature matrix, statistics are calculated indicator by indicator, including variance (measurement of data fluctuation), skewness (describing the asymmetry of data distribution) and kurtosis (reflecting the sharpness of data distribution). For example, when calculating the variance of the peak expression intensity sequence, a time window length of 30 time points is selected, and the skewness and kurtosis are calculated based on the statistical values ​​of the entire time series. By screening the thresholds of these statistical features (such as the variance threshold is set to 0.05, the skewness threshold is ±0.3, and the kurtosis threshold is 3), a set of paths with significant fluctuations is preliminarily screened out to facilitate the identification of potential abnormal dynamic responses.

[0021] Perform multi-scale time series analysis on the set of significant fluctuation paths, mark abnormal time series fluctuations, and obtain abnormal fluctuation path mark vectors; In this example, the time series data for each path in the selected set of significant fluctuation paths is fed into a multi-scale time series decomposition framework and decomposed into trend, cycle, and random fluctuation components. Anomaly detection is performed on the random component using a threshold rule (variance > 0.05 indicates significant fluctuation and dynamic instability), flagging abnormal time series fluctuation events. The anomaly flagging results are stored as a binary vector with a length equal to the length of the time series, where 1 represents an outlier and 0 represents a normal value. This step helps reveal sudden changes and anomalous behavior in path adaptability over time.

[0022] Filter the false positive paths from the abnormal fluctuation path marker vector to obtain the true variation path subset; In this example, a false-positive elimination mechanism was designed based on the obtained abnormal fluctuation path marker vectors, combined with common noise and model error characteristics found in historical simulation data. By setting a minimum consecutive abnormality length threshold (e.g., at least five consecutive abnormal points) and an abnormality frequency threshold, sporadic and non-persistent false-positive fluctuation paths were eliminated. Ultimately, a subset of pathways with true biological variation was screened, ensuring that subsequent analysis focused on reliable dynamic abnormal pathways.

[0023] Cluster analysis is performed on the subset of true variation paths, grouping them according to variation type and magnitude to form variation path clusters; In this example, a subset of true variation paths is grouped into path clusters using a feature similarity metric based on their variation type (e.g., sudden peak increase, continuous decrease, cyclical fluctuation, etc.) and variation amplitude. Each cluster contains paths with similar performance, facilitating analysis of their shared dynamic behavior patterns and underlying mechanisms. The clustering results are characterized by constructing a cluster center vector and an intra-cluster variance matrix. The cluster center vector represents the typical adaptive variation trend of the group of paths.

[0024] Combining the evolutionary trend and adaptability score of the mutation path cluster, we screen out the paths with significant mutations and optimized feasibility, and generate an optimizable screening path set.

[0025] In this example, a multidimensional screening strategy was implemented, combining analysis of the evolutionary trends of variant pathway clusters (e.g., trend persistence and fluctuation stability) with a comprehensive fitness score of the original pathways. The strategy retained pathways that exhibited both significant variation and optimization potential in simulation. Screening criteria included an evolutionary trend duration of at least 50 time steps and a fitness score above 0.8, ensuring that the pathways possessed biological significance and engineering feasibility in terms of functional expression and regulation. The resulting set of optimizable screening pathways was used for subsequent pathway integration and the construction of optimal expression configuration sets, facilitating the optimal design of protein regulatory strategies.

[0026] Optionally, step S1 specifically includes: Step S11: real-time proteomics data are collected through a mass spectrometer, a high-throughput liquid chromatography platform, and a cell expression database, and the real-time proteomics data are time-series aligned and data decoupled to obtain a raw protein expression data set; In this example, a Thermo Fisher Orbitrap Exploris 480 mass spectrometer coupled with an Agilent 1290 Infinity II high-throughput liquid chromatography platform was used to acquire multiple batches of protein samples from HeLa cells under different induced states. During mass spectrometry acquisition, a resolution of 120,000 was set, the full scan range was set to 350–1,800 m / z, and the liquid phase gradient elution time was set to 90 minutes. Publicly available protein expression data related to the selected cell state from cell expression databases (such as the Human Protein Atlas) were simultaneously imported. The acquired data were time-divided into 5-second intervals, and the experimental data were time-mapped and aligned with the static database data using a synchronous alignment module. After performing peak identification, isotope removal, and retention time recalibration on the acquired data, multi-source mapping was used to decouple the protein ID, expression abundance, and mass spectrometry signal-to-noise ratio fields into a single output, generating a raw protein expression dataset.

[0027] It is worth noting that the Thermo Fisher Orbitrap Exploris 480 is a high-resolution mass spectrometer used to precisely measure the mass and relative abundance of molecular ions. It is commonly used in proteomics analysis because it provides very high mass accuracy and sensitivity, helping to identify and quantify proteins in complex biological samples. The Agilent 1290 Infinity II High-Throughput Liquid Chromatography Platform is an ultra-high-performance liquid chromatography (UHPLC) system used to separate samples before entering the mass spectrometer, separating the components of complex mixtures by chemical properties (such as hydrophobicity and polarity) for more accurate detection and analysis by the mass spectrometer. When used together, the liquid chromatography performs separation and the mass spectrometer performs detection and analysis, enabling high-throughput, precise molecular identification and quantification of complex biological samples.

[0028] Step S12: performing high-dimensional semantic feature mapping conversion on the original protein expression dataset to construct a protein expression semantic tensor; In this example, based on the obtained protein expression original dataset, each protein expression event is encoded as a five-tuple by setting a five-dimensional semantic space projection (protein function, cell location, regulatory state, biological process, expression abundance). Using the definitions in the GeneOntology database as a reference, a nested semantic mapping dictionary was constructed, and this dictionary was used to semantically replace and map each row in the original expression matrix. The semantic structure was represented by a third-order tensor with the dimensions (protein_index × time_point × semantic_dimension). For example, for a sample containing 500 proteins, 60 time points, and 5-dimensional semantic information, the following semantic tensor was constructed: Each element represents the semantic vector projection corresponding to the protein at a specified time point, for example, [nucleus, activation state, high expression, involved in transcriptional regulation, transcription factor] is mapped to [0.89, 0.74, 0.92, 0.81, 0.93]. This ultimately forms a semantic tensor that can be used for subsequent context modeling.

[0029] Step S13: performing polymorphic context nested modeling based on the protein expression semantic tensor and extracting the nested relationship of environmental regulatory factors, thereby constructing a context nested representation model; In this embodiment, dynamic context nesting modeling is performed based on the semantic tensor structure and the semantic similarity between proteins under regulatory states. The nesting relationship is constructed by constructing a three-layer nested structure of "state-factor-expression" to establish a multi-granular expression environment model. The model is expressed in tensor expansion form as: Context_Nesting_Model={context_layer_1: state vector ,context_layer_2: environmental factor nested set F_i={f_ij|j=1...k},context_layer_3: expression response nested matrix }; where fi_ij represents the state For example, in the TGF-β induction state, Corresponding to the induced state, Including pH, temperature, Environmental factors such as concentration, The specific regulatory responses of these factors to protein expression are obtained through tensor analysis and conditional screening, ultimately resulting in a set of structured nested context models for the dynamic interpretation of expression status.

[0030] Step S14: analyzing the expression behavior vector sequence of the protein in various states based on the context nested representation model, thereby forming an expression state dynamic matrix; In this example, the context nesting model is input into the expression behavior extraction module to generate a dynamic sequence of expression states by identifying the expression vector changes of proteins in specific nested contexts. For example, the expression of protein P12345 under different stress conditions is The expression behavior will show significant fluctuations, and its expression behavior sequence is presented in the form of a state vector as follows: ; Summarize the expression behaviors of all proteins and construct the structure: ;in is the amount of protein, The state label information (such as high expression, stress response, and inhibitory expression) is annotated to ensure that each expression behavior vector has traceable contextual meaning, which serves as the basis for subsequent modeling.

[0031] Step S15: Fuse the expression state dynamic matrix and perform low-rank semantic compression to construct the protein expression context matrix.

[0032] In this embodiment, the above expression state dynamic matrix is ​​subjected to low-rank representation compression processing to retain the main change trends and eliminate redundant features. The first k=20 principal components are retained by singular value screening to construct a low-rank expression state matrix: Based on low-rank semantics, the expression matrices in different time periods and environments are integrated to construct the protein expression context matrix , the matrix dimensions are: ;in The compressed semantic dimension, for example, d = 32, is used. Each row represents the contextual expression profile of a protein. Structurally, each element is a comprehensive expression that integrates expression state, nested environmental factors, and semantic attributes, forming a unified input foundation for subsequent structural analysis and functional reduction.

[0033] Optionally, step S13 is specifically as follows: Step S131: extracting state identification factors and experimental metadata from the protein expression semantic tensor to obtain a factor candidate set; In this embodiment, the content of a specific dimension is extracted from the constructed protein expression semantic tensor as the state identification factor and the experimental metadata input source. ,in is the amount of protein, For time point, The tensor is a semantic dimension (e.g., function, location, regulatory state, biological process, expression level). We extract identifiers from the regulatory state dimension of this tensor and experimental metadata fields (e.g., induction method, culture conditions, stress type, etc.), and associate them with fields such as acquisition batch, culture temperature, treatment duration, and cell type in the experimental metadata to form a joint candidate factor description set for subsequent pathway construction.

[0034] Step S132: constructing an initial nested path graph based on the factor candidate set and performing context relevance evaluation to obtain an initial nested path graph; In this embodiment, the candidate factor set F_candidate is used as a node to construct the initial nested path graph , where V represents the candidate factor and E represents the contextual semantic co-occurrence relationship between the two factors. Contextual relationships are constructed using two sources: 1) Co-occurrence frequency analysis: calculating the number of co-occurrences of each factor combination under different experimental conditions; and 2) Semantic similarity analysis: assessing the regulatory similarity between factors based on a semantic embedding space (e.g., the GO definition), using a vector angle of less than 30° as the threshold for establishing connections. Setting the minimum co-occurrence frequency threshold to 3 and the minimum semantic similarity to 0.7, we obtain the following initial nested path diagram connection diagram: The edge weights of the graph represent the strength of contextual semantic coupling, and the structure is used to support subsequent subgraph segmentation.

[0035] Step S133: dividing the initial nested path graph into context-coupled subgraphs to construct a set of factor-coupled subgraphs; In this embodiment, in the initial nested path graph The entire graph is segmented by context coupling strength, and pairs of factors with significant semantic coupling are grouped into the same subgraph. Based on the continuity of the context semantic propagation path, connected context factor sequences are extracted as factor coupling subgraphs. For example, the following coupling subgraph set is constructed through propagation path analysis: Each subgraph represents a potential nested expression pathway, indicating the sequential relationships between multiple regulatory factors in a specific experimental context. Each subgraph has a minimum node count of 2 and a maximum length of 5 to avoid isolated nodes or lengthy pathways.

[0036] Step S134: Map each path pair in the factor coupling subgraph back to the behavior channel in the protein expression semantic tensor, thereby obtaining a context nested tensor structure; In this example, the path in each factor coupling subgraph is mapped back to the original behavior channel in the protein expression semantic tensor T_expr. The mapping logic performs index reverse retrieval based on the index position and channel number of the semantic dimension (regulatory state) corresponding to the path factor in the tensor. For example, if the subgraph path is , then locate all protein behavior vector sequences in the tensor that match these two labels in the regulatory state dimension and integrate them to generate a tensor structure The resulting set of contextual nested tensor structures provides a structural mapping between regulatory pathways and protein expression behaviors for expression state modeling.

[0037] Step S135: Construct a context nested representation model based on the context nested tensor structure.

[0038] In this embodiment, based on the tensor structure set constructed by T_nested, a complete context nested representation model C_nested_model is established by fusing the behavior tensors corresponding to each subgraph path to capture the regulatory nested relationship of multi-path context on protein expression behavior. The form of the context nested representation model can be: C_nested_model={context_unit_i:{subgraph_id:i,factor_sequence:[ ],behavior_profile: ,interaction_strength_matrix: , }; The interaction_strength_matrix represents the strength of the joint influence between factors in the pathway, obtained by evaluating the consistency of expression responses within the tensor and represented by a standard deviation matrix or correlation matrix. The entire model can be used to support regulatory modeling, expression prediction, or functional linkage analysis of protein expression under the combined effects of multiple state factors, with good structural interpretability and nested logical integrity.

[0039] Optionally, step S14 is specifically as follows: Step S141: parsing the context nested representation model to extract the state master control factor sequence and response channel structure, thereby constructing a state channel index table; In this embodiment, the context nested representation model C_nested_model is structurally parsed to extract the state master factor sequence contained in each context unit and the response channel information mapped thereto. The specific form of the structure contained in each context unit in the model C_nested_model can be expressed as: context_unit_i={factor_sequence:[ ],behavior_profile ,interaction_strength_matrix: }; factor_sequence is the master control factor sequence, and the S' dimension in behavior_profile corresponds one-to-one to the tensor channel mapping relationship. By traversing all context units in the model, the master control factor sequence and its mapped behavior channel index in the protein expression semantic tensor T_expr are extracted, thereby establishing a state-channel correspondence index table. The specific structure of the state-channel correspondence index table can be expressed as follows: ; This index table is used to guide subsequent behavioral feature extraction and expression modeling steps to ensure that the extracted data is tightly bound to the state label and avoid behavioral channel mismatch.

[0040] Step S142: extracting multi-channel behavior data from the protein expression semantic tensor according to the state channel index table, constructing a protein expression behavior feature space, and performing state expression vector aggregation modeling to obtain an expression behavior time series group; In this embodiment, based on the state channel index table Index_table, the semantic tensor of protein expression is expressed Extract multiple channel indices corresponding to each state label and integrate them to form a multi-channel behavior sub-tensor group T_sub. For example, for "hypoxia_response", extract channels 3, 5, and 9 to construct the behavior tensor fragment The set of fragments of all states is constructed into a behavioral feature space, and the structure of the behavioral feature space is: ; Then, each state tensor fragment is aggregated and modeled in the time dimension, and the behavior change of a protein in a certain state is represented by the state expression vector v_state(t), which is defined as follows: ; Get the expression behavior time series group: ; Each time series is used to characterize the evolutionary behavior of protein expression under a specific regulatory state.

[0041] Step S143: Modeling expression vector change patterns for each expression behavior time series group at different preset time resolutions, and marking the regulation time windows to form a nested expression trajectory set; In this embodiment, for the expression behavior time series group S_behavior, three time resolution windows are set: 2 hours, 6 hours, and 24 hours. Time series segments are intercepted respectively to observe the changing trends of expression vectors at different scales. At each resolution, a sliding window vector comparison is performed to identify the expression change interval, and the mean change rate exceeding the set threshold (such as ±15%) is used as the basis for marking the state transition. For example, in the "oxidative_stress" expression sequence: 1) The expression vector change rate reaches 20% in the 12th to 16th hour period, which is marked as the regulatory time window W1=[12h,16h]; 2) The change in the 30th to 36th hour period is stable and unmarked. The change time window marks of all states are merged to construct a nested expression trajectory set: ; This set of trajectories is used to model the causal structure of subsequent states.

[0042] Step S144: constructing a state vector graph model based on the nested expression trajectory set to dynamically deduce the causal regulatory relationship between different expression states and output a state regulation response graph; In this embodiment, in this step, a state vector graph model G_state=(V, E, W) is constructed based on the nested expression trajectory set. Where: V is the state factor set, such as ; E represents a directed edge with a regulatory relationship between states; W is the edge weight matrix, which represents the regulatory weight between different states. Weight deduction is based on the expression change response relationship within the time window. For example, if the expression change of "hypoxia_response" occurs within 2 hours before "oxidative_stress" and the similarity is higher than the set value (such as cosine similarity > 0.85), then the edge is established: ( ); the graph structure is as follows: This diagram expresses the regulatory causal structure of protein expression driven by multiple state factors.

[0043] Step S145: The time evolution trajectory of the nested expression trajectory set is integrated with the causal regulation weight in the state regulation response graph, and the temporal weight normalization reconstruction is performed to obtain the expression state dynamic matrix.

[0044] In this embodiment, the time evolution information of the nested expression trajectory set is integrated with the edge weights in the state causal graph to establish the final expression state dynamic matrix The fusion method is as follows: for each state v_i's expression vector sequence, a weighted combination is performed based on its incoming edge relationship in the state vector graph. To prevent weight imbalance, normalization is performed: 1) The control weight vector W[:,i] of each state is normalized to the interval [0,1]; 2) The normalized fusion value is z-score normalized to make the dynamic matrix expression of each state comparable. Ultimately, the expression state dynamic matrix M_dyn expresses the aggregate response behavior of each protein under the regulation of different state factors and changes over time, and is in the following form: The first row corresponds to the aggregated expression intensity of the "hypoxia_response" state at different time points, the second row corresponds to the "oxidative_stress" state, and so on. This matrix can serve as the core input for subsequent protein expression behavior reconstruction, functional mapping, and regulatory strategy deduction.

[0045] Optionally, step S2 is specifically: Step S21: performing nested extraction of expression semantic nodes on the protein expression context matrix to construct a cross-domain feature relationship graph; In this embodiment, based on the obtained protein expression context matrix (in is the number of proteins, is the time point, (where is the number of contextual semantic channels) and semantic nesting extraction is performed on the expression nodes in each semantic channel. The nesting extraction operation is based on contextual semantic similarity. The expression node pairs with semantic cosine similarity greater than 0.85 between contextual channels are selected to form a semantic nesting pair set: ; Then, the feature entity node graph G_feat=(V,E,A) is constructed based on the nested pair set, where: 1) V is the protein expression semantic node set; 2) E is the semantic nested edge; 3) The semantic feature vector for each node is constructed with a dimension d of 128. Furthermore, to enhance the interaction between expression behaviors across contextual channels, a cross-channel edge weight correction mechanism is introduced into the graph structure. This mechanism adds edges to nodes with consistent expression trends, with a threshold of a direction angle of less than 20° within the sliding window. The resulting cross-domain feature relationship graph serves as the basis for subsequent structural analysis.

[0046] Step S22: performing graph sparsification reconstruction on the cross-domain feature graph of the expression behavior to obtain a structure compression representation graph; In this embodiment, the constructed feature relationship graph G_feat is structurally sparse and reconstructed to reduce redundant edges and excessive connections. The specific processing method is as follows: first, the edge connectivity of each node is calculated, and its local edge weight aggregation distribution is statistically analyzed; edges with connectivity less than the mean μ and edge weight lower than the lower quartile Q1 are removed; at the same time, for some node pairs that show strong correlation but no structural connection, if their semantic cosine similarity is greater than 0.9 and the Euclidean distance of the time behavior curve is less than 0.9, the edge is removed. (default ), then add new edges. After the above pruning and filling mechanisms, the structural compression graph is reconstructed , where the number of nodes remains unchanged, but the number of edges is reduced by about 35% on average compared to the original graph. This graph structure serves as the input for domain mapping, facilitating subsequent spatial alignment operations.

[0047] Step S23: connecting to a standard protein domain database, mapping the behavior nodes in the structure compression representation graph to the domain coordinate system in the standard protein domain database, and performing chimeric modeling to generate a behavior-structure chimeric graph; In this embodiment, the behavior nodes in the compressed graph G_sparse are mapped to a standard protein domain database (such as the Pfam database or the SCOPe database). The standard domain model provided by the database includes protein families, functional regions, and three-dimensional coordinate fragments. The standard domain model of a specific database can be expressed as: The behavioral nodes in the map are mapped to protein position indexes, and the nodes are mapped to the corresponding domain coordinates by parsing the proteins, time points, and semantic labels. Chimera modeling uses structural alignment evaluation indicators such as structural center angle and residue overlap to select the optimal chimera pair. A behavior-structure chimera graph is formed. ,in: is a set of behavior nodes; is the set of structure domain nodes; It maps edges between behavior nodes and structural domains, and the edge weight is the embedding fit (0~1).

[0048] Step S24: performing path-aware extended modeling on the behavior-structure chimeric graph to generate a response path graph group; In this embodiment, based on the behavior-structure mosaic graph , perform path-aware extension operations. First, define the path extension window (the default length is 5 steps), starting from each chimeric behavior node, combine its behavior trend direction with the physical contact topology of the domain to construct a local response path: 1) If the expression increase rate of the node in 3 consecutive steps is >20%, and its domain belongs to the same functional module (such as kinase structure module), then the path is extended to a potential response chain; 2) If there is a reversal of the behavior trend (such as recovery after a decrease in expression), a branch path is constructed and marked as a turning node. All paths are combined to form a response path graph group , each graph corresponds to a protein regulatory response process.

[0049] Step S25: performing multi-scale graph convolution feature aggregation on the response path graph group to extract the causal response weight matrix; In this embodiment, for the response path diagram group Each path subgraph in the graph performs multi-scale graph convolutional feature aggregation modeling. The context behavior vector of each node in the graph is As the initial features, the scale range is set to {1,2,4} (representing the 1st, 2nd and 4th order neighborhoods), and the graph structure-guided feature propagation and weighted integration are performed. After aggregation, the global expression vector is extracted for each path graph. , calculate the behavioral response correlation between the graph and the state label, and form a causal response weight matrix , where: the row and column indices are the response path numbers; the values ​​are the response coupling degrees between the paths (such as the cross-regulation response strength, ranging from [0,1]).

[0050] Step S26: Based on the causal significance and path density threshold in the causal response weight matrix, weighted reconstruction and screening are performed on each path in the response path graph group to construct a multi-scale causal structure map.

[0051] In this embodiment, each path in the response path diagram group is weightedly reconstructed and screened based on two core indicators: 1) causal significance threshold : Weight value from M_causal, set the threshold as , path pairs below this value are eliminated; 2) Path density threshold : It is defined as the ratio of the number of edges in the path graph to the maximum possible number of edges, set to =0.45. Values ​​below this are considered weak behavioral connections and are not included in the final graph. The selected pathways are fused to construct a multi-scale causal structure graph G_causal_final=(V,E,W). Edge weights in the graph inherit the filtered weights from the causal-response matrix, while node attributes retain the contextual behavioral vectors inherited from the original graph. This graph provides a structural foundation for subsequent protein regulatory mechanism inference, expression prediction, or target identification.

[0052] Optionally, step S23 is specifically as follows: Step S231: connecting to a standard protein domain database, and semantically aligning the semantic features of each behavior node in the structure compression representation graph with the domain description vector of the standard protein domain database to obtain a semantic index mapping table; In this embodiment, the standard protein domain database (such as Pfam-A or SCOPe) is docked to extract its built-in domain description information and construct a domain semantic description vector set D_struct={ , ,..., }, where each Represents the semantic description vector of the structural domain, with a dimension of 256. At the same time, the semantic vector of the behavior node in the structural compression representation graph is (Inherited from the previous graph) is normalized. On this basis, cosine similarity matching is performed on all behavior nodes and structural domain description vectors, and the matching threshold is set to 0.7 to construct a semantic index mapping table. , where each behavior node can be mapped to at most three structural domains, sorted by similarity. This mapping table serves as the basis for subsequent spatial projection and positioning, ensuring the accuracy of semantic matching of structural domains.

[0053] Step S232: performing domain space projection on each behavior node in the structure compression representation graph based on the semantic index mapping table, and performing coordinate encoding in combination with the domain coordinate system of the standard protein domain database to obtain a domain positioning tensor; In this embodiment, each behavior node v_i in the diagram is mapped to the three-dimensional coordinate system corresponding to the structural domain d_j according to the generated semantic index mapping table T_map. The standard structural domain database provides the structural domain coordinate origin, orientation matrix and functional block coordinates. Taking the Pfam structural domain PF00001 as an example, the description form of its structural domain PF00001 includes: Centroid: (x=12.4, y=8.3, z=5.1), Orientation: [0.6, 0.2, 0.7], FunctionalBlocks: { :[15-25], :[55-75]}; Combine the time index t_i and channel coordinate c_i of the behavior node in the expression context matrix, map it to the relative offset vector in the structure domain coordinate space, and then use the standard structure domain orientation transformation matrix for coordinate encoding to obtain the structure domain positioning tensor ,in is the number of behavioral nodes, 3 is the dimension of the three-dimensional space, and K is the number of candidate structural domains. This tensor is used for spatial neighborhood structure matching and context coupling verification.

[0054] Step S233: performing context-coupled retrieval and neighborhood cross-validation on the structural domain positioning tensor to obtain a structural matching enhancement matrix; In this embodiment, the domain positioning tensor Execute context-coupled retrieval. Specifically, for any two behavior nodes , if they are mapped to the same domain If the distance between their positioning vectors in Euclidean space is less than 5, they constitute a candidate pair of spatial coupling. At the same time, cross-validation is performed in combination with the semantic similarity of expression (the threshold is set to 0.8). All node pairs that meet the conditions are constructed into an enhanced matrix , each element in the matrix Represents the structural matching enhancement score of the node pair, with a value range of [0,1]. This matrix provides a weight basis for the subsequent construction of the structural mapping relationship.

[0055] Step S234: constructing an initial behavior-structure mapping diagram based on the structure matching enhancement matrix; In this embodiment, based on the enhancement matrix M_enhance, an initial behavior-structure mapping diagram is constructed. , where: 1) V is the set of behavior nodes; 2) E is the set of nodes that pass The edge set determined by the score greater than 0.7; 3) W is the edge weight set, each edge The weight of The graph excludes structural domain nodes; only the structural correlation connections between behavioral nodes are retained. To ensure the graph's structural connectivity, nodes with low edge density are supplemented with virtual edges connecting them to the center of their most similar structural neighborhood. This initial graph is used for subsequent structural-functional clustering and chimeric modeling.

[0056] Step S235: performing chimeric metric modeling using functional clustering information and structural category similarity of domains in a standard protein domain database, and outputting a chimeric association tensor; In this embodiment, the domain function clusters (such as "kinase class", "transporter class") and their structural category labels (such as fold, Bucket structure, etc.), and the category similarity score is given for the domain mapped by each behavior node. For example, if two behavior nodes are mapped to PF00069 and PF07714 respectively, their functional clustering is consistent (both are Ser / Thr kinases), and the category similarity score can reach 0.9. Constructing a chimeric association tensor , where the dimensions are as follows: the first and second dimensions are behavior node indices; the first channel of the third dimension is the functional clustering consistency score, and the second channel is the category similarity score. The values ​​in this tensor serve as a quantitative indicator of the degree of chimeric compatibility between pairs of behavior nodes.

[0057] Step S236: Use the chimeric association tensor to assign chimeric strength labels to the edges in the initial behavior-structure mapping graph, and perform graph edge recalibration and connection optimization to construct a behavior-structure chimeric graph.

[0058] In this embodiment, the score value in the embedding association tensor T_embed is projected onto the edge weight attribute in the initial behavior-structure mapping graph G_init, and each edge e_{ij} is reassigned a comprehensive embedding strength score. Specifically, the following weighted combination is used: ;in , , Used to adjust the weights of the three chimera influencing factors. After recalibrating all edges, perform connection optimization: delete edges with chimera strength lower than 0.5; if the connectivity of a node is lower than 2, automatically connect it to the node with the highest chimera score. Finally, the optimized behavior-structure chimera graph is obtained. ,This graph has the ability of semantic-structural dual alignment, and can be used as the basic graph structure for regulatory ,pathway modeling and causal analysis.

[0059] Optionally, the protein function expression function reduction in step S3 is specifically as follows: Perform causal chain expansion on each expression path in the multi-scale causal structure graph, set the maximum path depth to 6, extract the function expression sequence in each expression path, and construct a path expression function set; In this embodiment, when processing the expression path in the multi-scale causal structure graph G_causal=(V,E,W), for each path, it is expanded in the direction of causal weight from the starting node, and recursively traverses up to 6 layers of connection depth. Each node contains an expression function label f_i, for example, f_i=tanh(Wx+b), which corresponds to the activation function expression. For any path , extract the function expression sequence on it , records the nesting levels between functions and the flow of input and output variables. The expression function sequences of all paths are uniformly encapsulated into a set , which serves as the basic function set for subsequent structural modeling. To control the extraction process of redundant paths, only paths with a sum of weights greater than 2.5 are retained.

[0060] The sparsity control parameter is set to 0.03, and the path expression function set is symbolically sparsely modeled to generate a function sparse structure tensor; In this embodiment, for the constructed expression function set F_all, the sparse control parameter is set to 0.03 to limit the symbol density in each expression function sequence. In the modeling process, each function expression is first symbolized into a standard infix expression (such as sigmoid(Wx+b) is converted to sigmoid,+,W,x,b). Sparse modeling aims to control the complexity of the structure. After converting the function representation into a symbol chain form, the symbols with a frequency below the threshold are removed. , an operator, variable, or subexpression, where is the total number of all symbol instances in F_all. The final output is a sparse structure tensor ,in is the number of paths, l is the maximum function length (taken as 16), is the symbol dimension (encoded as a 128-dimensional vector). This tensor reflects the functional symbolic structure pattern of each path and is used for subsequent similarity clustering.

[0061] Perform local sensitive hashing on the function expressions in the function sparse structure tensor, set the hash distance threshold to 0.15, and use the intra-cluster function similarity ≥ 85% as the aggregation condition to output the expression function aggregation tensor; In this embodiment, Each function expression vector in is considered as a high-dimensional sparse vector and mapped to a low-dimensional hash space through locality sensitive hashing (LSH). The hash distance threshold is set to 0.15, and four groups of parallel hash function families are used (20 hash functions in each group) to aggregate function expressions with similar symbolic structures. The aggregation criterion is: if the average Hamming distance of two function expressions in all hash function groups does not exceed 0.15 and the matching degree at the symbolic structure level exceeds 85%, they are considered to be members of the same cluster. The final output is the expression function aggregation tensor , where K is the number of function clustering clusters, is the number of functions per cluster (dynamically changing), is the length of the symbol vector (128 dimensions), and this structure is used to identify function clusters with homogeneity in expression functions.

[0062] Based on the gradient propagation flow graph and causal control weight graph of each expression function cluster in the expression function aggregation tensor, sub-functions with causal control weight ≥ 0.6 and average gradient weight ≥ 0.65 are set as core sub-functions, and low-weight functions in the expression function cluster are pruned to output a function pruned set; In this embodiment, based on the function clustering clusters in T_agg, the gradient propagation graph of each cluster is extracted. , where the edge weight is the local gradient value and the node is the sub-function identifier. At the same time, combined with the causal control graph The causal edge weights in the causal edge weights are set as follows to filter sub-functions: If a sub-function is The average gradient propagation weight in ≥ 0.65; and its The causal control weight in is ≥0.6; it is marked as a core sub-function. The low-weight sub-functions that do not meet the above conditions are removed from The function expressions in each cluster are removed, their positions and structural expressions are recorded, and a function pruning set F_pruned = {f_i} is constructed for subsequent topological stability modeling. This pruning process effectively controls the complexity of function expressions in each cluster, with an average pruning ratio of approximately 30%.

[0063] Perform structural transformation deduction on the sub-functions in the function clipping set and match them with the stable function patterns in the preset expression topology stability rule library. Set the matching tolerance threshold to 0.1 to obtain the expression topology stability map. In this embodiment, for each sub-function structure expression in F_pruned, a transformation deduction process is performed, including three types of structural operations: variable renaming, hierarchical expansion, and nested reconstruction. Convert to Then, the function templates in the expression topology stability rule library are matched in the form of a structural graph (nodes are operators, edges are data flows). The library contains about 200 typical stable function patterns, and the matching tolerance is set to 0.1 (i.e., structural similarity ≥ 0.9). The successfully matched function expression is added to the expression topology stability graph. , used to identify functional nodes and flow directions that have a stabilizing effect under multi-path conditions.

[0064] Based on the expression topology stability map and function clipping set, sub-functions are reorganized, connected and sorted to obtain the minimum combination unit of functional expression.

[0065] In this embodiment, based on the obtained G_stable and F_pruned, the minimum functional expression unit reorganization process is executed. First, all core sub-functions are divided into functional hierarchies (such as activation class, gating class, normalization class), and node connections are combined according to the topological connection direction to construct an expression subgraph. The nodes in each subgraph are rearranged according to the topological sorting to ensure the continuity of the data dependency chain, and the complexity of its combination expression (such as the total number of symbols, function depth, etc.) is calculated. After removing redundant and repeated paths, the minimum functional expression combination unit set is output. , each unit represents a functional combination subgraph with stable structure and clear expression function, which can be used to regulate the underlying interpretation of expression mechanisms and build a modular expression system.

[0066] Optionally, the structure-function transformation rule extraction in step S3 is specifically as follows: Taking the minimum combination unit set of functional expression as input, the symbolic function topology, path link structure and causal response factor are extracted to construct the expression-function structure map; In this embodiment, the minimum combination unit set is expressed by the generated function Based on this, we extract the symbolic function sequence from each unit u_i and mark its topological nested relationship in the graph structure. Each expression unit is represented as a directed structure graph. , where nodes V_ui represent symbolic functions (such as ReLU, Sigmoid, multiplication, etc.), and edges E_ui represent the input-output relationship between functions. All G_ui are uniformly encoded to construct the full graph , and fuse the causal response factor tensor (representing the response intensity of each expression unit to m functional indicators), the response factor is annotated on the edge of the structure diagram as a weight vector. Finally, an expression-function structure map is formed for subsequent topological feature induction and functional interpretation.

[0067] Perform topological rearrangement modeling on the expression-function structure map and extract the structural graph fingerprint of each expression unit in the topological rearrangement modeling result to construct a functional structure fingerprint set; In this embodiment, a topological rearrangement model is performed on G_expr-func, and the graph is reordered using the principle of minimizing structural entropy, giving priority to expression units with higher connection centrality and response factor values. During the rearrangement process, the nodes of each expression subgraph G_ui are renumbered, and its standardized topological arrangement vector is output. , to unify the logical execution order. Then, based on the topological adjacency matrix A_ui and the node function label sequence , constructing a structural graph fingerprint , where d is the function embedding dimension (set to 64), and each function label is embedded in the function semantic vector by looking up the table. The structural fingerprints of all expression units constitute the functional structure fingerprint set .

[0068] The structural graph fingerprints in the functional structure fingerprint set are compared with the structure-function mapping correspondence in the topological stability rule base for vector similarity, and symbolic function mapping and local subgraph identification are performed to obtain the structure-function symbol correspondence matrix. In this embodiment, the fingerprint F_ui in the functional structure fingerprint set is compared with the standard structure template in the expression topology stability rule library. Perform vector similarity comparison. Each standard template is composed of a function label sequence and a standard topology graph, which is represented as a triple ( ). Use cosine similarity to compare the similarity between F_ui and each S_i fingerprint vector, set the similarity threshold to 0.85, and map the successful matching results into a structure-function symbol correspondence matrix ,in Indicates that there is a symbolic correspondence between the expression unit u_i and the template structure S_j. This matrix can serve as a semantic guide for subsequent local graph recognition and expression reconstruction.

[0069] Based on the symbolic relationship in the structure-function symbol correspondence matrix, rule reduction and expression-driven recoding are performed to obtain a set of rule expression substructures; In this embodiment, based on the matching relationship in M_sym, a set of expression units with similar expression functions but redundant topological structures is extracted. For the expression units in this set, their common sub-expression structures (such as shared activation function sequences, input and output links, etc.) are extracted, and redundant paths, redundant connections and repeated sub-functions are eliminated to generate a set of regular expression sub-structures. Each substructure r_i is represented as a directed subgraph with a function expression label and a response value interval. For example, the substructure It can be constructed from three function nodes [Sigmoid→Multiply→Add], with edge weights attached with response weights in the range [0.4, 0.8]. Through structural normalization and symbol relabeling, the driver recoding of the expression units is completed, ensuring that the key functional paths are preserved while reducing the complexity of the graph structure.

[0070] Based on the rule expression substructure set, directed graph modeling and path scoring are performed to obtain a structure-function transformation path diagram; In this embodiment, all the rule expression substructures in R_expr are used to construct a structure-function transformation path diagram , where each node represents an expression substructure and the edge represents a feasible path from one structure to another. The edge weight is calculated using a comprehensive scoring method: based on the response weight increase ratio (ΔC); the expression complexity decrease (ΔN); and the topological similarity increase with the structural template (ΔS); the three are weighted into a path score value S Perform path traversal on G_path, record all reachable paths and their scores, and finally output the structure-function transformation path graph and its path score table , used to identify the optimal functional structure combination path.

[0071] Highly robust structural transformation paths are screened from the structure-function transformation path diagram to form a set of functional mapping units.

[0072] In this example, based on the G_path scoring table P_scores, we screened out paths with a score of no less than 0.85, and further evaluated their structural robustness changes in perturbation experiments (e.g., functional perturbation ±10%), setting the robustness score threshold to 0.8. For paths that meet the conditions, we extracted the minimum expression structure link between the start and end nodes and constructed it as a functional mapping unit. The final output is the functional mapping unit set , where each Mi represents a type of structural-functional bridging unit with high stability and high matching between topological changes and functional expression, which is used to guide subsequent expression reconstruction and functional modular integration.

[0073] Optionally, the unit configuration path validity and structural stability evaluation in step S4 is specifically as follows: Based on the preliminary scheduling scheme from the functional mapping unit set and protein scheduling simulation results, a directed graph structure is constructed to represent the dependency relationship and execution order between the functional mapping units in the scheduling path, and a unit scheduling graph is generated; In this embodiment, after obtaining the function mapping unit set Preliminary scheduling scheme based on protein scheduling simulation results Based on this, for each Mi, its upstream and downstream dependencies, function triggering boundaries, and time sequence are extracted, and a directed graph G_sched=(V_sched,E_sched) is constructed. Node V_sched represents the function mapping unit in the schedule, and edges E_sched represent the inter-function order and causal dependency chain. Each node Mi is assigned its execution time window [t_start, t_end] and required resource identifiers, and conditional trigger labels (such as those expressing gating signals and energy input) are attached to the edges. The resulting unit scheduling graph supports modeling the parallelism and dependencies of multipath scheduling structures, providing structural support for subsequent conflict detection.

[0074] Perform bottleneck and feedback closed-loop modeling on the unit scheduling graph, and extract path timing conflict features based on the bottleneck and feedback closed-loop modeling results to obtain a path conflict vector group; In this embodiment, a structural backtracking analysis is performed on the scheduling graph G_sched to identify all the ring subgraphs with feedback paths. , and aggregate modeling is performed on the cumulative execution time and state response interval of each node in each loop. If the execution delay difference in a closed-loop path is greater than the set threshold ΔT_max=0.12s, it is marked as a timing bottleneck feedback loop. At the same time, the re-entry points and shared resource nodes on all paths are collected to construct a path overlap matrix , count the number of timing conflicts between each pair of paths. Based on R_conflict and the bottleneck closed loop link, extract the path conflict vector group ,Each vector represents the key resource competition indexes such as the path’s conflict density, feedback intensity, timing conflict intensity, conflict frequency and conflict duration.

[0075] Model the connection tightness, energy level transition distribution and cooperative configuration pressure of the nodes in the unit scheduling graph to generate a structural energy distribution map; In this embodiment, the structural characteristics of each node in G_sched are analyzed to extract its topological connection tightness (based on the connectivity of the node and the local average clustering coefficient), energy level transition frequency (reflecting the number of intermediate energy state changes required from the input state to the excited state), and local synergistic pressure (based on the tensor distribution of the force matrix between adjacent units). The above three parameters are combined to form the node energy vector , and construct the node energy tensor on the entire graph ,Finally, a structural energy distribution map is drawn to reflect the ,load hotspots and activation centers of each region in the unit ,scheduling diagram during the execution process for use in ,execution efficiency evaluation.

[0076] Evaluate the execution efficiency of scheduling paths based on the unit scheduling graph and the path conflict vector group; In this embodiment, based on the combination of the scheduling graph G_sched and the path conflict vector group C_vec, the actual scheduling efficiency of each complete path P_i from the starting point to the end point is evaluated. The efficiency score uses the following formula group: benchmark execution time ; Actual delay time ; Path span length L_i; Unit execution efficiency Δ_conflict (path timing conflict delay) represents the cumulative execution delay caused by scheduling conflicts in resource usage, input / output dependencies, and other aspects of the key nodes in the path. It is calculated as follows: for each path P_i in the scheduling graph G_sched, assume there are m conflict points (including resource sharing conflicts, data dependency conflicts, etc.), and the delay corresponding to each conflict point is δ_j; the conflict points are obtained through the path conflict vector group C_vec; ; where δ_j ( ) can be estimated based on parameters such as resource queue waiting time, idle time after conflict conditions are triggered, or conflict mitigation scheduling delay specified by the system. ) represents the logic delay and state synchronization lag introduced to meet the feedback control consistency when there is a feedback structure path in the scheduling graph (for example, function A depends on B, and B indirectly depends on A). Its estimation is based on the following structure: closed-loop path ;The longest logical waiting time in the path τ_max ( );The difference between the synchronous trigger signals τ_sync ( ); Status refresh frequency f_update in the same cycle ( ); then we can calculate: If there are multiple feedback loops, the maximum or average value of all loops involved in path P_i is used for evaluation, or simulation and measurement are performed based on the modeling accuracy. Evaluation parameter records are attached to the output of each path, including the delay source node number, conflict vector dimension, feedback impact factor, etc. This forms a path efficiency report table. , providing a quantitative reference for path optimization and configuration screening.

[0077] The path performance score table is obtained by using the dual indicators of execution efficiency and stability of the scheduling path and the structure energy distribution map to quantify the scores. In this embodiment, the execution efficiency Eff_i and the average energy level value E_avg of the corresponding path in the structure energy tensor T_energy are input into the dual-index scoring function The full score is set to 1.0, and the performance score of each path is calculated and a score table is formed. High scores indicate high execution efficiency and low structural energy consumption, making them valuable for retention. Furthermore, for paths with scores below 0.6, the report records the location of structural bottlenecks and inefficiency factors, providing guidance for refined path optimization. The score sheet forms the input for subsequent hierarchical clustering.

[0078] Based on the path performance score table, hierarchical clustering and association mapping are performed on each scheduling path in the unit scheduling diagram, and evolution trend simulation is performed to obtain the path evolution decision diagram; In this embodiment, the path performance score table Score_path is clustered among multiple indicators, and the structural similarity metric Sim(P_i, P_j) is used to construct a path similarity matrix based on the score differences. The clustering threshold θ = 0.75 is used to divide the path into multiple functional behavior clusters. Score trend modeling is performed within each cluster to construct the path score evolution trajectory diagram. Each edge represents the feasibility of a natural transition between paths in the direction of score improvement (e.g., local reconstruction from an inefficient path to an efficient path). The final output is the path evolution decision graph G_decision, where each subgraph represents a possible path evolution strategy that can be used to support the upgrade and improvement of scheduling paths or the reduction of redundancies.

[0079] The paths with scores ≥ 0.85 and the paths with average structural energy levels of path nodes ≤ 0.25 in the path evolution decision graph are subjected to redundancy pruning and structural verification to screen out a set of stable configuration paths. The stable configuration path sets are dynamically spliced ​​to form a configuration decision graph.

[0080] In this embodiment, a path set P_stable with all path scores s_i≥0.85 and the average structural energy level of nodes in its path E_avg≤0.25 is screened out from G_decision. Node redundancy detection is performed on each path, including: repeated calculation functions; redundant nested loops; empty operation nodes. The above redundancy is trimmed, and the structural matching interface and stability rule library are called to verify the logical integrity of the trimmed path. Under the premise of ensuring functional continuity and stability, the screened path set is dynamically spliced. The splicing principle is based on input and output compatibility, structural continuity and the average response time difference does not exceed 0.03s. The splicing result forms a configuration decision diagram G_config=(V_conf,E_conf), which provides an efficient and stable graphics-level execution model for function deployment.

[0081] It is particularly important to evaluate the execution efficiency of the scheduling path as follows: Extract the path length, dependency depth, node connection mode, and number of feedback loops in each scheduling path from the unit scheduling graph, and associate the timing conflict information in the path conflict vector group to generate a path execution feature set; In this embodiment, based on the unit scheduling graph G_sched=(V,E), all valid scheduling path sets are identified from it. ,For each path P_i, the following four types of structural features are extracted in turn: 1) Path length: record the total number of nodes in the path, set as , the unit is the number of hops; 2) Dependency depth: Based on the hierarchical structure of dependency edges in the path, extract the length of the longest directed dependency chain 3) Node connection mode: Count the distribution of connection types between nodes in the path (e.g., “sequential,” “parallel,” “backflow”), and mark the connection dimension vector of each node; 4) Number of feedback loops: Extract the number of feedback substructures based on whether there are self-loops or cross-layer loops on the path ; Then, a mapping relationship is established between the path P_i and the timing conflict record C_vec(i) in the path conflict vector group, from which the conflict type corresponding to the path is extracted (startup conflicts, resource contention conflicts, etc.), conflict density (Number of conflict events per unit time) and duration of conflict , and finally generate a path execution feature set with the following structure: .

[0082] Calculate the path execution complexity based on the dependency hierarchy, conflict density and feedback structure distribution in the path execution feature set; In this embodiment, the path execution feature set in the previous step is used as input to construct the execution complexity index for each path. The specific implementation method is: based on the dependency hierarchy Setting dependency coefficient weights =0.35, according to the number of feedback loops Assigning closed-loop factor weights =0.25, and the tension tensor is determined by the connection method. The structural perturbation value (such as the maximum connectivity difference) sets the connection complexity weight =0.4. The execution complexity score of this path is calculated as follows: ; Among them, the tension tensor The distribution matrix representing the connection weights of each node in the path. The more dispersed the value, the greater the connection instability during the scheduling process.

[0083] Based on the path execution complexity, a multi-index fusion model is constructed by combining the temporal conflict intensity, conflict frequency and conflict duration characteristics in the path conflict vector group. In this embodiment, the timing conflict characteristic values ​​in the path conflict vector group, such as the conflict intensity , conflict frequency and the average duration of conflicts , together with the execution complexity index C_exec(i) to build a multi-index fusion model. Set the model input weight ratio as: complexity item =0.4, conflict intensity =0.3, conflict frequency =0.2, conflict duration =0.1, then the model fusion score is: in, The unit is normalized conflict tension, is the number of conflict events per unit time, is the average duration in seconds.

[0084] A multi-index fusion model is used to perform weighted scoring, and the weighted scoring results are transformed into nonlinear mapping to generate the scheduling path execution efficiency.

[0085] In this embodiment, the fusion score As input, we transform it into path execution efficiency by constructing a set of nonlinear mapping formulas based on empirical functions. . Assume that the nonlinear mapping adopts a logarithmic decay scoring function: ;in is the fusion value with the highest score in the current; the conversion result The range is limited to the interval [0,1], and the closer the value is to 1, the higher the scheduling efficiency. This mapping function can effectively compress the weight of high-complexity task paths, making their weight lower in the scoring system, which is conducive to scheduling priority sorting.

[0086] It should be noted that the English phrases contained in the parameters in the embodiments of this application are named according to the data type or data meaning, and this application does not impose any restrictions on the naming of these parameters.

[0087] Optionally, this specification further provides a biological information processing system applied to synthetic biology, for executing the biological information processing method applied to synthetic biology as described above, the biological information processing system applied to synthetic biology comprising: A context decoding module is used to acquire real-time proteomics data and perform polymorphic context decoding on the real-time proteomics data to obtain a protein expression context matrix; The interaction modeling module is used to deconstruct and reconstruct the cross-domain graph structure of the behavioral features in the protein expression context matrix to obtain a reconstructed protein matrix; based on the reconstructed protein matrix, the biological response pathway-protein domain interaction is analyzed to obtain a multi-scale causal structure map; Functional transformation analysis module, used to reduce protein functional expression functions of path nodes in multi-scale causal structure maps to obtain the minimum functional expression module combination unit; based on the minimum functional expression module combination unit, structure-function transformation rules are extracted to obtain a functional mapping unit set; The protein scheduling simulation module is used to perform protein scheduling simulation using the functional mapping unit set as input, and evaluate the effectiveness of the unit configuration path and the structural stability based on the protein scheduling simulation results, thereby dynamically assembling and forming a configuration decision diagram; The virtual response simulation module is used to perform multi-scenario virtual response simulation on the path in the configuration decision diagram, and to verify and score the path adaptability of the multi-scenario virtual simulation results to obtain a path adaptability feedback table; The variability screening module is used to perform multidimensional clustering and variability screening on the path adaptability feedback table to obtain the optimal expression configuration set.

[0088] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.

[0089] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.

Claims

1. A biological information processing method applied to synthetic biology, characterized in that: The following steps are involved: Step S1: Acquire real-time proteomics data and perform polymorphic context decoding on the real-time proteomics data to obtain a protein expression context matrix; Step S2: Deconstruct and reconstruct the behavioral features in the protein expression context matrix across domains to obtain a reconstructed protein matrix; analyze the biological response pathway-protein domain interactions based on the reconstructed protein matrix to obtain a multi-scale causal structure map; Step S3: performing protein function expression function reduction on the path nodes in the multi-scale causal structure map to obtain the minimum module combination unit of function expression; performing structure-function transformation rule extraction based on the minimum module combination unit of function expression to obtain a function mapping unit set; Step S4: Using the functional mapping unit set as input, perform protein scheduling simulation, and evaluate the effectiveness of the unit configuration path and the structural stability based on the protein scheduling simulation results, thereby dynamically assembling to form a configuration decision diagram; Step S5: Perform multi-scenario virtual response simulation on the paths in the configuration decision diagram, and perform path adaptability verification and scoring on the multi-scenario virtual simulation results to obtain a path adaptability feedback table; Step S6: Perform multidimensional clustering and variability screening on the path adaptability feedback table to obtain the optimal expression configuration set.

2. The biological information processing method applied to synthetic biology according to claim 1, characterized in that: Step S1 is specifically as follows: Step S11: real-time proteomics data are collected through a mass spectrometer, a high-throughput liquid chromatography platform, and a cell expression database, and the real-time proteomics data are time-series aligned and data decoupled to obtain a raw protein expression data set; Step S12: performing high-dimensional semantic feature mapping conversion on the original protein expression dataset to construct a protein expression semantic tensor; Step S13: performing polymorphic context nested modeling based on the protein expression semantic tensor and extracting the nested relationship of environmental regulatory factors, thereby constructing a context nested representation model; Step S14: analyzing the expression behavior vector sequence of the protein in various states based on the context nested representation model, thereby forming an expression state dynamic matrix; Step S15: Fuse the expression state dynamic matrix and perform low-rank semantic compression to construct the protein expression context matrix.

3. The biological information processing method applied to synthetic biology according to claim 2, characterized in that: Step S13 is specifically as follows: Step S131: extracting state identification factors and experimental metadata from the protein expression semantic tensor to obtain a factor candidate set; Step S132: constructing an initial nested path graph based on the factor candidate set and performing context relevance evaluation to obtain an initial nested path graph; Step S133: dividing the initial nested path graph into context-coupled subgraphs to construct a set of factor-coupled subgraphs; Step S134: Map each path pair in the factor coupling subgraph back to the behavior channel in the protein expression semantic tensor, thereby obtaining a context nested tensor structure; Step S135: Construct a context nested representation model based on the context nested tensor structure.

4. The biological information processing method applied to synthetic biology according to claim 2, characterized in that: Step S14 is specifically as follows: Step S141: parsing the context nested representation model to extract the state master control factor sequence and response channel structure, thereby constructing a state channel index table; Step S142: extracting multi-channel behavior data from the protein expression semantic tensor according to the state channel index table, constructing a protein expression behavior feature space, and performing state expression vector aggregation modeling to obtain an expression behavior time series group; Step S143: Modeling expression vector change patterns for each expression behavior time series group at different preset time resolutions, and marking the regulation time windows to form a nested expression trajectory set; Step S144: constructing a state vector graph model based on the nested expression trajectory set to dynamically deduce the causal regulatory relationship between different expression states and output a state regulation response graph; Step S145: The time evolution trajectory of the nested expression trajectory set is integrated with the causal regulation weight in the state regulation response graph, and the temporal weight normalization reconstruction is performed to obtain the expression state dynamic matrix.

5. The biological information processing method applied to synthetic biology according to claim 1, characterized in that: Step S2 is specifically as follows: Step S21: performing nested extraction of expression semantic nodes on the protein expression context matrix to construct a cross-domain feature relationship graph; Step S22: performing graph sparsification reconstruction on the cross-domain feature graph of the expression behavior to obtain a structure compression representation graph; Step S23: connecting to a standard protein domain database, mapping the behavior nodes in the structure compression representation graph to the domain coordinate system in the standard protein domain database, and performing chimeric modeling to generate a behavior-structure chimeric graph; Step S24: performing path-aware extended modeling on the behavior-structure chimeric graph to generate a response path graph group; Step S25: performing multi-scale graph convolution feature aggregation on the response path graph group to extract the causal response weight matrix; Step S26: Based on the causal significance and path density threshold in the causal response weight matrix, weighted reconstruction and screening are performed on each path in the response path graph group to construct a multi-scale causal structure map.

6. The biological information processing method applied to synthetic biology according to claim 5, characterized in that: Step S23 is specifically as follows: Step S231: connecting to a standard protein domain database, and semantically aligning the semantic features of each behavior node in the structure compression representation graph with the domain description vector of the standard protein domain database to obtain a semantic index mapping table; Step S232: performing domain space projection on each behavior node in the structure compression representation graph based on the semantic index mapping table, and performing coordinate encoding in combination with the domain coordinate system of the standard protein domain database to obtain a domain positioning tensor; Step S233: performing context-coupled retrieval and neighborhood cross-validation on the structural domain positioning tensor to obtain a structural matching enhancement matrix; Step S234: constructing an initial behavior-structure mapping diagram based on the structure matching enhancement matrix; Step S235: performing chimeric metric modeling using functional clustering information and structural category similarity of domains in a standard protein domain database, and outputting a chimeric association tensor; Step S236: Use the chimeric association tensor to assign chimeric strength labels to the edges in the initial behavior-structure mapping graph, and perform graph edge recalibration and connection optimization to construct a behavior-structure chimeric graph.

7. The biological information processing method applied to synthetic biology according to claim 1, characterized in that: The protein function expression function reduction in step S3 is specifically as follows: Perform causal chain expansion on each expression path in the multi-scale causal structure graph, set the maximum path depth to 6, extract the function expression sequence in each expression path, and construct a path expression function set; The sparsity control parameter is set to 0.03, and the path expression function set is symbolically sparsely modeled to generate a function sparse structure tensor; Perform local sensitive hashing on the function expressions in the function sparse structure tensor, set the hash distance threshold to 0.15, and use the intra-cluster function similarity ≥ 85% as the aggregation condition to output the expression function aggregation tensor; Based on the gradient propagation flow graph and causal control weight graph of each expression function cluster in the expression function aggregation tensor, sub-functions with causal control weight ≥ 0.6 and average gradient weight ≥ 0.65 are set as core sub-functions, and low-weight functions in the expression function cluster are pruned to output a function pruned set; Perform structural transformation deduction on the sub-functions in the function clipping set and match them with the stable function patterns in the preset expression topology stability rule library. Set the matching tolerance threshold to 0.1 to obtain the expression topology stability map. Based on the expression topology stability map and function clipping set, sub-functions are reorganized, connected and sorted to obtain the minimum combination unit of functional expression.

8. The biological information processing method applied to synthetic biology according to claim 1, characterized in that: The structure-function transformation rule extraction in step S3 is specifically as follows: Taking the minimum combination unit set of functional expression as input, the symbolic function topology, path link structure and causal response factor are extracted to construct the expression-function structure map; Perform topological rearrangement modeling on the expression-function structure map and extract the structural graph fingerprint of each expression unit in the topological rearrangement modeling result to construct a functional structure fingerprint set; The structural graph fingerprints in the functional structure fingerprint set are compared with the structure-function mapping correspondence in the topological stability rule base for vector similarity, and symbolic function mapping and local subgraph identification are performed to obtain the structure-function symbol correspondence matrix. Based on the symbolic relationship in the structure-function symbol correspondence matrix, rule reduction and expression-driven recoding are performed to obtain a set of rule expression substructures; Based on the rule expression substructure set, directed graph modeling and path scoring are performed to obtain a structure-function transformation path diagram; Highly robust structural transformation paths are screened from the structure-function transformation path diagram to form a set of functional mapping units.

9. The biological information processing method applied to synthetic biology according to claim 1, characterized in that: The unit configuration path validity and structural stability evaluation in step S4 are specifically as follows: Based on the preliminary scheduling scheme from the functional mapping unit set and protein scheduling simulation results, a directed graph structure is constructed to represent the dependency relationship and execution order between the functional mapping units in the scheduling path, and a unit scheduling graph is generated; Perform bottleneck and feedback closed-loop modeling on the unit scheduling graph, and extract path timing conflict features based on the bottleneck and feedback closed-loop modeling results to obtain a path conflict vector group; Model the connection tightness, energy level transition distribution and cooperative configuration pressure of the nodes in the unit scheduling graph to generate a structural energy distribution map; Evaluate the execution efficiency of scheduling paths based on the unit scheduling graph and the path conflict vector group; The path performance score table is obtained by using the dual indicators of execution efficiency and stability of the scheduling path and the structure energy distribution map to quantify the scores. Based on the path performance score table, hierarchical clustering and association mapping are performed on each scheduling path in the unit scheduling diagram, and evolution trend simulation is performed to obtain the path evolution decision diagram; The paths with scores ≥ 0.85 and the paths with average structural energy levels of path nodes ≤ 0.25 in the path evolution decision graph are subjected to redundancy pruning and structural verification to screen out a set of stable configuration paths. The stable configuration path sets are dynamically spliced ​​to form a configuration decision graph.

10. A biological information processing system applied to synthetic biology, characterized in that: For executing the biological information processing method applied to synthetic biology according to claim 1, the biological information processing system applied to synthetic biology comprises: A context decoding module is used to acquire real-time proteomics data and perform polymorphic context decoding on the real-time proteomics data to obtain a protein expression context matrix; The interaction modeling module is used to deconstruct and reconstruct the cross-domain graph structure of the behavioral features in the protein expression context matrix to obtain a reconstructed protein matrix; based on the reconstructed protein matrix, the biological response pathway-protein domain interaction is analyzed to obtain a multi-scale causal structure map; Functional transformation analysis module, used to reduce protein functional expression functions of path nodes in multi-scale causal structure maps to obtain the minimum functional expression module combination unit; based on the minimum functional expression module combination unit, structure-function transformation rules are extracted to obtain a functional mapping unit set; The protein scheduling simulation module is used to perform protein scheduling simulation using the functional mapping unit set as input, and evaluate the effectiveness of the unit configuration path and the structural stability based on the protein scheduling simulation results, thereby dynamically assembling and forming a configuration decision diagram; The virtual response simulation module is used to perform multi-scenario virtual response simulation on the path in the configuration decision diagram, and to verify and score the path adaptability of the multi-scenario virtual simulation results to obtain a path adaptability feedback table; The variability screening module is used to perform multidimensional clustering and variability screening on the path adaptability feedback table to obtain the optimal expression configuration set.

Citation Information

Patent Citations

  • Gene design method and platform based on AI improvement

    CN118212972A

  • Gene fusions and gene variants associated with cancer

    CN118910253A

  • Mining method and system for synthetic biological functional element and storage medium

    CN119673283A

  • Product and methods useful for modulating and evaluating immune responses

    US20220397568A1

  • Artificial intelligence-simulation based science

    US20240370608A1

Cited By

  • Human cell viability data analysis system based on big data

    CN121789830A

  • Personal password credibility evaluation method fusing biological characteristics and rule engine

    CN122133121A