Decoding method for predicting response information on the basis of response data of cell to external dynamic signal

Through machine learning methods, formalize the cell response problem to dynamic signals, establish a dynamic model and train a classifier model, which solves the problem that traditional methods are difficult to deeply interpret the dynamic characteristics of cell responses, and realizes the revelation of quantitative analysis and decision-making rules.

WO2025123414A1PCT designated stage expired Publication Date: 2025-06-19SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Patent Information

Application Number
PCT/CN2023/141314
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-11
Filing Date
2023-12-23
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Traditional methods are difficult to deeply interpret the dynamic characteristics of cells' response to environmental signals, cannot fully reflect the true nature and details of environmental information signals, and cannot realize real-time single-cell monitoring.

Method used

Using machine learning methods, dynamic signal recognition problems are formalized into classification tasks. By establishing dynamic models, generating simulation training data, extracting feature values ​​and training classifier models, the cell response information to external dynamic signals is predicted.

Benefits of technology

Quantitative analysis and in-depth interpretation of cell dynamic response data is realized, revealing the decision rules of cells in the process of identifying and responding to environmental signals, and can more accurately predict and judge cell response behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023141314_19062025_PF_FP_ABST
    Figure CN2023141314_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a decoding method for predicting response information on the basis of response data of a cell to an external dynamic signal. The decoding method of the present invention formalizes a dynamic signal recognition problem into a classification task, uses simulation data samples to train a model, and then uses the trained model to predict and determine dynamic response data in a real experiment, so as to acquire response information of a dynamic signal. The present invention uses a machine learning method to predict and determine dynamic response data of cells, so as to provide deeper insights for researchers, helping the researchers understand how cells extract information from environments and make adaptive responses.
Need to check novelty before this filing date? Find Prior Art

Description

A decoding method for predicting response information based on cell response data to external dynamic signals Technical Field

[0001] The present invention belongs to the field of biotechnology, and in particular relates to a decoding method for predicting response information based on cell response data to external dynamic signals. Background Art

[0002] Microbial survival requires rapid responses to environmental changes. Changes in environmental information are highly complex across time and space, and the impacts of environmental factors are often interdependent and additive. Microbial cells must efficiently and accurately perceive various signals from the external environment and adjust their life activities accordingly to adapt to environmental changes or utilize resources within it.

[0003] Any changes in the intracellular and extracellular environments can affect vital life processes such as cell survival, proliferation, and physiological functions. For example, changes in temperature and nutrients can directly or indirectly affect various physical and chemical reactions within cells, thereby affecting cell structure and function. Furthermore, the infectious capacity of pathogenic microorganisms can vary under different environmental conditions, which can also have a significant impact on host cell survival.

[0004] Therefore, microbial cells must recognize the importance and urgency of sensing changes in the external environment so they can promptly implement appropriate physiological regulatory strategies. This is the significance of expressing external information within cells. Microorganisms achieve this crucial function by transmitting and conveying signals within the cell through second messenger molecules.

[0005] Second messenger molecules, crucial for cell signaling, are transported across the cytoplasm, mediating intercellular and intracellular signaling. They convey environmental information and regulate gene expression and biological responses. Dynamic changes in the concentration and structure of second messengers reflect the real-time response of cells to external stimuli and are the primary means by which cells express environmental information.

[0006] However, because biochemical reactions within cells are random events, individual cells vary in gene expression levels, metabolic processes, and other aspects. This results in individual cells exhibiting varying second messenger concentration response curves under the same environmental stimulus. However, biologists have traditionally used qualitative analytical methods to study cellular responses to environmental signals. This approach lacks quantitative metrics and is unable to provide a deep understanding of the decision-making rules.

[0007] In addition, simply observing the concentration or structural level of the second messenger in the cell at a single time point without considering the dynamic characteristics of the cell response is insufficient to fully reflect the true nature and details of the environmental information signal, and cannot fully reflect the characteristics of the environmental signal. Therefore, it is necessary to comprehensively observe and analyze the evolution of the entire second messenger concentration response curve in order to better distinguish the essential characteristics of environmental information, which is especially important for studying how cells learn and perceive external environmental information. Tracking the entire response process using traditional methods is complex and tedious, and real-time single-cell monitoring is almost impossible using conventional methods.

[0008] Summary of the Invention

[0009] To address these issues, the present invention utilizes machine learning methods to formalize the dynamic signal recognition problem as a classification task, thereby predicting and determining the dynamic response data of cells. Specifically, this study formalizes the dynamic signal recognition problem as a classification task and trains a model using simulated data samples. This trained model is then used to predict and determine dynamic response data from real experiments, thereby obtaining dynamic signal response information.

[0010] In one aspect, the present invention provides a decoding method for predicting response information based on cell response data to external dynamic signals, comprising the following steps:

[0011] S1) establishing a dynamic model of the target system's dynamic response;

[0012] S2) obtaining simulated training data using a random simulation algorithm;

[0013] S3) segmenting the obtained training data;

[0014] S4) extracting feature values ​​from the segmented data;

[0015] S5) using a machine learning method to perform training using the feature values ​​obtained in step S4) to obtain a classifier model;

[0016] S6) Using the classifier model obtained in step S5), the corresponding response information is predicted based on the actual response data of the cells to the external dynamic signal.

[0017] Furthermore, establishing a dynamic response kinetic model in S1) is a process of describing the behavior of the target system as an ODE.

[0018] Furthermore, S2) simulates training data using a stochastic simulation algorithm. This algorithm can introduce randomness into the established kinetic model to simulate the effects of biochemical reactions or other random phenomena within the system. This algorithm introduces noise into the simulation process and generates noisy kinetic response data, more realistically reflecting the actual system behavior. Simulating the light on and off states at different time intervals generates simulated training data.

[0019] Furthermore, S3) segments the acquired training data. For subsequent processing and analysis, the generated noisy dynamic response data requires appropriate segmentation. Segmenting the data can be done based on different time periods to better understand the system's cyclical behavior and to perform data alignment to ensure effective feature extraction and machine learning model application.

[0020] Furthermore, the segmentation is performed by dividing each period of turning on and off the light source into small segments. Since different periods of turning on and off the light source are used, each period of turning on and off the light source is aligned after segmentation. Furthermore, S4) feature extraction is performed on the segmented data. Feature extraction is performed on the segmented data in order to extract meaningful information from the data.

[0021] Furthermore, eigenvalues ​​include classification features such as standard deviation, variance, skewness, correlation coefficient, energy, etc. For specific systems, other specific feature extraction methods can also be considered to better capture the behavior and characteristics of the target system.

[0022] Furthermore, in step S5), a machine learning method is used to train the feature values ​​obtained in step S4) to obtain a classifier model. The feature values ​​extracted in step S4) are used as input, and the corresponding light source on and off is dynamically encoded as output. This is used as training data to train the machine learning model to obtain a classifier model.

[0023] Furthermore, by using neural network simulation and training the above training data, a neural network classifier model capable of classification is obtained.

[0024] Furthermore, during the training process, techniques such as cross-training are further used to evaluate and select the best model.

[0025] Furthermore, for different systems, suitable machine learning algorithms and model architectures can be selected, such as decision trees, support vector machines, neural networks, naive Bayes, random forests, and logistic regression to achieve accurate model training.

[0026] Furthermore, S6) the classifier model obtained in step S5) predicts corresponding response information based on the actual response data of the cells to the external dynamic signal.

[0027] Furthermore, the target system is cAMP, factors affecting cAMP concentration, and downstream effector proteins activated by cAMP; further, the factors affecting cAMP concentration are bPAC and CpdA, and the activation of bPAC is affected by external dynamic signals; the downstream effector protein can be an active protein or a fluorescent protein.

[0028] Another aspect of the present invention provides a decoding device for predicting response information based on cell response data to external dynamic signals, the device comprising: a modeling module, a module for obtaining simulated training data, a training data segmentation module, a feature value extraction module, a machine learning training module, and a decoding module;

[0029] The modeling module is used to establish a dynamic model of the target system dynamic response;

[0030] The module for obtaining simulated training data is used for obtaining simulated training data by a random simulation algorithm;

[0031] The training data segmentation module is used to segment the obtained training data;

[0032] The feature value extraction module is used to extract feature values ​​from the segmented data;

[0033] The machine learning training module is used to adopt a machine learning method to train the feature values ​​obtained by the feature value extraction module to obtain a classifier model;

[0034] The decoding module is used to predict corresponding response information based on the real response data of cells to external dynamic signals using the classifier model obtained by the machine learning training module.

[0035] Yet another aspect of the present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method described above is implemented.

[0036] In another aspect, the present invention provides a computer storage medium having a computer program stored thereon, wherein the computer program implements the above method when executed by a processor. Beneficial effects

[0037] This invention uses machine learning as a tool, with broad applications in big data classification and recognition. Through machine learning, researchers can gain a more comprehensive understanding of the dynamic responses of cells to environmental signals. This overcomes the existing problem of observing cells at a single time point, which fails to capture the dynamic characteristics of cellular responses and fully reflects the true nature and details of environmental information signals.

[0038] The method of the present invention can provide quantitative indicators and deeper interpretations, helping to reveal the decision-making rules of cells in the process of recognizing and responding to environmental signals.

[0039] In addition, machine learning can process large amounts of data to more accurately predict and judge the response behavior of cells.

[0040] Using machine learning methods to predict and judge the dynamic response data of cells can provide researchers with deeper insights, helping them understand how cells extract information from the environment and respond adaptively.

[0041] By transferring traditional human judgment to machine model judgment, judgment decoding of large amounts of data can be performed. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1: Flowchart of the present invention.

[0043] Figure 2: Modeling diagram.

[0044] Figure 3: Random simulation data with a 300s interval period.

[0045] Figure 4: Schematic diagram of machine learning model training and its truth table.

[0046] Figure 5: Experimental data decoding. DETAILED DESCRIPTION

[0047] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below, but it should not be understood as limiting the scope of implementation of the present invention.

[0048] As shown in FIG1 , the present invention provides a decoding method for predicting response information based on cell response data to external dynamic signals.

[0049] In this solution, the decoding method includes the following steps:

[0050] S1) establishing a dynamic model of dynamic response;

[0051] S2) obtaining simulated training data using a random simulation algorithm;

[0052] S3) segmenting the obtained training data;

[0053] S4) extracting feature values ​​from the segmented data;

[0054] S5) using a machine learning method to perform training using the feature values ​​obtained in step S4) to obtain a classifier model;

[0055] S6) Using the classifier model obtained in step S5), the corresponding response information is predicted based on the actual response data of the cells to the external dynamic signal.

[0056] As shown in Figures 1-5, the present invention first simulates the dynamic response of a second messenger system when receiving control signals. It then uses the Stochastic Simulation Algorithm (SSA) to simulate the system's dynamics in the presence of noise. Next, by simulating random modulation signals under various conditions, the simulated dynamic process data is segmented and feature values ​​extracted. This is followed by machine learning model training to generate a classifier model. Finally, the trained classifier is applied to experimental data on the dynamic response.

[0057] In some specific embodiments, the second messenger system in the cell is studied and analyzed as the response regulatory system, and it can be expected that the analysis method of the present invention can also be applied to other cell regulatory systems that can respond to external dynamic information stimulation.

[0058] S1) (See Figure 2) Building a dynamic response kinetic model involves describing the behavior of the target system as an ODE. For research scenarios where the second messenger system is a response-regulating system, the ODE is derived based on the second messenger system's intracellular chemical reactions and system characteristics.

[0059] In the physical model of the second messenger system, light source stimulation is used as the input signal, and whether the fluorescent protein in the cell emits light is used as the output signal. The core of the intracellular response regulation system is cAMP, and the concentration of cAMP is affected by bPAC and CpdA. The synthesis is affected by bPAC protein, while the decomposition is affected by CpdA. bPAC is obtained in the cell through bPAC gene expression; whether the output signal fluorescent protein emits light is affected by cAMP.

[0060] Therefore, based on the above physical model, a chemical reaction network model of this physical model can be constructed. The on and off of the light source stimulation signal can be represented by 1 and 0, respectively, and the chemical reaction can be decomposed into r1) photons bind to bPAC to form bPAC-pho; r2) bPAC-pho converts to the activated state bPAC*; r3) the activated state bPAC* regulates cAMP and then converts to the inactivated state bPAC; the inactivated state bPAC participates in the r1) photon-bPAC binding reaction; r4) the activated state bPAC* regulates cAMP; r5) cAMP binds to CpdA to form CpdA-cAMP; r6) CpdA-cAMP decomposes cAMP and releases CpdA; r7) cAMP affects the luminescence of the fluorescent protein Rflamp.

[0061] Construct a kinetic model of the dynamic response based on the reactions of the chemical reaction network.

[0062] Chemical reaction equation:

[0063] Differential equations:

[0064] S2) Simulated training data is obtained using a stochastic simulation algorithm. This algorithm can introduce randomness into the established kinetic model to simulate the effects of biochemical reactions or other random phenomena within the system. This algorithm introduces noise into the simulation process and generates noisy kinetic response data, more realistically reflecting the actual system behavior. Simulated lighting is switched on and off at different time intervals to generate simulated training data.

[0065] Figure 3 shows 3000 seconds of noisy kinetic response data generated by a random simulation algorithm at 300-second intervals. Figure 2 shows the cAMP concentration at 1 and 0, i.e., after light stimulation and after the light is turned off. It can be seen that the cAMP concentration fluctuates over the 3000-second period, i.e., during the five cycles of on and off the light. If only the cAMP concentration at a single moment is used as the indicator of the cell's response to light regulation, it is difficult to accurately determine the cell's regulatory response. For example, the cAMP value at 2200 seconds, when the light is off, is higher than the 200-second period when the light is on.

[0066] S3) The obtained training data is segmented. In order to carry out subsequent processing and analysis, the generated kinetic response data with noise needs to be appropriately segmented. The segmented data can be aligned according to different time periods in order to better understand the periodic behavior of the system to ensure the effective application of subsequent feature extraction and machine learning models. As shown in Figure 4, each cycle of turning on and off the light source is divided into small segments. Since different cycles of turning on and off the light source are used, each cycle of turning on and off the light source is aligned after segmentation. If the data source is experimental data rather than simulation data, it is necessary to preprocess according to the experimental results, that is, convert the data of the cAMP probe obtained in the experiment into the concentration of cAMP molecules.

[0067] S4) Extracting features from the segmented data. Extracting features from the segmented data is done to extract meaningful information from the data.

[0068] In some specific implementation schemes, the characteristic values ​​include classification features such as standard deviation, variance, skewness, correlation coefficient, energy, etc. For specific systems, other specific feature extraction methods can also be considered to better capture the behavior and characteristics of the target system.

[0069] S5) Using machine learning to train the feature values ​​obtained in step S4) to obtain a classifier model. The feature values ​​extracted in step S4) are used as input, and the corresponding light source on and off is dynamically encoded as output. This is used as training data to train the machine learning model to obtain a classifier model.

[0070] In some specific implementation schemes, a neural network classifier model capable of classification is obtained by training the above training data using neural network simulation.

[0071] In some specific implementation schemes, during the training process, cross-training and other techniques are further used to evaluate and select the best model.

[0072] In some specific implementation schemes, suitable machine learning algorithms and model architectures can be selected for different systems, such as decision trees, support vector machines, neural networks, naive Bayes, random forests, and logistic regression to achieve accurate model training.

[0073] S6) Using the classifier model obtained in step S5, the classifier model predicts corresponding response information based on the actual cell response data to external dynamic signals. As shown in FIG5 , in some specific embodiments, the following method is used for experimentation and verification: the bacteria are regulated to turn a light source on and off, and simulated signals 1 and 0 are input; simultaneously, data regarding the dynamic response of cAMP is obtained using a cAMP probe expressed in the bacteria. The cAMP dynamic response data is segmented as described in S3) and feature value extracted as described in S4). The extracted feature values ​​are used as input and analyzed using the classifier model obtained in S5), resulting in predicted simulated signals 1 and 0 corresponding to the on and off states of the light source.

[0074] By comparing the predicted results with the actual results, the decoding capability and accuracy of the system of the present invention can be determined. As shown in Figure 5, the truth table shows that the decoding capability and accuracy of the present invention are high. This indicates that the method of the present invention can be used to decode and obtain information about the cell's response to an environmental signal, in this example, the on and off of a light source.

[0075] In some specific implementation schemes, the solution of the present invention is not only applicable to the second messenger system, but also can be used in other response regulation systems using similar methods to decode the response information.

[0076] Specifically, for other response control systems, when establishing a kinetic model of dynamic response in S1), a kinetic model can be constructed based on the dynamic response characteristics and behavior, and the physical or chemical reactions of the dynamic response process to simulate the dynamic response of the system under different conditions.

[0077] In some specific embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0078] In some specific implementation schemes, a computer-readable storage medium is also provided, and its implementation principle and technical effects are similar to those of the above-mentioned method implementation scheme, which will not be repeated here.

[0079] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be read by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital versatile disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).

[0080] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A decoding method for predicting response information based on cell response data to external dynamic signals, characterized in that, It includes the following steps: S1) Establish a kinetic model for the dynamic response of the target system; S2) Obtain simulated training data using a stochastic simulation algorithm; S3) Split the obtained training data; S4) Extract eigenvalue from the split data; S5) Use machine learning to train with the eigenvalues obtained in step S4) to obtain a classifier model; S6) Use the classifier model obtained in step S5) to predict the corresponding response information for the true response data of the cell to the external dynamic signal; Preferably, the target system is cAMP, the factors affecting the cAMP concentration, and the downstream effector proteins activated by cAMP; Further, the factors affecting the cAMP concentration are bPAC and CpdA, and the activation of bPAC is affected by the external dynamic signal; The downstream effector protein can be an active protein or a fluorescent protein.

2. The decoding method according to claim 1, characterized in that, The establishment of the kinetic model for dynamic response in S1) is a process of describing the behavior of the target system as a differential equation (ODE).

3. The decoding method according to claim 1, characterized in that, In S2), simulated training data is obtained using a stochastic simulation algorithm. Noise is introduced during the simulation process by the stochastic simulation algorithm, and kinetic response data with noise is generated to simulate the influence of biochemical reactions or other random phenomena within the system.

4. The decoding method according to claim 1, characterized in that, S3) Split the obtained training data; The data is split and aligned according to different time periods to ensure the effective application of subsequent feature extraction and machine learning models; Preferably, the data is split by dividing each period of turning on and off the light source into small segments. Due to the use of different periods of turning on and off the light source, each period of turning on and off the light source is aligned after splitting.

5. The decoding method according to claim 1, characterized in that, S4) Extract eigenvalue from the split data. Extracting eigenvalue from the split data is to extract meaningful information from the data; Preferably, the eigenvalue is a classification feature, and the classification feature is selected from standard deviation, variance, skewness, correlation coefficient, energy.

6. The decoding method according to claim 1, characterized in that, S5) Use machine learning to train with the eigenvalues obtained in step S4) to obtain a classifier model. Use the eigenvalues extracted in S4) as input, and use the corresponding dynamic encoding of turning on and off the light source as output, and use this as training data to train the machine learning model to obtain a classifier model; Preferably, it is simulated by a neural network. Through the training of the above training data, a neural network classifier model capable of classification is obtained; Preferably, during the training process, techniques such as cross-training are further used to evaluate and select the best model; Preferably, for different systems, suitable machine learning algorithms and model architectures can be selected to achieve accurate model training, and the machine learning algorithms and model architectures are selected from decision tree, support vector machine, neural network, naive Bayes, random forest, logistic regression.

7. The decoding method according to claim 1, characterized in that, S6) Use the classifier model obtained in step S5) to predict the corresponding response information for the true response data of the cell to the external dynamic signal.

8. A decoding device for predicting response information based on cell response data to external dynamic signals, characterized in that, The device includes: a modeling module, a module for obtaining simulated training data, a training data splitting module, an eigenvalue extraction module, a machine learning training module, and a decoding module; The modeling module is used to establish a dynamic response kinetic model of the target system; The module for obtaining simulated training data is used to obtain simulated training data by a stochastic simulation algorithm; The training data segmentation module is used to segment the obtained training data; The eigenvalue extraction module is used to extract eigenvalues from the segmented data; The machine learning training module is used to train by machine learning using the eigenvalues obtained by the eigenvalue extraction module to obtain a classifier model; The decoding module is used to predict corresponding response information for the true response data of cells to external dynamic signals by using the classifier model obtained by the machine learning training module.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1-7.

10. A computer storage medium, characterized in that,A computer program is stored thereon, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Deep learning based long-chain non-coding RNA subcellular position prediction algorithm

    CN107577924A

  • Cell discrimination model construction method and device, electronic equipment and storage medium

    CN116486403A

  • Cell type prediction model training method, cell type prediction method and device

    CN117037917A

  • Cell detection model training method, device and system and electronic device

    CN117058484A

  • Machine learning models for cell localization and classification learned using repel coding

    US20230186659A1

Cited By

  • Multi-factor environment simulation system for high-speed train passenger comfort research

    CN120656354A