Disease marker structural dynamics characteristic determination method

Through omics technology screening and deep learning models combined with atomic force microscopy and isothermal titration calorimeter analysis, the problem of accurate characterization of the dynamic characteristics of disease markers structures is solved, and accurate description of the dynamic changes of marker networks is achieved.

CN120279999APending Publication Date: 2025-07-08QINGDAO RAISECARE BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510423414.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to accurately characterize the structural dynamic characteristics of disease markers, especially the dynamic changes of markers in complex physiological environments, and the existing network analysis methods have not fully considered the impact of marker structural characteristics on their functions.

Method used

Omics technology is used to screen disease-related markers, establish a structural functional database, and build a dynamic action relationship network through principal component decomposition calculation and graph neural network model. Time series analysis is performed by combining atomic force microscopy and isothermal titration calorimeter, and input into deep learning models for training to achieve accurate characterization of the dynamic characteristics of the structure of disease markers.

Benefits of technology

The accurate characterization of the dynamic characteristics of the structure of disease markers is achieved, the analysis accuracy is improved, the dynamic changing characteristics of the marker network are revealed, and the accuracy of the disease prediction model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279999A_ABST
    Figure CN120279999A_ABST
Patent Text Reader

Abstract

The invention provides a method for determining structural dynamic characteristics of a disease marker, which belongs to the technical field of disease markers, and comprises the following steps of: screening the disease marker through an omics technology, establishing a structural function database, carrying out standardization processing on data, and then carrying out principal component decomposition to obtain stable and variable structural domain components; and constructing a marker action relation graph by using a protein interaction network, calculating network topology parameters and a shortest path, and inputting a path matrix into a graph neural network to obtain a marker representation vector. Meanwhile, carrying out morphology analysis by adopting an atomic force microscope to obtain time sequence characteristic data, and measuring binding parameters of the marker and the ligand by utilizing an isothermal titration calorimetry. And finally, inputting multi-dimensional data such as the stable structural domain component, the variable structural domain component, the marker characterization vector, the morphology feature data and the combination parameter matrix into a deep learning model, carrying out training optimization through cross validation, and establishing a disease prediction model, thereby determining the structural dynamic features of the disease-related markers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of disease markers, and more specifically, relates to a method for determining the structural dynamics characteristics of disease markers. Background Art

[0002] Disease markers are important molecular indicators for diagnosing and predicting the development of diseases. With the rapid development of biomedical technologies, disease marker screening methods based on proteomics, metabolomics, and transcriptomics have emerged continuously, providing important bases for the early diagnosis and prognosis evaluation of diseases. In the prior art, the analysis of disease markers mainly focuses on the study of static structural characteristics and expression levels, and common analysis methods include techniques such as mass spectrometry analysis, immunohistochemical detection, and Western blot. Currently, the structural analysis of disease markers mainly uses means such as X-ray crystallography and nuclear magnetic resonance spectroscopy to obtain their three-dimensional structure information. Although these methods can provide static structure information of the markers, they cannot reflect the dynamic change characteristics of the markers in the physiological environment. During the development of diseases, the structures of the markers often change dynamically, and these changes are closely related to the progression of diseases. In the prior art, molecular dynamics simulation is mainly used to study the structural dynamics characteristics of the markers, but due to the limitations of computing resources and the lack of simulation accuracy, it is difficult to accurately describe the structural change rules of the markers in a complex physiological environment. In terms of the analysis of the interaction network of markers, the prior art mainly relies on database mining and statistical analysis methods to construct a static protein interaction network. This method ignores the dynamic characteristics of the interactions between the markers and is difficult to reflect the dynamic changes of the marker network during the development of diseases. In addition, existing network analysis methods often simplify the markers as network nodes and fail to fully consider the influence of the structural characteristics of the markers on their functions. In terms of the study of marker functions, traditional methods mainly rely on biochemical experiments and cell experiments to speculate on their functions by observing the expression changes of the markers. This method is time-consuming, costly, and difficult to reveal the molecular mechanism of the functional changes of the markers. Although various high-throughput detection technologies have been developed in recent years, these technologies mainly focus on the changes in the expression levels of the markers and lack in-depth research on their structural dynamics characteristics.

[0003] That is to say, the prior art often only focuses on the structural characteristics of disease markers and fails to accurately characterize the structural dynamics characteristics of disease markers. Summary of the Invention

[0004] In view of this, the present invention provides a method for determining the structural dynamics characteristics of disease markers, which can solve the technical problem that the prior art fails to accurately characterize the structural dynamics characteristics of disease markers.

[0005] The present invention is implemented as follows:

[0006] The present invention provides a method for determining the structural dynamics characteristics of disease markers, including: screening disease-related markers using omics technology, obtaining structural and functional data of disease-related markers, establishing a structural and functional database, performing standardization and global normalization processing on the structural and functional data to obtain a threshold of the structural and functional data; performing principal component decomposition calculation on the structural and functional data of disease-related markers to obtain stable domain components and variable domain components; constructing a network of interaction relationships of disease-related markers; constructing a marker action pathway matrix using the shortest path algorithm; inputting the marker action pathway matrix into a graph neural network model to obtain a marker representation vector; performing time-series topography analysis using an atomic force microscope; measuring binding parameters using an isothermal titration calorimeter to construct a binding parameter matrix; inputting the above data into a deep learning model; training the model using a cross-validation method to determine the structural dynamics characteristics of disease-related markers.

[0007] Specifically, it includes the following steps:

[0008] S10. Screen disease-related markers using omics technology, obtain the structural and functional data of the disease-related markers, establish a structural and functional database, perform standardization and global normalization processing on the structural and functional data to obtain a threshold of the structural and functional data;

[0009] S20. Perform principal component decomposition calculation on the structural and functional data of the disease-related markers to obtain stable domain components and variable domain components, and calculate the fluctuation thresholds of the stable domain components and the variable domain components;

[0010] S30. Use a protein interaction network to construct a network of interaction relationships of the disease-related markers, use the names of the disease-related markers as network nodes, use the interactions between the disease-related markers as network edges, and calculate the topological parameters of the protein interaction network;

[0011] S40. Based on the network of interaction relationships, use the shortest path algorithm to calculate the action pathways between the network nodes, construct a marker action pathway matrix, and obtain the key path weights;

[0012] S50. Input the marker action pathway matrix into a graph neural network model, perform quantitative characterization on the disease-related markers, obtain a marker representation vector, and generate the node embedding vector;

[0013] S60. Perform time-series topography analysis on the disease-related markers using an atomic force microscope to obtain topography feature data, calculate the surface roughness change rate of the topography feature data and the time-series change data of the topography feature, and establish a correspondence relationship between the topography feature data and the marker representation vector;

[0014] S70. Use an isothermal titration calorimeter to measure the binding parameters of the disease-related biomarker and the ligand, construct a binding parameter matrix, calculate the stability coefficient of the binding parameter matrix, and measure the dynamic change values of the binding constant and the binding enthalpy change;

[0015] S80. Input the stable domain component, the variable domain component, the biomarker characterization vector, the morphological feature data, and the binding parameter matrix into a deep learning model;

[0016] S90. Use a cross-validation method to train and optimize the deep learning model, establish a disease prediction model, and determine the structural dynamics characteristics of the disease-related biomarker.

[0017] Among them, the use of omics technology to screen disease-related biomarkers specifically refers to using proteomics, metabolomics, and transcriptomics technologies to perform high-throughput detection and identification of the disease-related biomarkers in biological samples, qualitatively and quantitatively analyze the disease-related biomarkers through a mass spectrometer, and establish the molecular weight data, abundance data, modification site data, and expression level data of the disease-related biomarkers;

[0018] The structural and functional database specifically is a multi-dimensional relational database composed of a primary structure information data table, a secondary structure information data table, a functional domain data table, and a modification site data table of the disease-related biomarker, including the amino acid sequence, molecular weight, isoelectric point, hydrophobicity index, secondary structure conformation, functional domain, post-translational modification site, and structural stability parameters of the disease-related biomarker;

[0019] Perform standardization and global normalization processing on the structural and functional data. Specifically, perform normalization processing on the molecular weight data using the maximum-minimum normalization method, perform normalization processing on the abundance data using the logarithmic normalization method, perform normalization processing on the expression level data using the median normalization method, and perform binary normalization processing on the modification site data to generate the threshold of the structural and functional data;

[0020] Perform principal component decomposition calculation on the structural and functional data of the disease-related biomarker. Specifically, use the principal component analysis method to perform dimensionality reduction processing on the structural and functional data, extract the main domain features, use the singular value decomposition method to decompose the structural and functional data into the stable domain component and the variable domain component, perform time series analysis on the stable domain component and the variable domain component, and calculate the fluctuation threshold, where the stable domain component represents the conservative structural features of the disease-related biomarker, and the variable domain component represents the changing structural features of the disease-related biomarker;

[0021] Construct a relationship graph network of the disease-related markers using a protein interaction network. Specifically, establish the connection relationships between the disease-related markers based on the physical interactions, functional correlations, and expression quantity correlations among the disease-related markers. Use the disease-related markers as the network nodes and the interaction intensity between the disease-related markers as the weight value of the network edges to construct a weighted directed graph network. Calculate the degree distribution, clustering coefficient, and centrality measure of the network to obtain the topological parameters of the protein interaction network.

[0022] Use the shortest path algorithm to calculate the action pathways between the network nodes and construct the marker action pathway matrix. Specifically, use the Dijkstra algorithm to calculate the shortest action path between any two network nodes, construct a distance matrix based on the weight values of the shortest action paths, and transform the distance matrix into the standardized marker action pathway matrix to obtain the key path weights.

[0023] The graph neural network model is specifically a deep learning architecture based on a graph convolutional neural network, and its structure includes a graph convolutional layer, a pooling layer, a fully connected layer, and an output layer. Among them, the graph convolutional layer is used to extract the local features of the network nodes, the pooling layer is used to reduce the feature dimension, the fully connected layer is used for feature fusion, and the output layer is used to generate the marker representation vector and the node embedding vector.

[0024] Quantitatively represent the disease-related markers. Specifically, use the graph neural network model to extract features and reduce the dimension of the network nodes, transform the high-dimensional structural and functional features into low-dimensional vector representations, and generate the marker representation vectors.

[0025] Use an atomic force microscope to perform time-series topography analysis on the disease-related markers. Specifically, use the tapping mode to perform scanning imaging of the disease-related marker samples at continuous time points under room temperature conditions, obtain the three-dimensional topography images of the disease-related markers, measure the length, width, height, and surface roughness parameters of the disease-related markers, calculate the surface roughness change rate, and record the time-series change data of the topography features.

[0026] Establish the correspondence between the topography feature data and the marker representation vectors. Specifically, use a multi-layer perceptron to establish a mapping function from the topography feature data to the marker representation vectors, use the topography feature data as the input layer, use the marker representation vectors as the output layer, and optimize the parameters of the mapping function through the backpropagation algorithm.

[0027] The binding parameters of the disease-related biomarker and the ligand are determined using an isothermal titration calorimeter. Specifically, the ligand solution is added dropwise to the disease-related biomarker solution under constant temperature conditions, and the heat released or absorbed during the binding process is measured. The binding constant, binding enthalpy change, and binding entropy change of the disease-related biomarker and the ligand are calculated by fitting the heat curve. The dynamic change values of the binding constant and the binding enthalpy change are obtained by recording the measurement results at different time points, and the binding parameter matrix is constructed. The stability coefficient of the binding parameter matrix is calculated;

[0028] The deep learning model is specifically a deep neural network based on the encoder network and decoder network structures. The structure is to set a mathematical model module between the encoder network and the decoder network. The mathematical model module includes a structural dynamics feature extraction function, a multi-scale feature fusion function, and a dynamic feature mapping function;

[0029] The structural dynamics feature extraction function is used to extract key dynamic feature representations from the biomarker structural dynamics data. The inputs include the fluctuation thresholds of the stable domain component and the variable domain component, the surface roughness change rate in the morphological feature data, the dynamic change values of the binding constant and the binding enthalpy change in the binding parameter matrix, and the output is the feature vector of the biomarker structural dynamics;

[0030] The multi-scale feature fusion function is used to integrate and fuse multi-scale feature data. The inputs include the topological parameters of the protein interaction network, the key path weights, the node embedding vectors, and the structural dynamics feature vectors, and the output is the comprehensive representation of the multi-scale features;

[0031] The dynamic feature mapping function is used to map the comprehensive representation to the dynamic prediction space. The inputs include the structural function data threshold, the comprehensive representation of the multi-scale features, the binding parameter stability coefficient, and the time series change data of the morphological features, and the output is the time series prediction result of the biomarker structural dynamics features;

[0032] The specific structure of the deep learning model is composed of a 5-layer encoder network, the mathematical model module, and a 5-layer decoder network connected in series. Each layer of the network adopts a batch normalization and residual connection structure;

[0033] The 5-layer encoder network adopts a hierarchical fully connected neural network structure. The input dimension of the first layer is the original feature dimension, and the output dimensions of the first layer to the fifth layer decrease successively at a ratio of 30%. Each layer of the network uses a leaky rectified linear unit activation function with a slope of 0.02. A batch normalization layer and a dropout layer are added after each layer of the network. The momentum parameter of the batch normalization layer is set to 0.9, and the dropout rate of the dropout layer is set to 0.3;

[0034] The 5-layer decoder network adopts a fully connected neural network structure symmetric to the encoder network, where the output dimensions of the first layer to the fifth layer increase successively at a ratio of 40%, and the hyperbolic tangent function is used as the activation function between each layer of the network. A batch normalization layer and a dropout layer are added after each layer of the network. The momentum parameter of the batch normalization layer is set to 0.95, and the dropout rate of the dropout layer is set to 0.25. The output dimension of the last layer is the same as the original feature dimension.

[0035] Compared with the prior art, a method for determining the structural dynamics characteristics of disease markers provided by the present invention realizes the accurate characterization of the structural dynamics characteristics of disease markers by innovatively integrating multi-omics data, structural function information, and network analysis methods. The method has the following remarkable technical effects:

[0036] First of all, the present invention establishes a multi-dimensional structural function database. Through standardization and normalization processing, the problem of integrating different types of data is effectively solved. Through principal component decomposition calculation, the structural characteristics of the marker are successfully decomposed into a stable domain and a variable domain, enabling the dynamic change characteristics of the marker structure to be accurately captured.

[0037] Secondly, the present invention innovatively combines the protein-protein interaction network with the graph neural network to construct a dynamic marker interaction relationship network. Through the shortest path algorithm and critical path weight analysis, not only the functional connections between markers are revealed, but also the accurate description of the dynamic changes of the network is realized. The application of the graph neural network model enables the effective extraction of the network characteristics of the marker and their conversion into low-dimensional vector representations.

[0038] Thirdly, the present invention uses atomic force microscopy for time-series topography analysis and combines it with isothermal titration calorimetry to obtain direct experimental evidence of the structural changes of the marker. By establishing the correspondence between the topography feature data and the marker characterization vector, the effective correlation between the microscopic structural changes and the macroscopic functional characteristics is realized. This multi-scale characterization method greatly improves the analysis accuracy of the structural dynamics characteristics of the marker.

[0039] In addition, the deep learning model designed by the present invention realizes the effective integration of multi-source heterogeneous data through the encoder-decoder structure and special mathematical model modules. The design of the structural dynamics feature extraction function, multi-scale feature fusion function, and dynamic feature mapping function enables the model to accurately capture the dynamic change characteristics of the marker. The application of the batch normalization and residual connection structures significantly improves the training stability and prediction accuracy of the model.

[0040] In summary, the present invention solves the technical problem that the prior art fails to accurately characterize the structural dynamics characteristics of disease markers. Brief Description of the Drawings

[0041] Figure 1 It is a flowchart of the method provided by the present invention;

[0042] Figure 2 It is a scatter plot of principal component analysis of the EGFR domain characteristics in Example 2;

[0043] Figure 3 It is a visualization graph of the EGFR protein interaction network in Example 2;

[0044] Figure 4 It is a surface plot of the dynamic change of the three-dimensional structure of EGFR in Example 2. Detailed Embodiments

[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0046] As Figure 1 shown, it is a flowchart of a method for determining the structural dynamics characteristics of a disease biomarker provided by the present invention. This method includes the following steps:

[0047] S10. Use omics technology to screen disease-related biomarkers, obtain the structural and functional data of disease-related biomarkers, establish a structural and functional database, perform standardization and global normalization processing on the structural and functional data, and obtain the structural and functional data threshold;

[0048] S20. Perform principal component decomposition calculation on the structural and functional data of disease-related biomarkers, obtain the stable domain component and the variable domain component, and calculate the fluctuation threshold of the stable domain component and the variable domain component;

[0049] S30. Use the protein interaction network to construct a network of interaction relationships of disease-related biomarkers. Use the names of disease-related biomarkers as network nodes and the interactions between disease-related biomarkers as network edges, and calculate the topological parameters of the protein interaction network;

[0050] S40. Based on the interaction relationship graph network, use the shortest path algorithm to calculate the interaction pathways between network nodes, construct a biomarker interaction pathway matrix, and obtain the key path weights;

[0051] S50. Input the biomarker interaction pathway matrix into a graph neural network model, perform quantitative characterization on disease-related biomarkers, obtain biomarker characterization vectors, and generate node embedding vectors;

[0052] S60. Use an atomic force microscope to perform time-series topography analysis on disease-related markers, obtain topographical feature data, calculate the surface roughness change rate of the topographical feature data and the time-series change data of the topographical features, and establish the corresponding relationship between the topographical feature data and the marker characterization vector;

[0053] S70. Use an isothermal titration calorimeter to measure the binding parameters between disease-related markers and ligands, construct a binding parameter matrix, calculate the stability coefficient of the binding parameter matrix, and measure the dynamic change values of the binding constant and the binding enthalpy change;

[0054] S80. Input the stable domain component, variable domain component, marker characterization vector, topographical feature data, and binding parameter matrix into a deep learning model;

[0055] S90. Use the cross-validation method to train and optimize the deep learning model, establish a disease prediction model, and determine the structural dynamics characteristics of disease-related markers.

[0056] The following is a detailed description of the specific implementation manners of the above steps:

[0057] The specific implementation manner of step S10 is to obtain and process multi-omics data of disease-related markers. First, use proteomics technology to screen for markers, use a liquid chromatography tandem mass spectrometer to separate and detect biological samples, and obtain the molecular weight data and abundance data of the markers; at the same time, use metabolomics technology to analyze the metabolites of the markers, and use a gas chromatography mass spectrometer to measure the content changes of metabolites; then use transcriptomics technology to analyze the gene expression of the markers, and use a high-throughput sequencing platform to detect the gene transcription level. Then establish a structure-function database, including a primary structure information data table of the marker, recording the amino acid sequence, molecular weight, and isoelectric point; a secondary structure information data table, recording conformational information such as α-helix and β-sheet; a functional domain data table, recording conserved domains and functional sites; a modification site data table, recording the positions of post-translational modifications. Standardize the obtained data. For the molecular weight data, use the maximum-minimum normalization method to normalize the data to between 0 and 1; for the abundance data, use the logarithmic normalization method to eliminate the magnitude difference of the data; for the expression level data, use the median normalization method to reduce the influence of outliers; for the modification site data, use the binary normalization method to convert the modification state to 0 or 1. Finally, determine the data thresholds. The molecular weight data threshold is set to 0.5, the abundance data threshold is set to 2-fold change, the expression level data threshold is set to 3-fold standard deviation, and the modification site data threshold is set to 0.8. The purpose of this step is to obtain the multi-dimensional data characteristics of the markers and provide a data basis for subsequent analysis.

[0058] The specific implementation of step S20 is to perform dimensionality reduction and decomposition analysis on the marker structure - function data. First, the principal component analysis method is used to reduce the dimensionality of the data, calculate the eigenvalues and eigenvectors, and select the principal components with a cumulative contribution rate reaching 85% as the main features. Then, the singular value decomposition method is used to decompose the data matrix into a left - singular vector matrix, a singular value diagonal matrix, and a right - singular vector matrix. The main components are screened according to the singular value size, and a stable domain component matrix is constructed. The variable domain component matrix is obtained by subtracting the stable domain component matrix from the original data matrix. Time - series analysis is performed on the stable domain components and variable domain components, and the mean and standard deviation of the time series are calculated. The fluctuation threshold is set as the mean plus 2 times the standard deviation. The purpose of this step is to separate the conservative structure features and variable structure features of the markers and determine the structure fluctuation range.

[0059] The specific implementation of step S30 is to construct a protein - protein interaction network and analyze the network characteristics. First, connection relationships are established based on the physical interactions, functional correlations, and expression correlations between markers. The markers are used as network nodes, and the interaction strength is used as the weight of the edges. Then, the topological parameters of the network are calculated, including degree distribution, clustering coefficient, and centrality measure. The degree distribution represents the distribution of the number of node connections, the clustering coefficient represents the connection density between node neighbors, and the centrality measure represents the importance of nodes in the network. Finally, the modular structure of the network is analyzed to identify functionally related protein complexes. This step uses complex network analysis methods, and the purpose is to reveal the interaction relationships and network structure characteristics between markers.

[0060] The specific implementation of step S40 is to calculate the shortest action paths between network nodes. First, the Dijkstra algorithm is used to calculate the shortest path between any two nodes, considering the weight values of the edges, and find the path with the minimum sum of weights. Then, a distance matrix is constructed according to the weight values of the shortest paths, and the shortest distances between node pairs are recorded. The distance matrix is standardized to obtain a standardized marker action path matrix. Finally, the weights of the critical paths are calculated, and the weight threshold is set as 0.7 to screen important action paths. The purpose of this step is to determine the signal transmission paths between markers and identify critical action paths.

[0061] The specific implementation of step S50 is to use a graph neural network model for feature extraction. First, a graph convolutional neural network model is constructed, which includes multiple graph convolutional layers, pooling layers, fully connected layers, and output layers. The graph convolutional layer uses Chebyshev polynomial approximation for feature extraction, and the output dimension of each layer is set to 0.7 times the input dimension. The pooling layer adopts max pooling operation to compress the feature dimension. The fully connected layer uses the hyperbolic tangent function as the activation function, and the dropout rate is set to 0.3. The output layer generates a biomarker representation vector and a node embedding vector. Then, the biomarker action pathway matrix is input into the model, and a low-dimensional feature representation is obtained through forward propagation calculation. The purpose of this step is to obtain a vectorized representation of the biomarker for subsequent analysis.

[0062] The specific implementation of step S60 is to perform biomarker morphology analysis. First, an atomic force microscope is used to observe the biomarker sample at room temperature, and continuous scanning is carried out in tapping mode to obtain a three-dimensional morphology image. Then, parameters such as the length, width, and height of the biomarker are measured, and the surface roughness is calculated. The morphological features at different time points are analyzed, the change rate of the surface roughness is calculated, and the time series change data is recorded. Finally, a mapping relationship between the morphological feature data and the biomarker representation vector is established, and a multi-layer perceptron is used for feature mapping. The number of neurons in the hidden layer is set to 1.5 times the input dimension. The purpose of this step is to obtain the microscopic structural features of the biomarker and its dynamic change law.

[0063] The specific implementation of step S70 is to determine the interaction parameters between the biomarker and the ligand. First, an isothermal titration calorimeter is used for determination at a constant temperature. The ligand solution is gradually added to the biomarker solution, and the heat change is recorded. Then, the binding constant, binding enthalpy change, and binding entropy change are calculated through heat curve fitting. Measurements are carried out at different time points to obtain the dynamic change data of the binding parameters. A binding parameter matrix is constructed, and the stability coefficient is calculated. The binding constant threshold is set to 105, and the binding enthalpy change threshold is set to 10 kcal / mol. The purpose of this step is to determine the binding characteristics and stability between the biomarker and the ligand.

[0064] The specific implementation of step S80 is to construct a deep learning model for feature integration. First, a deep neural network based on an encoder-decoder structure is designed, and a mathematical model module is added between the two networks. The encoder network adopts a 5-layer fully connected structure, and the output dimension of each layer decreases by 30%. The leaky rectified linear unit is used as the activation function. The decoder network adopts a symmetric structure, and the output dimension of each layer increases by 40%. The hyperbolic tangent function is used as the activation function. The mathematical model module includes a structural dynamics feature extraction function, a multi-scale feature fusion function, and a dynamic feature mapping function. Then, the stable domain component, variable domain component, biomarker representation vector, morphological feature data, and binding parameter matrix are input into the model. The purpose of this step is to integrate multi-dimensional feature data and construct a prediction model.

[0065] The specific implementation of step S90 is to train and optimize the deep learning model. First, the data set is divided into a training set, a validation set, and a test set according to a ratio of 7:2:1; the 5-fold cross-validation method is used to evaluate the model performance, and the prediction accuracy, sensitivity, and specificity are calculated; then the model parameters are optimized, the initial learning rate is set to 0.001, the batch size is set to 32, and the number of training epochs is set to 100; the early stopping strategy is used to prevent overfitting, and the training is stopped when the validation set loss does not decrease for 5 consecutive epochs; finally, a disease prediction model is established to determine the structural dynamics characteristics of the biomarker. The purpose of this step is to optimize the model parameters and improve the prediction accuracy.

[0066] The following will describe in detail the calculation process involved in the present invention.

[0067] 1. Standardization and normalization of structural and functional data:

[0068] Formula for maximum-minimum normalization of molecular weight data:

[0069] In the formula, X is the original molecular weight data; X min is the minimum value of the data; X max is the maximum value of the data; X norm is the normalized data value.

[0070] Formula for logarithmic normalization of abundance data:

[0071] In the formula, Y is the original abundance data; Y ref is the reference abundance value; Y norm is the normalized data value.

[0072] Formula for median normalization of expression level data:

[0073] In the formula, Z is the original expression level data; Z median is the median; MAD is the median absolute deviation; Z norm is the normalized data value.

[0074] Formula for binary normalization of modification site data:

[0075]

[0076] In the formula, W is the original modification site data; W threshold is the threshold; W norm is the normalized data value.

[0077] 2. Principal component decomposition calculation:

[0078] Principal component analysis dimensionality reduction formula: T = XP;

[0079] Where, X is the original data matrix; P is the eigenvector matrix; T is the principal component score matrix.

[0080] Singular value decomposition formula: X = U∑V T +E;

[0081] Where, X is the original data matrix; U is the left singular vector matrix; ∑ is the singular value diagonal matrix; V is the right singular vector matrix; E is the error matrix.

[0082] Calculation formula for the stable domain component:

[0083] Where, U k is the first k left singular vectors; ∑ k is the first k singular values; V k is the first k right singular vectors; S is the stable domain component matrix.

[0084] Calculation formula for the variable domain component: D = X - S;

[0085] Where, X is the original data matrix; S is the stable domain component matrix; D is the variable domain component matrix.

[0086] Calculation formula for the fluctuation threshold: θ = μ + ασ;

[0087] Where, μ is the mean of the time series; σ is the standard deviation; α is the adjustment factor; θ is the fluctuation threshold.

[0088] 3. Calculation of protein-protein interaction network topological parameters:

[0089] Calculation formula for degree distribution:

[0090] Where, n k is the number of nodes with degree k; N is the total number of nodes; P(k) is the degree distribution probability.

[0091] Calculation formula for clustering coefficient:

[0092] Where, L i is the actual number of connections between the neighbors of node i; k i is the degree of node i; C i is the clustering coefficient of node i.

[0093] Calculation formula for centrality measure:

[0094] Where, σ st is the number of shortest paths between nodes s and t; σst (v) is the number of shortest paths passing through node v; BC(v) is the betweenness centrality of node v.

[0095] 4. Shortest path algorithm calculation:

[0096] Dijkstra's algorithm distance calculation formula: d(i,j) = min{d(i,k) + w(k,j)};

[0097] In the formula, d(i,j) is the shortest distance from node i to j; w(k,j) is the weight from node k to j; k is the intermediate node.

[0098] Normalized distance matrix calculation formula:

[0099] In the formula, D is the original distance matrix; D min is the minimum distance; D max is the maximum distance; D norm is the normalized distance matrix.

[0100] 5. Graph neural network node representation calculation:

[0101] Graph convolution layer calculation formula:

[0102] In the formula, H (l) is the node feature matrix of the l-th layer; is the adjacency matrix with self-loops added; is the degree matrix; W (l) is the weight matrix; σ is the activation function.

[0103] Pooling layer calculation formula:

[0104] In the formula, is the feature after pooling of node i; h j is the feature of the neighbor nodes; N(i) is the set of neighbors of node i.

[0105] Node embedding vector calculation formula:

[0106] In the formula, is the node feature of the last layer; f MLP is the multi-layer perceptron function; e i is the node embedding vector.

[0107] 6. Topography feature analysis calculation:

[0108] Surface roughness calculation formula:

[0109] In the formula, y iis the height measurement value; is the average height; n is the number of measurement points; R a is the arithmetic mean roughness.

[0110] Calculation formula for surface roughness change rate:

[0111] In the formula, R t is the roughness at time t; R0 is the initial roughness; Δt is the time interval; ΔR is the roughness change rate.

[0112] 7. Calculation of combined parameters:

[0113] Calculation formula for combined constant:

[0114] In the formula, [PL] is the complex concentration; [P] is the protein concentration; [L] is the ligand concentration; K a is the binding constant.

[0115] Calculation formula for binding enthalpy change:

[0116] In the formula, Q is the heat change; n is the number of binding sites; [P] t is the total protein concentration; ΔH is the binding enthalpy change.

[0117] Calculation formula for stability coefficient:

[0118] In the formula, σ i is the standard deviation of parameter i; μ i is the mean of parameter i; N is the number of parameters; S is the stability coefficient.

[0119] 8. Calculation of deep learning model:

[0120] Function for extracting structural dynamics characteristics:

[0121]

[0122] In the formula, F dyn is the m-dimensional dynamic feature vector; is the fluctuation threshold of the i-th stable domain component, obtained by the aforementioned singular value decomposition; is the fluctuation threshold of the i-th variable domain component, obtained by the aforementioned singular value decomposition; R i is the surface roughness change rate at the i-th time point, calculated by atomic force microscopy measurement; is the binding constant at the i-th time point, determined by isothermal titration calorimetry; ΔH i is the binding enthalpy change at the i-th time point, determined by isothermal titration calorimetry; w iis the weight coefficient of each feature, in the range [0, 1], optimized by cross-validation; α i , β i , γ i , δ i , η i , λ i are trainable parameters, with initial values randomly sampled from [-1, 1]; ε is a random noise term, following a normal distribution with a mean of 0 and a variance of 0.1; n is the length of the time series.

[0123] Multi-scale feature fusion function:

[0124]

[0125] In the formula: F multi is a p-dimensional multi-scale feature vector; is the degree distribution feature at the k-th scale, and the calculation formula is: where d i is the degree of node i; is the clustering coefficient feature at the k-th scale, and the calculation formula is: where C i is the clustering coefficient of node i; is the centrality feature at the k-th scale, and the calculation formula is: where BC(i) is the betweenness centrality of node i; is the path length feature at the k-th scale, and the calculation formula is: where d ij is the shortest path length from node i to j; is the path weight feature at the k-th scale, and the calculation formula is: where w ij is the sum of weights on the shortest path from node i to j; is the local embedding feature at the k-th scale, obtained by attention aggregation on the k-th order neighbors of the node; is the global embedding feature at the k-th scale, compressed by the graph pooling layer; is the dynamic feature at the k-th scale; v k is the weight coefficient at the k-th scale, in the range [0, 1], optimized by cross-validation; φ k , ψ k , ω k , τ k are trainable parameters, with initial values randomly sampled from [-1, 1]; ξ is a random noise term, following a normal distribution with a mean of 0 and a variance of 0.1; K is the number of scales, usually taking 3 - 5; N is the total number of nodes.

[0126] Dynamic feature mapping function:

[0127]

[0128] where: Y pred is a q-dimensional predicted result vector; is the structural function data threshold vector at time t, obtained through the aforementioned normalization process; is the maximum value of the threshold vector; is the multi-scale feature vector at time t; is the maximum value of the multi-scale feature vector; S t is the binding parameter stability coefficient at time t; S max is the maximum value of the stability coefficient; T t is the morphological feature time series data at time t; T max is the maximum value of the time series data; u t is the weight coefficient at time point t, in the range [0, 1], optimized through cross-validation; ρ t , σ t , μ t , ν t are trainable parameters, with initial values randomly sampled between [-1, 1]; ζ is a random noise term, following a normal distribution with a mean of 0 and a variance of 0.1; T is the prediction time window length.

[0129] The construction processes of the three functions in the mathematical model module are specifically described as follows:

[0130] In the construction process of the structural dynamics feature extraction function, the multi-dimensional characteristics of the data are first considered. In terms of the structural domain features, the stable structural domain components are obtained through singular value decomposition and the variable structural domain components are obtained by subtracting the stable components from the original data and the weight coefficients α i , β i are introduced to balance the contributions of the two structural domains, forming the structural domain feature term For the surface morphology features, the surface roughness change rate R i is introduced and the quadratic term is used to describe the non-linear change, and the coefficients γ i , δ i are used to adjust the proportion of the linear and non-linear terms, constituting the morphology feature term In terms of the combined dynamics features, the binding constant and the binding enthalpy change ΔH i are obtained through isothermal titration calorimetry measurement, and the coefficients η i , λ i are used to balance the effects of the two thermodynamic parameters, forming the binding feature term Finally, these three features are combined in the form of a weighted sum, and the time-point weight w is introduced i and a random noise term ε to obtain a complete structural dynamics feature extraction function

[0131] The construction process of the multi-scale feature fusion function fully considers multiple levels of the network structure. In terms of topological features, the degree distribution feature clustering coefficient feature and centrality feature are defined respectively to describe the connection pattern, local connection density and importance of nodes. In terms of path features, the overall connection characteristics of the network are described by the path length and path weight For the embedding feature, the attention mechanism is used to calculate the local embedding and the global embedding is obtained through graph pooling to effectively extract local and global features. Finally, different-scale features are fused by weighted average to obtain a complete multi-scale feature fusion function

[0132] The construction process of the dynamic feature mapping function focuses on the standardization and time-series characteristics of data. First, the structural function data multi-scale features and stability coefficient are normalized to eliminate the influence of dimensions. Then, the changes at adjacent time points are calculated to capture time-series features. By introducing a square transformation to enhance the non-linear expression ability and using the time weight coefficient u t to adjust the contributions of different time points, the dynamic feature mapping function is finally obtained

[0133] The parameter optimization of these three functions adopts a multi-level optimization strategy. For the optimization of the weight coefficients (w i , v k , u t ), by dividing the data set into a training set, a validation set and a test set, the model is trained on the training set, the performance is evaluated on the validation set, and the optimal weights are found through grid search. For the optimization of trainable parameters, the backpropagation algorithm is used. First, the parameters are initialized to random values, then the gradient of the loss function is calculated, and the parameter values are updated. The optimization is iterated until convergence. In terms of hyperparameter tuning, the number of scales K is determined to be optimal at 3 - 5 through experiments, the time window length T is selected according to the data characteristics, and the variance of the noise term is empirically set to 0.1. These optimization strategies ensure the effectiveness and robustness of the model and can accurately capture the structural dynamics features of the biomarkers.

[0134] Specifically, the principle of the present invention is as follows: At the level of data acquisition and processing, the present invention uses multi-omics technologies to comprehensively screen disease markers. This method can simultaneously obtain the structural, functional, and expression information of the markers. By establishing a multi-dimensional relational database, the effective integration of different types of data is achieved. Standardization and normalization processing ensure the comparability of data from different sources, laying a foundation for subsequent analysis. The application of principal component decomposition conforms to the biological principle of protein domain organization and can effectively distinguish conserved structural features and variable structural features.

[0135] At the level of network analysis, the present invention constructs a protein interaction network based on physical interaction, functional correlation, and expression correlation. This method fully considers the multi-level relationships between markers. Analyzing the network structure through the shortest path algorithm conforms to the actual characteristics of signal transmission in biomolecular networks. The application of graph neural networks can effectively handle the non-Euclidean characteristics of network data, and its hierarchical feature extraction process is consistent with the organizational principle of biomolecular networks.

[0136] At the level of experimental verification, the present invention uses atomic force microscopy for morphological analysis. This method can directly observe the structural changes of markers under near-physiological conditions. The application of isothermal titration calorimetry provides thermodynamic parameters for the interaction between markers and ligands, which directly reflect the energy changes during the molecular binding process. The combination of these two experimental methods obtains both structural information and energy information, providing a complete experimental basis for understanding the kinetic characteristics of markers.

[0137] At the level of model construction, the deep learning model designed by the present invention adopts an encoder-decoder structure, which can effectively handle the dimensionality reduction and reconstruction problems of high-dimensional data. The design of the mathematical model module considers three key links: feature extraction, feature fusion, and dynamic mapping, which conforms to the logical sequence of data analysis. The application of batch normalization and residual connections solves the problem of gradient disappearance in the training of deep neural networks, improving the training efficiency and stability of the model.

[0138] A specific embodiment 1 of the present invention is provided below. The specific implementation of each step in this embodiment 1 is described in detail as follows: The specific implementation of step S10 is to screen disease-related markers using omics technology. First, technologies such as proteomics, metabolomics, and transcriptomics are applied to perform high-throughput detection and identification of disease-related markers in biological samples. A mass spectrometer can be used to qualitatively and quantitatively analyze these markers to obtain their molecular weight data X, abundance data Y, modification site data W, and expression level data Z. These data are integrated to establish a structural and functional database of disease-related markers. This database includes information such as the amino acid sequence, molecular weight, isoelectric point, hydrophobicity index, secondary structure conformation, functional domain, post-translational modification site, and structural stability parameters of the markers.

[0139] Next, these structural and functional data are subjected to standardization and global normalization processing. The molecular weight data X is normalized using the minimum-maximum normalization method: where X min and X max are the minimum and maximum values of the data respectively. The abundance data Y is normalized using the logarithmic normalization method: where Y ref is the reference abundance value. The expression level data Z is normalized using the median normalization method: Znorm = where Z median is the median and MAD is the median absolute deviation. The modification site data W is normalized using the binary normalization method: where W threshold is the threshold. Through these standardization and normalization steps, structural and functional data thresholds for subsequent analysis can be generated.

[0140] The specific implementation of step S20 is to perform principal component decomposition calculation on the structural and functional data of disease-related markers. First, the principal component analysis method is applied to reduce the dimensionality of the structural and functional data X: T = XP; where T is the principal component score matrix and P is the eigenvector matrix. Then, the singular value decomposition method is used to decompose the structural and functional data X into a stable domain component S and a variable domain component D: X = U∑V T + E; where D is the left singular vector matrix, ∑ is the singular value diagonal matrix, V is the right singular vector matrix, and E is the error matrix. The stable domain component The variable domain component D = X - S; where U k , ∑ k and V k are the first k left singular vectors, singular values, and right singular vectors respectively.

[0141] Perform time series analysis on these two components and calculate their fluctuation threshold θ: θ = μ + ασ; where μ is the mean of the time series, σ is the standard deviation, and α is the adjustment factor. The stable domain component S represents the conservative structural characteristics of the biomarker, and the variable domain component D represents the changing structural characteristics of the biomarker.

[0142] The specific implementation of step S30 is to construct a relationship graph network of disease-related biomarkers using a protein-protein interaction network. First, establish the connection relationship between them based on the physical interaction, functional correlation, and expression level correlation between biomarkers. Take the biomarkers as network nodes v and the interaction strength between biomarkers as the weight value w(k,j) of the network edges to construct a weighted directed graph network.

[0143] Then, calculate the topological parameters of this network. The degree distribution calculation formula is: where n k is the number of nodes with degree k, and N is the total number of nodes. The clustering coefficient calculation formula is: where L i is the actual number of connections between the neighbors of node i, and k i is the degree of node i. The centrality measure uses betweenness centrality, and the calculation formula is: where σ st is the number of shortest paths between nodes s and t, and σ st (v) is the number of shortest paths passing through node v.

[0144] The specific implementation of step S40 is to calculate the action pathways between network nodes based on the relationship graph network using the shortest path algorithm and construct a biomarker action pathway matrix. Apply Dijkstra's algorithm to calculate the shortest action path between any two network nodes i and j, and the distance calculation formula is: d(i,j) = min{d(i,k) + w(k,j)}; where d(i,j) is the shortest distance from node i to j, and w(k,j) is the weight from node k to j. Construct a distance matrix D with the weight values of the shortest action paths and standardize it to obtain the standardized biomarker action pathway matrix D norm : where D min and D max are the minimum distance and the maximum distance respectively. Thus, the key path weight is obtained.

[0145] The specific implementation of step S50 is to quantitatively characterize disease-related biomarkers using a graph neural network model. This model adopts a deep learning architecture based on graph convolutional neural networks. The calculation formula of the graph convolutional layer is: where H (l)is the node feature matrix of the l-th layer, is the adjacency matrix with self-loops added, is the degree matrix, W (l) is the weight matrix, and σ is the activation function. The calculation formula of the pooling layer is: where is the feature of node i after pooling, h j is the feature of neighbor nodes, and N(i) is the neighbor set of node i.

[0146] Finally, generate the biomarker characterization vector e i and the node embedding vector: where is the node feature of the last layer, f MLP is the multi-layer perceptron function. Through the feature extraction and dimensionality reduction processing of these network layers, the high-dimensional structural and functional features can be transformed into low-dimensional vector representations.

[0147] The specific implementation of step S60 is to perform time-series topography analysis on disease-related biomarkers using an atomic force microscope. Measure the length, width, height, and surface roughness parameter y i , calculate the surface roughness R a and the surface roughness change rate ΔR: where is the average height, n is the number of measurement points, R t and R0 are the roughness at time t and the initial roughness respectively, and Δt is the time interval. There is a corresponding relationship between these topographical feature data and the aforementioned biomarker characterization vector, and a mapping function can be established using a multi-layer perceptron for mapping.

[0148] The specific implementation of step S70 is to measure the binding parameters of disease-related biomarkers and ligands using an isothermal titration calorimeter. The binding constant K a and the binding enthalpy change ΔH of the biomarker and ligand can be calculated: where [PL] is the complex concentration, [P] is the protein concentration, [L] is the ligand concentration, Q is the heat change, n is the number of binding sites, and [P] t is the total protein concentration. By measuring the binding parameters at different time points, the dynamic change values of the binding constant and the binding enthalpy change can be obtained. Construct these binding parameters into a matrix S and calculate its stability coefficient: where σ i is the standard deviation of parameter i, μ i is the mean of parameter i, and N is the number of parameters.

[0149] The specific implementation of step S80 is to input the above various structural dynamics characteristic data into a deep learning model. The model adopts an architecture based on an encoder network and a decoder network, and a mathematical model module is set between the two. The mathematical model module includes a structural dynamics feature extraction function F dyn 、a multi-scale feature fusion function F multi and a dynamic feature mapping function Y pred .

[0150] The input of the structural dynamics feature extraction function includes the fluctuation threshold of the stable domain component and the variable domain component , the surface roughness change rate R i , the binding constant and the dynamic change value of the binding enthalpy change ΔH i , etc. Its calculation formula is: where w i is the feature weight coefficient, α, β, γ, δ, η, λ are trainable parameters, ε is a random noise term, and n is the time series length.

[0151] The input of the multi-scale feature fusion function includes the topological parameters T deg 、T clu 、T cen of the protein interaction network, the path length and weight feature P len 、P wei , the local and global embedding features E loc 、E glo , and the structural dynamics features , etc. Its calculation formula is: where v k is the scale weight coefficient, φ, ψ, ω, τ are trainable parameters, ξ is a random noise term, and K is the number of scales.

[0152] The input of the dynamic feature mapping function includes the structural function data threshold multi-scale features the binding parameter stability coefficient S t and the morphological feature time series T t , etc. Its calculation formula is: where u t is the weight coefficient at time point t, ρ, σ, μ, v are trainable parameters, ζ is a random noise term, and T is the prediction time window length.

[0153] The specific implementation of step S90 is to use the cross-validation method to train and optimize the deep learning model. The model consists of a 5-layer encoder network, a mathematical model module, and a 5-layer decoder network. The encoder network adopts a fully connected neural network structure, and the leaky rectified linear unit activation function σ(x) = max(0, x) is used between each layer of the network. A batch normalization layer and a dropout layer are added after each layer of the network to improve the generalization ability of the network. The decoder network adopts a fully connected neural network structure symmetric to the encoder, and the hyperbolic tangent function is used as the activation function. Similarly, a batch normalization layer and a dropout layer are added after each layer of the network. Finally, a disease prediction model can be established to determine the structural and dynamic characteristics of disease-related markers.

[0154] To better understand and implement the present invention, Example 2 of a specific application scenario of the present invention is provided below: A research center has established a structural and functional database of dozens of tumor markers. These data include information such as the amino acid sequence, molecular weight, isoelectric point, hydrophobicity index, secondary structure conformation, functional domain, post-translational modification site, and structural stability parameter of the markers. To eliminate technical biases during the testing process, the research team standardized and normalized these raw data. The maximum-minimum normalization was used for the molecular weight data, and it was normalized to the range of 0-1; the logarithmic normalization was used for the abundance data, and it was presented in the form of binary logarithm; the median normalization was used for the expression data, and it was normalized with the median deviation; the binary normalization was used for the modification site data, and 0 / 1 was used to indicate whether the site had a modification. Through these preprocessing steps, the center established a high-quality structural and functional database of tumor markers.

[0155] Next, the research team performed principal component analysis on these data and extracted the stable domain features and variable domain features of the markers. Taking the lung adenocarcinoma marker EGFR as an example, its primary structure includes 1210 amino acids, with a molecular weight of 133 kDa, an isoelectric point of 7.6, and a hydrophobicity index of -0.37. The secondary structure of the EGFR protein is mainly composed of α-helices and β-sheets, and the functional domains include an extracellular ligand-binding region, a transmembrane region, and an intracellular tyrosine kinase region. In addition, there are multiple post-translational modification sites in EGFR, such as serine phosphorylation at position 728 and tyrosine phosphorylation at position 845.

[0156] Through principal component analysis, the research team decomposed the structural and functional data of EGFR into stable domain components and variable domain components. The stable domain components mainly include conserved structural features such as α-helix structures, extracellular ligand-binding regions, and transmembrane regions, which remain relatively stable in different tumor samples. The variable domain components, on the other hand, reflect the more tumorigenesis-sensitive structural features such as the intracellular tyrosine kinase region and serine phosphorylation sites of EGFR, and their time series fluctuations reflect the dynamic changes of EGFR during the tumor process. As Figure 2 shown, it demonstrates the distribution characteristics of the stable and variable domains of EGFR in the principal component space. The blue dots represent the stable domains (such as α-helix structures, extracellular ligand-binding regions, etc.), and the red dots represent the variable domains (such as the intracellular tyrosine kinase region, phosphorylation sites, etc.). From the distribution, it can be seen that the stable domains have better aggregation, while the variable domains have a larger dispersion degree.

[0157] To further analyze the role of EGFR in the tumor signaling pathway, the research team constructed an EGFR interaction relationship network based on protein-protein interactions. This network contains more than 300 protein nodes that interact with EGFR and its upstream and downstream, and the connection relationships between the nodes are determined by physical interactions, functional correlations, and expression quantity correlations, etc. By calculating the topological parameters of the network, the team found that EGFR is in a central hub position in this network, with a power-law distribution characteristic of degree distribution and a relatively high clustering coefficient, indicating that it plays a key hub role in the tumor cell signal regulation network.

[0158] In addition, applying the shortest path algorithm, the research team also analyzed the interaction pathways between EGFR and other key signaling pathway proteins. The results showed that EGFR regulates the proliferation, survival, and migration behaviors of tumor cells through classical signaling pathways such as phosphatidylinositol 3-kinase (PI3K) and mitogen-activated protein kinase (MAPK). For these key interaction pathways, the team constructed an EGFR signaling pathway matrix, providing important topological information for subsequent structural dynamics analysis. As Figure 3 shown, it demonstrates the interaction relationship network of EGFR and its related proteins. The central red node represents EGFR, and the surrounding blue nodes represent the proteins that interact with it. The arrows indicate the direction of interaction, and the node sizes represent the importance in the network. It can be seen that EGFR is in the core position of the network and has direct interactions with multiple PI3K and MAPK pathway proteins.

[0159] In order to more comprehensively characterize the structure-function relationship of EGFR, the research team combined atomic force microscopy technology to conduct a time series morphological analysis. Under room temperature, the team used tapping mode to scan and image the EGFR protein sample for 24 hours and obtained its three-dimensional morphological image. By measuring the length, width, height and other parameters of EGFR, the rate of change of its surface roughness over time was calculated. The results showed that with the development of tumor progression, the surface roughness of EGFR showed a dynamic change trend of first increasing and then decreasing, reflecting the remodeling process of its structure on a time scale. Figure 4 As shown in the figure, the dynamic evolution of EGFR structural changes in time and space is shown. The X-axis represents time, the Y-axis represents the spatial scale, the Z-axis represents the amplitude of structural changes, and the color represents the intensity of structural changes. Through this figure, we can intuitively observe the fluctuation characteristics of EGFR structure over time and space.

[0160] At the same time, the research team also used an isothermal titration calorimeter to determine the binding kinetic parameters of EGFR and specific ligands. Under constant temperature conditions, the ligand solution was added to the EGFR protein solution, and the heat changes during the binding process were monitored in real time. Through thermal curve fitting, the binding constant and binding enthalpy change of EGFR and ligand can be calculated. The results showed that with the progression of the tumor, the binding constant of EGFR showed a dynamic change trend of first decreasing and then increasing, while the binding enthalpy change showed the opposite trend. These results reveal the dynamic change characteristics of EGFR activity regulation and its ability to bind to ligands during tumor progression.

[0161] The above multi-scale EGFR structural dynamics feature data were input into the deep learning model for training and optimization, and the research team built a prediction model for lung adenocarcinoma diagnosis. The model includes three modules: structural dynamics feature extraction, multi-scale feature fusion, and dynamic feature mapping.

[0162] In the structural dynamics feature extraction module, the team designed a weighted summation function that comprehensively considers indicators such as the fluctuation threshold of the stable domain of EGFR, the fluctuation threshold of the variable domain, the rate of change of surface roughness, the dynamic change of the binding constant, and the dynamic change of the binding enthalpy change:

[0163]

[0164] Among them, w i is the weight coefficient of each feature, α, β, γ, δ, η, λ are trainable parameters, and ε is a random noise term. Through this feature extraction function, key predictive factors can be extracted from the structural dynamics data of EGFR.

[0165] In the multi-scale feature fusion module, the team considered multiple aspects of information, such as the topological parameters of the EGFR interaction network, the weights of key interaction pathways, the node embedding vectors, and the aforementioned structural dynamics feature vectors:

[0166]

[0167] Among them, v k is the weight coefficient at different scales, φ, ψ, ω are trainable parameters, and ξ is a random noise term. By fusing features at different scales, the mechanism of EGFR in tumorigenesis can be characterized more comprehensively.

[0168] Finally, in the dynamic feature mapping module, the team designed a prediction function based on time series, taking into account information such as the threshold of standardized structural function data, the multi-scale feature fusion results, the combination parameter stability, and the time series of morphological features:

[0169]

[0170] Among them, u t is the weight coefficient at different time points, ρ, σ, μ, v are trainable parameters, and ζ is a random noise term. Through this dynamic feature mapping function, the trend of structural dynamics changes of EGFR in the next year can be predicted, providing an important basis for the diagnosis and treatment of lung adenocarcinoma.

[0171] Table 1 Results of standardization processing of EGFR structural function data

[0172] Feature Normalized value Molecular weight 0.792 Isoelectric point 0.682 Hydrophobicity index -0.374 Secondary structure (α-helix) 0.615 Secondary structure (β-sheet) 0.472 Number of post-translational modification sites 1 Structural stability 0.853

[0173] Table 2 Topological parameters of the EGFR interaction network

[0174] Parameter Calculation result Degree distribution index 2.38 Average clustering coefficient 0.714 Average shortest path length 3.86 Betweenness centrality 0.423

[0175] Table 3 Trend of changes in EGFR structural dynamics characteristics

[0176] Feature Early stage of tumor Middle stage of tumor Advanced stage of tumor Fluctuation threshold of stable domain 0.35 0.41 0.47 Fluctuation threshold of variable domain 0.52 0.68 0.82 Change rate of surface roughness 0.028 0.034 0.025 Binding constant <![CDATA[1.8×10 8 M -1 > <![CDATA[9.2×10 7 M -1 > <![CDATA[2.1×10 8 M -1 > Enthalpy change of binding -52 kJ / mol -41 kJ / mol -48 kJ / mol

[0177] Through systematic analysis of the structural dynamics of EGFR, the research team found that it showed an obvious dynamic change trend during the occurrence and development of lung adenocarcinoma. In the early stage of the tumor, the stable domain of EGFR was relatively conservative, the variable domain was more active, and the surface roughness changed greatly, indicating that it was in the process of structural remodeling and functional activation. At this time, the binding activity of EGFR with ligands was strong, the binding constant was high, and the binding enthalpy change was large. In the middle stage of the tumor, the fluctuation of the variable domain of EGFR increased, the change of surface roughness tended to be gentle, the binding activity with ligands decreased, the binding constant decreased, and the binding enthalpy change decreased. In the late stage of the tumor, the fluctuation of the stable domain of EGFR increased, the variable domain also tended to be stable, the change of surface roughness increased again, and the binding activity recovered somewhat.

[0178] It can be inferred that these characteristics of structural dynamics changes are highly correlated with the mechanism of action of EGFR in the occurrence and development of lung adenocarcinoma. In the early stage of the tumor, EGFR is highly active, regulating the proliferation and survival of tumor cells through classical signaling pathways such as MAPK and PI3K, making the tumor in a stage of rapid progression. With the development of the tumor, the structure of EGFR begins to show a certain degree of denaturation and inactivation, reducing its binding ability with ligands. In the late stage, the fluctuation of the stable domain of EGFR increases, and it may participate in the invasion and metastasis of tumor cells through other signaling pathways.

[0179] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Tables 4 and 5 below.

[0180] Table 4 Variable Explanation Table (First Part of Variables)

[0181]

[0182] Table 5 Variable Explanation Table (Second Part of Variables)

[0183]

[0184]

[0185] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.

Claims

1. A method for determining the structural dynamics characteristics of a disease biomarker, characterized in that, It includes the following steps: Using omics technology to screen disease-related markers, obtaining the structural and functional data of disease-related markers, establishing a structural and functional database, performing standardization and global normalization on the structural and functional data to obtain the threshold of the structural and functional data; performing principal component decomposition calculation on the structural and functional data of disease-related markers to obtain the stable domain component and the variable domain component; constructing a network of the interaction relationship graph of disease-related markers; constructing a marker action pathway matrix using the shortest path algorithm; inputting the marker action pathway matrix into a graph neural network model to obtain the marker characterization vector; performing time series topography analysis using atomic force microscopy; measuring the binding parameters using isothermal titration calorimetry to construct a binding parameter matrix; inputting the above data into a deep learning model; training the model using the cross-validation method to determine the structural dynamics characteristics of disease-related markers.

2. The method for determining the structural dynamics characteristics of a disease biomarker according to claim 1, wherein Using omics technology to screen disease-related markers, specifically, using proteomics, metabolomics and transcriptomics technologies to perform high-throughput detection and identification of disease-related markers in biological samples, performing qualitative and quantitative analysis of disease-related markers through a mass spectrometer, and establishing the molecular weight data, abundance data, modification site data and expression level data of disease-related markers.

3. The method for determining the structural dynamics characteristics of a disease biomarker according to claim 2, wherein The structural and functional database is specifically a multi-dimensional relational database composed of a primary structure information data table, a secondary structure information data table, a functional domain data table and a modification site data table of disease-related markers, including the amino acid sequence, molecular weight, isoelectric point, hydrophobicity index, secondary structure conformation, functional domain, post-translational modification site and structural stability parameter of disease-related markers.

4. The method for determining the structural dynamics characteristics of a disease biomarker according to claim 3, wherein Performing standardization and global normalization on the structural and functional data, specifically, performing normalization on the molecular weight data using the maximum-minimum normalization method, performing normalization on the abundance data using the logarithmic normalization method, performing normalization on the expression level data using the median normalization method, and performing normalization on the modification site data using the binary normalization method to generate the threshold of the structural and functional data.

5. The method for determining the structural dynamics characteristics of a disease biomarker according to claim 4, characterized in that Performing principal component decomposition calculation on the structural and functional data of disease-related markers, specifically, using the principal component analysis method to perform dimensionality reduction processing on the structural and functional data, extracting the main domain features, using the singular value decomposition method to decompose the structural and functional data into a stable domain component and a variable domain component, performing time series analysis on the stable domain component and the variable domain component, and calculating the fluctuation threshold.

6. The method for determining the structural dynamics characteristics of a disease biomarker according to claim 5, characterized in that, Constructing a network of the interaction relationship graph of disease-related markers, specifically, establishing the connection relationship between disease-related markers based on the physical interaction, functional correlation and expression level correlation between disease-related markers, using disease-related markers as network nodes, using the interaction intensity between disease-related markers as the weight value of the network edge, constructing a weighted directed graph network, and calculating the degree distribution, clustering coefficient and centrality measure of the network.

7. The method for determining the structural dynamics characteristics of a disease biomarker according to claim 6, wherein Construct a marker action pathway matrix using the shortest path algorithm. Specifically, use Dijkstra's algorithm to calculate the shortest action path between any two network nodes, construct a distance matrix based on the weight values of the shortest action paths, transform the distance matrix into a standardized marker action pathway matrix, and obtain the key path weights.

8. The method for determining the structural dynamics characteristics of a disease biomarker according to claim 7, wherein, The graph neural network model is specifically a deep learning architecture based on graph convolutional neural networks. The structure includes a graph convolutional layer, a pooling layer, a fully connected layer, and an output layer. Among them, the graph convolutional layer is used to extract the local features of network nodes, the pooling layer is used to reduce the feature dimension, the fully connected layer is used for feature fusion, and the output layer is used to generate marker representation vectors and node embedding vectors.

9. The method for determining the structural dynamics characteristics of a disease biomarker according to claim 8, wherein Perform time-series topography analysis using an atomic force microscope. Specifically, use the tapping mode to perform scanning imaging of disease-related marker samples at continuous time points under room temperature conditions, obtain three-dimensional topography images of disease-related markers, measure the length, width, height, and surface roughness parameters of disease-related markers, calculate the surface roughness change rate, and record the time-series change data of topographical features.

10. The method for determining the structural dynamics characteristics of a disease biomarker according to claim 9, characterized in that, Determine the binding parameters using an isothermal titration calorimeter. Specifically, measure the binding constant and binding enthalpy change between disease-related markers and ligands, construct a binding parameter matrix, calculate the stability coefficient of the binding parameter matrix, analyze the dynamic changes of the binding constant and binding enthalpy change, and evaluate the thermodynamic properties of the binding between disease-related markers and ligands.