Disease marker expression network reconstruction feature key point determination method
Through wavelet decomposition and spectral clustering methods, combined with multi-scale network reconstruction and stability analysis, the inaccuracy problem of determining the key points of the reconstruction feature of the disease marker expression network in the prior art is solved, and more accurate disease marker network feature recognition and key node group recognition are achieved.
Patent Information
- Application Number
- CN202510438393.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-08
AI Technical Summary
The existing disease marker expression network reconstruction and key node identification methods have insufficient accuracy in data collection, feature extraction, network reconstruction and key point recognition, which is difficult to reflect the complexity of the biomolecular network and limit its application value in disease diagnosis and treatment.
Wavelet decomposition technology is used to process the multi-dimensional data of markers, and the frequency domain correlation matrix and multi-scale weight are constructed. Combined with the network center degree and spectral clustering methods, key node groups are identified and feature key points are determined. The identification accuracy is improved through multi-scale topology and stability analysis.
Multidimensional feature fusion and multi-scale network reconstruction of marker expression networks are realized, data comparability and explanatory network model are improved, and important and stable key node groups are identified, providing a reliable basis for disease diagnosis and treatment.
Smart Images

Figure CN120279988A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of disease markers, and more specifically, relates to a method for determining key points of a disease marker expression network reconstruction feature. Background Art
[0002] Analysis of biomarker expression regulatory networks is an important direction in current molecular medicine research. Traditional biomarker research is mainly based on single biological indicators or pairwise interactions between molecules, making it difficult to reflect the complex regulatory relationships in the living system. In recent years, with the rapid development of molecular biology and bioinformatics, network biology analysis methods based on omics data have gradually received attention. Such methods can depict the mutual regulatory mechanisms between biomarkers from a systematic level, providing a new perspective for deeply understanding the molecular basis of disease occurrence and development.
[0003] The existing disease biomarker expression network reconstruction and key node identification mainly include the following aspects:
[0004] 1. Biomarker expression data acquisition method: Currently, it mainly relies on immunohistochemical staining or immunofluorescence staining techniques to obtain the expression information of biomarkers in tissue sections. These methods can only provide static expression intensity data and cannot reflect the dynamic interactions between biomarkers. At the same time, due to the lack of standardization in image acquisition and analysis, the data comparability between different batches and different samples is poor.
[0005] 2. Biomarker expression feature extraction method: Existing methods mainly focus on the analysis of the expression intensity of a single biomarker, ignoring the interaction relationship between biomarkers. The description of spatial distribution features is limited to simple density statistics and cannot depict complex cell community structures. In addition, there is a lack of analysis of the dynamic change rules of biomarker expression.
[0006] 3. Network reconstruction method: Existing network reconstructions are mostly based on simple statistical indicators such as correlation coefficients, making it difficult to reflect the complexity of biological regulation. The network topology is relatively single, and it is difficult to discover hidden regulatory patterns. At the same time, the insufficient discrimination of connections with different intensities also limits the interpretability of the network model.
[0007] 4. Key node identification method: Existing methods mainly rely on a single centrality index, such as degree centrality, betweenness centrality, etc., lacking comprehensive consideration of the multi-dimensional features of nodes. In addition, these methods ignore the importance differences of nodes at different scales and are difficult to comprehensively evaluate their key roles in the network.
[0008] Generally speaking, existing methods for analyzing marker expression networks have some limitations in key aspects such as data collection, feature extraction, network reconstruction, and key point identification. It is difficult to fully reflect the complexity of biomolecular networks, thus limiting their application value in disease diagnosis and treatment. That is to say, there is an inaccurate technical problem in the method for determining the key points of the reconstructed features of the disease marker expression network in the prior art. Summary of the Invention
[0009] In view of this, the present invention provides a method for determining key points of reconstructed features of a disease marker expression network, which can solve the inaccurate technical problem in the method for determining key points of reconstructed features of a disease marker expression network in the prior art. The method includes the following operation steps:
[0010] Process the tissue sample to be tested to obtain marker multi-dimensional data; perform wavelet decomposition on the marker multi-dimensional data to obtain marker multi-scale coefficients, where the marker multi-scale coefficients include a reference component, a high-frequency component, a medium-frequency component, and a low-frequency component; construct a frequency-domain correlation matrix based on the marker multi-scale coefficients, where the frequency-domain correlation matrix characterizes the correlation strength between different frequency components; construct a network reconstruction model according to the frequency-domain correlation matrix and multi-scale weights, where the multi-scale weights include a reference weight, a high-frequency weight, a medium-frequency weight, and a low-frequency weight; calculate the network centrality based on the network reconstruction model, where the network centrality includes connection centrality, path centrality, and structural centrality; identify a key node group and determine the key points according to the network centrality; calculate the multi-scale importance of the key points in the multi-scale topological structure and output a sorting result.
[0011] Among them, processing the tissue sample to be tested includes: performing preparation processing on the tissue sample to be tested, including tissue sectioning, fixation, dehydration, and embedding; performing marker staining processing on the tissue sample to be tested by immunohistochemical staining.
[0012] Among them, obtaining the marker multi-dimensional data includes: collecting a fluorescence image of the tissue sample to be tested through a fluorescence microscope, where the fluorescence image includes multiple fluorescence channels; performing image segmentation processing on the fluorescence image to extract the target cell region.
[0013] Among them, the marker multi-dimensional data includes: marker expression intensity data obtained by measuring the fluorescence intensity of the marker in the target cell region; marker spatial distribution data obtained by recording the spatial position of the target cell region in the fluorescence image; marker co-expression data obtained by analyzing the expression patterns of different markers in the same target cell region; marker time-series data obtained by collecting the marker expression intensity data at time intervals.
[0014] Among them, wavelet decomposition of the marker multi-dimensional data includes: dividing the marker multi-dimensional data into a reference component, a high-frequency component, a medium-frequency component, and a low-frequency component according to the energy distribution of the marker multi-scale coefficients.
[0015] Among them, constructing the frequency-domain correlation matrix includes: calculating the frequency-domain correlation degrees between the reference component and the high-frequency component, the medium-frequency component, and the low-frequency component; calculating the multi-scale weights of network nodes based on the frequency-domain correlation matrix, where the multi-scale weights include a reference weight, a high-frequency weight, a medium-frequency weight, and a low-frequency weight.
[0016] Among them, constructing the network reconstruction model includes: constructing a multi-scale topological structure of the marker expression network according to the frequency-domain correlation matrix and the multi-scale weights.
[0017] Among them, the network centrality includes connection centrality, path centrality, and structural centrality; constructing a node feature matrix based on the network centrality, and using the spectral clustering method to identify key node groups in the network reconstruction model.
[0018] Among them, calculating the multi-scale importance of the characteristic key points includes: calculating the internal connectivity and external connectivity of the key node groups in the multi-scale topological structure; performing network stability analysis on the key point data.
[0019] Among them, sorting the characteristic key points according to the multi-scale importance and outputting the sorting result.
[0020] The effect of the present invention is as follows:
[0021] The present invention proposes a new method for reconstructing a disease marker expression network and identifying key nodes, which has made remarkable improvements on the basis of the existing technology.
[0022] First, in terms of data collection and preprocessing, the method of the present invention establishes a standardized fluorescence intensity correction process, introduces compensation for reference intensity and background noise, and improves the comparability of data between different batches and different samples. At the same time, an adaptive image segmentation algorithm is adopted, combined with morphological processing, which can more accurately extract the target cell region and lay a foundation for subsequent feature analysis. In addition, the present invention realizes the collaborative analysis of multi-channel fluorescence images and can simultaneously obtain the distribution information of multiple markers in tissues.
[0023] Secondly, in terms of biomarker expression feature extraction, the present invention proposes a multi-dimensional feature fusion framework, which integrates various biological features such as the expression intensity, spatial distribution, co-expression relationship, and temporal variation of biomarkers, and can characterize the complex expression patterns of biomarkers more comprehensively than existing methods. In particular, in co-expression analysis, the present invention develops a local-global correlation evaluation method to reveal the regulatory relationships at different scales. At the same time, temporal data analysis is introduced to discover the laws of dynamic changes in biomarker expression.
[0024] Thirdly, in terms of network reconstruction, the present invention establishes a multi-scale network model based on wavelet analysis, which can distinguish the contributions of different frequency components to biological regulation and better reflect complex molecular interactions. When calculating the node weights, the present invention comprehensively considers the local correlation degree, global energy distribution, and frequency characteristics, making the network structure closer to biological reality. In addition, the multi-scale network representation constructed by the present invention can deeply characterize the interactions between biomarkers at different intensity levels.
[0025] Finally, in terms of key node identification, the present invention develops a multi-dimensional node importance evaluation system, which comprehensively considers the connection centrality, path centrality, and structural centrality of nodes. At the same time, a spectral clustering method is introduced to identify key node groups, improving the systematicness of key point screening. In addition, a network stability analysis link is added to ensure that the identified key nodes have an important and stable position in the overall network.
[0026] Generally speaking, the method of the present invention has made important innovations in key aspects such as data processing, feature extraction, network reconstruction, and key point identification, and is significantly superior to the prior art in terms of accuracy, reliability, and biological relevance. It solves the technical problem that the method for determining the key feature points of the disease biomarker expression network reconstruction in the prior art is not accurate enough. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a flowchart of the method provided by the present invention;
[0028] Figure 2 is a fluorescence intensity distribution diagram of the biomarker in Example 2;
[0029] Figure 3 is a temporal variation diagram of the spatial distribution entropy in Example 2;
[0030] Figure 4 is a wavelet energy distribution diagram in Example 2;
[0031] Figure 5 is a key node stability index diagram in Example 2. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0033] As Figure 1 shown, the present invention includes the following operating steps:
[0034] S01. Prepare the tissue sample to be tested, including tissue sectioning, fixation, dehydration, and embedding;
[0035] S02. Perform marker staining on the tissue sample to be tested using immunohistochemical staining;
[0036] S03. Collect fluorescence images of the tissue sample to be tested through a fluorescence microscope, where the fluorescence images include multiple fluorescence channels;
[0037] S04. Perform image segmentation on the fluorescence images to extract the target cell region;
[0038] S05. Measure the fluorescence intensity of the marker within the target cell region to obtain marker expression intensity data;
[0039] S06. Record the spatial position of the target cell region in the fluorescence image to obtain marker spatial distribution data;
[0040] S07. Analyze the expression patterns of different markers within the same target cell region to obtain marker co-expression data;
[0041] S08. Collect the marker expression intensity data at time intervals to obtain marker time series data;
[0042] S09. Combine the marker expression intensity data, marker spatial distribution data, marker co-expression data, and marker time series data to form marker multi-dimensional data;
[0043] S10. Perform wavelet decomposition on the marker multi-dimensional data to obtain marker multi-scale coefficients;
[0044] S11. Divide the marker multi-dimensional data into a reference component, a high-frequency component, a medium-frequency component, and a low-frequency component according to the energy distribution of the marker multi-scale coefficients;
[0045] S12. Calculate the frequency domain correlation degrees between the reference component and the high-frequency component, medium-frequency component, and low-frequency component, and construct a frequency domain correlation matrix;
[0046] S13. Calculate the multi-scale weights of network nodes based on the frequency domain correlation matrix, where the multi-scale weights include a reference weight, a high-frequency weight, a medium-frequency weight, and a low-frequency weight;
[0047] S14. Construct a multi-scale topological structure of the biomarker expression network based on the frequency-domain correlation matrix and the multi-scale weights to obtain a network reconstruction model;
[0048] S15. Calculate the network centrality of each node in the network reconstruction model, where the network centrality includes connection centrality, path centrality, and structural centrality;
[0049] S16. Construct a node feature matrix based on the network centrality, and use the spectral clustering method to identify key node groups in the network reconstruction model;
[0050] S17. Calculate the internal connectivity and external connectivity of the key node groups in the multi-scale topological structure, determine the characteristic key points to obtain key point data; perform network stability analysis on the key point data, calculate the multi-scale importance of the characteristic key points, sort the characteristic key points according to the multi-scale importance, and output the sorting result.
[0051] The following describes in detail the specific implementation manners of the above steps.
[0052] The specific implementation manner of step S01 is to perform preparation processing on the tissue sample to be measured. First, tissue sectioning is carried out to cut the sample to be measured into thin slices; then a fixing solution is used to fix the sections to maintain the integrity of the cell structure; subsequently, dehydration treatment is performed to remove the moisture in the sections; finally, through embedding technology, the dehydrated sections are embedded in a specific supporting material for subsequent staining and observation. This series of preparation processes aims to ensure that the tissue sample to be measured can be effectively labeled by the subsequent immunohistochemical staining method.
[0053] The specific implementation manner of step S02 is to perform biomarker staining on the tissue sample to be measured by immunohistochemical staining. This method utilizes the high-affinity interaction between specific antibodies and biomarker molecules to mark the distribution of the target biomarker in the tissue sample to be measured through an immune reaction. Specifically, first, the prepared tissue sections are incubated with a specific antibody solution to make the antibody specifically bind to the biomarker in the tissue; then a secondary antibody labeled with a fluorescent substance or an enzyme is added to form a complex with the primary antibody; finally, through microscopic observation or enzymatic color reaction, the spatial distribution information of the biomarker in the tissue can be obtained. The purpose of this step is to obtain qualitative and quantitative information of various biomarkers in the tissue sample to be measured.
[0054] The specific implementation of step S03 is to collect fluorescence images of the tissue sample to be tested through a fluorescence microscope. After completing immunohistochemical staining, the stained sections are placed under a fluorescence microscope for imaging. By setting different fluorescence channels, fluorescence signals of different types of markers can be captured respectively. For example, the FITC channel can be set to detect green fluorescence, the Texas Red channel to detect red fluorescence, etc., so as to obtain multi-channel fluorescence images containing information of multiple markers. The purpose of this step is to lay a foundation for subsequent image analysis and provide data support for quantitatively evaluating the expression levels of various markers.
[0055] The specific implementation of step S04 is to perform image segmentation processing on the obtained fluorescence images to extract the target cell regions. First, an adaptive threshold segmentation method is used to binarize the images. This method automatically determines the optimal segmentation threshold according to the gray-scale statistical characteristics of different regions in the images, and can well handle the problem of uneven illumination during the imaging process. Then, morphological processing means such as opening operation and closing operation are combined to remove noise interference in the images and retain the complete target cell regions. The purpose of this step is to accurately extract the cell regions to be analyzed from the complex fluorescence images and provide a reliable data source for subsequent quantitative expression of markers.
[0056] The specific implementation of step S05 is to measure the fluorescence intensity of the markers within the target cell regions to obtain marker expression intensity data. First, the fluorescence intensity of each pixel point is corrected, taking into account the influence of background noise and reference intensity to improve the comparability of the data. Specifically, normalization is performed using the maximum reference intensity, and the corrected intensity value reflects the relative expression level of the marker at that pixel point. Then, the average fluorescence intensity of the markers within the entire target cell region is statistically calculated in a weighted average manner, and the weight coefficient is related to the pixel area size to consider the influence of the region size on the expression intensity. The purpose of this step is to obtain the quantitative expression levels of each marker within the target cell regions and lay a foundation for subsequent multi-dimensional feature extraction and network reconstruction analysis.
[0057] The specific implementation of step S06 is to record the spatial positions of the target cell regions in the fluorescence images to obtain the spatial distribution data of the markers. First, a spatial adjacency matrix between cell regions is constructed, and the adjacent relationship is defined by setting an Euclidean distance threshold, which reflects the local structural characteristics of the target cells in space. Secondly, the spatial distribution entropy is calculated to describe the spatial distribution uniformity of the entire cell population, and the aggregation or dispersion patterns of the cells can be identified. The purpose of this step is to obtain the distribution of the markers in the tissue space and provide spatial topological information for subsequent multi-scale network reconstruction.
[0058] The specific implementation of step S07 is to analyze the expression patterns of different markers within the same target cell region to obtain marker co-expression data. First, the local co-expression index is calculated, which takes into account the correlation of marker expression within the spatial neighborhood and can capture the co-expression patterns at the cellular level. Then, all local information is integrated through weighted averaging to construct a global co-expression matrix, which reflects the overall co-expression relationship between markers. The purpose of this step is to extract the co-expression characteristics between markers and provide correlation information for subsequent network construction.
[0059] The specific implementation of step S08 is to collect marker expression intensity data according to time intervals to obtain marker time-series data. First, the original time-series data is standardized, and a noise compensation term is introduced to enhance the data stability. Then, the autocorrelation function of the time series is calculated, and the periodic change patterns existing in the marker expression can be found. The purpose of this step is to obtain the dynamic information of marker expression in the time dimension and provide support for multi-dimensional feature extraction and time-series analysis.
[0060] The specific implementation of step S09 is to combine the marker expression intensity data, marker spatial distribution data, marker co-expression data, and marker time-series data to form marker multi-dimensional data. Specifically, a feature vector containing 4 dimensions is designed, which respectively reflects the expression intensity, spatial distribution, co-expression relationship, and time-series change characteristics of the markers. To eliminate the influence of dimensional differences, the feature data of each dimension is normalized. The purpose of this step is to construct a multi-dimensional data set that comprehensively depicts the marker expression characteristics and provide basic data support for subsequent network reconstruction and key point identification.
[0061] The specific implementation of step S10 is to perform wavelet decomposition on the marker multi-dimensional data to obtain marker multi-scale coefficients. First, the input multi-dimensional feature data is decomposed using two-dimensional wavelet transform to obtain wavelet coefficients at different scales. Wavelet transform can achieve time-frequency joint analysis, thereby extracting multi-scale feature information. Then, the energy distribution of the wavelet coefficients at each scale is calculated for subsequent frequency band division. The purpose of this step is to convert the original multi-dimensional feature data into a multi-scale feature representation, laying the foundation for multi-scale analysis of network reconstruction.
[0062] The specific implementation of step S11 is to divide the multi-dimensional data of the marker into a reference component, a high-frequency component, a medium-frequency component, and a low-frequency component according to the energy distribution of the multi-scale coefficient of the marker. Specifically, first calculate the energy proportion of each scale wavelet coefficient as the basis for frequency band division. Then set the high and low frequency thresholds, divide those with energy distribution greater than the high frequency threshold into high-frequency components, those between the high and low frequency thresholds into medium-frequency components, and those less than the low frequency threshold into low-frequency components. The purpose of this step is to extract information of different frequency components from multi-dimensional features, laying a foundation for subsequent multi-scale network analysis.
[0063] The specific implementation of step S12 is to calculate the frequency domain correlation degrees between the reference component and the high-frequency, medium-frequency, and low-frequency components, and construct a frequency domain correlation matrix. First, calculate the correlation degree between different frequency bands through the frequency domain coherence analysis method, which takes into account the influence of power spectral density and cross spectral density. Then arrange the calculated correlation coefficients in the form of a block matrix, that is, the frequency domain correlation matrix. The purpose of this step is to depict the interaction relationship between the multi-scale features of the marker, providing a basis for subsequent multi-scale network reconstruction.
[0064] The specific implementation of step S13 is to calculate the multi-scale weights of the network nodes according to the frequency domain correlation matrix. First, calculate the local weight and global weight of the nodes based on local correlation degree and global energy distribution respectively. Then perform a weighted combination of the two weights with an exponential decay term related to frequency characteristics to obtain the comprehensive multi-scale weights. The purpose of this step is to balance the contributions of different scale features to the importance of nodes, providing a weight basis for constructing the multi-scale network topology.
[0065] The specific implementation of step S14 is to construct the multi-scale topology of the marker expression network according to the frequency domain correlation matrix and multi-scale weights, and obtain a network reconstruction model. First, calculate the connection strength between nodes, taking into account the comprehensive influence of node weights, correlation degrees, and spatial distances. Then construct the network adjacency matrix corresponding to the corresponding scale according to different strength thresholds, realizing the multi-scale network representation. The purpose of this step is to establish a network model reflecting the multi-dimensional regulatory relationship of the marker, providing a basis for subsequent key point identification.
[0066] The specific implementation of step S15 is to calculate the network centrality of each node in the network reconstruction model. First, use eigenvector centrality to analyze the importance of nodes in the overall network. Secondly, introduce an improved betweenness centrality, taking into account the influence of the shortest path length. The purpose of this step is to evaluate the importance indicators of network nodes from multiple perspectives, providing a basis for subsequent key point screening.
[0067] The specific implementation of step S16 is to construct a node feature matrix based on network centrality and use spectral clustering method to identify key node groups in the network reconstruction model. First, construct a Laplacian matrix according to the adjacency matrix, and then perform eigen-decomposition on it. By analyzing the spectral gap, determine the optimal number of clusters. Finally, apply the spectral clustering algorithm to group the network nodes, so as to identify the set of nodes that play key roles in the overall network. The purpose of this step is to discover important node groups that play key roles in the biomarker expression regulation network from the system level.
[0068] The specific implementation of step S17 is to calculate the internal connectivity and external connectivity of the key node group in the multi-scale topological structure, determine the characteristic key points, and conduct network stability analysis. First, evaluate the internal tightness and external connectivity of the key node group at different scales as the basis for judging the importance of nodes. Then conduct perturbation sensitivity analysis to examine the impact of nodes on the properties of the overall network. Finally, calculate the multi-scale stability index of the nodes by integrating internal stability and external connection ability. By sorting the multi-scale importance of the characteristic key points, the final key point recognition result can be obtained. The purpose of this step is to ensure that the identified key nodes have an important and stable position in the network, providing a reliable reference basis for subsequent biological applications.
[0069] The following describes the specific calculation process involved in the present invention in detail.
[0070] First, the image processing related calculations of S01 - S04:
[0071] The image segmentation process in S04 uses an adaptive threshold method to determine the optimal threshold through the maximum inter-class variance criterion, which can effectively handle the situation of uneven illumination; the morphological processing combines opening operation and closing operation, which can effectively remove noise and maintain the integrity of the target area, and is specifically expressed as follows:
[0072]
[0073] In the formula, P(x, y) is the segmentation result; f(x, y) is the gray value of the original image; T is the segmentation threshold, which can be obtained by the maximum inter-class variance method:
[0074] T = argmax t {ω1(t)ω2(t)[μ1(t) - μ2(t)] 2};
[0075] In the formula, ω1(t) and ω2(t) are the pixel ratios of the foreground and background regions respectively; μ1(t) and μ2(t) are the average gray values of the foreground and background regions respectively.
[0076] The morphological processing of the cell region is specifically expressed as follows:
[0077]
[0078] Where M is the processed mask; I is the original binary image; B1 and B2 are morphological operators; represent dilation and erosion operations respectively.
[0079] The process of measuring the fluorescence intensity of the marker in S05 is described as follows:
[0080] The fluorescence intensity correction of a single pixel takes into account background noise and reference intensity, introduces the maximum reference intensity for normalization, and improves the comparability between different samples;
[0081]
[0082] Where I c (x, y) is the corrected intensity; I r (x, y) is the original intensity; I b (x, y) is the background intensity; is the maximum reference intensity; is the average background intensity.
[0083] The regional fluorescence intensity statistics adopt a weighted average method, taking into account the influence of pixel area:
[0084]
[0085] Where I m,n is the average fluorescence intensity of the nth marker in the mth target cell region.
[0086] The extraction of spatial distribution characteristics in S06 is specifically expressed as follows:
[0087] The spatial adjacency matrix defines the connection relationship between cells through the Euclidean distance threshold and can reflect the local spatial structure:
[0088]
[0089] Where A i,j is the adjacency relationship; d i,j is the Euclidean distance between cell regions i and j; d threshold is the adjacency threshold.
[0090] The spatial distribution entropy measures the uniformity of cell distribution and can identify aggregation or dispersion patterns:
[0091]
[0092] Where p i,j is the probability of the standardized spatial distance distribution; N is the number of cell regions.
[0093] The co-expression analysis of markers in S07 is described as follows:
[0094] The local co-expression index considers the expression correlation within the spatial neighborhood and can capture the local co-expression pattern:
[0095]
[0096] In the formula, L i,j (k) is the local co-expression index of markers i and j within the neighborhood of region k; N k is the neighborhood set of region k.
[0097] The global co-expression matrix integrates all local information through weighted averaging and reflects the overall co-expression law:
[0098]
[0099] In the formula, G i,j is the global co-expression intensity of markers i and j; w k is the weight coefficient of region k.
[0100] The processing of time-series data in S08 is specifically represented as follows:
[0101] The time-series preprocessing adopts the local standardization method and introduces a noise compensation term, enhancing the stability of the data:
[0102]
[0103] In the formula, X t is the time-series data after standardization; x t is the original data; μ t , σ t are the local mean and standard deviation respectively; η t is the noise compensation term.
[0104] The time-series correlation analysis can discover the periodic expression pattern:
[0105]
[0106] In the formula, R(τ) is the autocorrelation function at time delay τ; N is the time-series length.
[0107] The multi-dimensional data fusion in S09 is specifically represented as follows:
[0108] The eigenvector design integrates the information of four dimensions: intensity, space, co-expression, and time-series:
[0109] V i =[I i , S i, G i , T i T ;
[0110] In the formula, V i is the feature vector of the i-th region; I i , S i , G i , T i are intensity, space, co-expression, and time-series features, respectively.
[0111] Feature normalization:
[0112]
[0113] In the formula, is the normalized feature vector; V min , V max are the minimum and maximum values of each dimension, respectively.
[0114] The wavelet decomposition process in S10 is described as follows:
[0115] The multi-dimensional wavelet transform realizes time-frequency joint analysis and can extract multi-scale features:
[0116]
[0117] In the formula, W j,k,l is the two-dimensional wavelet coefficient; V n,m is the input data; ψ j,k (n) and φ j,l (m) are the wavelet function and the scaling function, respectively. The energy distribution calculation and frequency band division can distinguish the contributions of different frequency components:
[0118]
[0119] In the formula, E j is the energy of the j-th scale; K, L are the number of coefficients at the corresponding scale.
[0120] The multi-scale component division in S11 is specifically expressed as follows:
[0121] Energy ratio threshold:
[0122]
[0123] In the formula, θ j is the energy proportion of the j-th scale; J is the total number of scales.
[0124] Frequency band division determination:
[0125]
[0126] In the formula, F j is the frequency band type; τ h , τ l are the high and low frequency thresholds, with ranges of [0.6, 0.8] and [0.2, 0.4] respectively.
[0127] The frequency domain correlation analysis in S12 is described as follows:
[0128] The frequency domain coherence analysis considers the interaction between different frequency bands:
[0129]
[0130] In the formula, C i,j (f) is the coherence coefficient at frequency f; S xy (f) is the cross-spectral density; S xx (f), S yy (f) are the power spectral densities.
[0131] The construction of the correlation matrix adopts a block matrix form, which is convenient for subsequent network reconstruction:
[0132]
[0133] In the formula, R ij is the correlation degree between different frequency bands.
[0134] The multi-scale weight calculation in S13 is described as follows: The weight calculation comprehensively considers the local correlation degree, global energy distribution and frequency characteristics, and introduces an exponential decay term to smooth the frequency transition:
[0135] Local weight:
[0136]
[0137] In the formula, is the local weight based on the correlation degree.
[0138] Global weight:
[0139]
[0140] In the formula, is the global weight based on energy.
[0141] Comprehensive weight:
[0142]
[0143] In the formula, γ1, γ2, γ3 are weight coefficients; f i is the center frequency; f0 is the reference frequency; λ is the attenuation coefficient.
[0144] The network reconstruction in S14 is specifically represented as follows:
[0145] The topological connection strength combines node weights, association degrees, and spatial distances, comprehensively describing the interaction between nodes:
[0146]
[0147] In the formula, S ij is the connection strength between nodes i and j; d ij is the node distance; α and β are adjustment parameters.
[0148] The multi-scale adjacency matrix realizes the network representation at different scales:
[0149]
[0150] In the formula, A m is the adjacency matrix at the m-th scale; H(·) is the step function; θ m is the connection threshold.
[0151] The centrality calculation in S15 is described as follows:
[0152] The eigenvector centrality reflects the importance of a node in the overall network:
[0153] λv = Av;
[0154] In the formula, λ is the eigenvalue; v is the eigenvector; A is the adjacency matrix.
[0155] The improved betweenness centrality takes into account the influence of path length:
[0156]
[0157] In the formula, B i is the betweenness centrality of node i; d st is the shortest path length; D is the network diameter; δ is the distance weight.
[0158] The spectral clustering in S16 is specifically represented as follows:
[0159] Based on the eigen-decomposition of the Laplacian matrix, the community structure of the network can be effectively discovered:
[0160] L = D - W;
[0161] In the formula, L is the Laplacian matrix; D is the degree matrix; W is the weight matrix.
[0162] Eigen-decomposition:
[0163] Lv k = λ k v k, k = 1, 2, …, K;
[0164] where, v k is the k-th eigenvector, and the selection of eigenvectors takes into account the spectral gap; K is the number of clusters.
[0165] The stability analysis in S17 is described as follows:
[0166] The perturbation sensitivity analysis evaluates the impact of nodes on the overall properties of the network:
[0167]
[0168] where, S i is the sensitivity of node i; F j is the j-th property of the network; x i is the state of node i.
[0169] The stability index balances the internal stability and external connection ability of nodes:
[0170]
[0171] where, SI i is the stability index of node i; β is the sensitivity weight; α is the external connectivity weight.
[0172] The key points for determining the characteristics of traditional disease biomarker expression network reconstruction mainly include the following aspects:
[0173] 1. Traditional biomarker expression data collection methods:
[0174] Mainly rely on single immunohistochemical staining or immunofluorescence staining, and can only obtain static biomarker expression intensity information; there is a lack of standardized calibration methods in the image acquisition process, resulting in poor data comparability between different batches and different samples; image segmentation mainly uses fixed threshold methods, which are difficult to adapt to complex imaging conditions.
[0175] 2. Traditional expression feature extraction methods:
[0176] Only focus on the expression intensity of a single biomarker, ignoring the interaction relationship between biomarkers; the spatial distribution characteristics mainly rely on simple density statistics and cannot describe complex spatial structure patterns; there is a lack of collection and analysis of temporal expression data and cannot reflect the dynamic change law of biomarker expression.
[0177] 3. Traditional network reconstruction methods:
[0178] The co-expression network is mainly constructed based on the correlation coefficient, without considering the multi-scale characteristics of biomarker expression; the network connections are mainly determined according to a single threshold, lacking the distinction of different intensity connections; the network topology is relatively simple and it is difficult to reflect complex biological regulatory relationships.
[0179] 4. Traditional key point recognition methods:
[0180] The importance of nodes is mainly evaluated based on a single network centrality index (such as degree centrality, betweenness centrality, etc.); the comprehensive consideration of multi-dimensional features of nodes is lacking; the importance differences of nodes at different scales are not considered.
[0181] In contrast, the method of the present invention has the following remarkable advantages and effects:
[0182] 1. In terms of data collection and preprocessing:
[0183] (1) A standardized fluorescence intensity correction method is established, introducing reference intensity and background correction, improving the comparability of data;
[0184] (2) An adaptive threshold segmentation method is adopted, combined with morphological processing, improving the accuracy of target area extraction;
[0185] (3) The collaborative analysis of multi-channel fluorescence images is realized, and the expression information of multiple biomarkers can be obtained simultaneously.
[0186] 2. In terms of feature extraction:
[0187] (1) A multi-dimensional feature fusion framework is proposed, integrating the expression intensity, spatial distribution, co-expression relationship and temporal variation characteristics of biomarkers;
[0188] (2) The local-global co-expression analysis method is developed, which can discover the expression regulation patterns at different spatial scales;
[0189] (3) Temporal data analysis is introduced to reveal the dynamic evolution law of biomarker expression.
[0190] 3. In terms of network reconstruction:
[0191] (1) A multi-scale analysis method based on wavelet decomposition is established, which can distinguish the contributions of different frequency components;
[0192] (2) A comprehensive weight calculation method is designed to balance the influences of local correlation degree, global energy distribution and frequency characteristics;
[0193] (3) A multi-scale network representation is constructed, which can reflect the interactions of biomarkers at different intensity levels.
[0194] 4. In terms of key point recognition:
[0195] (1) A multi-dimensional node importance evaluation system is developed, comprehensively considering connection centrality, path centrality, and structural centrality;
[0196] (2) The spectral clustering method is introduced to identify key node groups, improving the systematicness of key point screening;
[0197] (3) The stability analysis link is added to ensure the reliability of the identification results.
[0198] A specific embodiment 1 of the present invention is provided below. The specific implementation manners of each step in this embodiment 1 are described in detail as follows: The specific implementation manner of step S01 is to perform preparation processing on the tissue sample to be measured. First, the tissue sample to be measured is sectioned. The purpose of sectioning is to cut the tissue sample into thin slices to facilitate subsequent processing and observation. Then, fixation processing is performed. Fixation processing is to use chemical reagents to fix and keep the cell structure intact to prevent deformation or damage during subsequent operations. Commonly used fixatives include formaldehyde, ethanol, etc. Next is dehydration processing. Dehydration is to remove the water in the section to prepare for subsequent embedding. Usually, gradient ethanol solutions are used for dehydration. Finally, embedding processing is performed. Embedding is to embed the dehydrated section into a specific supporting material, such as paraffin, resin, etc., to facilitate subsequent cutting and staining. This series of preparation processes aims to ensure that the tissue sample to be measured can be effectively labeled by subsequent immunohistochemical staining methods.
[0199] The specific implementation manner of step S02 is to perform marker staining processing on the tissue sample to be measured using immunohistochemical staining. First, the prepared tissue section is incubated with a specific antibody solution. The specific antibody mentioned here refers to an antibody that can specifically bind to the marker molecule to be measured. Through the immune reaction, the antibody will bind to the target marker molecule in the tissue. Then, a secondary antibody labeled with a fluorescent substance or an enzyme is added. The secondary antibody will form a complex with the previously bound primary antibody. Finally, through microscopic observation or enzymatic color reaction, the spatial distribution information of the marker in the tissue can be obtained. Immunohistochemical staining uses the specificity of the antigen-antibody reaction to effectively label the distribution of the target marker in the tissue section.
[0200] The specific implementation of step S03 is to collect the fluorescence image of the tissue sample to be tested through a fluorescence microscope. First, place the stained tissue section under the fluorescence microscope for imaging. The fluorescence microscope can excite the fluorescent labeling substances in the section and capture the emitted fluorescence signals. By setting different fluorescence channels, such as the FITC channel, Texas Red channel, etc., the fluorescence images of different types of markers can be obtained respectively. For example, the FITC channel is used to detect green fluorescence, and the Texas Red channel is used to detect red fluorescence. Through multi-channel fluorescence imaging, fluorescence image data containing information of multiple markers can be obtained. The purpose of this step is to provide basic data for subsequent image analysis and quantitative expression of markers.
[0201] The specific implementation of step S04 is to perform image segmentation processing on the obtained fluorescence image to extract the target cell region. First, use the adaptive threshold segmentation method for image binarization. The adaptive threshold segmentation method automatically determines the optimal segmentation threshold T according to the statistical characteristics of different regions in the image:
[0202] T = argmax t {ω1(t)ω2(t)[μ1(t) - μ2(t)] 2};
[0203] In the formula, ω1(t) and ω2(t) are the pixel ratios of the foreground and background regions respectively, and μ1(t) and μ2(t) are the average gray values of the foreground and background regions respectively. This adaptive threshold method can well handle the problem of uneven brightness in different regions. Then, combined with morphological processing means such as opening operation and closing operation, remove the noise interference in the image and retain the complete target cell region:
[0204]
[0205] In the formula, M is the processed mask, I is the original binary image, B1 and B2 are morphological operators, and represent dilation and erosion operations respectively. The purpose of this step is to accurately extract the cell region to be analyzed from the complex fluorescence image.
[0206] The specific implementation of step S05 is to measure the fluorescence intensity of the marker within the target cell region to obtain the marker expression intensity data. First, correct the fluorescence intensity of each pixel point. The correction takes into account the influence of background noise and reference intensity, and uses the maximum reference intensity for normalization:
[0207]
[0208] In the formula, I c (x,y) is the corrected intensity, Ir (x, y) is the original intensity, I b (x, y) is the background intensity, is the maximum reference intensity, is the average background intensity. Then, the average fluorescence intensity of the markers within the entire target cell region is statistically calculated using the weighted average method:
[0209]
[0210] In the formula, I m,n is the average fluorescence intensity of the nth marker in the mth target cell region, M i,j is the segmentation mask. The purpose of this step is to obtain the quantitative expression levels of each marker within the target cell region.
[0211] The specific implementation of step S06 is to record the spatial position of the target cell region in the fluorescence image to obtain the spatial distribution data of the markers. First, construct the spatial adjacency matrix A between cell regions:
[0212]
[0213] In the formula, A i,j indicates whether the ith and jth cell regions are adjacent, d i,j is the Euclidean distance between the two regions, d threshold is the adjacency threshold. Then, calculate the spatial distribution entropy H s to describe the distribution uniformity of the entire cell population:
[0214]
[0215] In the formula, p i,j is the normalized spatial distance distribution probability, and N is the number of cell regions. Through the two indicators of spatial adjacency relationship and distribution entropy, the distribution characteristics of target cells in the tissue space can be characterized.
[0216] The specific implementation of step S07 is to analyze the expression patterns of different markers within the same target cell region to obtain the co-expression data of the markers. First, calculate the local co-expression index L i,j (k):
[0217]
[0218] In the formula, L i,j (k) represents the local co-expression index of markers i and j in region k, N k is the neighborhood set of region k, I m,i and I m,jThey are the expression intensities of marker i and marker j in the m-th region respectively. Then, the global co-expression matrix G is calculated by weighted average:
[0219]
[0220] In the formula, w k is the weight coefficient of region k. The local co-expression index and the global co-expression matrix jointly characterize the co-expression relationship between markers.
[0221] The specific implementation of step S08 is to collect marker expression intensity data according to the time interval to obtain marker time series data. First, the original time series data is standardized:
[0222]
[0223] In the formula, X t is the standardized time series data, x t is the original data, μ t and σ t are the local mean and standard deviation respectively, and η t is the noise compensation term. Then, the autocorrelation function R(τ) of the time series is calculated:
[0224]
[0225] In the formula, τ is the time delay and N is the time series length. Autocorrelation analysis helps to discover the periodic change law in marker expression. Through the preprocessing and correlation analysis of the time series data, the dynamic change characteristics of marker expression can be obtained.
[0226] The specific implementation of step S09 is to combine the marker expression intensity data, marker spatial distribution data, marker co-expression data, and marker time series data to form marker multi-dimensional data. First, a feature vector V i :
[0227] V i =[I i ,S i ,G i ,T i T ;
[0228] In the formula, I i is the marker expression intensity feature, S i is the spatial distribution feature, G i is the co-expression feature, and T i is the time series feature. Then, the feature data of each dimension is normalized:
[0229]
[0230] In the formula, V min and V max are respectively the minimum and maximum values of the feature in each dimension. Through the design and normalization of the feature vectors, multiple expression features are fused into a unified multi-dimensional data representation.
[0231] The specific implementation of step S10 is to perform wavelet decomposition on the multi-dimensional data of the marker to obtain the multi-scale coefficients of the marker. First, the two-dimensional discrete wavelet transform is used to decompose the input multi-dimensional feature data V n,m :
[0232]
[0233] In the formula, W j,k,l is the wavelet coefficient at the j-th scale, ψ j,k (n) and φ j,l (m) are the wavelet function and the scaling function respectively. Then, calculate the energy E j of the wavelet coefficient at each scale:
[0234]
[0235] In the formula, K and L are the number of coefficients at the corresponding scale. Wavelet decomposition can extract multi-scale feature information, laying a foundation for subsequent frequency band division and network reconstruction.
[0236] The specific implementation of step S11 is to divide the multi-dimensional data of the marker into a reference component, a high-frequency component, an intermediate-frequency component, and a low-frequency component according to the energy distribution of the multi-scale coefficients of the marker. First, calculate the energy ratio θ j of each scale:
[0237]
[0238] In the formula, J is the total number of scales. Then, set the frequency band division thresholds τ h and τ l , τ h ranges from [0.6, 0.8], τ l ranges from [0.2, 0.4], and determine the frequency band type F j :
[0239]
[0240] In this way, the original multi-dimensional features can be divided into four different frequency band components: reference, high-frequency, intermediate-frequency, and low-frequency.
[0241] The specific implementation of step S12 is to calculate the frequency-domain correlation degrees between the reference component and the high-frequency, medium-frequency, and low-frequency components, and construct a frequency-domain correlation matrix R. First, the correlation coefficient C between different frequency bands is calculated using frequency-domain coherence analysis i,j (f):
[0242]
[0243] In the formula, S xy (f) is the cross-spectral density, and S xx (f) and S yy (f) are the power spectral densities respectively. Then, these frequency-domain correlation coefficients are organized in the form of a block matrix:
[0244]
[0245] The frequency-domain correlation matrix R depicts the interaction relationship between the characteristics of different frequency bands.
[0246] The specific implementation of step S13 is to calculate the multi-scale weights of network nodes according to the frequency-domain correlation matrix. First, the local weight of a node is calculated based on the local correlation degree
[0247]
[0248] Then, the global weight of a node is calculated based on the energy distribution of wavelet coefficients
[0249]
[0250] Finally, the local weight, global weight, and frequency characteristics are weighted and combined to obtain the comprehensive multi-scale weight w i :
[0251]
[0252] In the formula, γ1, γ2, and γ3 are weight coefficients, f i is the central frequency of the node, f0 is the reference frequency, and λ is the attenuation coefficient. The multi-scale weight can balance the contributions of different characteristics to the importance of the node.
[0253] The specific implementation of step S14 is to construct the multi-scale topological structure of the biomarker expression network according to the frequency-domain correlation matrix and the multi-scale weights. First, the connection strength S between nodes is calculated ij :
[0254]
[0255] In the formula, d ijis the distance between nodes, and α and β are adjustment parameters. The connection strength integrates node weight, correlation degree, and spatial distance information. Then, according to different strength thresholds θ m construct a multi-scale adjacency matrix A m :
[0256]
[0257] where H(·) is the step function. In this way, a multi-scale network representation reflecting connections of different strengths can be obtained.
[0258] The specific implementation of step S15 is to calculate the network centrality of each node in the network reconstruction model. First, use eigenvector centrality to analyze the importance of nodes in the overall network:
[0259] λv = Av;
[0260] where λ is the eigenvalue, v is the eigenvector, and A is the adjacency matrix. Then, introduce the improved betweenness centrality B i :
[0261]
[0262] where σ st (i) is the number of shortest paths passing through node i, σ st is the total number of shortest paths, d st is the shortest path length, D is the network diameter, and δ is the distance weight. Eigenvector centrality and betweenness centrality evaluate the importance of nodes in the network from different perspectives.
[0263] The specific implementation of step S16 is to construct a node feature matrix based on network centrality and use the spectral clustering method to identify key node groups in the network reconstruction model. First, construct the Laplacian matrix L according to the adjacency matrix:
[0264] L = D - W;
[0265] where D is the degree matrix and W is the weight matrix. Then, perform eigenvalue decomposition on the Laplacian matrix:
[0266] Lv k = λ k v k , k = 1, 2, …, K;
[0267] where v k is the k-th eigenvector and K is the number of clusters. Determine the optimal number of clusters by analyzing the spectral gap. Finally, apply the spectral clustering algorithm to divide the network nodes into different key node groups. This step can discover the node groups that play key roles in the overall network from the system level.
[0268] The specific implementation of step S17 is to calculate the internal connectivity and external connectivity of the key node group in the multi-scale topological structure, determine the characteristic key points, and perform network stability analysis. First, evaluate the internal connectivity CI i and the external connectivity CE i . Then, conduct perturbation sensitivity analysis S i :
[0269]
[0270] In the formula, F j is the j-th property of the network, and x i is the state of node i. Finally, comprehensively considering the internal stability and external connectivity ability, calculate the multi-scale stability index SI i :
[0271]
[0272] In the formula, β and α are weight coefficients. Through stability analysis, the characteristic key points with important and stable effects in the network can be determined.
[0273] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Tables 1 and 2 below.
[0274] Table 1 Variable Explanation Table (First Part)
[0275]
[0276] Table 2 Variable Explanation Table (Second Part)
[0277]
[0278]
[0279] Next, an example of a specific application scenario of the present invention is provided: A medical research team is conducting in-depth research on the molecular mechanism of lung cancer. The researchers collected 20 lung cancer tissue specimens and 20 normal lung tissue specimens, and hoped to find out the key regulatory molecules by analyzing the expression characteristics of lung cancer markers in these samples, so as to provide a basis for subsequent diagnosis and treatment.
[0280] First, the team prepared the collected tissue specimens (step S01). The researchers sliced the samples, fixed them with formaldehyde solution, then dehydrated them through an ethanol gradient, and finally embedded them in paraffin blocks. The sections processed in this way can maintain the integrity of the cell morphological structure and create good conditions for subsequent labeling and observation.
[0281] Next, the research team used immunohistochemical staining (Step S02) to stain the prepared tissue sections for markers. The researchers selected 5 biomarkers closely related to lung cancer, including epidermal growth factor receptor (EGFR), cytokeratin 19 (CK19), Ki-67, p53, and VEGF. These markers all play important regulatory roles in the occurrence and development of lung cancer. The researchers first incubated the sections with the corresponding primary antibody solution to specifically bind the antibody to the marker molecules, and then added the secondary antibody labeled with a fluorescent dye. By observing under a fluorescence microscope, the distribution of these markers in the tissue sections could be obtained.
[0282] After completing the immunofluorescence staining of the tissue sections, the research team used a fluorescence microscope (Step S03) to perform multi-channel imaging on the samples. The researchers set up the FITC channel (green fluorescence) to detect CK19 and Ki-67, the Texas Red channel (red fluorescence) to detect EGFR and p53, and the Cy5 channel (dark red fluorescence) to detect VEGF. Through multi-channel imaging, the fluorescence signals of these 5 markers in the same tissue section could be obtained simultaneously. To eliminate environmental interference during imaging, the researchers also performed background correction and brightness adjustment on the images.
[0283] Next, the research team processed and analyzed the obtained fluorescence images (Step S04). First, the researchers used the adaptive threshold segmentation method to perform binary processing on the images. This method automatically determines the optimal threshold based on local statistical features and can well segment the cell regions in different brightness areas. Then, the researchers combined morphological processing means such as opening operation and closing operation to remove noise interference in the images and retain the complete target cell regions. Through this series of image preprocessing, the research team finally obtained a clear cell region mask, laying a foundation for subsequent expression analysis.
[0284] Next, the research team measured the fluorescence intensity of each marker in the target cell region (Step S05). To correct for background noise and brightness differences between different sections, the researchers introduced the maximum reference intensity for normalization:
[0285]
[0286] where I c (x,y) is the corrected fluorescence intensity of a single pixel, I r (x,y) is the original intensity, I b (x,y) is the background intensity, is the maximum reference intensity, is the average background intensity. Then, the researchers used the weighted average method to statistically calculate the average fluorescence intensity of each marker in the entire target cell region:
[0287]
[0288] Among them, I m,n is the average fluorescence intensity of the nth biomarker in the mth target cell region, and M i,j is the segmentation mask. The quantitative expression data of the 5 biomarkers obtained in this way lay the foundation for subsequent network construction and key node identification.
[0289] As Figure 2 shown, the normalized fluorescence intensity comparison of 5 key biomarkers (EGFR, CK19, Ki-67, p53, and VEGF) in lung cancer tissues and normal tissues is presented. In the figure, it is shown in the form of a bar chart, with orange representing lung cancer tissues and blue representing normal tissues. It can be clearly seen from the figure that the expressions of EGFR and CK19 in lung cancer tissues are significantly higher than those in normal tissues, while the expression levels of p53 and VEGF are higher in normal tissues.
[0290] The research team then analyzed the spatial distribution characteristics of the target cell regions in the images (Step S06). First, the researchers constructed a spatial adjacency matrix A between cell regions:
[0291]
[0292] where d i,j is the Euclidean distance between the ith and jth cell regions, and d threhsold is the set adjacency threshold, with a value of 50 μm. Through this definition of adjacency relationship, the local distribution structure of cells in space can be characterized. Then, the research team calculated the spatial distribution entropy H s of the entire cell population:
[0293]
[0294] where p i,j is the normalized spatial distance distribution probability, and N is the total number of cell regions. The spatial distribution entropy can reflect the uniformity of cell distribution and help identify aggregated or dispersed patterns. Through the two indicators of spatial adjacency relationship and distribution entropy, the research team has a preliminary understanding of the spatial distribution characteristics of cells in the sample sections.
[0295] As Figure 3 shown, the changing trends of the spatial distribution entropy of lung cancer tissues and normal tissues during the 72-hour observation period are depicted. The red curve represents lung cancer tissues, and the blue curve represents normal tissues. The figure shows that the overall spatial distribution entropy of lung cancer tissues is higher than that of normal tissues and exhibits more obvious periodic fluctuation characteristics.
[0296] In addition, the research team also analyzed the co-expression relationship of different markers within the same target cell region (step S07). First, the researchers calculated the local co-expression index L i,j (k):
[0297]
[0298] where N k is the neighborhood set of region k, and I m,i and I m,j are the expression intensities of marker i and marker j within the m-th region, respectively. The local co-expression index reflects the correlation of marker expressions in spatially adjacent regions. Then, the research team calculated the global co-expression matrix G by weighted averaging:
[0299]
[0300] where w k is the weight coefficient of region k. The global co-expression matrix G synthesizes all local information and represents the co-expression relationship between marker i and marker j as a whole. Through local-global co-expression analysis, the research team found that in lung cancer tissues, EGFR and CK19 showed a strong co-expression relationship, while in normal lung tissues, p53 and VEGF had a high co-expression level. This co-expression pattern between markers provided important clues for subsequent network construction.
[0301] Meanwhile, the research team also collected the expression data of these markers at different time points (step S08) in order to discover their dynamic change rules. The researchers first standardized the original time-series data:
[0302]
[0303] where X t is the standardized time-series data, x t is the original data, μ t and σ t are the local mean and standard deviation, respectively, and η t is the noise compensation term. Then, the research team calculated the autocorrelation function R(τ) of the time series:
[0304]
[0305] where τ is the time delay and N is the time-series length. Through autocorrelation analysis, the researchers found that EGFR and CK19 showed periodic expression changes in lung cancer tissues, while p53 and VEGF had relatively stable time characteristics in normal lung tissues. These time-series characteristics laid the foundation for subsequent network dynamics analysis.
[0306] Integrate the obtained expression intensity, spatial distribution, co-expression relationship, and temporal variation characteristics of the markers together (step S09), and the research team constructs a feature vector V containing four dimensions
[0307] : i :
[0308] V i = [I i , S i , G i , T i T ;
[0309] where I i is the expression intensity feature, S i is the spatial distribution feature, G i is the co-expression feature, and T i is the temporal feature. To eliminate the influence of dimensional differences, the researchers normalized the feature data of each dimension:
[0310]
[0311] where V min and V max are the minimum and maximum values of the features of each dimension respectively. Through this multi-dimensional feature fusion, the research team obtained a dataset that comprehensively characterizes the marker expression features, which serves as the basis for subsequent network reconstruction and key node identification.
[0312] Next, the research team performed wavelet decomposition on these multi-dimensional feature data (step S10) to extract feature information at different scales. The researchers used two-dimensional discrete wavelet transform:
[0313]
[0314] where W j,k,l is the wavelet coefficient at the j-th scale, and ψ j,k (n) and φ j,l (m) are the wavelet function and the scaling function respectively. Then, the research team calculated the energy E j of the wavelet coefficients at each scale:
[0315]
[0316] Among them, k and L are the number of coefficients at the corresponding scales. Through wavelet decomposition and energy calculation, the researchers found that in lung cancer tissues, the expression characteristics of CK19 and EGFR are mainly concentrated in the middle and high frequency bands, while in normal lung tissues, the characteristics of p53 and VEGF are more reflected in the low frequency band. This difference in multi-scale characteristics provides an important clue for subsequent network reconstruction analysis.
[0317] Such as Figure 4 The figure shows the energy distribution characteristics of lung cancer tissues and normal tissues at different wavelet scales. The red curve represents lung cancer tissues, and the blue curve represents normal tissues. The figure shows that lung cancer tissues have a higher energy proportion in the middle and high frequency scales (2-3).
[0318] Based on the multi-scale characteristics obtained by wavelet analysis, the research team divided the biomarker data into frequency bands (step S11). First, the researchers calculated the energy proportion θ j :
[0319]
[0320] where J is the total number of scales. Then, according to the energy proportion, the high and low frequency thresholds τ h and τ l were set, with values of 0.7 and 0.3 respectively, to determine the frequency band type F j :
[0321]
[0322] Through this method, the research team divided the original multi-dimensional feature data into four parts: the reference component, the high frequency component, the middle frequency component, and the low frequency component.
[0323] Next, the research team began to construct the biomarker expression regulation network (steps S12 - S14). First, the researchers calculated the coherence C i,j (f):
[0324]
[0325] where S xy (f) is the cross-spectral density, and S xx (f) and S yy (f) are the power spectral densities respectively. Through frequency domain coherence analysis, the research team found that in lung cancer tissues, there is a strong correlation between EGFR and CK19 in the high frequency band, while in normal lung tissues, there is a significant correlation between p53 and VEGF in the low frequency band.
[0326] Based on the frequency domain coherence, the research team constructed the frequency domain correlation matrix R:
[0327]
[0328] Among them, R ij represents the correlation degree between different frequency bands.
[0329] Next, the research team calculated the multi-scale weights of the network nodes (step S13). First, the local weights of the nodes were calculated based on the local correlation degree
[0330]
[0331] Then, the global weights of the nodes were calculated based on the energy distribution of the wavelet coefficients
[0332]
[0333] Finally, the local weights, global weights, and frequency characteristics were weighted and combined to obtain the comprehensive multi-scale weight w i :
[0334]
[0335] Among them, γ1 = 0.4, γ2 = 0.4, γ3 = 0.2, λ = 0.05, f0 = 0.5.
[0336] Based on the frequency domain correlation matrix R and the multi-scale weight w i , the research team constructed the multi-scale topological structure of the biomarker expression regulation network (step S14). First, the researchers calculated the connection strength S between nodes ij :
[0337]
[0338] Among them, α = 0.5, β = 0.02, d ij is the Euclidean distance between nodes i and j. Then, according to different intensity thresholds θ m (θ1 = 0.8, θ2 = 0.6, θ3 = 0.4, θ4 = 0.2), four-scale network adjacency matrices A were constructed m :
[0339]
[0340] Among them, H(·) is the step function. Through this multi-scale network representation, the research team can more precisely characterize the interactions of biomarkers at different intensity levels.
[0341] After constructing the multi-scale network model, the research team began to identify key marker nodes (Steps S15 - S17). First, the researchers calculated the eigenvector centrality and betweenness centrality of each node in the overall network:
[0342] λv = Av;
[0343]
[0344] where δ = 0.1. These two centrality metrics reflect the importance of nodes in the network from different perspectives.
[0345] Then, the research team constructed a node feature matrix based on network centrality and used spectral clustering method (Step S16) to identify 3 key node groups:
[0346] EGFR and CK19;
[0347] p53 and VEGF;
[0348] Ki-67.
[0349] Finally, the researchers calculated the internal connectivity CI i and external connectivity CE i in the multi-scale topological structure, and carried out perturbation sensitivity analysis S i , and comprehensively evaluated the multi-scale stability index SI i (Step S17):
[0350]
[0351] where β = 0.5, α = 0.5. Through this series of analyses, the research team finally determined 3 node groups that play key roles in the entire marker regulatory network: EGFR / CK19, p53 / VEGF, and Ki-67. These key nodes provide important clues for the discovery of subsequent diagnostic and therapeutic targets.
[0352] As Figure 5 shown, the comparison of the multi-scale stability indices of the three key node groups (EGFR / CK19, p53 / VEGF, and Ki-67) in lung cancer tissues and normal tissues is presented. Orange represents lung cancer tissues and blue represents normal tissues. The figure shows that EGFR / CK19 has the highest stability in lung cancer tissues.
[0353] Generally speaking, this embodiment demonstrates the application of the method of the present invention in actual biomedical research. The research team adopted technologies such as multi-dimensional feature fusion, multi-scale network reconstruction, and comprehensive node evaluation proposed by the present invention, and successfully identified biomarkers that play a key role in the occurrence and development of lung cancer and their regulatory network structure. Compared with traditional single-index analysis and simple network construction methods, the method of the present invention can more comprehensively and accurately depict the complexity of molecular regulation, providing strong support for the exploration of new targets for disease diagnosis and treatment.
[0354] Technical principle
[0355] The method of the present invention can effectively solve the core problems existing in the prior art, mainly due to the following innovations and breakthroughs:
[0356] 1. Multi-dimensional feature fusion: The present invention designs a fusion framework that includes various biological features such as expression intensity, spatial distribution, co-expression relationship, and temporal variation. The comprehensive description of such multi-dimensional features is closer to the actual molecular regulation mechanism, laying a solid foundation for subsequent network construction and key point identification. Especially the introduction of local-global correlation evaluation in co-expression analysis can reflect the co-expression patterns at different scales, which has been significantly improved compared with single correlation analysis. In addition, the introduction of temporal data also enables the dynamic change characteristics of biomarker expression to be reflected.
[0357] 2. Multi-scale network reconstruction: The multi-scale network reconstruction method based on wavelet analysis can distinguish the contributions of different frequency components to biological regulation. This multi-scale network representation can more finely depict the interaction relationships between biomolecules compared with the existing single-threshold methods. When calculating the node weights, the present invention comprehensively considers local correlation degree, global energy distribution, and frequency characteristics, making the importance evaluation of each node more in line with biological reality. In addition, this network construction method based on multi-scale features provides richer topological information for subsequent key node identification.
[0358] 3. Comprehensive node evaluation: The present invention develops a multi-dimensional key node identification system, which not only considers the importance of nodes in the overall network, such as connection centrality and betweenness centrality, but also introduces the structural centrality of nodes at different scales. This comprehensive evaluation method more comprehensively depicts the key role of nodes in the network. At the same time, the method of spectral clustering is used to identify key node groups, and from the overall level, it can discover the set of biomarkers that play a key role in the regulatory network. In addition, the added stability analysis link ensures that the identified key nodes have an important and stable position in the network, improving the reliability of the results.
[0359] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.
Claims
1. A method for determining key points of a disease biomarker expression network reconstruction feature, characterized in that, Including: Processing the tissue sample to be measured to obtain multi-dimensional biomarker data; Performing wavelet decomposition on the multi-dimensional biomarker data to obtain multi-scale biomarker coefficients, where the multi-scale biomarker coefficients include a reference component, a high-frequency component, a medium-frequency component, and a low-frequency component; Constructing a frequency-domain correlation matrix based on the multi-scale biomarker coefficients, where the frequency-domain correlation matrix characterizes the correlation strength between different frequency components; constructing a network reconstruction model according to the frequency-domain correlation matrix and multi-scale weights, where the multi-scale weights include a reference weight, a high-frequency weight, a medium-frequency weight, and a low-frequency weight; calculating network centrality based on the network reconstruction model, where the network centrality includes connection centrality, path centrality, and structural centrality; identifying a key node group and determining characteristic key points according to the network centrality; calculating the multi-scale importance of the characteristic key points in the multi-scale topological structure and outputting a sorting result.
2. The method for determining key feature points of disease marker expression network reconstruction according to claim 1, wherein Processing the tissue sample to be measured includes: performing preparation processing on the tissue sample to be measured, including tissue sectioning, fixation, dehydration, and embedding; performing biomarker staining processing on the tissue sample to be measured by using immunohistochemical staining.
3. The method for determining the key feature points of the reconstruction of the disease marker expression network according to claim 1, wherein Obtaining the multi-dimensional biomarker data includes: collecting a fluorescence image of the tissue sample to be measured through a fluorescence microscope, where the fluorescence image includes multiple fluorescence channels; performing image segmentation processing on the fluorescence image to extract a target cell region.
4. The method for determining the key feature points of the reconstruction of the disease marker expression network according to claim 3, characterized in that The multi-dimensional biomarker data includes: biomarker expression intensity data obtained by measuring the fluorescence intensity of biomarkers in the target cell region; biomarker spatial distribution data obtained by recording the spatial position of the target cell region in the fluorescence image; biomarker co-expression data obtained by analyzing the expression patterns of different biomarkers in the same target cell region; biomarker time-series data obtained by collecting the biomarker expression intensity data at time intervals.
5. The method for determining the key feature points of the reconstruction of the disease biomarker expression network according to claim 1, wherein Performing wavelet decomposition on the multi-dimensional biomarker data includes: dividing the multi-dimensional biomarker data into a reference component, a high-frequency component, a medium-frequency component, and a low-frequency component according to the energy distribution of the multi-scale biomarker coefficients.
6. The method for determining the key feature points of the disease marker expression network reconstruction according to claim 5, characterized in that Constructing the frequency-domain correlation matrix includes: calculating the frequency-domain correlation degrees between the reference component and the high-frequency component, the medium-frequency component, and the low-frequency component; calculating the multi-scale weights of network nodes based on the frequency-domain correlation matrix, where the multi-scale weights include a reference weight, a high-frequency weight, a medium-frequency weight, and a low-frequency weight.
7. The method for determining the key points of the disease marker expression network reconstruction feature according to claim 6, wherein Constructing the network reconstruction model includes: constructing a multi-scale topological structure of a biomarker expression network according to the frequency-domain correlation matrix and the multi-scale weights.
8. The method for determining the key feature points of the reconstruction of the disease marker expression network according to claim 1, wherein The network centrality includes connection centrality, path centrality, and structural centrality; constructing a node feature matrix based on the network centrality and using a spectral clustering method to identify key node groups in the network reconstruction model.
9. The method for determining the key feature points of the reconstruction of the disease marker expression network according to claim 1, wherein Calculating the multi-scale importance of the characteristic key points includes: calculating the internal connectivity and external connectivity of the key node group in the multi-scale topological structure; performing network stability analysis on the key point data.
10. The method for determining key points of disease marker expression network reconstruction features according to claim 9, characterized in that, Sorting the characteristic key points according to the multi-scale importance and outputting a sorting result.
Citation Information
Cited By
Product sampling detection accuracy evaluation method and system based on AI
CN120976181A
AI-based product sampling detection accuracy evaluation method and system
CN120976181B