Data Processing Method and System for Tumor Cell Detection

Through the microfluidic chip combining gene expression and physical microenvironment data, pseudo-nearest neighbor method and autoencoder technology are used to solve the information loss problem caused by dimensionality reduction in high-dimensional data in tumor cell detection, and accurate dimensionality reduction and subtype classification are achieved, improving the accuracy of biological analysis and subtype distinction ability.

CN119889463BActive Publication Date: 2025-07-25ZHEJIANG RUISHENG MEDICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510362386.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-25
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

In the prior art In tumor cell detection, high-dimensional data dimensionality reduction methods may lead to the loss of key biological information, especially the annihilation of tumor stem cell marker characteristics, and the autoencoder cannot effectively distinguish cell subtypes during compression, resulting in blurred subtype boundaries.

Method used

Single cells were captured through microfluidic chips, combined with gene expression, metabolic oscillation and physical microenvironment data, and the embedding dimension was determined by pseudo-nearest neighbor method, a protective autoencoder was constructed, the bottleneck layer dimension was dynamically adjusted, key biomarker information was extracted, and the subtype classification boundaries were constructed through feature fusion.

Benefits of technology

It realizes accurate dimensionality reduction of tumor cell detection, retains key biological information, improves the accuracy and depth of biological analysis, can accurately distinguish between aggressive and inert subtypes, and supports subsequent biological analysis and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119889463B_ABST
    Figure CN119889463B_ABST
Patent Text Reader

Abstract

The present invention discloses a data processing method and system for tumor cell detection, relating to the technical field of data processing, including: sequencing cells and monitoring cell metabolism; performing time-delay reconstruction according to the embedding dimension and delay time, mapping gene expression data to the reconstructed phase space, and calculating the chaos degree of gene expression; establishing a biomarker whitelist, inputting the chaos degree into an adaptive weight allocator, setting feature retention weights for the biomarkers in the whitelist, constructing a protective autoencoder, dynamically adjusting the dimension of the bottleneck layer, and forcibly constraining when detecting the biomarkers in the whitelist; extracting coupling features from metabolic oscillation data and physical microenvironment data, calculating the gradient sensitivity of shear stress to metabolic oscillation, performing feature fusion on the coupling features and gene expression features, combining the gradient sensitivity, and constructing a subtype classification boundary for the fused features through feature vector selection, thus realizing precise data dimensionality reduction and in-depth analysis of biological characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and specifically to a data processing method and system for tumor cell detection. Background Art

[0002] In biomedical research, especially in the fields of oncology and cell biology, the processing and analysis of high-dimensional data play a crucial role. However, in practical applications, due to the excessively high data dimension, dimensionality reduction or feature compression is often required for more efficient data storage, transmission, and analysis. However, this process may be accompanied by a significant loss of key biological information, posing challenges to biological analysis.

[0003] Traditional dimensionality reduction methods, such as principal component analysis (PCA), linear discriminant analysis (LDA), etc., although able to reduce the data dimension to a certain extent, may ignore the non-linear relationships and dynamic features in the data, resulting in the loss of key biological information. For example, in tumor research, tumor stem cell markers such as CD133 are of great significance for understanding the occurrence, development, and treatment of tumors. However, during the dimensionality reduction process, if the method is inappropriate, the expression characteristics of these key markers may be annihilated, thus weakening the biological analysis ability of the data.

[0004] As a commonly used feature compression method, autoencoders learn the low-dimensional representation of data by training a neural network. However, when autoencoders are over-compressed, key biological information may also be lost. In addition, autoencoders may not be able to effectively distinguish cell subtypes with distinct biological characteristics during the compression process, such as invasive subtypes and indolent subtypes. These subtypes may be misdisplayed and mixed in the low-dimensional space, resulting in blurred or even overlapping boundaries between subtypes, bringing great difficulties to subsequent biological analysis and understanding. Summary of the Invention

[0005] (I) Technical Problems to be Solved

[0006] Aiming at the deficiencies of the prior art, the present invention provides a data processing method and system for tumor cell detection. By capturing single cells through a microfluidic chip, combining gene expression, metabolic oscillation, and physical microenvironment data, using the pseudo-nearest neighbor method to determine the embedding dimension, constructing a protective autoencoder to retain key biomarker information, and finally fusing features to construct subtype classification boundaries, precise dimensionality reduction of the data and in-depth analysis of biological characteristics are achieved.

[0007] (II) Technical Solutions

[0008] To achieve the above objectives, the present invention is realized through the following technical solutions: A data processing method for tumor cell detection, including:

[0009] Sequence the cells, monitor cell metabolism, obtain gene expression and metabolic oscillation data, quantify the physical microenvironment of the cells, and obtain physical microenvironment data;

[0010] Determine the embedding dimension by the false nearest neighbor method, determine the delay time by the first zero crossing of the autocorrelation function, perform time-delay reconstruction according to the embedding dimension and delay time, map the gene expression data to the reconstructed phase space, and calculate the chaos degree of gene expression;

[0011] Establish a white list of biomarkers, input the chaos degree into the adaptive weight allocator, set the feature retention weights for the biomarkers in the white list, construct a protective autoencoder, dynamically adjust the dimension of the bottleneck layer, input the gene expression data into the protective autoencoder, extract gene expression features, and enforce constraints when detecting the biomarkers in the white list;

[0012] Extract coupling features from the metabolic oscillation data and physical microenvironment data, calculate the gradient sensitivity of shear stress to metabolic oscillation, perform feature fusion on the coupling features and gene expression features, combine the gradient sensitivity, and construct a subtype classification boundary for the fused features through feature vector selection.

[0013] Furthermore, obtain physical microenvironment data, including: scanning the cell surface to obtain three-dimensional stiffness matrix data of the cells, and calculating the shear stress received by the cells in the flow field based on particle image velocimetry in the microfluidic chip: , where, represents the shear stress on the cell surface, μ represents the medium viscosity, v represents the velocity, x represents the position along the flow direction.

[0014] Furthermore, map the gene expression data to the reconstructed phase space, and the reconstruction formula:

[0015] , where, F(t) represents the reconstructed gene expression data, G(t) represents the gene expression data, t represents the time, represents the delay time, m represents the embedding dimension;

[0016] In the reconstructed phase space, calculate the distance evolution model for each pair of adjacent points: , where d represents the distance, t represents the time, represents the initial distance, λ represents the slope;

[0017] Take the logarithm of the distance evolution model: , the slope of the distance evolution model of adjacent point pairs is linearly fitted, and the average value of the slopes of all adjacent point pairs is used as the chaos degree.

[0018] Furthermore, the embedding dimension is determined by the false nearest neighbor method, including:

[0019] Obtain gene expression data; select a small initial embedding dimension m; calculate the nearest neighbor point of each data point at the current embedding dimension m; increase the embedding dimension m by 1, construct a new high-dimensional space, and map the gene expression data into this high-dimensional space; in the newly constructed high-dimensional space, calculate the nearest neighbor point of each data point. If the nearest neighbor point changes after the dimension increase, determine that the nearest neighbor point is a false nearest neighbor point, and calculate the false nearest neighbor point ratio; if the change rate of the false nearest neighbor point ratio exceeds the preset ratio threshold, continue to increase the embedding dimension m until the change rate of the false nearest neighbor point ratio does not exceed the preset ratio threshold; select the embedding dimension m corresponding to when the change rate of the false nearest neighbor point ratio does not exceed the preset ratio threshold as the final embedding dimension.

[0020] Furthermore, the delay time is determined by the first zero crossing of the autocorrelation function, including:

[0021] For gene expression data, calculate its autocorrelation function: , where R( ) represents the autocorrelation function, represents the delay time, G(t) represents the gene expression data, n represents the number of data points;

[0022] Find the position of the first zero crossing on the autocorrelation function, and take the value corresponding to the position of the first zero crossing as the delay time.

[0023] Furthermore, construct an adaptive weight allocator, input the chaos degree into the adaptive weight allocator, and calculate the feature retention threshold: , where T represents the feature retention threshold, σ represents the sigmoid function, represents the chaos degree, represents the reference chaos degree, represents the initial threshold; set the feature retention weight for the markers in the whitelist: , where ω represents the feature retention weight.

[0024] Furthermore, construct a protective autoencoder, set the dimension of the bottleneck layer as an adaptive variable, input the feature retention weight and the chaos degree into the bottleneck layer, and calculate the dynamic compression rate: , where DCR represents the dynamic compression ratio; calculating the dimension of the bottleneck layer: , where represents the bottleneck layer dimension value, represents the reference dimension; inputting the gene expression data into the protective autoencoder, when the whitelist markers are detected, forcing the constraint .

[0025] Furthermore, performing a tensor product operation on the metabolic oscillation data and the physical microenvironment data to extract coupling features: , where represents the metabolic oscillation value at the i th time point, is the three-dimensional stiffness matrix, j and k represent the rows and columns of the three-dimensional stiffness matrix respectively, represents the coupling feature;

[0026] Calculating the gradient sensitivity of shear stress to metabolic oscillation: , where γ represents the gradient sensitivity, represents the cell surface shear stress, v represents the velocity, x represents the position along the flow direction, M represents the metabolic oscillation value, and t represents time.

[0027] Furthermore, performing feature fusion on the coupling features and the gene expression features, and constructing a subtype classification boundary for the fused features through feature vector selection: , where X represents feature vector selection, γ represents the gradient sensitivity, represents the maximum shear stress difference, represents the metabolic response delay time.

[0028] A data processing system for tumor cell detection, including:

[0029] A data acquisition module that sequences cells, monitors cell metabolism, obtains gene expression and metabolic oscillation data, and quantifies the cell physical microenvironment to obtain physical microenvironment data;

[0030] A gene expression analysis module that determines the embedding dimension by the pseudo nearest neighbor method, determines the delay time by the first zero crossing of the autocorrelation function, performs time delay reconstruction according to the embedding dimension and the delay time, maps the gene expression data to the reconstructed phase space, and calculates the chaos degree of the gene expression;

[0031] An auto-encoding module that establishes a whitelist of biomarkers, inputs the chaos degree into an adaptive weight allocator, sets the feature retention weights for the biomarkers in the whitelist, constructs a protective auto-encoder, dynamically adjusts the dimension of the bottleneck layer, inputs gene expression data into the protective auto-encoder, extracts gene expression features, and enforces constraints when detecting the biomarkers in the whitelist;

[0032] A feature selection module that extracts coupled features through metabolic oscillation data and physical microenvironment data, calculates the gradient sensitivity of shear stress to metabolic oscillation, performs feature fusion on the coupled features and gene expression features, and constructs a subtype classification boundary for the fused features through feature vector selection in combination with the gradient sensitivity.

[0033] (III) Beneficial effects

[0034] The present invention provides a data processing method and system for tumor cell detection, which have the following beneficial effects:

[0035] (1) Through high-resolution acquisition of gene expression data, real-time monitoring of cell metabolism, and quantification of the cell physical microenvironment, etc., it is possible to more comprehensively understand the biological characteristics of tumor cells, including information such as their gene expression, metabolic state, and physical microenvironment, which helps to reduce the loss of key biological information during the process of dimensionality reduction or feature compression, and improve the accuracy and depth of biological analysis.

[0036] (2) Through the method of pseudo nearest neighbor and the method of the first zero crossing of the autocorrelation function, it is possible to accurately determine the embedding dimension and delay time, and then accurately map the gene expression data into the reconstructed phase space, so as to more truly reflect the dynamic behavior and potential non-linear characteristics of the data. The calculation of chaos degree helps to understand the complexity and instability of gene expression.

[0037] (3) Dynamically adjust the feature retention weights through an adaptive weight allocator, and then dynamically adjust the dimension of the bottleneck layer of the auto-encoder, which can flexibly adjust the dimensionality reduction strategy according to the complexity of the data and the distribution of key information, ensure the priority retention of key biological markers such as tumor stem cells, avoid the loss of key information that may occur during the dimensionality reduction or feature compression process, input the gene expression environment data into the protective auto-encoder, and when detecting the biomarkers in the whitelist, the forced constraint further enhances the accuracy and reliability of data processing.

[0038] (4) By extracting metabolic-physical coupled features and calculating the gradient sensitivity of shear stress to metabolic oscillation, a deeper understanding of the biological characteristics of tumor cells is achieved; feature fusion and feature vector selection further improve the accuracy of subtype classification, help to accurately distinguish key subtypes such as invasive subtypes and indolent subtypes, and provide strong support for subsequent biological analysis and treatment. Description of the drawings

[0039] Figure 1 Schematic diagram of the steps of the data processing method for tumor cell detection according to the present invention;

[0040] Figure 2 Schematic diagram of the autoencoder network structure according to the present invention;

[0041] Figure 3 Schematic diagram of the structure of the data processing system for tumor cell detection according to the present invention. Detailed implementation manners

[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0043] Please refer to Figure 1 - Figure 2 , the present invention provides a data processing method for tumor cell detection, including the following steps:

[0044] Step 1: Sequence the cells and monitor cell metabolism to obtain gene expression and metabolic oscillation data, and quantify the physical microenvironment of the cells to obtain physical microenvironment data;

[0045] The said Step 1 includes the following contents:

[0046] Step 101: Use a microfluidic chip (channel width is 20μm) to capture single cells. The microfluidic chip precisely controls the fluid flow to make the cells pass through the narrow channel one by one and be captured. The capture process needs to ensure the survival rate of the cells and avoid cell damage during the capture process;

[0047] Step 102: Set the lysis buffer pulse trigger frequency, such as triggering the lysis buffer pulse once every 30 seconds (lasting for 5ms), to lyse the captured single cells, and use the SMART-seq2 technology to generate full-length cDNA to ensure that the sequencing depth of each cell reaches or exceeds 50000 reads / cell, and the time resolution reaches the minute level, so as to provide high-resolution gene expression data;

[0048] Step 103: Integrate a fluorescence resonance energy transfer (FRET) biosensor (such as iGlucose) to monitor the glucose uptake rate in cells in real time, and set the sampling frequency, such as 100Hz, to capture fast metabolic oscillation data. The collected data is denoised by Kalman filtering to improve the accuracy of the data;

[0049] Step 104: Scan the cell surface using an atomic force microscope (AFM, with a probe stiffness of 0.1 N / m) to obtain three-dimensional stiffness matrix data of the cell. Based on the particle image velocimetry (PIV) technology in the microfluidic chip, calculate the shear stress exerted on the cell in the flow field: , where represents the shear stress on the cell surface, μ represents the medium viscosity, v represents the velocity, x represents the position along the flow direction;

[0050] Step 105: Generate an encrypted timestamp to ensure that gene expression, metabolism, and physical data are aligned to microsecond-level precision in time. Perform eigenvalue analysis on the three-dimensional stiffness matrix. Eigenvalue analysis includes, but is not limited to, the definition method, the characteristic polynomial method, or the iterative method. If the eigenvalue is greater than a characteristic threshold, such as 15 kPa, trigger the weighted filtering process for gene expression data. The weighted filtering process includes, but is not limited to, linear weighted filtering, exponential weighted filtering, or weighted median filtering; the characteristic threshold is the basis for determining whether to trigger the weighted filtering process, and the characteristic threshold is determined according to cell type, experimental conditions, and data analysis requirements.

[0051] When in use, combine the content of Steps 101 to 105:

[0052] Through high-resolution acquisition of gene expression data, real-time monitoring of cell metabolism, and quantification of the cell physical microenvironment, etc., it is possible to more comprehensively understand the biological characteristics of tumor cells, including information on their gene expression, metabolic state, and physical microenvironment, etc., which helps to reduce the loss of key biological information during the process of dimensionality reduction or feature compression, and improve the accuracy and depth of biological analysis.

[0053] Step Two: Determine the embedding dimension by the false nearest neighbor method, determine the delay time by the first zero-crossing of the autocorrelation function, perform time-delay reconstruction according to the embedding dimension and the delay time, map the gene expression data to the reconstructed phase space, and calculate the chaos degree of the gene expression;

[0054] The above Step Two includes the following content:

[0055] Step 201: Determine the embedding dimension by the false nearest neighbor method, determine the delay time by the first zero-crossing of the autocorrelation function, perform time-delay reconstruction according to the embedding dimension and the delay time, map the gene expression data to the reconstructed phase space, and the reconstruction formula:

[0056] , where F(t) represents the reconstructed gene expression data, G(t) represents the gene expression data, t represents the time moment, represents the delay time,m denotes the embedding dimension;

[0057] Among them, determining the embedding dimension by the false nearest neighbor method includes:

[0058] Data preparation: Obtain gene expression data;

[0059] Initial embedding dimension selection: Select a small initial embedding dimension m, usually starting from 2 or 3;

[0060] Calculate the nearest neighbor points: Under the current embedding dimension m, calculate the nearest neighbor points of each data point, which is achieved by calculating the Euclidean distance or other distance metrics;

[0061] Construct a high-dimensional space: Increase the embedding dimension m by 1, construct a new high-dimensional space, and map the original data into this high-dimensional space;

[0062] Calculate the proportion of false nearest neighbor points: In the newly constructed high-dimensional space, for each data point, calculate its nearest neighbor point. If the nearest neighbor point changes after the dimension increase, that is, it is not the same as the nearest neighbor point in the previous dimension, then judge this nearest neighbor point as a false nearest neighbor point, and calculate the proportion of false nearest neighbor points;

[0063] Judge whether the embedding dimension is sufficient: If the change rate of the proportion of false nearest neighbor points exceeds the preset proportion threshold, it indicates that the current embedding dimension m may not be sufficient to fully capture the dynamic behavior of the data. Therefore, it is necessary to continue to increase the embedding dimension m and repeat the above steps. If the change rate of the proportion of false nearest neighbor points does not exceed the preset proportion threshold, then it can be considered that the current embedding dimension m is sufficient;

[0064] In the false nearest neighbor method, the change rate of the proportion of false nearest neighbor points refers to the change amount of the proportion of false nearest neighbor points relative to the previous embedding dimension as the embedding dimension increases. This change rate is used to judge whether the current embedding dimension is large enough to fully capture the dynamic behavior in the data. The proportion threshold is obtained by gradually increasing the embedding dimension and calculating the change rate of the proportion of false nearest neighbor points, observing the trend of the change rate as the embedding dimension increases, and gradually adjusting the threshold according to the trend and the preliminary estimated threshold range until a suitable value is found;

[0065] Determine the final embedding dimension: Select the embedding dimension m corresponding to when the change rate of the proportion of false nearest neighbor points does not exceed the preset proportion threshold as the final embedding dimension;

[0066] Determine the delay time through the first zero crossing of the autocorrelation function, including:

[0067] For gene expression data, calculate its autocorrelation function: , where, R( )denotes the autocorrelation function, denotes the delay time, G(t) denotes the gene expression data, n denotes the number of data points;

[0068] On the graph of the autocorrelation function, find the position of the first zero-crossing, and take the value corresponding to the position of the first zero-crossing as the delay time;

[0069] Step 202: In the reconstructed phase space, calculate the distance evolution model for each pair of adjacent points: , where d represents the distance, t represents the time, denotes the initial distance, λ denotes the slope;

[0070] Step 203: Take the logarithm of the distance evolution model: , and after linearly fitting the slope of the distance evolution model of adjacent point pairs, take the average value of the slopes of all adjacent point pairs as the chaos degree ;

[0071] The chaos degree is obtained by calculating the average value of the slopes of the distance evolution models of adjacent point pairs (i.e., λ ), and this average value provides an overall measure of the chaotic behavior in the entire dataset. It should be noted that the chaos degree is a relative concept, which reflects the degree of sensitivity of the system to the initial conditions;

[0072] When in use, combine the content of Step 201 to Step 203:

[0073] Through the method of pseudo nearest neighbors and the method of the first zero-crossing of the autocorrelation function, the embedding dimension and the delay time can be accurately determined, and then the gene expression data can be accurately mapped into the reconstructed phase space, so as to more truly reflect the dynamic behavior and potential non-linear characteristics of the data. The calculation of the chaos degree helps to understand the complexity and instability of gene expression.

[0074] Step Three: Establish a biomarker whitelist, input the chaos degree into the adaptive weight allocator, set the feature retention weights for the biomarkers in the whitelist, construct a protective autoencoder, dynamically adjust the dimension of the bottleneck layer, input the gene expression data into the protective autoencoder, extract gene expression features, and enforce constraints when detecting the biomarkers in the whitelist;

[0075] The said Step Three includes the following content:

[0076] Step 301: Establish a biomarker whitelist database (including tumor stem cell markers such as CD133), extract gene expression features using a bidirectional long short-term memory (LSTM) network, set the number of memory units to 128, and the sliding step size of the time window to 5 minutes;

[0077] Step 302: Construct an adaptive weight allocator, input the chaos degree into the adaptive weight allocator, and calculate the feature retention threshold: , where T represents the feature retention threshold, σ represents the sigmoid function, represents the chaos degree, represents the reference chaos degree, represents the initial threshold. The reference chaos degree is a preset value used to standardize the chaos degree, and the initial threshold is a parameter adjusted according to experimental data and requirements; Set the feature retention weight for the markers in the whitelist: , where ω represents the feature retention weight;

[0078] Step 303: Construct a protective autoencoder network, set the dimension of the bottleneck layer as an adaptive variable, input the feature retention weight and the chaos degree into the bottleneck layer control module, and calculate the dynamic compression rate: , where DCR represents the dynamic compression rate; Calculate the dimension of the bottleneck layer: , where represents the value of the bottleneck layer dimension, represents the reference dimension. The reference dimension is a preset dimension reference value (such as 32 dimensions);

[0079] Step 304: Input the gene expression data into the protective autoencoder. When a marker in the whitelist is detected, enforce the constraint ;

[0080] When in use, combine the content of Steps 301 to 304:

[0081] Dynamically adjust the feature retention weight through the adaptive weight allocator, and then dynamically adjust the dimension of the bottleneck layer of the autoencoder, which can flexibly adjust the dimensionality reduction strategy according to the complexity of the data and the distribution of key information, ensure the priority retention of key biological markers such as cancer stem cells, avoid the loss of key information that may occur during dimensionality reduction or feature compression, input the gene expression environment data into the protective autoencoder, and when a marker in the whitelist is detected, the enforcement of the constraint further enhances the accuracy and reliability of data processing.

[0082] Step Four: Extract coupling features from the metabolic oscillation data and the physical microenvironment data, calculate the gradient sensitivity of shear stress to metabolic oscillation, perform feature fusion on the coupling features and the gene expression features, and combine the gradient sensitivity to construct a subtype classification boundary for the fused features through feature vector selection.

[0083] The above Step Four includes the following content:

[0084] Step 401: Establish a metabolic-physical coupling tensor: Perform a tensor product operation on the metabolic oscillation data and the physical microenvironment data to extract metabolic-physical coupling features: , where represents the metabolic oscillation value at the i th time point, is a three-dimensional stiffness matrix, j and k represent the rows and columns of the three-dimensional stiffness matrix respectively, represents the coupling feature;

[0085] Step 402: Calculate the gradient sensitivity of shear stress to metabolic oscillation: , where γ represents the gradient sensitivity, represents the cell surface shear stress, v represents the velocity, x represents the position along the flow direction, M represents the metabolic oscillation value, and t represents time;

[0086] Step 403: Perform feature fusion on the metabolic-physical coupling features and gene expression features. The feature fusion is performed by the splicing method, and a subtype classification boundary is constructed for the fused features through feature vector selection: , where X represents feature vector selection, γ represents the gradient sensitivity, represents the maximum shear stress difference (per unit length along the flow direction), represents the metabolic response delay time (the time lag from the change in shear stress to the metabolic oscillation response);

[0087] When in use, combine the content of Step 401 to Step 402:

[0088] By extracting the metabolic-physical coupling features and calculating the gradient sensitivity of shear stress to metabolic oscillation, a deeper understanding of the biological characteristics of tumor cells is achieved; feature fusion and feature vector selection further improve the accuracy of subtype classification, which helps to accurately distinguish key subtypes such as invasive subtypes and indolent subtypes, providing strong support for subsequent biological analysis and treatment.

[0089] Please refer to Figure 3 , the present invention also provides a data processing system for tumor cell detection, including: a data acquisition module, a gene expression analysis module, an autoencoder module, and a feature selection module; where

[0090] The data acquisition module sequences cells and monitors cell metabolism to obtain gene expression and metabolic oscillation data, and quantifies the cell physical microenvironment to obtain physical microenvironment data;

[0091] The gene expression analysis module determines the embedding dimension by the pseudo nearest neighbor method, determines the delay time by the first zero crossing of the autocorrelation function, performs time-delay reconstruction according to the embedding dimension and the delay time, maps the gene expression data to the reconstructed phase space, and calculates the chaos degree of the gene expression;

[0092] The autoencoder module establishes a white list of biomarkers, inputs the chaos degree into the adaptive weight allocator, sets the feature retention weights for the biomarkers in the white list, constructs a protective autoencoder, dynamically adjusts the dimension of the bottleneck layer, inputs the gene expression data into the protective autoencoder, extracts the gene expression features, and enforces constraints when detecting the biomarkers in the white list;

[0093] The feature selection module extracts the coupling features through the metabolic oscillation data and the physical microenvironment data, calculates the gradient sensitivity of the shear stress to the metabolic oscillation, performs feature fusion on the coupling features and the gene expression features, combines the gradient sensitivity, and constructs a subtype classification boundary for the fused features through feature vector selection.

[0094] In the application, several formulas involved are dimensionless and their numerical values are calculated. The formula is obtained by software simulation of a large amount of collected data to approximate the real situation as closely as possible, and the coefficients in the formula are set by those skilled in the art according to the actual situation.

[0095] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution.

[0096] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0097] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application.

Claims

1. A data processing method for tumor cell detection, characterized in that: Including: Sequencing cells, monitoring cell metabolism, obtaining gene expression and metabolic oscillation data, quantifying the physical microenvironment of cells, and obtaining physical microenvironment data; Determining the embedding dimension by the false nearest neighbor method, determining the delay time by the first zero crossing of the autocorrelation function, performing time-delay reconstruction according to the embedding dimension and the delay time, and mapping the gene expression data to the reconstructed phase space. Reconstruction formula: , where F(t) represents the reconstructed gene expression data, G(t) represents the gene expression data, t represents the time, represents the delay time, m represents the embedding dimension; In the reconstructed phase space, calculate the distance evolution model for each pair of adjacent points: , where d represents the distance, represents the initial distance, λ represents the slope; Take the logarithm of the distance evolution model: Linear fit the slope of the distance evolution model of adjacent point pairs, and take the average value of the slopes of all adjacent point pairs as the chaos degree; Establish a biomarker whitelist, construct an adaptive weight allocator, input the chaos degree into the adaptive weight allocator, and calculate the feature retention threshold: , where T represents the feature retention threshold, σ represents the sigmoid function, represents the chaos degree, represents the reference chaos degree, represents the initial threshold; Set the feature retention weight for the biomarkers in the whitelist: , where ω represents the feature retention weight; Construct a protective autoencoder, input the feature retention weight and chaos degree into the bottleneck layer, and calculate the dynamic compression rate: , where DCR represents the dynamic compression rate; set the dimension of the bottleneck layer as an adaptive variable, and calculate the dimension of the bottleneck layer: , where represents the dimension value of the bottleneck layer, represents the reference dimension; input the gene expression data into the protective autoencoder, and when the whitelist markers are detected, enforce the constraint ; Extracting coupling features from the metabolic oscillation data and the physical microenvironment data, calculating the gradient sensitivity of shear stress to metabolic oscillation, performing feature fusion on the coupling features and the gene expression features, and combining the gradient sensitivity to construct a subtype classification boundary for the fused features by feature vector selection.

2. The data processing method for tumor cell detection according to claim 1, wherein: Obtain physical microenvironment data, including: scanning the cell surface to obtain three-dimensional stiffness matrix data of the cells, and calculating the shear stress exerted on the cells in the flow field based on particle image velocimetry in a microfluidic chip: , where represents the shear stress on the cell surface, μ represents the medium viscosity, v represents the velocity, x represents the position along the flow direction.

3. The data processing method for tumor cell detection according to claim 1, wherein: Determining the embedding dimension by the false nearest neighbor method, including: Obtaining gene expression data; selecting an initial embedding dimension m; calculating the nearest neighbor points of each data point at the current embedding dimension m; increasing the embedding dimension m by 1, constructing a new high-dimensional space, and mapping the gene expression data into this high-dimensional space; in the new high-dimensional space, calculating the nearest neighbor points of each data point. If the nearest neighbor points change after the dimension increase, determine that the nearest neighbor points are false nearest neighbor points, and calculate the false nearest neighbor point ratio; if the change rate of the false nearest neighbor point ratio exceeds a preset ratio threshold, continue to increase the embedding dimension m until the change rate of the false nearest neighbor point ratio does not exceed the preset ratio threshold; select the embedding dimension m corresponding to when the change rate of the false nearest neighbor point ratio does not exceed the preset ratio threshold as the final embedding dimension.

4. The data processing method for tumor cell detection according to claim 1, wherein: Determining the delay time by the first zero crossing of the autocorrelation function, including: For gene expression data, calculate its autocorrelation function: , where R( ) represents the autocorrelation function, n represents the number of data points; Find the position of the first zero crossing on the autocorrelation function, and use the value corresponding to the position of the first zero crossing as the delay time.

5. The data processing method for tumor cell detection according to claim 1, wherein: Perform a tensor product operation on the metabolic oscillation data and the physical microenvironment data to extract coupling features: , where represents the metabolic oscillation value at the i -th time point, is a three-dimensional stiffness matrix, j and k represent the rows and columns of the three-dimensional stiffness matrix respectively, represents the coupling feature; Calculating the gradient sensitivity of shear stress to metabolic oscillations: , where γ represents the gradient sensitivity, M represents the metabolic oscillation value, t represents time.

6. The data processing method for tumor cell detection according to claim 5, wherein: Fuse the coupling features and gene expression features, and construct a subtype classification boundary for the fused features by feature vector selection: , where X represents feature vector selection, γ represents gradient sensitivity, represents the maximum shear stress difference, represents the metabolic response delay time.

7. A data processing system for tumor cell detection, which is used to implement the method described in any one of claims 1 to 6, characterized in that: Including: A data acquisition module that sequences cells, monitors cell metabolism, obtains gene expression and metabolic oscillation data, quantifies the physical microenvironment of cells, and obtains physical microenvironment data; A gene expression analysis module that determines the embedding dimension by the false nearest neighbor method, determines the delay time by the first zero crossing of the autocorrelation function, performs time-delay reconstruction according to the embedding dimension and the delay time, maps the gene expression data to the reconstructed phase space, and calculates the chaos degree of gene expression; An autoencoder module that establishes a biomarker whitelist, inputs the chaos degree into an adaptive weight allocator, sets feature retention weights for the biomarkers in the whitelist, constructs a protective autoencoder, dynamically adjusts the dimension of the bottleneck layer, inputs the gene expression data into the protective autoencoder, extracts gene expression features, and enforces constraints when detecting the biomarkers in the whitelist; Feature selection module, which extracts coupled features from metabolic oscillation data and physical microenvironment data, calculates the gradient sensitivity of shear stress to metabolic oscillation, performs feature fusion on the coupled features and gene expression features, and combines the gradient sensitivity to construct a subtype classification boundary for the fused features through feature vector selection.

Citation Information

Patent Citations

  • Method for visualization of developmental landscapes from single-cell multimodal data

    EP4270398A1

  • Topological features and time-bandwidth signature of heart signals as biomarkers to detect deterioration of a heart

    US20200155015A1