A cell state recognition method based on multi-modal image deep learning
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHANGCHUN UNIV OF SCI & TECH
- Filing Date
- 2026-04-07
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]现有细胞状态识别方法多依赖单一显微图像特征或人工提取特征进行分析,难以同时兼顾细胞形貌信息与力学信息的联合表征,导致对细胞复杂状态差异的刻画不够全面
1、本发明通过对原子力显微镜获取的力曲线进行反演处理,构建高度图、杨氏模量图、粘附力图和参考力图,并形成统一的多通道输入数据,使细胞的形貌信息与力学信息得到协同表征,较单一图像或单一参数更有利于完整刻画细胞状态差异,从而提高细胞状态识别的准确性与稳定性。
Smart Images

Figure CN122200651B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical imaging, specifically a cell state recognition method based on multimodal image deep learning. Background Technology
[0002] Cell state identification is a crucial technological foundation for cell biology research, drug screening, disease mechanism analysis, and clinical testing. By identifying morphological changes, mechanical responses, and surface interaction characteristics of cells under different treatment conditions, it can provide a basis for assessing cell activity, determining drug efficacy, analyzing pathological states, and monitoring the treatment process.
[0003] Existing cell state recognition methods mostly rely on single microscopic image features or manually extracted features for analysis, making it difficult to simultaneously consider the joint representation of cell morphology and mechanical information, resulting in an incomplete characterization of complex cell state differences. When cells have similar appearances but different internal mechanical responses, traditional methods are prone to insufficient recognition accuracy and poor stability. At the same time, cell samples have strong uncertainties in imaging posture, local region distribution, and noise perturbation, which further affect the final discrimination results, thus failing to meet the practical needs of fine cell state recognition. Summary of the Invention
[0004] The purpose of this invention is to provide a cell state recognition method based on multimodal image deep learning to solve the problems raised in the prior art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: A cell state recognition method based on multimodal image deep learning, the method comprising the following steps: Two-dimensional data of the cell under test are obtained by atomic force microscopy. The two-dimensional data includes a height map and a force curve. A parameter map characterizing the mechanical properties of the cell is generated based on the force curve. The atomic force microscope (AFM) is a scanning detection device that characterizes surface morphology and measures mechanical properties based on the interaction force between the probe and the sample surface. It can move the nanoscale probe at the front end of the microcantilever beam close to the sample surface and control the probe to move along a preset direction at each scanning position on the sample surface. Furthermore, the process of acquiring two-dimensional data of the cells to be tested includes: The surface of the cell to be tested is scanned pixel by pixel using atomic force microscopy in quantitative imaging mode. The pixel-by-pixel scanning process includes: The probe of the atomic force microscope is controlled to move vertically toward the surface of the cell to be tested at each scanning position. After the probe moves to the target position, the probe is controlled to move in the opposite direction to the termination position, and the displacement information and force information corresponding to each scanning position during the movement are recorded. Force curves for each scanning position are constructed based on the displacement information and the force information; Specifically, the cells to be tested are fixed on the detection carrier and placed within the detection area of the atomic force microscope. Based on the size of the cells to be tested, the target resolution, and the detection accuracy requirements, the scanning area, scanning range, scanning step size, and number of scanning positions are set so that the scanning area covers the target surface area of the cells to be tested. The probe of the atomic force microscope is controlled to move sequentially to each scanning position, and a complete displacement scan is performed in the vertical direction at each scanning position to obtain the displacement information and force information corresponding to that scanning position. At a single scanning location, the probe is controlled to move vertically toward the surface of the cell being tested; As the probe continues to approach the surface of the cell to be tested, the interaction between the probe and the surface of the cell to be tested gradually changes. When the probe reaches the target position, it is controlled to move in the opposite direction to the termination position, thereby completing one round-trip detection process at that scanning position. During this round-trip detection process, the displacement changes of the probe in the vertical direction and the corresponding force changes are continuously recorded. The displacement sequence and force sequence recorded at the same scanning position are correlated to construct a force curve for that scanning position. The force curve is a data curve reflecting the relationship between the probe displacement change and the force change at a single scanning position. This force curve can characterize the changes in the mechanical response of the probe and the surface of the cell to be tested before contact, after contact, and during the separation process. Since each scanning position corresponds to a force curve, after completing pixel-by-pixel scanning in the entire scanning area, a set of force curves corresponding to all scanning positions can be obtained. When the applied force reaches a preset value, the displacement information at that moment is used as the height value of the corresponding scanning position. These height values are then arranged according to the two-dimensional coordinates of each scanning position to obtain the height map. The height map reflects the surface morphology undulations of the cell being tested. Furthermore, the process of generating a parametric map characterizing the mechanical properties of the cell based on the force curve includes: The force curve segment from the point where the probe contacts the cell to be tested is extracted from each scanning position as the approach segment; specifically, the contact point can be defined as the position in the force curve where the force information begins to deviate continuously from the baseline force. Based on the proximity segment in the force curve corresponding to each scanning position, the Young's modulus value of each scanning position is obtained by inversion through the contact mechanics model; A contact mechanics model is used to describe the relationship between force and deformation during the probe's indentation into the surface of the cell under test. Based on displacement and force change data in the adjacent segments of each scanning position, the local indentation response of the probe on the surface of the cell under test at that scanning position is determined. This local indentation response is then substituted into the preset contact mechanics model for fitting calculation to obtain the Young's modulus value corresponding to that scanning position. The Young's modulus value is used to characterize the local stiffness characteristics of the cell under test at the corresponding scanning position. A larger Young's modulus value indicates that the area corresponding to that scanning position is less prone to deformation under force; a smaller Young's modulus value indicates that the area corresponding to that scanning position is more prone to deformation. By performing inversion calculations on all scanning positions separately, a set of Young's modulus values covering the entire scanning area can be obtained. Specifically, taking the Hertzian contact model as an example, let the force information of the probe at a certain scanning position be F, the indentation depth be δ, the equivalent radius of the probe be R, and the Poisson's ratio of the cell to be measured be ν, then the Hertzian contact model can be expressed as: ; Among them, E * The equivalent elastic modulus, E, is the equivalent elastic modulus when the probe is considered a rigid body relative to the cell being tested. * The Young's modulus E satisfies: ; Substituting the above equation into the Hertzian contact model, we obtain Young's modulus E: ; The adhesion force value at each scanning position is extracted based on the proximity segment, and a reference force value is preset for each scanning position; Specifically, the adhesion force value is determined based on the mechanical response characteristics corresponding to the contact between the probe and the surface of the cell to be tested and the establishment of an interfacial interaction in the approach segment; in the approach segment of each scanning position, the characteristic force value reflecting the adsorption interaction between the probe and the surface of the cell to be tested is read, and the characteristic force value is used as the adhesion force value of that scanning position; the adhesion force value is used to characterize the strength of the adhesion response when the surface of the cell to be tested interacts with the probe at the corresponding scanning position, and the difference in adhesion force value at different scanning positions can reflect the spatial distribution difference in the interfacial interaction characteristics of the local area of the surface of the cell to be tested; Specifically, the reference force value is a parameter used to characterize the force response at each scanning position under a unified evaluation benchmark. A unified preset displacement position, a unified preset indentation depth position, or a unified preset force determination rule is used to read the force curve corresponding to each scanning position, thereby determining the reference force value for each scanning position. For example, the force value corresponding to the same indentation depth can be selected in the proximal segment of each scanning position as the reference force value for that scanning position; alternatively, the force value can be read at the same displacement position as the reference force value. Since the benchmark for reading the reference force value remains consistent across all scanning positions, the reference force values at each scanning position are comparable and can reflect the differences in force response of different local regions of the tested cell under the same evaluation conditions. The Young's modulus value, adhesion force value, and reference force value are arranged according to the pixel coordinates corresponding to each scanning position to obtain the Young's modulus map, adhesion force map, and reference force map for each scanning position. The parameter diagrams include Young's modulus diagram, adhesion force diagram, and reference force diagram.
[0006] The height map and parameter map are preprocessed, and the preprocessed height map and parameter map are spatially aligned and channel constructed to obtain the model input tensor; the model input tensor is divided into image blocks, and the image blocks are converted into embedding sequences; Furthermore, the preprocessing of the height map and parameter map includes: The height map and parameter map are respectively flattened, corrected for background, denoised and normalized to obtain preprocessed height map and preprocessed parameter map; the preprocessed height map and preprocessed parameter map have the same size.
[0007] Flattening correction is used to eliminate the overall tilt or bending trend introduced by the tilt of the scanning substrate, uneven placement of the sample, or low-frequency drift of the scanning device in the height map and parametric map, so that the base surface of the corrected image tends to be flat. Background correction is used to correct background responses in an image that are irrelevant to the intrinsic features of the cells being tested, so that the baseline values of the image return to a uniform reference level. Denoising is used to suppress isolated noise points, high-frequency fluctuations, and local artifacts in images caused by random disturbances of the detection system, environmental vibrations, electrical signal fluctuations, or scanning errors, in order to preserve effective structural information related to the state of the cells under test. Normalization refers to mapping pixel values in images of different modalities to a uniform numerical scale range, making height maps and parametric maps comparable in numerical distribution. Since height maps and parametric maps represent different physical quantities, their original value ranges and dimensions are usually different. If channel stitching is performed directly, a certain modality may dominate the model learning process due to its larger numerical range. Normalization can eliminate dimensional and amplitude differences between different modalities, making each modal image contribute more evenly to subsequent joint input, which is beneficial for the model to learn cell morphology and mechanical features simultaneously. Specifically, all pixel values in the height map or parametric map are read, and the maximum and minimum values among the pixel values are used as statistical benchmarks. Each pixel value in the image is linearly mapped according to its relative position between the minimum and maximum values, so that all pixel values fall into the preset range. The dimensions of the preprocessed height map and the preprocessed parametric map are unified to S×S; Furthermore, the process of performing the spatial alignment includes: The preprocessed height map is used as the reference image, and the preprocessed parametric map is mapped to coordinates to make the pixel positions of the preprocessed parametric map consistent with the pixel positions of the preprocessed height map. Specifically, the coordinate information corresponding to the position of each pixel in the preprocessed height map is read, and a reference coordinate system is established; Read the coordinate information corresponding to the position of each pixel in the preprocessed parametric image and establish the parametric image coordinate system; Obtain the coordinate offset, rotation, and scaling of the preprocessed parametric map relative to the preprocessed height map; A coordinate mapping relationship is constructed based on the coordinate offset, rotation, and scaling; according to the coordinate mapping relationship, the original coordinates of each pixel position in the preprocessed parameter map are converted into mapped coordinates; Based on the mapped coordinates, the pixel values in the preprocessed parameter map are redistributed to the corresponding target pixel positions; Interpolation calculations are performed on the target pixel positions that exceed the original sampling positions of the preprocessed parameter map to obtain the corresponding pixel values; the parameter map after pixel value redistribution is output as an aligned parameter map, so that the pixel positions of the aligned parameter map are consistent with the pixel positions of the preprocessed height map; The process of constructing the channels and obtaining the model input tensor includes: The spatially aligned height map and parameter map are concatenated along the channel dimension to obtain the model input tensor.
[0008] Specifically, the process of stitching along the channel dimension includes: The spatially aligned height map and parametric map are concatenated to form the model input tensor X∈R. H×W×C Where X represents the model input tensor, R represents the real number space, H and W are the input resolutions, and C is the number of channels; Furthermore, the process of dividing the model input tensor into image patches and converting the image patches into embedding sequences includes: The model input tensor is divided into several image blocks according to a preset block size; Each image patch is flattened and converted into an embedding vector through linear mapping; The embedding vectors are arranged according to the positional order of the image blocks to obtain the embedding sequence, which also includes a preset classification label and positional code.
[0009] Specifically, the model input tensor X is divided into N = (H / P) × (W / P) image blocks according to P × P, where P represents the side length of the image block, N represents the number of image blocks, and H and W are the input resolutions. When H or W is not divisible by P, the input tensor is first padded with zeros, cropped, or scaled until it is divisible by P before being divided. The i-th image patch is flattened as The D-dimensional embedding vector Z is obtained through linear mapping. i =X i ×W e +b e ;where X i Z represents the vector after the i-th image patch is flattened. i Let R represent the ith embedding vector corresponding to the ith image patch, C be the number of channels, and W be the number of channels. e and b e These represent the weight matrix and bias term of the linear embedding layer, respectively. The embedding sequence is Z0=[Z cls Z1, ..., Z N ]+E pos Z cls Z represents the classification label. N Let E represent the Nth embedding vector. pos Indicates positional encoding; During the encoding process, the classification label and the embedding vector corresponding to each image patch jointly participate in the attention calculation and continuously absorb the feature information in each image patch; after completing multi-layer encoding, the output vector corresponding to the classification label represents the global semantic information of the entire model input tensor and serves as the input for subsequent classification layers to determine cell state. Position encoding is used to attach the position information of each image patch in the original model input tensor to the embedding vector corresponding to each image patch. Since the model input tensor is divided into image patches and converted into embedding vectors, each embedding vector itself only represents the content features of the corresponding image patch and cannot directly reflect the position of the image patch in the original image. Therefore, position encoding needs to be set and added to the corresponding embedding vector so that the visual Transformer model can distinguish the spatial position relationship of different image patches when performing attention calculation.
[0010] The embedded sequence is input into a visual Transformer model to obtain a global feature vector representing the cell state; based on the global feature vector, the cell state category and probability are output.
[0011] Furthermore, the visual Transformer model is trained. The visual Transformer model includes several encoder blocks, each of which includes a multi-head self-attention layer and a multi-layer perceptron layer. The embedded sequence is input into each encoder block for feature extraction, and the global feature vector is output.
[0012] Specifically, the embedded sequence Z0 is input into a Transformer model consisting of L stacked encoder layers; The layer comprises multi-head self-attention and multi-layer perceptron sublayers, and employs residual connections and layer normalization. Its computation can be expressed as: ; ; Wherein, LN represents layer normalization, which is used to normalize each feature dimension of the input feature sequence in order to stabilize the feature distribution; Indicates the encoder layer index. Indicates the first The embedding sequence of the layer encoder, Indicates the first The intermediate feature sequence of the layer encoder, Indicates the first Output sequence features of the layer encoder; MSA stands for Multi-Head Self-Attention, which is used to model the correlation between vectors in the input feature sequence. It obtains the correlation features between different image patches and between the classification label and each image patch through parallel computation of multiple attention heads. Its output is the attention feature sequence corresponding to the dimension of the input sequence. MLP stands for Multilayer Perceptron, which is used to perform nonlinear transformations on the feature sequence after attention updates, and to further map and enhance the features, including fully connected layers, activation functions, and fully connected layers again; Specifically, firstly, the first Layer Embedded Sequence Perform layer normalization and input it into the multi-head self-attention module; then combine the resulting multi-head self-attention output with the original input feature sequence. Add them together to obtain the intermediate output feature sequence. Then output the intermediate feature sequence. Perform layer normalization and input it into the multilayer perceptron module; then compare the resulting multilayer perceptron output with the intermediate output feature sequence. Add them together to get the first one. The final output feature sequence of the layer encoder ; Repeat the above calculation process until all encoder layers have been calculated, and use the feature sequence output by the last encoder layer as the global feature vector h = [h1, h2, ..., h2]. D ], where D represents the feature dimension, h D This represents the value of the global feature vector in the Dth feature dimension; Furthermore, during the training process of the visual Transformer model, the model input tensor corresponding to the labeled cell samples is input into the visual Transformer model, a loss function is constructed based on the output category results and the true labels, and the model parameters of the visual Transformer model are updated according to the loss function.
[0013] Specifically, the parameters are updated through backpropagation using the cross-entropy loss function, which is: ; Where y represents the ID of the true category, p y L represents the probability value of the predicted probability vector p corresponding to the true class y. CE Represents the cross-entropy loss function; The AdamW optimizer and cosine annealing learning rate scheduling are preferred, and hybrid precision training and gradient pruning can be used to improve training stability.
[0014] Furthermore, the process of outputting the cell state category and probability based on the global feature vector includes: The global feature vector is input into the visual Transformer model to obtain the classification score corresponding to each cell state. The scores for each category are normalized to obtain the probability corresponding to each cell state; The cell state with the highest probability is taken as the category result.
[0015] After completing the feature extraction of the aforementioned encoder block, the global feature vector that can characterize the overall features of the current cell under test is read, and the global feature vector is input into the classification layer of the visual Transformer model to obtain the score results of the current cell under test belonging to each cell state. The classification layer can be implemented using a linear mapping method, that is, based on the mapping relationship between each feature component in the global feature vector and each cell state, a classification score corresponding to each cell state is output; the classification score is used to characterize the degree of matching between the current cell to be tested and the corresponding cell state. The larger the classification score, the stronger the tendency for the current cell to be tested to be identified as the corresponding cell state. Specifically, the global feature vector is linearly transformed through the classification layer to obtain the classification score vector S=W. s ×h+b s W s and b s ... Furthermore, the classification score vector S = [s1, s2, ..., s...] m ], where s m This represents the classification score corresponding to the state of the m-th cell.
[0016] The normalization of the classification score vector can be achieved using the softmax function, which involves performing an exponential transformation on the classification score corresponding to each category and dividing it by the sum of the exponential transformation results of the classification scores corresponding to all categories, thereby obtaining the probability value corresponding to each cell state. After this processing, the output result corresponding to each cell state is transformed from the original score into a comparable probability distribution form. Using the cell state with the highest probability as the category result means that after obtaining the probability corresponding to each cell state, the probability values in the probability vector are compared, and the cell state corresponding to the highest probability value is selected as the final identification result of the current cell to be tested; in other words, the current cell to be tested is determined to be the cell state most likely to correspond to the cell state indicated by the highest probability. Furthermore, the process of obtaining the category results and probabilities also includes: The input tensor of the model is rotated, flipped, clipped, and scaled to obtain multiple enhanced tensors. Each enhancement tensor is input into the visual Transformer model to obtain several probabilities, which are then weighted and fused to obtain the category result and probability.
[0017] Specifically, after obtaining the model input tensor, instead of performing a single inference based solely on the model input tensor, multiple geometric transformations are performed on the model input tensor to generate multiple augmented tensors corresponding to the original input. These geometric transformations include rotation, flipping, clipping, and scaling. Let T be the transformation function corresponding to the k-th augmentation operation. k Then the k-th augmentation tensor can be represented as: X k =T k (X), where X represents the model input tensor; X k Represents the k-th augmentation tensor; K is an integer greater than or equal to 2; The rotation is used to rotate and transform the model input tensor according to a preset angle; The flipping is used to perform a mirror transformation along the horizontal or vertical axis; The cropping is used to extract a local area and then adjust it to the input size; The scaling is used to enlarge or shrink the input area and then restore it to a uniform size; After the above processing, K enhancement tensors corresponding to the same test cell but with different spatial expression forms were obtained; For each augmentation tensor, it is input into the visual Transformer model for independent inference; that is, the same processing flow as the original input is performed on each augmentation tensor, including image patch segmentation, embedding sequence construction, encoder feature extraction, and classification output; the prediction probability vector corresponding to the k-th augmentation tensor is: ; Where m represents the number of cell state categories. Let represent the probability that the k-th augmentation tensor belongs to the m-th cell state; after performing the above reasoning process on all K augmentation tensors, K sets of predicted probabilities can be obtained; the K sets of predicted probabilities are weighted and fused to obtain the fused cell state probabilities. Specifically, the cell state refers to the comprehensive phenotypic category of the surface morphology and local mechanical characteristics exhibited by the test cells under the current culture conditions, drug treatment conditions, or control conditions. It can be an untreated state, a drug-responsive state, an inhibited state, or other state categories classified according to experimental labeling rules. Furthermore, the global feature vector is input into the classification layer and the classification score, probability, and category result are output, which reflects the classification output process under a single inference condition and is used to complete a single classification judgment. The model input tensor is rotated, flipped, clipped, and scaled, and then inferred separately before probability fusion is performed. This reflects the fusion output process based on multiple enhanced inputs during the testing phase, and is used to fuse multiple prediction results based on a single classification judgment. There is also a hierarchical relationship between the two: The process of obtaining classification scores, probabilities, and category results from global feature vectors is the basic output process necessary for the model to complete a recognition. The process of rotating, flipping, cropping, and scaling the model input tensor and then performing weighted fusion adds two steps to the basic output process: enhanced input and probability fusion. That is, each augmentation tensor still needs to undergo classification score calculation and probability normalization before the probability result used for fusion can be obtained.
[0018] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention constructs height maps, Young's modulus maps, adhesion force maps, and reference force maps by inverting force curves obtained from atomic force microscopy, and forms unified multi-channel input data, so that the morphological and mechanical information of cells can be synergistically characterized. Compared with a single image or a single parameter, this is more conducive to completely depicting the differences in cell state, thereby improving the accuracy and stability of cell state identification.
[0019] 2. This invention inputs multi-channel cell characterization data into a visual Transformer model and utilizes its ability to extract global features and cross-regional correlation features to achieve fine classification of cell states. It can more effectively distinguish cell samples that look similar but have different internal mechanical responses, reduce the interference of local noise and local deformation on the discrimination results, and improve the model's generalization ability and classification reliability.
[0020] 3. This invention enhances the model's adaptability to pose changes, scale changes, and local perturbations by performing enhancement processing on the model input tensor through rotation, flipping, pruning, and scaling, and by fusing the predicted probabilities corresponding to multiple enhancement results. This reduces the impact of random errors in single inference on the final result, making the output cell state category results and probabilities more robust and facilitating subsequent analysis and application. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating a cell state recognition method based on multimodal image deep learning according to the present invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] Example: Figure 1 As shown, the present invention provides a technical solution: a cell state recognition method based on multimodal image deep learning, the method comprising the following steps: Two-dimensional data of the cell under test are obtained by atomic force microscopy. The two-dimensional data includes a height map and a force curve. A parameter map characterizing the mechanical properties of the cell is generated based on the force curve. The process of acquiring two-dimensional data of the cells to be tested includes: The surface of the cell to be tested is scanned pixel by pixel using atomic force microscopy in quantitative imaging mode. The pixel-by-pixel scanning process includes: The probe of the atomic force microscope is controlled to move vertically toward the surface of the cell to be tested at each scanning position. After the probe moves to the target position, the probe is controlled to move in the opposite direction to the termination position, and the displacement information and force information corresponding to each scanning position during the movement are recorded. Force curves for each scanning position are constructed based on the displacement information and the force information; In this embodiment, glioma cells are used as the test cells; Glioma cells were cultured adherently on coverslips, and single-cell data were acquired using atomic force microscopy in quantitative imaging mode. The pixel resolution was set to 256×256, the scanning frequency to 2Hz, the peak force to 1nN, the scanning speed to 70μm / s, and the Z-axis range to 2.6μm to fully cover the morphological undulations of the cell surface. When the force information reaches the preset force, the displacement information at this time is used as the height value of the corresponding scanning position; and the height value is arranged according to the two-dimensional coordinates of each scanning position to obtain the height map.
[0024] The process of generating a parametric map characterizing the mechanical properties of cells based on the force curve includes: The force curve segment from the moment the probe contacts the cell to be tested is extracted as the approach segment for each scanning position. Based on the proximity segment in the force curve corresponding to each scanning position, the Young's modulus value of each scanning position is obtained by inversion through the contact mechanics model; The Young's modulus value, adhesion force value, and reference force value are arranged according to the pixel coordinates corresponding to each scanning position to obtain the Young's modulus map, adhesion force map, and reference force map for each scanning position. The parameter diagrams include Young's modulus diagram, adhesion force diagram, and reference force diagram.
[0025] In this embodiment, Young's modulus map, adhesion force map, and reference force map are generated based on force curve inversion. After flattening, denoising, and background correction of the image, the cell region is cropped and uniformly scaled to 384×384. Pixel values are normalized. During the training phase, data augmentation is performed by random rotation up to 20° to keep the pose perturbation within a reasonable range without damaging cell features. Random horizontal flipping probability is 0.5 to balance the data orientation distribution and improve generalization ability. Random cropping and scaling scale range [0.6, 1.0] is used to preserve the cell core discrimination region without losing key information. The height map and parameter map are preprocessed, and the preprocessed height map and parameter map are spatially aligned and channel constructed to obtain the model input tensor; the model input tensor is divided into image blocks, and the image blocks are converted into embedding sequences; The preprocessing process for the height map and parameter map includes: The height map and parameter map are respectively flattened, corrected for background, denoised and normalized to obtain preprocessed height map and preprocessed parameter map; the preprocessed height map and preprocessed parameter map have the same size.
[0026] The process of performing the spatial alignment includes: The preprocessed height map is used as the reference image, and the preprocessed parametric map is mapped to coordinates to make the pixel positions of the preprocessed parametric map consistent with the pixel positions of the preprocessed height map. The process of constructing the channels and obtaining the model input tensor includes: The spatially aligned height map and parameter map are concatenated along the channel dimension to obtain the model input tensor.
[0027] The process of dividing the model input tensor into image patches and converting the image patches into embedding sequences includes: The model input tensor is divided into several image blocks according to a preset block size; Each image patch is flattened and converted into an embedding vector through linear mapping; The embedding vectors are arranged according to the positional order of the image blocks to obtain the embedding sequence, which also includes a preset classification label and positional code.
[0028] In this embodiment, the 384×384 input tensor is divided into image blocks of 16×16, and the number of image blocks N=(384 / 16)×(384 / 16)=576. After flattening the image blocks, a 768-dimensional embedding vector is obtained by linear mapping. The embedded sequence is input into a visual Transformer model to obtain a global feature vector representing the cell state; based on the global feature vector, the cell state category and probability are output.
[0029] The visual Transformer model is trained. The visual Transformer model includes several encoder blocks. Each encoder block includes a multi-head self-attention layer and a multi-layer perceptron layer. The embedded sequence is input into each encoder block for feature extraction, and the global feature vector is output.
[0030] During the training of the visual Transformer model, the model input tensor corresponding to the labeled cell samples is input into the visual Transformer model. A loss function is constructed based on the output category results and the true labels, and the model parameters of the visual Transformer model are updated according to the loss function.
[0031] In this embodiment, the visual Transformer model adopts a 12-layer encoder stacked structure. Each encoder layer contains multi-head self-attention and multi-layer perceptron sub-layers. The number of attention heads is set to 12, and the MLP hidden dimension is set to 3072 to broaden the feature mapping space and enhance the representation capability. Dropout is set to 0.1 to suppress model overfitting in small sample scenarios. Encoder calculation follows the formula, and residual connections and layer normalization ensure the stability of the training process. The process of outputting the cell state category and probability based on the global feature vector includes: The global feature vector is input into the visual Transformer model to obtain the classification score corresponding to each cell state. The scores for each category are normalized to obtain the probability corresponding to each cell state; The cell state with the highest probability is taken as the category result.
[0032] In this embodiment, the cross-entropy loss function is used to calculate the classification error, the AdamW optimizer is selected, and the learning rate is set to 1×10⁻⁶. -4 To balance convergence speed and parameter update accuracy; the weight decay is set to 0.05 to constrain parameter amplitude and reduce the risk of overfitting; The learning rate scheduling uses a cosine annealing strategy, which calculates the total number of complete cosine annealing iterations T. max Setting eta to 100 sets the lower bound of the learning rate during the annealing process. min Set to 1×10 -6 To avoid parameter oscillations caused by a sudden increase in the learning rate in the later stages, the number of training rounds was set to 10. To prevent overtraining on small sample datasets, the batch size was set to 64 to balance memory usage and gradient calculation stability. Mixed precision training was adopted and the gradient clipping threshold was set to 1.0 to prevent gradient explosion. After model inference and probability fusion, the state classification probabilities of the cells to be tested are: untreated state 0.023, mild response state 0.057, and significant response state 0.920. The model output class result is the significant response state, and the high confidence probability meets the requirements for accurate determination of cell state.
[0033] The process of obtaining the category results and probabilities also includes: The input tensor of the model is rotated, flipped, clipped, and scaled to obtain multiple enhanced tensors. Each enhancement tensor is input into the visual Transformer model to obtain several probabilities, which are then weighted and fused to obtain the category result and probability.
[0034] For example, if the number of iterations K is set to 5, performing random rotation of 15°, flipping, and slight affine transformation on the same image yields 5 enhanced versions. After inference, the probabilities are weighted and fused to obtain the unprocessed state (0.891), the mildly responsive state (0.085), and the significantly responsive state (0.024). The model outputs the unprocessed state as the category result. In this embodiment, 70 cell samples were selected for performance verification. The sample composition included 20 control samples, 30 samples in the early stage of drug action, and 20 samples in the stage of significant drug inhibition. The balanced sample ratio ensured the objectivity of the classification and evaluation results. The overall identification accuracy of the method reached 97.14%. The precision, recall, and F1 score of the control sample were 1.0000, 0.9000, and 0.9474. All three indicators of the early stage of drug action were 1.0000. The precision, recall, and F1 score of the stage of significant drug inhibition were 0.9091, 1.0000, and 0.9524. The predicted distribution of multiple cell sample states was statistically analyzed and the inhibition evaluation indicators were calculated to complete the ranking and screening of candidate inhibitors.
[0035] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A cell state recognition method based on multimodal image deep learning, characterized in that: The method includes the following steps: Two-dimensional data of the cell under test are obtained by atomic force microscopy. The two-dimensional data includes a height map and a force curve. A parameter map characterizing the mechanical properties of the cell is generated based on the force curve. The process of acquiring two-dimensional data of the cells to be tested includes: The surface of the cell to be tested is scanned pixel by pixel using atomic force microscopy in quantitative imaging mode. The pixel-by-pixel scanning process includes: The probe of the atomic force microscope is controlled to move vertically toward the surface of the cell to be tested at each scanning position. After the probe moves to the target position, the probe is controlled to move in the opposite direction to the termination position, and the displacement information and force information corresponding to each scanning position during the movement are recorded. Force curves for each scanning position are constructed based on the displacement information and the force information; When the force information reaches the preset force, the displacement information at this time is used as the height value of the corresponding scanning position; and the height values are arranged according to the two-dimensional coordinates of each scanning position to obtain the height map; The process of generating a parametric map characterizing the mechanical properties of cells based on the force curve includes: The force curve segment from the moment the probe contacts the cell to be tested is extracted as the approach segment for each scanning position. Based on the proximity segment in the force curve corresponding to each scanning position, the Young's modulus value of each scanning position is obtained by inversion through the contact mechanics model; the adhesion force value of each scanning position is extracted according to the proximity segment, and the reference force value of each scanning position is preset. The Young's modulus value, adhesion force value, and reference force value are arranged according to the pixel coordinates corresponding to each scanning position to obtain the Young's modulus map, adhesion force map, and reference force map for each scanning position. The parameter diagrams include Young's modulus diagram, adhesion force diagram, and reference force diagram; The height map and parameter map are preprocessed, and the preprocessed height map and parameter map are spatially aligned and channel constructed to obtain the model input tensor; the model input tensor is divided into image blocks, and the image blocks are converted into embedding sequences; The embedded sequence is input into a visual Transformer model to obtain a global feature vector representing the cell state; based on the global feature vector, the cell state category and probability are output.
2. The cell state recognition method based on multimodal image deep learning according to claim 1, characterized in that: The process of performing the spatial alignment includes: The preprocessed height map is used as the reference image, and the preprocessed parametric map is mapped to coordinates to make the pixel positions of the preprocessed parametric map consistent with the pixel positions of the preprocessed height map. The process of constructing the channels and obtaining the model input tensor includes: The spatially aligned height map and parameter map are concatenated along the channel dimension to obtain the model input tensor.
3. The cell state recognition method based on multimodal image deep learning according to claim 1, characterized in that, The process of dividing the model input tensor into image patches and converting the image patches into embedding sequences includes: The model input tensor is divided into several image blocks according to a preset block size; Each image patch is flattened and converted into an embedding vector through linear mapping; The embedding vectors are arranged according to the positional order of the image blocks to obtain the embedding sequence, which also includes a preset classification label and positional code.
4. The cell state recognition method based on multimodal image deep learning according to claim 1, characterized in that: The visual Transformer model is trained. The visual Transformer model includes several encoder blocks. Each encoder block includes a multi-head self-attention layer and a multi-layer perceptron layer. The embedded sequence is input into each encoder block for feature extraction, and the global feature vector is output.
5. The cell state recognition method based on multimodal image deep learning according to claim 4, characterized in that: During the training of the visual Transformer model, the model input tensor corresponding to the labeled cell samples is input into the visual Transformer model. A loss function is constructed based on the output category results and the true labels, and the model parameters of the visual Transformer model are updated according to the loss function.
6. The cell state recognition method based on multimodal image deep learning according to claim 1, characterized in that, The process of outputting the cell state category and probability based on the global feature vector includes: The global feature vector is input into the visual Transformer model to obtain the classification score corresponding to each cell state. The scores for each category are normalized to obtain the probability corresponding to each cell state; The cell state with the highest probability is taken as the category result.
7. The cell state recognition method based on multimodal image deep learning according to claim 1, characterized in that, The process of obtaining the category results and probabilities also includes: The input tensor of the model is rotated, flipped, clipped, and scaled to obtain multiple enhanced tensors. Each enhancement tensor is input into the visual Transformer model to obtain several probabilities, which are then weighted and fused to obtain the category result and probability.
8. The cell state recognition method based on multimodal image deep learning according to claim 1, characterized in that, The preprocessing process for the height map and parameter map includes: The height map and parameter map are respectively flattened, corrected for background, denoised and normalized to obtain preprocessed height map and preprocessed parameter map; the preprocessed height map and preprocessed parameter map have the same size.
Citation Information
Patent Citations
Noninvasive sperm cell death and viability detection method
CN120142258A
Biological cell image classification method and system and storage medium
CN120564186A