Circulating tumor cell recognition model training method and system based on deep learning
Through the cross-modal feature extraction of multi-channel pathological sample images and the deep convolutional network of multi-stage training strategies, the problem of insufficient feature misjudgment and behavioral pattern recognition capabilities in the recognition of circulating tumor cells in the prior art is solved, and higher recognition accuracy and clinical applicability are achieved.
Patent Information
- Application Number
- CN202510586266.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-08
AI Technical Summary
The prior art fails to effectively fuse cell morphological characteristics and microenvironment characteristics in the recognition of circulating tumor cells, resulting in misjudgment of characteristics when facing complex pathological scenarios. The traditional model lacks the ability to recognize behavioral patterns such as loose cell connections and matrix penetration, which affects the pathological interpretability of clinical deployment.
By acquiring multi-channel pathological sample images, cross-modal feature extraction is used to generate fusion feature vectors, and a deep convolutional network combined with a multi-stage training strategy includes cross-channel feature attention mechanism, adversarial sample generation and knowledge distillation to enhance the feature discriminant and compression capabilities of the model.
It significantly improved the accuracy and clinical applicability of circulating tumor cell recognition results, improved the model's ability to distinguish complex pathological scenarios, and maintained pathological explanatory.
Smart Images

Figure CN120107724A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning, and in particular to a circulating tumor cell recognition model training method and system based on deep learning. Background Art
[0002] In the field of medical image analysis, the identification of circulating tumor cells is of great significance for cancer metastasis monitoring. The current mainstream method extracts the morphological features of the nucleus based on single-channel pathological images, and realizes tumor cell detection by training the classification model through convolutional neural network. The existing technology focuses on the calculation of isolated morphological parameters such as the size of the nucleus and the deviation of the nuclear-cytoplasmic ratio, and directly outputs the classification results using an end-to-end deep learning model. It fails to effectively integrate the spatial interaction characteristics of tumor cells and the microenvironment, resulting in the model's feature misjudgment when facing complex pathological scenes such as dye differences and immune cell infiltration. Traditional training strategies are insufficient in recognizing behavioral patterns such as loose intercellular connections and matrix penetration that are unique to circulating tumor cells. In addition, the existing model compression methods mostly achieve lightweighting through global channel pruning, which destroys the local feature response path that is sensitive to tumor diagnosis and affects the pathological interpretability during clinical deployment. Based on this, the existing technology has problems such as large positioning deviation in real clinical samples, high rate of missed detection of undifferentiated cells, and difficulty in filtering false positive samples. Summary of the invention
[0003] The present invention provides a circulating tumor cell recognition model training method and system based on deep learning.
[0004] In a first aspect, an embodiment of the present invention provides a circulating tumor cell recognition model training method based on deep learning, the method comprising: obtaining a multi-channel pathological sample image set of circulating tumor cells, the multi-channel pathological sample image set comprising cell microscopic images processed with different stains, each sample image carrying corresponding cell type annotation information; performing cross-modal feature extraction on the multi-channel pathological sample images to generate a fusion feature vector comprising cell morphological features and microenvironmental features, wherein the microenvironmental features include matrix component distribution features and immune cell infiltration features; constructing a deep convolutional network with a multi-stage training strategy, wherein the first stage adopts a cross-channel feature attention mechanism to optimize convolution kernel parameters, the second stage enhances feature discriminability by adversarial sample generation, and the third stage adopts knowledge distillation to compress the network scale; adopting a hierarchical cross-validation mechanism to verify and adjust the network parameters of the deep convolutional network during the training process, wherein the hierarchical cross-validation mechanism is divided into multiple validation subsets according to the biological characteristics corresponding to the cell type annotation information, and the network parameters are adjusted based on the classification error of each validation subset; generating a final recognition model based on the verified and adjusted network parameters, and the final recognition model is configured to receive a clinical pathological image stream and output the positioning coordinates and classification confidence of circulating tumor cells.
[0005] In a second aspect, an embodiment of the present invention provides a computer system, comprising: a memory, in which a computer program is stored; and a processor, for loading the computer program to implement the deep learning-based circulating tumor cell recognition model training method as described above.
[0006] The circulating tumor cell recognition model training method and system provided by the present invention significantly improves the accuracy and clinical applicability of the recognition results by integrating the cell morphological features and microenvironmental features of multi-channel pathological images and adopting a multi-stage optimization strategy. A spatial position correlation alignment mechanism is established in cross-modal feature extraction, so that the cell nuclear atypia parameters and the matrix fiber orientation characteristics form a biomedical semantic association, effectively capturing the morphological-microenvironmental coordinated change law in the process of tumor cell invasion and metastasis; through adversarial samples, the co-occurrence pattern of the tightness of intercellular connections and the density of immune cell aggregation is generated to enhance the model's ability to discriminate cell variation characteristics in real complex pathological scenes; combined with knowledge distillation technology, the sensitive response path of the teacher network to the nuclear atypia index is inherited, and the key biological feature recognition capabilities are retained while compressing the model scale. The network parameters are further dynamically adjusted based on the hierarchical verification mechanism of the biological characteristics of the tumor to ensure the balance of the generalization performance of the model on epithelial, mesenchymal and undifferentiated tumor cell subsets. The final output positioning coordinates and classification confidence are deeply integrated with the immune infiltration density and cell connection density indicators, so that the recognition results are consistent with the spatial distribution of cells and have clinical pathological diagnostic interpretability, overcoming the problems of high false positive rate and insufficient cell subtype differentiation caused by ignoring microenvironment interactions in traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0008] Figure 1 This is a flowchart of a circulating tumor cell recognition model training method based on deep learning provided by an embodiment of the present invention.
[0009] Figure 2 It is a schematic diagram of the composition of a computer system provided by an embodiment of the present invention.
[0010] Reference numerals: 101 - processor; 102 - communication interface; 103 - memory. DETAILED DESCRIPTION
[0011] See also Figure 1 , Figure 1 A flowchart of a circulating tumor cell recognition model training method based on deep learning provided by an embodiment of the present invention, the circulating tumor cell recognition model training method based on deep learning can be executed by a computer system, and the circulating tumor cell recognition model training method based on deep learning may include the following steps:
[0012] Step S100: Acquire a multi-channel pathology sample image set of circulating tumor cells, wherein the multi-channel pathology sample image set includes cell microscopic images processed with different stains, and each sample image carries corresponding cell type annotation information.
[0013] In an embodiment of the present invention, a multi-channel pathology sample image set is composed of cell microscopic images processed with different stains. The staining agent treatment is to better display the different structures and components of cells, and different stains can highlight the specific characteristics of cells. Cell microscopic images are images obtained by observing and photographing cell samples under a microscope, which contain information such as cell morphology and structure. Cell type annotation information is a clear identification of the cell type in each sample image, such as epithelial tumor cells, mesenchymal tumor cells, etc. The annotation information is the basis for supervised learning in subsequent model training.
[0014] The present invention does not limit the source of samples in the multi-channel pathology sample image set. For example, under the permission of laws and regulations, circulating tumor cell samples treated with different stains are collected from the hospital's pathology database, and then these samples are observed and photographed using professional microscope equipment to obtain cell microscopic images. For each sample image, the cell type can be annotated by a field expert.
[0015] Step S200: performing cross-modal feature extraction on the multi-channel pathological sample image to generate a fusion feature vector including cell morphology features and microenvironment features, wherein the microenvironment features include matrix component distribution features and immune cell infiltration features.
[0016] Cross-modal feature extraction refers to features extracted from data of different modalities (i.e., multi-channel pathological sample images processed with different stains in the embodiment of the present invention). Cell morphological features are the morphological and structural features of the cell itself, such as the size, shape, nuclear-cytoplasmic ratio of the nucleus, etc. These features can reflect the biological characteristics and pathological state of the cell. Microenvironmental features are the characteristics of the surrounding environment in which the cell is located, including matrix component distribution characteristics and immune cell infiltration characteristics. Matrix component distribution characteristics describe the distribution of the extracellular matrix, such as the direction and density of matrix fibers; immune cell infiltration characteristics reflect the distribution and aggregation of immune cells in tumor tissues, which helps to understand the immune response and prognosis of the tumor. The fused feature vector is a vector obtained by integrating the cell morphological features and the microenvironmental features, which contains comprehensive information about the cell and its microenvironment.
[0017] The process of cross-modal feature extraction of multi-channel pathological sample images can be performed by, for example, using a deep learning-based convolutional neural network (CNN), such as ResNet, VGG, etc. Then, the extracted cell morphology features and microenvironment features are fused to generate a fused feature vector.
[0018] As an implementation manner, step S200, performing cross-modal feature extraction on a multi-channel pathological sample image to generate a fusion feature vector including cell morphology features and microenvironment features, may specifically include the following steps S210-S260:
[0019] Step S210: performing superpixel segmentation operation on the multi-channel pathological sample image to generate structural data covering the cell nucleus region, cytoplasm region and extracellular matrix, the structural data including cell nucleus morphology parameters, cell membrane curvature parameters and extracellular matrix distribution parameters of each region.
[0020] Superpixel segmentation is used to segment an image into several small regions with similar features, namely superpixels. In an embodiment of the present invention, a superpixel segmentation operation is performed on a multi-channel pathological sample image to divide the cell nucleus region, cytoplasm region and extracellular matrix region in the image to generate structural data.
[0021] The nucleus morphological parameters include the size, shape, nuclear-to-cytoplasmic ratio of the nucleus, etc. These parameters can be obtained by geometric analysis of the segmented nucleus region. For example, the area, perimeter, length of the major axis and minor axis of the nucleus are calculated, and then the nuclear-to-cytoplasmic ratio and other morphological indicators are calculated based on these parameters. The cell membrane curvature parameter describes the curvature of the cell membrane, which can be obtained by analyzing the curvature of the cell membrane boundary. The extracellular matrix distribution parameters include the direction and density of the matrix fibers, etc. These parameters can be obtained by analyzing the image features of the extracellular matrix region. For example, the direction and density of the matrix fibers are calculated using texture analysis methods.
[0022] There are many algorithms that can be used to perform superpixel segmentation, such as the Simple Linear Iterative Clustering (SLIC) algorithm. The SLIC algorithm divides an image into several superpixels by clustering in the color space and coordinate space of the image. For example, some cluster centers are first randomly initialized in the image, and then pixels are assigned to different clusters based on the distance between the pixels and the cluster centers, and the positions of the cluster centers are continuously updated until the convergence conditions are met.
[0023] Step S220: Perform dual-branch feature extraction based on the structural data, wherein the first branch extracts the nuclear atypia index and the cell-to-cell connection density index as cell morphological features through geometric topological analysis, and the second branch extracts the matrix fiber orientation features and the immune cell aggregation density as microenvironment features through spatial frequency transformation.
[0024] The dual-branch feature extraction is to input the structural data into two different branches for feature extraction, and each branch extracts different types of features. The first branch mainly focuses on the morphological characteristics of the cells themselves, and extracts the nuclear atypia index and the cell-to-cell connection tightness index through geometric topological analysis. In an embodiment of the present invention, the boundaries of the cell nucleus and the cell-to-cell connection area are analyzed to extract relevant feature indicators.
[0025] The nuclear atypia index is an indicator to measure the degree of abnormal nuclear morphology, which reflects the malignancy of tumor cells. The cell-to-cell connection tightness index describes the tightness of the connection between cells, which helps to understand the growth and migration characteristics of cells. The second branch focuses on the microenvironment characteristics of cells, and extracts the matrix fiber orientation characteristics and immune cell aggregation density through spatial frequency transformation. Spatial frequency transformation is used to convert images from the spatial domain to the frequency domain, which can reveal the distribution of different frequency components in the image. In an embodiment of the present invention, the direction of the matrix fibers and the aggregation density of immune cells are extracted by performing spatial frequency transformation on the image of the extracellular matrix region.
[0026] In the first branch, geometric topological analysis can adopt the following specific steps. First, perform edge detection operation on the structural data to generate the nuclear contour boundary data and the intercellular connection area boundary data. The boundary data contains pixel-level coordinate sequence and curvature change information. Edge detection can use the Canny algorithm, which finds the edge points in the image by calculating the gradient value of the pixels in the image. Then, perform polygon fitting operation on the nuclear contour boundary data to generate the polygon vertex distribution describing the nuclear morphology, calculate the nuclear cytoplasm ratio deviation, contour concave-convex index and main axis deflection angle based on the vertex distribution, and fuse these parameters to generate the nuclear atypia index. Polygon fitting can use the least squares method to approximate the outline of the cell nucleus by fitting the vertices of the polygon. For the intercellular connection area boundary data, perform topological structure analysis, extract the statistical distribution of the overlapping length and contact angle of the adjacent cell contact surface, and generate the intercellular connection tightness index by combining the area proportion of the connection area. Topological structure analysis can extract relevant features by analyzing the topological relationship of the boundary data, such as connectivity, number of holes, etc.
[0027] In the second branch, the spatial frequency transform can be performed using discrete Fourier transform (DFT) or fast Fourier transform (FFT). First, a spatial frequency transform operation is performed on the extracellular matrix region in the structural data to convert the pixel distribution into a frequency domain energy distribution, and the main frequency direction is extracted as the matrix fiber orientation feature. Then, a regional clustering analysis is performed on the frequency domain energy distribution to identify the spatial density distribution of the immune cell aggregation region, and the immune cell aggregation density is generated by combining the area of the clustering region with the contact boundary length of the adjacent matrix region. The regional clustering analysis can use the K-means clustering algorithm to identify the immune cell aggregation region by clustering the points in the frequency domain energy distribution into different regions. For example, for a multi-channel pathological sample image, after the structural data is obtained by superpixel segmentation, it is input into the first branch and the second branch for feature extraction. In the first branch, the Canny algorithm is used to detect the boundaries of the cell nucleus and the intercellular connection area, and then the least squares method is used for polygon fitting to calculate the cell nuclear atypia index and the intercellular connection tightness index. In the second branch, FFT was used to perform spatial frequency transformation on the extracellular matrix region to extract the matrix fiber orientation characteristics, and then the K-means clustering algorithm was used to identify the immune cell aggregation area and calculate the immune cell aggregation density.
[0028] As an implementation mode, in step S220, the first branch extracts the nuclear atypia index and the cell-to-cell connection density index as cell morphological features through geometric topological analysis, which may specifically include the following steps S221-S223:
[0029] Step S221: performing edge detection operation on the structural data to generate cell nucleus contour boundary data and intercellular connection area boundary data, the boundary data including pixel-level coordinate sequence and curvature change information.
[0030] In the embodiment of the present invention, the edge detection operation is performed on the structural data to extract the boundary information of the cell nucleus and the intercellular connection area. The pixel-level coordinate sequence, i.e., the coordinate value of each pixel on the boundary, can accurately describe the position of the boundary. The curvature change information reflects the curvature of the boundary, which helps analyze the morphology of the cell nucleus and the structure of the intercellular connection.
[0031] A variety of algorithms can be used to perform edge detection operations, such as the aforementioned Canny algorithm. During the execution process, the image is first Gaussian smoothed to reduce the noise in the image; then, the gradient value and gradient direction of each pixel in the image are calculated; then, the gradient value is non-maximum suppressed to refine the edge; finally, the edge is screened using the double threshold method to determine the true edge point. For example, for a structural data image obtained by superpixel segmentation, the Canny algorithm is used for edge detection. First, the image is Gaussian smoothed to remove the noise in the image, then the gradient value and gradient direction of each pixel are calculated, then non-maximum suppression is performed, and pixels whose gradient values are not local maxima are set to zero. Finally, the double threshold method is used to determine pixels with gradient values greater than the high threshold as edge points, and pixels with gradient values between the high threshold and the low threshold and connected to the edge point are also determined as edge points, thereby obtaining the cell nucleus contour boundary data and the cell-to-cell connection area boundary data, which contain pixel-level coordinate sequences and curvature change information.
[0032] Step S222: performing polygon fitting operation on the cell nucleus contour boundary data to generate polygon vertex distribution describing the cell nucleus morphology, calculating the nuclear-cytoplasmic ratio deviation, contour concavity index and main axis deflection angle based on the vertex distribution, and fusing the parameters to generate the cell nucleus atypia index.
[0033] Polygon fitting can approximate discrete point sets with polygons. In an embodiment of the present invention, a polygon fitting operation is performed on the nucleus contour boundary data, the purpose of which is to describe the morphology of the nucleus with the vertex distribution of the polygon. The nuclear cytoplasm ratio deviation, that is, the difference between the nuclear cytoplasm ratio of the nucleus and the nuclear cytoplasm ratio of normal cells, can reflect the abnormality of the nucleus. The contour concave-convex index describes the concave-convex situation of the nucleus contour and reflects the irregularity of the nucleus. The main axis deflection angle is the deflection angle of the main axis of the nucleus relative to a certain reference direction, which can reflect the orientation of the nucleus. Specifically, polygon fitting can use the least squares method to find the best function match for the data by minimizing the sum of squares of the error. In an embodiment of the present invention, the nuclear contour boundary data is fitted by the least squares method to determine a set of polygon vertices so that the error between the polygon formed by these vertices and the nucleus contour is minimized. When calculating the nuclear cytoplasm ratio deviation, the area of the nucleus and the area of the cytoplasm are first calculated, and then the nuclear cytoplasm ratio is calculated, and then compared with the nuclear cytoplasm ratio of normal cells to obtain the nuclear cytoplasm ratio deviation. The contour concavity index can be obtained by calculating the concavity of the polygon vertices, for example, counting the number of convex vertices and concave vertices in the polygon and calculating their ratio. The main axis deflection angle can be obtained by calculating the main axis direction of the polygon, for example, using the principal component analysis (PCA) method to calculate the main axis direction of the polygon. Finally, the nuclear cytoplasm ratio deviation, the contour concavity index and the main axis deflection angle are fused to generate the nuclear atypia index. The fusion method can be weighted summation, for example, assigning a weight to each parameter, and then multiplying and adding them to obtain the nuclear atypia index. For example, for a cell nucleus contour boundary data, the least squares method is used to perform polygon fitting to obtain a set of polygon vertex distributions. Then, the nuclear cytoplasm ratio deviation, the contour concavity index and the main axis deflection angle of the cell nucleus are calculated, and these parameters are weighted summed to obtain the nuclear atypia index.
[0034] Step S223: Perform topological structure analysis on the boundary data of the intercellular connection area, extract the statistical distribution of the overlapping length and contact angle of the adjacent cell contact surfaces, and generate an intercellular connection density index based on the area ratio of the connection area.
[0035] Topological structure analysis can reveal topological features such as the connectivity and number of holes of an object. In an embodiment of the present invention, topological structure analysis is performed on the boundary data of the intercellular connection area, with the purpose of extracting the statistical distribution of the overlap length and contact angle of the adjacent cell contact surfaces. The overlap length is the length of the adjacent cell contact surface, and the contact angle is the angle between the adjacent cell contact surfaces. The area ratio of the connection area is the ratio of the area of the intercellular connection area to the total area of the cell.
[0036] Topological structure analysis can be achieved by analyzing the topological relationship of boundary data. For example, the boundary data is represented as a graph using graph theory, where nodes represent points on the boundary and edges represent connections between adjacent nodes. By analyzing the connectivity of the graph and the distance between nodes, the statistical distribution of the overlap length and contact angle of adjacent cell contact surfaces can be extracted. Then, the cell-to-cell connection tightness index is generated by combining the area ratio of the connection area. The method for generating the cell-to-cell connection tightness index can be to comprehensively consider the overlap length, the statistical distribution of the contact angle and the area ratio of the connection area, for example, using a function to combine these parameters. For example, for a cell-to-cell connection area boundary data, the graph theory method is used to represent it as a graph, and the connectivity of the graph and the distance between nodes are analyzed to obtain the statistical distribution of the overlap length and contact angle of adjacent cell contact surfaces. Assuming that the average value of the overlap length is L, the standard deviation of the contact angle is A, and the area ratio of the connection area is S, a function f(L,A,S) is used to generate the cell-to-cell connection tightness index, where f can be a linear combination function, such as f(L,A,S)=0.5L+0.3(1 / A)+0.2S. It is understandable that before the combination, the dimensions of each parameter can be unified or eliminated by normalization or standardization (such as maximum-minimum normalization, normalizing the value to [0,1]), and then subsequent calculations can be performed. This is common knowledge and will not be described in detail. Similarly, for subsequent similar fusion calculations of parameters of different dimensions, those skilled in the art can unify or eliminate conflicts and align dimensions of related dimensions according to common knowledge to meet calculation needs.
[0037] In one embodiment, the second branch extracts matrix fiber orientation characteristics and immune cell aggregation density as microenvironment characteristics through spatial frequency transformation, which may specifically include the following steps S224-S226:
[0038] Step S224: performing a spatial frequency transformation operation on the extracellular matrix region in the structural data, converting the pixel distribution into a frequency domain energy distribution, and extracting the main frequency direction as a matrix fiber orientation feature.
[0039] Spatial frequency transform is used to convert an image from a spatial domain to a frequency domain, and can reveal the distribution of different frequency components in an image. In an embodiment of the present invention, a spatial frequency transform operation is performed on an extracellular matrix region in the structural data, and the purpose is to convert the pixel distribution in the extracellular matrix region into a frequency domain energy distribution. The frequency domain energy distribution describes the energy size of different frequency components in an image, and it can reflect the texture and structural information of the image. The dominant frequency direction is the direction with the highest energy in the frequency domain energy distribution, and it can represent the direction of the matrix fibers.
[0040] As mentioned above, the spatial frequency transform operation can be performed using a discrete Fourier transform (DFT) or a fast Fourier transform (FFT). The image of the extracellular matrix region is cropped and preprocessed, and then converted into a frequency domain energy distribution using the FFT algorithm. Next, the direction with the highest energy is determined in the frequency domain energy distribution and used as the main frequency direction, that is, the matrix fiber orientation feature. For example, for an image of an extracellular matrix region, the FFT algorithm is used to convert it into a frequency domain energy distribution. In the frequency domain energy distribution, the main frequency direction is determined by searching for the position of the maximum energy. Assuming that the frequency domain energy distribution is a two-dimensional matrix, the position of the maximum energy is determined by traversing each element in the matrix, and then the direction corresponding to the position is calculated as the matrix fiber orientation feature.
[0041] Step S225: Perform regional clustering analysis on the frequency domain energy distribution to identify the spatial density distribution of the immune cell aggregation region, and generate the immune cell aggregation density by combining the area of the cluster region and the contact boundary length of the adjacent matrix region.
[0042] Regional cluster analysis is used to divide data points into different clusters, and data points with similar characteristics can be grouped together. Regional cluster analysis is performed on the frequency domain energy distribution to identify the spatial density distribution of immune cell aggregation areas. The area of the cluster region is the area of the immune cell aggregation region, and the contact boundary length of the adjacent matrix region is the length of the contact boundary between the immune cell aggregation region and the adjacent matrix region.
[0043] Regional cluster analysis can use K-means clustering algorithm, by constantly updating the position of cluster center, data points are divided into K clusters. For example, K cluster centers are randomly initialized, and then the data points are assigned to different clusters according to the distance between the data points and the cluster centers, and then the position of the cluster centers is updated until the cluster centers no longer change or meet the convergence conditions. In an embodiment of the present invention, the data points in the frequency domain energy distribution are used as input, and the K-means clustering algorithm is used to divide them into different clusters to identify the immune cell aggregation area. Then, the contact boundary length of the cluster area and the adjacent matrix area is calculated, and the immune cell aggregation density is generated by combining these two parameters. The method for generating the immune cell aggregation density can be to comprehensively consider the contact boundary length of the cluster area and the adjacent matrix area, for example, using a function to combine these two parameters. For example, for a frequency domain energy distribution, the K-means clustering algorithm is used to divide it into 3 clusters to identify the immune cell aggregation area. Assuming that the area of the clustering region is A and the contact boundary length of the adjacent matrix region is B, the function g(A,B) is used to generate the immune cell aggregation density, where g can be a linear combination function, such as g(A,B)=0.6A+0.4B.
[0044] Step S226: combining the cell nuclear atypia index and the cell-cell connection tightness index into a cell morphology feature, and combining the matrix fiber orientation feature and the immune cell aggregation density into a microenvironment feature.
[0045] In an embodiment of the present invention, the nuclear atypia index and the cell-cell connection tightness index extracted above are combined to form a cell morphological feature. The cell morphological feature combines the degree of morphological abnormality of the cell nucleus and the tightness of the cell-cell connection, and can more comprehensively describe the morphological characteristics of the cell. The matrix fiber orientation characteristics and the immune cell aggregation density are combined to form a microenvironment feature. The combination method can be vector splicing. That is, the nuclear atypia index and the cell-cell connection tightness index are used as elements of the vector and spliced into a new vector as the cell morphological feature. Similarly, the matrix fiber orientation characteristics and the immune cell aggregation density are used as elements of the vector and spliced into a new vector as the microenvironment feature.
[0046] Step S230: Input the cell morphological features and microenvironmental features into the cross-modal alignment module. The module generates a feature alignment matrix by calculating the position correlation between the two types of features in the spatial coordinate system. Each element in the matrix represents the degree of matching between the cell morphological features and the microenvironmental features at the corresponding spatial position.
[0047] The cross-modal alignment module is used to handle the alignment problem between different modal features. In an embodiment of the present invention, the cell morphological features and microenvironmental features are input into the cross-modal alignment module to determine the positional correlation between the two types of features in the spatial coordinate system. The spatial coordinate system can be a pixel coordinate system of an image, and the correlation between the features is determined by associating the features with the pixel positions of the image. The feature alignment matrix is a two-dimensional matrix, and each element in the matrix represents the degree of matching between the cell morphological features and the microenvironmental features at the corresponding spatial position.
[0048] The cross-modal alignment module can be implemented, for example, using a neural network based on an attention mechanism. The attention mechanism focuses on the important parts of the features, thereby improving the accuracy of feature alignment. Specifically, the cell morphological features and microenvironmental features are encoded and converted into vectors of the same dimension. Then, the attention mechanism is used to calculate the correlation between the two types of features to generate a feature alignment matrix. For example, a multi-layer perceptron (MLP) is used to encode cell morphological features and microenvironmental features and convert them into 128-dimensional vectors. Then, the attention mechanism is used to calculate the correlation between the two vectors to generate a 128×128 feature alignment matrix. Each element in the matrix represents the degree of match between a certain dimension of the cell morphological feature and a certain dimension of the microenvironmental feature, and the larger the value, the higher the degree of match.
[0049] Step S240: A multi-level feature interaction mechanism is used to process the feature alignment matrix, wherein the primary interaction layer captures the morphology-microenvironment association pattern in the local area through convolution operations, and the advanced interaction layer establishes cross-regional feature dependencies through a self-attention mechanism.
[0050] The multi-level feature interaction mechanism is used to process the interaction between features and extract the correlation information between features at different levels. The primary interaction layer focuses on the feature interaction in the local area and captures the morphology-microenvironment correlation pattern in the local area through convolution operation. The convolution operation can extract local features in the image. The advanced interaction layer focuses on the cross-regional feature interaction and establishes the cross-regional feature dependency through the self-attention mechanism. The self-attention mechanism can automatically focus on the important parts of the features, thereby establishing long-distance dependencies between features.
[0051] The primary interaction layer can use a convolutional neural network (CNN) to gradually extract features in the image through the convolution operation of the convolution layer and the downsampling operation of the pooling layer. In an embodiment of the present invention, a CNN is used to process the feature alignment matrix to capture the morphology-microenvironment association pattern in the local area, and the advanced interaction layer can use a Transformer model. In an embodiment of the present invention, the feature alignment matrix is regarded as a sequence, and the Transformer model is used to process it to establish cross-regional feature dependencies. For example, for the feature alignment matrix, a CNN containing 3 convolutional layers and 2 pooling layers is first used to process it to capture the morphology-microenvironment association pattern in the local area. Then, the output of the CNN is input into the Transformer model, and the self-attention mechanism is used to establish cross-regional feature dependencies.
[0052] Step S250: Perform dynamic gating fusion on the interactively processed features to generate a four-dimensional feature tensor including cell nucleus level morphological features, intercellular connection morphological features, matrix component microenvironment features, and immune infiltration microenvironment features.
[0053] Dynamic gated fusion can dynamically adjust the weight of the feature according to the importance of the feature. In an embodiment of the present invention, dynamic gated fusion is performed on the interactively processed features, with the purpose of fusing the nuclear hierarchical morphological features, intercellular connection morphological features, matrix component microenvironment features, and immune infiltration microenvironment features to generate a four-dimensional feature tensor. The four-dimensional feature tensor contains comprehensive information about cells and their microenvironments, which helps to improve the recognition accuracy of subsequent models. Dynamic gated fusion can specifically use a gated recurrent unit (GRU) to dynamically adjust the gating signal according to the input features, thereby controlling the fusion process of the features. Specifically, the interactively processed features are first input into the GRU, and the GRU generates a gating signal according to the input features, and then the features are weighted and fused according to the gating signal to generate a four-dimensional feature tensor. For example, assuming that the interactively processed features include nuclear hierarchical morphological features, intercellular connection morphological features, matrix component microenvironment features, and immune infiltration microenvironment features, these features are input into a GRU. GRU generates a gating signal based on the input features. For example, for the nuclear level morphological features, a weight coefficient of 0.4 is generated, for the intercellular connection morphological features, a weight coefficient of 0.3 is generated, for the matrix component microenvironment features, a weight coefficient of 0.2 is generated, and for the immune infiltration microenvironment features, a weight coefficient of 0.1 is generated. Then, the features are weighted and fused according to these weight coefficients to generate a four-dimensional feature tensor.
[0054] Step S260: Input the four-dimensional feature tensor into the compression excitation network, strengthen the morphological and microenvironmental feature channels that are sensitive to diagnosis through the channel attention mechanism, suppress redundant feature channels, and finally output the fusion feature vector after dimension compression.
[0055] The compression excitation network is a network for feature compression and feature selection. It strengthens the feature channels sensitive to diagnosis and suppresses redundant feature channels through the channel attention mechanism. The channel attention mechanism is a mechanism that automatically pays attention to the important parts of the feature channel, and can dynamically adjust the weight of the channel according to the importance of the feature channel. In an embodiment of the present invention, the four-dimensional feature tensor is input into the compression excitation network for the purpose of compressing and selecting the features and outputting the fused feature vector after dimension compression. The compression excitation network can be specifically implemented by the existing Squeeze-and-Excitation (SE) module. The SE module consists of two parts: Squeeze and Excitation. The squeezing part compresses the spatial dimension of the four-dimensional feature tensor into a scalar through global average pooling to obtain the global feature information of each channel. The excitation part generates the weight coefficient of each channel according to the global feature information through a fully connected layer and a Sigmoid activation function. Then, the weight coefficient is multiplied by the original feature channel to strengthen the feature channel sensitive to diagnosis and suppress the redundant feature channel. Finally, the processed features are dimensionally compressed and the fused feature vector is output. For example, for a four-dimensional feature tensor, first use global average pooling to compress its spatial dimension into a scalar to obtain the global feature information of each channel. Then, the global feature information is input into a fully connected layer and the weight coefficient of each channel is obtained through the Sigmoid activation function.
[0056] Step S300: Construct a deep convolutional network with a multi-stage training strategy, in which the first stage uses a cross-channel feature attention mechanism to optimize the convolution kernel parameters, the second stage enhances feature discriminability by generating adversarial samples, and the third stage uses knowledge distillation to compress the network scale.
[0057] A deep convolutional network is a deep learning model based on a convolutional neural network that can automatically learn features in an image. A multi-stage training strategy divides the training process into multiple stages, and each stage uses a different training method to achieve different training goals. In an embodiment of the present invention, a deep convolutional network with a multi-stage training strategy is constructed. In the first stage, a cross-channel feature attention mechanism is used to optimize the convolution kernel parameters, with the aim of making the convolution kernel pay more attention to feature channels that are sensitive to diagnosis. In the second stage, feature discriminability is enhanced by generating adversarial samples. Adversarial samples are samples that can mislead the model after perturbation. By generating adversarial samples, the robustness and feature discrimination ability of the model can be improved. In the third stage, knowledge distillation is used to compress the network scale. Knowledge distillation is a method of transferring knowledge from a large model to a small model. Knowledge distillation can reduce the scale of the model without losing too much performance.
[0058] The cross-channel feature attention mechanism can be implemented in a similar way to the SE module. By calculating the spatial correlation between each feature channel and the nuclear atypia index, a channel enhancement coefficient matrix is generated, and then the convolution kernel is weighted channel by channel according to the channel enhancement coefficient matrix to optimize the convolution kernel parameters. Adversarial sample generation can adopt a gradient-based method, such as the fast gradient sign method (FGSM), to generate adversarial samples by adding perturbations to the original samples. Knowledge distillation can adopt the teacher-student model method, with a large model as the teacher model and a small model as the student model. By letting the student model learn the output of the teacher model, knowledge transfer is achieved. For example, a deep convolutional network containing multiple convolutional layers and fully connected layers is constructed. In the first stage, the cross-channel feature attention mechanism is used to optimize the convolution kernel parameters so that the convolution kernel pays more attention to the feature channels that are sensitive to diagnosis. In the second stage, the FGSM method is used to generate adversarial samples, and the adversarial samples and original samples are input into the network for training together to enhance feature discriminability. In the third stage, the large model is used as the teacher model and the small model is used as the student model. The knowledge of the teacher model is transferred to the student model using the knowledge distillation method to compress the network size.
[0059] As an implementation method, the above multi-stage training strategy may specifically include the following steps S310 to S340:
[0060] Step S310: Input the initial parameters of the deep convolutional network into the first-stage channel optimization module, generate a channel enhancement coefficient matrix by calculating the spatial correlation between the activation map of each feature channel and the nuclear atypia index, weight the convolution kernel channel by channel based on the channel enhancement coefficient matrix, and output the first-stage optimization parameters.
[0061] In an embodiment of the present invention, the initial parameters of the deep convolutional network are input into the first-stage channel optimization module, and the channel enhancement coefficient matrix is generated by calculating the spatial correlation between the activation map of each feature channel and the nuclear atypia index. The feature channel activation map is the feature map output by the convolution layer, which reflects the activation of each feature channel on the image. The nuclear atypia index is an index extracted previously for measuring the degree of abnormal nuclear morphology. The channel enhancement coefficient matrix is a two-dimensional matrix, and each element in the matrix represents the enhancement weight of the corresponding feature channel. The spatial correlation between the activation map of each feature channel and the nuclear atypia index can be calculated using the mutual information method. Mutual information is an index used to measure the statistical dependence between two random variables. In an embodiment of the present invention, the spatial correlation between the feature channel activation map and the nuclear atypia index is measured by calculating the mutual information value of the feature channel activation map and the nuclear atypia index. The specific implementation process includes the following steps: First, the activation map set output by each feature channel of the deep convolutional network is obtained, and the activation map set contains the feature response strength of different convolutional layers at different spatial positions. Then, the spatial distribution map corresponding to the nuclear atypia index is extracted, and each pixel position of the spatial distribution map stores the atypia index value of the nuclear at that position. Next, a spatial normalization operation is performed on the activation map set, and the activation map and spatial distribution map of each feature channel are adjusted to the same size and a pixel-level position mapping relationship is established. After that, the mutual information value of each feature channel in the spatial dimension is calculated based on the normalized activation map and spatial distribution map. The mutual information value represents the statistical dependence of the activation intensity and heteromorphism index of the channel. A nonlinear transformation operation is performed on the mutual information value, and the mutual information value with an absolute value lower than the set threshold is set to zero and the positive value response is retained to generate the preliminary weight coefficient of each feature channel. For the preliminary weight coefficient, a smoothing filter is performed according to the difference in weight coefficients of adjacent feature channels to eliminate isolated noise points and enhance spatial coherence, and the optimized weight coefficient is obtained. Based on the optimized weight coefficient, the initial channel enhancement coefficient matrix is generated, and each element in the initial channel enhancement coefficient matrix corresponds to the enhancement weight of the feature channel. Finally, the coefficient values of the elements in the initial channel enhancement coefficient matrix that exceed the preset interval are clipped to the interval boundary and rescaled to the target numerical range to obtain the channel enhancement coefficient matrix. For example, for a deep convolutional network, obtain the activation map set output by each feature channel and the spatial distribution map corresponding to the nuclear atypia index. Perform spatial normalization on the activation map set to make it the same size as the spatial distribution map and establish a pixel-level position mapping relationship. Then, calculate the mutual information value of each feature channel. Assuming that the mutual information value of a feature channel is 0.6, after nonlinear transformation and threshold processing, the initial weight coefficient is 0.6. Smooth filter the initial weight coefficient to obtain an optimized weight coefficient of 0.65.The initial channel enhancement coefficient matrix is generated based on the optimized weight coefficient. Assuming that an element in the matrix is 1.2, which exceeds the preset interval [0,1], it is clipped to 1 and rescaled to the target numerical range to obtain the channel enhancement coefficient matrix. Finally, the convolution kernel is weighted channel by channel according to the channel enhancement coefficient matrix, and the first stage optimization parameters are output. Based on the above process, the first stage channel optimization module may include a feature extraction unit, a correlation calculation unit, a weight generation unit and a convolution kernel weighting unit. The feature extraction unit completes the acquisition of the activation map set and the extraction of the spatial distribution map. The correlation calculation unit completes the spatial normalization and mutual information calculation. The weight generation unit completes the nonlinear transformation, smoothing filtering and matrix generation adjustment. The convolution kernel weighting unit weights the convolution kernel channel by channel according to the channel enhancement coefficient matrix and outputs the first stage optimization parameters. For the actions performed by each unit, those skilled in the art can select a specific execution structure according to actual needs. For example, feature extraction can select a convolution layer, normalization can select bilinear interpolation, smoothing filtering can select Gaussian filtering, etc., and there is no specific limitation.
[0062] As an implementation manner, step S310, calculating the spatial correlation between each characteristic channel activation map and the nuclear atypia index to generate a channel enhancement coefficient matrix, may specifically include the following steps S311-S318:
[0063] Step S311: Obtain a set of activation maps output by each feature channel of the deep convolutional network, where the activation map set contains the feature response strengths of different convolutional layers at different spatial positions.
[0064] The activation map is a feature map output by the convolution layer, which reflects the activation of each feature channel on the image. The activation maps of different convolutional layers contain feature information at different levels. The activation map of the shallow convolutional layer mainly reflects the local features of the image, and the activation map of the deep convolutional layer mainly reflects the global features of the image. In an embodiment of the present invention, a set of activation maps output by each feature channel of the deep convolutional network is obtained, with the purpose of collecting the feature response strengths of different convolutional layers at different spatial positions, in preparation for the subsequent calculation of spatial correlation.
[0065] The method for obtaining the activation map set can be to add an output layer after each convolution layer of the deep convolution network, save the output of each convolution layer, and form an activation map set. For example, for a deep convolution network containing 5 convolution layers, add an output layer after each convolution layer. When an image is input, the output of each convolution layer is recorded to obtain 5 activation maps, which constitute the activation map set. Each activation map is a three-dimensional tensor, which contains three dimensions: feature channel, height and width. The feature channel represents different feature types, the height and width represent the spatial position of the image, and each element in the activation map represents the feature response intensity of the corresponding feature channel at the corresponding spatial position.
[0066] Step S312: extracting the spatial distribution map corresponding to the cell nuclear atypia index, and storing the atypia index value of the cell nucleus at each pixel position in the spatial distribution map.
[0067] The nuclear atypia index is an index extracted above for measuring the degree of abnormal nuclear morphology. The spatial distribution map is a two-dimensional image, in which each pixel position stores the atypia index value of the nucleus at that position. In the embodiment of the present invention, the spatial distribution map corresponding to the nuclear atypia index is extracted, the purpose of which is to associate the nuclear atypia index with the spatial position of the image, and to provide a basis for the subsequent calculation of spatial correlation.
[0068] The method for extracting the spatial distribution map can be to map the nuclear atypia index to each pixel position of the image based on the previously extracted nuclear atypia index and the pixel position information of the image. For example, for a multi-channel pathological sample image, the atypia index of each cell nucleus has been calculated through the aforementioned steps, and the position information of each cell nucleus has been determined. The atypia index of each cell nucleus is assigned to the spatial distribution map of the pixel position where it is located. For the pixel position without a cell nucleus, its atypia index value can be set to 0. In this way, the spatial distribution map corresponding to the nuclear atypia index is obtained.
[0069] Step S313: performing a spatial normalization operation on the activation map set, adjusting the activation map and the spatial distribution map of each feature channel to the same size and establishing a pixel-level position mapping relationship.
[0070] The spatial normalization operation is an operation that adjusts images of different sizes to the same size, which can ensure that the activation map and the spatial distribution map of each feature channel are consistent in spatial position. In an embodiment of the present invention, the spatial normalization operation is performed on the activation map set, the purpose of which is to adjust the activation map and the spatial distribution map of each feature channel to the same size and establish a pixel-level position mapping relationship to facilitate the subsequent calculation of spatial correlation.
[0071] As mentioned above, the spatial normalization operation can use the bilinear interpolation method to calculate the grayscale value of the scaled pixel by linearly interpolating the grayscale values of adjacent pixels. Specifically, for each activation map and spatial distribution map, the scaling ratio is calculated according to the target size, and then the activation map and the spatial distribution map are adjusted to the target size using the bilinear interpolation method. At the same time, a pixel-level position mapping relationship is established, that is, the correspondence between the pixel positions in each activation map and the spatial distribution map is recorded. Exemplarily, for an activation map and a spatial distribution map, the target size is 256×256 pixels. Calculate the scaling ratio of the activation map and the spatial distribution map, assuming that the original size of the activation map is 512×512 pixels and the scaling ratio is 0.5, and use the bilinear interpolation method to adjust the activation map to 256×256 pixels. Similarly, the spatial distribution map is also adjusted to 256×256 pixels. At the same time, a correspondence between the pixel positions in the activation map and the spatial distribution map is established, such that the pixel at position (100,100) in the activation map corresponds to the pixel at position (100,100) in the spatial distribution map.
[0072] Step S314: Based on the normalized activation map and the spatial distribution map, the mutual information value of each feature channel in the spatial dimension is calculated, and the mutual information value represents the statistical dependence of the activation intensity of the channel and the heteromorphism index.
[0073] Mutual information is an indicator for measuring the statistical dependence between two random variables. In an embodiment of the present invention, the mutual information value of each feature channel in the spatial dimension is calculated based on the normalized activation map and the spatial distribution map, with the purpose of measuring the statistical dependence between the activation intensity of each feature channel and the nuclear atypia index. The larger the mutual information value, the stronger the correlation between the activation intensity of the feature channel and the nuclear atypia index.
[0074] The method for calculating the mutual information value can be to use the mutual information formula in information theory (general formula, not repeated here). Specifically, the normalized activation map and spatial distribution map are regarded as two random variables, their pixel values are discretized into several intervals, and then the number of pixels in each interval is counted, and the joint probability distribution and marginal probability distribution are calculated. Finally, the mutual information value is calculated according to the mutual information formula.
[0075] Step S315: Perform a nonlinear transformation operation on the mutual information value, set the mutual information value whose absolute value is lower than the set threshold to zero and retain the positive value response, and generate the preliminary weight coefficient of each feature channel.
[0076] Nonlinear transformation operations can change the distribution and characteristics of data. In the embodiment of the present invention, nonlinear transformation operations are performed on the mutual information values, the purpose of which is to remove noise information in the mutual information values whose absolute values are lower than a set threshold, retain positive responses that are important for diagnosis, and generate preliminary weight coefficients for each feature channel.
[0077] Nonlinear transformation operations can be performed using threshold processing, for example. A threshold is set, and the part of the mutual information value whose absolute value is lower than the threshold is set to zero, and the positive value response is retained. For example, if the threshold is set to 0.1, the mutual information value of a feature channel is 0.05, and it is set to zero; for the feature channel with a mutual information value of 0.2, its value is retained. After threshold processing, the preliminary weight coefficients of each feature channel are obtained.
[0078] Step S316: For the preliminary weight coefficients, smoothing filtering is performed according to the differences in weight coefficients of adjacent feature channels to eliminate isolated noise points and enhance spatial coherence, thereby obtaining optimized weight coefficients.
[0079] Smoothing filtering is used to remove noise from data, and can make data smoother by averaging or weighted averaging adjacent data points. In an embodiment of the present invention, for the preliminary weight coefficient, smoothing filtering is performed according to the difference in weight coefficients of adjacent feature channels, the purpose is to eliminate isolated noise points in the preliminary weight coefficient, enhance spatial coherence, and obtain optimized weight coefficients. Smoothing filtering can use the Gaussian filtering method to perform weighted averaging according to the distance and weight of adjacent data points. Specifically, for the preliminary weight coefficient of each feature channel, the weighted average is calculated using a Gaussian function according to the difference in weight coefficients of its adjacent feature channels to obtain the optimized weight coefficient.
[0080] Step S317: Generate an initial channel enhancement coefficient matrix based on the optimized weight coefficients, where each element in the initial channel enhancement coefficient matrix corresponds to the enhancement weight of the feature channel.
[0081] In an embodiment of the present invention, an initial channel enhancement coefficient matrix is generated based on the optimized weight coefficients, the purpose of which is to organize the optimized weight coefficients into a matrix. The method for generating the initial channel enhancement coefficient matrix may be to arrange the optimized weight coefficients into a matrix in the order of feature channels. For example, assuming that the deep convolutional network has 100 feature channels, the optimized weight coefficients are w1, w2, ..., w100, and these weight coefficients are arranged into a 1×100 matrix as the initial channel enhancement coefficient matrix.
[0082] Step S318: The coefficient values of the elements in the initial channel enhancement coefficient matrix that exceed the preset interval are clipped to the interval boundary and rescaled to the target value range to obtain the channel enhancement coefficient matrix.
[0083] The preset interval is used to limit the value of the channel enhancement coefficient. In an embodiment of the present invention, the element coefficient values in the initial channel enhancement coefficient matrix that exceed the preset interval are clipped to the interval boundary and rescaled to the target numerical range, in order to ensure that the element coefficient values in the channel enhancement coefficient matrix are within a reasonable range, and to avoid the occurrence of too large or too small coefficient values affecting the weighted effect of the convolution kernel. The method of clipping and rescaling can be: first determine the preset interval, such as [0, 1]. For the element coefficient values in the initial channel enhancement coefficient matrix that exceed the interval, clip them to the interval boundary, such as clipping element coefficient values greater than 1 to 1, and clipping element coefficient values less than 0 to 0. Then, rescale the clipped matrix to the target numerical range, such as linearly mapping the element coefficient values in the matrix to the range of [0.1, 0.9]. For example, for an initial channel enhancement coefficient matrix, where a certain element coefficient value is 1.2, it is clipped to 1. Assuming the target numerical range is [0.1, 0.9], the element coefficient values in the clipped matrix are linearly mapped to the range to obtain the channel enhancement coefficient matrix.
[0084] Step S320: Input the first-stage optimization parameters into the second-stage adversarial training module, and generate a perturbation sample set based on the optimized convolution kernel. The perturbation samples in the perturbation sample set are constructed by adjusting the spatial co-occurrence pattern of the intercellular connection density index and the immune cell aggregation density, and maintaining the continuity constraint of the matrix fiber orientation characteristics.
[0085] The second-stage adversarial training module is a module for generating adversarial samples and performing adversarial training. In an embodiment of the present invention, the first-stage optimization parameters are input into the second-stage adversarial training module, and a set of perturbed samples is generated based on the optimized convolution kernel. Perturbed samples are samples that can mislead the model after perturbation. By generating perturbed samples, the robustness and feature discrimination ability of the model can be improved.
[0086] Adjusting the spatial co-occurrence pattern of the intercellular connection density index and the immune cell aggregation density means adjusting the spatial distribution of the intercellular connection density index and the immune cell aggregation density in the image to make it different from the original sample. Maintaining the continuity constraint of the matrix fiber orientation feature means maintaining the continuity of the matrix fiber orientation feature in the process of generating the perturbation sample to avoid unreasonable changes.
[0087] The method for generating perturbation samples can be a gradient-based method, such as the fast gradient sign method (FGSM). Specifically, the original sample is first input into a deep convolutional network to calculate the gradient of the loss function to the input sample. Then, according to the sign of the gradient and the preset perturbation intensity, the original sample is perturbed to generate a perturbation sample. During the perturbation process, the spatial co-occurrence pattern of the cell-to-cell connection density index and the immune cell aggregation density is adjusted, and the continuity constraint of the matrix fiber orientation feature is maintained. For example, for an original sample, it is input into a deep convolutional network to calculate the gradient of the loss function to the input sample. Assuming that the perturbation intensity is 0.01, the original sample is perturbed according to the sign of the gradient, and the spatial co-occurrence pattern of the cell-to-cell connection density index and the immune cell aggregation density is adjusted, and the continuity constraint of the matrix fiber orientation feature is maintained to generate a perturbation sample. Repeat this process to generate a set of perturbation samples.
[0088] It can be understood that, based on the above description, those skilled in the art can select the corresponding implementation method to realize the architecture of the second-stage adversarial training module according to actual needs. For example, a generator-discriminator structure can be adopted. The generator generates a set of perturbation samples based on the convolution kernel optimized in the first stage by using the gradient ascent method, adjusts the spatial co-occurrence pattern of the intercellular connection density index and the immune cell aggregation density, and maintains the continuity of the matrix fiber orientation characteristics; the discriminator is a deep convolutional network, which updates the fully connected layer parameters by comparing the distribution offset of the original sample and the perturbation sample in the microenvironment feature space using the stochastic gradient descent method.
[0089] Step S330: Alternately input the perturbed sample set and the original sample into the deep convolutional network, update the fully connected layer parameters by comparing the distribution offset of the original features and the perturbed features in the microenvironment feature space, and output the second stage optimization parameters.
[0090] In an embodiment of the present invention, the perturbed sample set and the original sample are alternately input into the deep convolutional network, the purpose of which is to allow the model to learn the features of the original sample and the perturbed sample at the same time, thereby improving the robustness and feature discrimination ability of the model. The distribution offset of the original feature and the perturbed feature in the microenvironment feature space is compared, that is, the distance between the feature vectors of the original sample and the perturbed sample in the microenvironment feature space is calculated, and the distribution offset of the feature is measured by this distance.
[0091] The method for updating the parameters of the fully connected layer can be to use the stochastic gradient descent (SGD) algorithm to gradually reduce the loss function by continuously updating the parameters of the model. Specifically, the perturbed sample set and the original sample are first input alternately into the deep convolutional network to calculate the distribution offset of the original features and the perturbed features in the microenvironment feature space. Then, the loss function is calculated based on the distribution offset, and the fully connected layer parameters are updated using the SGD algorithm. For example, for a deep convolutional network, the perturbed sample set and the original sample are input alternately into the network to calculate the distribution offset of the original features and the perturbed features in the microenvironment feature space. Assuming that the distribution offset is 0.5, the loss function is calculated based on this offset, and the fully connected layer parameters are updated using the SGD algorithm. Repeat this process until the loss function converges and output the second stage optimization parameters.
[0092] Step S340: Input the second-stage optimization parameters into the third-stage compression module, match the feature activation path differences between the teacher network and the student network in the process of nuclear atypia index extraction, use the channel importance ranking algorithm to remove redundant channels, and output the lightweight third-stage optimization parameters.
[0093] The third stage compression module is used to compress the network scale. In an embodiment of the present invention, the second stage optimization parameters are input into the third stage compression module. By matching the feature activation path differences between the teacher network and the student network in the process of extracting the nuclear atypia index, the channel importance ranking algorithm is used to remove redundant channels, with the purpose of reducing the scale of the network without losing too much performance. The teacher network is a large deep convolutional network with high performance. The student network is a small deep convolutional network with a small scale. The channel importance ranking algorithm is an algorithm for evaluating the importance of channels, and channels can be ranked according to their contribution to the model performance, thereby finding redundant channels.
[0094] The method of matching the difference in feature activation paths between the teacher network and the student network in the process of extracting the nuclear atypia index can be to calculate the difference between the feature activation maps of the two networks under the same input. The channel importance ranking algorithm can adopt a gradient-based method, such as calculating the gradient of each channel to the loss function, and ranking the channels according to the size of the gradient. Specifically, first, the second-stage optimization parameters are loaded into the teacher network and the student network respectively, the same sample is input, and the difference between the feature activation maps of the two networks in the process of extracting the nuclear atypia index is calculated. Then, the channels of the student network are sorted using the channel importance ranking algorithm to find redundant channels. Finally, the redundant channels are removed and the lightweight third-stage optimization parameters are output. For example, for a teacher network and a student network, the second-stage optimization parameters are loaded into the two networks respectively, a multi-channel pathological sample image is input, and the difference between the feature activation maps of the two networks in the process of extracting the nuclear atypia index is calculated. The channels of the student network are sorted using the gradient-based channel importance ranking algorithm. Assuming that the last 20% of the channels contribute less to the model performance after sorting, these channels are removed and the lightweight third-stage optimization parameters are output. It is understandable that, based on the above description, those skilled in the art can select a corresponding implementation method to implement the architecture of the third-stage compression module according to actual needs, such as using a teacher-student network structure.
[0095] The above three stages adopt a progressive parameter transfer mechanism, in which the convolution kernel of the second stage inherits the channel weighted result of the optimized parameters in the first stage, the fully connected layer of the third stage inherits the spatial distribution pattern of the optimized parameters in the second stage, and the training samples in each stage dynamically adjust the sampling ratio according to the recognition error of the cell membrane curvature parameters in the previous stage.
[0096] The progressive parameter transfer mechanism is to transfer the optimized parameters of the previous stage to the next stage in the multi-stage training process as the initial parameters of the next stage training. In the embodiment of the present invention, the convolution kernel of the second stage inherits the channel weighted result of the optimized parameters of the first stage, that is, the convolution kernel of the second stage uses the convolution kernel parameters optimized in the first stage. The fully connected layer of the third stage inherits the spatial distribution pattern of the optimized parameters of the second stage, that is, the fully connected layer of the third stage uses the fully connected layer parameters optimized in the second stage.
[0097] The sampling ratio of training samples in each stage is dynamically adjusted according to the recognition error of cell membrane curvature parameters in the previous stage. That is, according to the recognition error of cell membrane curvature parameters by the model in the previous stage, the sampling ratio of different types of samples in the training samples in the current stage is adjusted. If the recognition error of cell membrane curvature parameters of a certain type of sample in the previous stage is large, the sampling ratio of such samples in the training samples in the current stage is increased to improve the model's recognition ability of such samples. For example, after the first stage of training, it was found that the model had a large recognition error in the cell membrane curvature parameters of epithelial tumor cells. In the second stage of training, the sampling ratio of epithelial tumor cell samples in the training samples was increased.
[0098] As an implementation method, the process of generating the adversarial sample may specifically include the following steps S21-S26:
[0099] Step S21: extracting the characteristic gradient distribution of the deep convolutional network on the multi-channel pathological sample image set, and identifying the cross-modal gradient association pattern of the cell nuclear atypia index and the immune cell aggregation density based on the characteristic gradient distribution.
[0100] The characteristic gradient distribution is the gradient distribution of the loss function to the input image when the deep convolutional network inputs a set of multi-channel pathological sample images. The cross-modal gradient association pattern is the association pattern between the nuclear atypia index and the immune cell aggregation density in the characteristic gradient distribution. In an embodiment of the present invention, the characteristic gradient distribution of the deep convolutional network on the multi-channel pathological sample image set is extracted, and the purpose is to identify the cross-modal gradient association pattern between the nuclear atypia index and the immune cell aggregation density by analyzing the characteristic gradient distribution.
[0101] Specifically, extracting the characteristic gradient distribution can be inputting a set of multi-channel pathological sample images into a deep convolutional network, calculating the loss function, and then using an automatic derivation tool to calculate the gradient of the loss function to the input image. Based on the characteristic gradient distribution, the cross-modal gradient association pattern is identified by analyzing the relationship between the gradient values corresponding to the nuclear atypia index and the immune cell aggregation density. For example, by calculating the correlation between the gradient values corresponding to the nuclear atypia index and the immune cell aggregation density, the association pattern between them is found.
[0102] Step S22: generating a disturbance direction template according to the cross-modal gradient association pattern, wherein the disturbance direction template includes a pixel offset direction that produces a coordinated gradient change in the cell nucleus morphological parameters and the matrix fiber orientation characteristics.
[0103] The perturbation direction template is a template used to guide the generation of perturbation samples, and includes the pixel offset direction that produces a synergistic gradient change in the cell nucleus morphology parameters and the matrix fiber orientation characteristics. In an embodiment of the present invention, the perturbation direction template is generated according to the cross-modal gradient association pattern, with the purpose of providing guidance for the subsequent generation of perturbation samples, so that the perturbation samples can maintain the synergistic change relationship between the cell nucleus morphology parameters and the matrix fiber orientation characteristics while changing them.
[0104] The method for generating a perturbation direction template can be to find the pixel position and offset direction that produce a synergistic gradient change in the cell nuclear morphology parameters and the matrix fiber orientation characteristics according to the cross-modal gradient association pattern. Specifically, the cross-modal gradient association pattern is analyzed to find the pixel positions where the gradient values corresponding to the cell nuclear morphology parameters and the matrix fiber orientation characteristics increase or decrease at the same time, and the offset directions of these pixels are recorded to generate a perturbation direction template. For example, for a cross-modal gradient association pattern, when it is found that the gradient value of a certain pixel affects the cell nuclear morphology parameters and the matrix fiber orientation characteristics at the same time, the offset direction of the pixel is recorded, and this process is repeated to generate a perturbation direction template.
[0105] Step S23: Input the multi-channel pathological sample image into the deep convolutional network, extract the spatial gradient response map of the fused feature vector, and generate an initial perturbation mask based on the distribution similarity between the spatial gradient response map and the perturbation direction template.
[0106] The spatial gradient response map is the gradient distribution of the fused feature vector in the spatial dimension. The initial perturbation mask is a binary mask used to indicate which pixels need to be perturbed. In an embodiment of the present invention, a multi-channel pathological sample image is input into a deep convolutional network, the spatial gradient response map of the fused feature vector is extracted, and an initial perturbation mask is generated based on the distribution similarity between the spatial gradient response map and the perturbation direction template, in order to determine which pixels need to be perturbed to generate effective perturbation samples.
[0107] The method for extracting the spatial gradient response map may be to use an automatic differentiation tool, such as an autograd block. Specifically, a multi-channel pathological sample image is input into a deep convolutional network, a fused feature vector is calculated, and then the gradient of the fused feature vector to the input image is calculated using an automatic differentiation tool to obtain a spatial gradient response map. Based on the distribution similarity between the spatial gradient response map and the perturbation direction template, by calculating the correlation between them, determine which pixels have a gradient distribution similar to the perturbation direction template, set the positions corresponding to these pixels to 1, and set other positions to 0 to generate an initial perturbation mask. For example, for a spatial gradient response map and a perturbation direction template, calculate the correlation between them, assume that the correlation threshold is 0.8, set the pixel positions with a correlation greater than 0.8 to 1, and set other positions to 0 to generate an initial perturbation mask.
[0108] Step S24: Perform a morphological boundary constraint operation on the initial disturbance mask to eliminate isolated disturbance points that have no intersection with the cytoplasm region boundary, and retain the disturbance region that overlaps with the cell nucleus contour boundary data space.
[0109] The morphological boundary constraint operation is used to process the binary mask, and the shape and boundary of the mask are adjusted by morphological transformation. In an embodiment of the present invention, the morphological boundary constraint operation is performed on the initial perturbation mask to eliminate isolated perturbation points that have no intersection with the cytoplasm region boundary, retain the perturbation area that overlaps with the nuclear contour boundary data space, and make the perturbation sample more reasonable and effective. The morphological boundary constraint operation can use expansion and corrosion operations. The expansion operation can expand the white area in the binary mask, and the corrosion operation can reduce the white area. Specifically, the initial perturbation mask is firstly expanded to expand the perturbation area, and then the intersection operation is performed with the cytoplasm region boundary to eliminate the isolated perturbation points that have no intersection with the cytoplasm region boundary. Then, the processed mask is corroded to restore the perturbation area to its original size. Finally, the intersection operation is performed with the nuclear contour boundary data to retain the perturbation area that overlaps with the nuclear contour boundary data space. For example, for an initial perturbation mask, the perturbation area is expanded using the expansion operation, and then the intersection operation is performed with the cytoplasm region boundary to eliminate the isolated perturbation points. Next, the perturbed area is restored to its original size using the corrosion operation, and finally an intersection operation is performed with the cell nucleus contour boundary data to retain the perturbed area that overlaps with the cell nucleus contour boundary data space.
[0110] Step S25: superimpose the retained disturbance area onto the original multi-channel pathological sample image, generate a disturbance image and input it into a deep convolutional network, and calculate the consistency index between the change direction of the nuclear atypia index before and after the disturbance and the change direction of the immune cell aggregation density.
[0111] The consistency index is the degree of consistency between the change direction of the nuclear atypia index before and after the disturbance and the change direction of the immune cell aggregation density. In this embodiment of the present invention, the retained disturbance area is superimposed on the original multi-channel pathological sample image, the disturbance image is generated and input into the deep convolutional network, and the consistency index between the change direction of the nuclear atypia index before and after the disturbance and the change direction of the immune cell aggregation density is calculated, with the purpose of evaluating the effectiveness of the disturbance sample and ensuring that the disturbance sample can have the expected impact on the model.
[0112] The method of superimposing the retained disturbance region onto the original multi-channel pathology sample image is to adjust the pixel values of these pixels on the original image according to the pixel positions corresponding to the retained disturbance region. For example, if the pixel values of the pixel positions corresponding to the disturbance region need to be increased by a fixed offset, the pixel values of these pixels in the original image are added with the offset to generate a disturbed image.
[0113] The perturbed image is input into the deep convolutional network to extract the nuclear atypia index and immune cell aggregation density again. The consistency index of the change direction of the nuclear atypia index before and after the perturbation and the change direction of the immune cell aggregation density can be calculated in the following way. First, the difference of the nuclear atypia index before and after the perturbation is calculated. The difference between the density of immune cells and .like If it is positive, it means that the nuclear atypia index increases; if it is negative, it means that it decreases. There are also corresponding positive and negative values to indicate the increase or decrease in the density of immune cell aggregation. The consistency index can be defined as: and If the two numbers have the same sign (i.e. both are positive or both are negative), the consistency index is 1; if they have different signs, it is -1; if one of the differences is 0, a smaller consistency index value, such as 0, can be set according to the specific situation.
[0114] Step S26: Dynamically adjust the intensity distribution of the perturbation direction template based on the consistency index to generate an adversarial sample set, in which the perturbation intensity of each sample in the adversarial sample set is inversely correlated with the local change rate of the cell membrane curvature parameter.
[0115] Dynamically adjusting the intensity distribution of the perturbation direction template is to make the generated perturbation samples more effective and better able to mislead the model. Based on the consistency index, if the consistency index is high, it means that the current perturbation direction template can better change the nuclear atypia index and immune cell aggregation density, and the change direction of the two is consistent. At this time, the intensity of the perturbation direction template can be appropriately increased; if the consistency index is low, it is necessary to reduce the intensity of the perturbation direction template and adjust its direction.
[0116] The specific method of adjusting the intensity distribution of the disturbance direction template can be to set a basic intensity value S 0 , calculate the adjusted strength value S according to the consistency index C. For example, a linear adjustment method can be used, S=S 0 +k×C, where k is an adjustment factor. If C=1, the intensity is increased; if C=-1, the intensity is decreased.
[0117] When generating an adversarial sample set, the perturbation intensity of each sample is inversely correlated with the local change rate of the cell membrane curvature parameter. The local change rate of the cell membrane curvature parameter can be obtained by calculating the curvature difference between adjacent pixels of the cell membrane. For areas with a large local change rate, it means that the morphological changes of the cell membrane are more drastic. At this time, the perturbation intensity should be reduced to avoid excessive perturbation and sample distortion; for areas with a small local change rate, the perturbation intensity can be appropriately increased. Multiply the adjusted perturbation intensity by the perturbation direction template and superimpose it on the original sample to generate an adversarial sample. Repeat this process until a sufficient number of adversarial samples are generated to form an adversarial sample set.
[0118] As an implementation manner, the above-mentioned knowledge distillation and compression process may specifically include the following steps S31 to S36:
[0119] Step S31: Use the convolutional layer parameters of the deep convolutional network as the teacher network parameters, and initialize the convolutional layer parameters of the student network so that its channel dimension matches the teacher network parameters but the number of channels is reduced.
[0120] The convolutional layer parameters of the deep convolutional network are used as the parameters of the teacher network, that is, the weights and bias parameters of the convolutional layer of the trained deep convolutional network are saved as the parameters of the teacher network. When initializing the convolutional layer parameters of the student network, its channel dimension should match the parameters of the teacher network, which means that the student network and the teacher network are consistent in the dimension of the feature channel, but the number of channels of the student network should be reduced, which can reduce the size of the model to a certain extent. The method of initializing the convolutional layer parameters of the student network can use random initialization, such as using the Xavier initialization method, which can make the network have better convergence in the early stage of training.
[0121] Step S32: extracting a set of intermediate feature activation maps of the teacher network on the multi-channel pathological sample image, where the set of intermediate feature activation maps includes the spatial response distribution of cell morphological features and microenvironment features in different convolutional layers.
[0122] The intermediate feature activation map is the feature map output by the teacher network at different convolutional layers. These feature maps reflect the spatial response distribution of cell morphological features and microenvironmental features in different convolutional layers. The purpose of extracting the intermediate feature activation map set is to provide a learning target for the student network, so that the student network can learn the feature representation of the teacher network when processing multi-channel pathological sample images. The process of extracting the intermediate feature activation map set is to input the multi-channel pathological sample image into the teacher network and record the feature map at the output of each convolutional layer. Since the feature maps of different convolutional layers have different levels and semantic information, the feature map of the shallow convolutional layer mainly reflects the local features of the image, such as edges, textures, etc.; the feature map of the deep convolutional layer mainly reflects the global features of the image, such as the shape and category of the object. Therefore, the intermediate feature activation map set contains the spatial response distribution of cell morphological features and microenvironmental features at different levels. For example, for a teacher network, a multi-channel pathological sample image is input into the network, and the output feature map of each convolutional layer is recorded respectively. These feature maps constitute the intermediate feature activation map set.
[0123] Step S33: perform cross-layer spatial alignment on the intermediate feature activation map of the student network and the intermediate feature activation map of the teacher network, compensate for the resolution difference between the two through dilated convolution and generate a feature similarity distribution map.
[0124] Cross-layer spatial alignment is to align the intermediate feature activation maps output by the student network and the teacher network at different convolutional layers in spatial positions in order to compare their similarities. Since the structures of the student network and the teacher network may be different, the resolution of the intermediate feature activation maps may also be different, so it is necessary to compensate for this resolution difference through dilated convolution.
[0125] Atrous convolution is a convolution operation that introduces a hole in the convolution kernel. It can expand the receptive field of the convolution kernel without increasing the number of parameters, thereby improving the resolution of the feature map. Specifically, the intermediate feature activation maps of the student network and the teacher network are matched in the order of the convolution layers. For feature maps with different resolutions, the feature map of the student network is upsampled using atrous convolution to make its resolution the same as that of the feature map of the teacher network. Then, the similarity between the corresponding feature maps of the student network and the teacher network is calculated to generate a feature similarity distribution map. The similarity can be calculated by the mean square error (MSE), that is, calculating the average of the squares of the difference between the corresponding pixel values of the two feature maps. For example, for the feature map F output by the student network and the teacher network at a certain convolution layer, s and F t , calculate their mean square error , where N is the number of pixels in the feature map. The mean square error of each pixel is taken as the value of the corresponding pixel in the feature similarity distribution map.
[0126] Step S34: Dynamically adjust the channel attention coefficient of the student network based on the feature similarity distribution map, suppress channels with significant differences in response patterns from the teacher network and enhance the weights of spatially consistent channels.
[0127] The channel attention coefficient is used to adjust the importance of different feature channels in the student network. The channel attention coefficient of the student network is dynamically adjusted based on the feature similarity distribution map. The purpose is to make the student network pay more attention to the channels that are consistent with the response mode of the teacher network and suppress the channels that have significant differences from the response mode of the teacher network.
[0128] The process of dynamically adjusting the channel attention coefficient is, for example, to first calculate the average similarity score of each channel based on the feature similarity distribution map. For each channel, the pixel values of the feature similarity distribution map on the channel are averaged to obtain the average similarity score of the channel. Then, the channel attention coefficient is adjusted according to the average similarity score. For channels with higher average similarity scores, it means that the response mode of the channel is more consistent with that of the teacher network, and its channel attention coefficient is enhanced; for channels with lower average similarity scores, it means that there is a significant difference in the response mode of the channel and the teacher network, and its channel attention coefficient is suppressed. For example, a threshold T is set, and for channels with an average similarity score greater than T, its channel attention coefficient is multiplied by a coefficient greater than 1, such as 1.2; for channels with an average similarity score less than T, its channel attention coefficient is multiplied by a coefficient less than 1, such as 0.8.
[0129] Step S35: Input the compressed feature vector of the student network into the classifier, calculate the difference loss between it and the classification probability distribution of the teacher network, and generate an adaptive distillation loss coefficient in combination with the cell type annotation information.
[0130] The compressed feature vector of the student network is input into the classifier to obtain the classification probability distribution of the student network. The classification probability distribution of the teacher network is obtained by inputting the same samples into the classifier of the teacher network.
[0131] The difference loss between the classification probability distribution of the student network and the teacher network can be calculated using the KL divergence (Kullback-Leibler divergence), which represents the information loss of one probability distribution relative to another probability distribution. For example, for the classification probability distribution P of the student network s And the classification probability distribution P of the teacher network t ,calculate .
[0132] The purpose of generating the adaptive distillation loss coefficient by combining the cell type annotation information is to make the distillation loss more reasonable. The cell type annotation information is the true category label of each sample. The weight of the distillation loss can be dynamically adjusted according to the difference between the sample category and the classification probability distribution. For example, for samples that are misclassified, the weight of the distillation loss is increased; for samples that are correctly classified, the weight of the distillation loss is reduced. The adaptive distillation loss coefficient can be calculated by a function g, for example, , a is the preset adjustment coefficient.
[0133] Step S36: Iteratively update the fully connected layer parameters of the student network according to the adaptive distillation loss coefficient until the performance difference between the comprehensive performance index of the student network and the teacher network under the layered cross-validation mechanism is stabilized within a preset range, and the compression process is terminated.
[0134] Iteratively updating the parameters of the fully connected layer of the student network is achieved through optimization algorithms, such as stochastic gradient descent (SGD), adaptive moment estimation (Adam), etc. In each iteration, the loss function is calculated based on the adaptive distillation loss coefficient, and then the optimization algorithm is used to update the weights and biases of the fully connected layer.
[0135] The stratified cross-validation mechanism is used to evaluate model performance. The dataset is stratified by category, and then divided into multiple subsets. Each time, one subset is used as a validation set, and the remaining subsets are used as training sets. This process is repeated multiple times to obtain the performance indicators of the model on different subsets. Finally, these indicators are combined to obtain the comprehensive performance indicators of the model.
[0136] The preset range is a pre-set performance difference threshold. When the comprehensive performance index of the student network under the stratified cross-validation mechanism is stable within this range compared with the performance difference of the teacher network, it means that the student network has learned most of the knowledge of the teacher network, and the compression process can be terminated at this time. For example, the preset range can be set to a difference of less than 5% between the accuracy of the student network and the accuracy of the teacher network. After each iteration of updating the fully connected layer parameters of the student network, the performance of the student network is evaluated using the stratified cross-validation mechanism and compared with the performance of the teacher network. When the performance difference between the two is stable within the preset range, the training is stopped and the knowledge distillation compression process is completed.
[0137] Step S400: Use a stratified cross-validation mechanism to verify and adjust the network parameters of the deep convolutional network during the training process, wherein the stratified cross-validation mechanism divides the biological characteristics corresponding to the cell type annotation information into multiple validation subsets, and adjusts the network parameters based on the classification error of each validation subset.
[0138] The Stratified Cross-Validation mechanism can fully consider the category distribution of the data and avoid evaluation bias caused by data imbalance. In an embodiment of the present invention, the biological characteristics corresponding to the cell type annotation information are divided into multiple validation subsets because different types of cells have different biological characteristics, such as epithelial tumor cells and mesenchymal tumor cells. There are differences in morphology, function, etc. Stratifying the data according to cell type can make the data in each validation subset have similar biological characteristics, thereby more accurately evaluating the performance of the model on different types of cells.
[0139] The purpose of adjusting network parameters based on the classification error of each validation subset is to make the model achieve better performance on different types of cells. The classification error is the difference between the model's classification result of the sample and the true label. By calculating the classification error of each validation subset, we can understand the performance of the model on different types of cells. If the classification error of a validation subset is large, it means that the model has poor performance on this type of cell, and the network parameters need to be adjusted to improve the model's ability to recognize this type of cell.
[0140] As an implementation mode, step S400 uses a layered cross-validation mechanism to verify and adjust the network parameters of the deep convolutional network during training, which may specifically include the following steps S410-S450:
[0141] Step S410: dividing the multi-channel pathology sample image set into an epithelial tumor cell subset, a mesenchymal tumor cell subset, and a tumor cell subset with an undetermined degree of differentiation according to the biological characteristics corresponding to the cell type annotation information.
[0142] Epithelial tumor cells are tumor cells that originate from epithelial tissues and have specific morphology and biological characteristics, such as regular cell arrangement and obvious polarity. Mesenchymal tumor cells are tumor cells that originate from mesenchymal tissues, and their morphology and biological characteristics are different from those of epithelial tumor cells, such as diverse cell morphology and rich mesenchymal components. Tumor cells of undetermined differentiation are those whose degree of differentiation cannot be clearly determined, and their biological characteristics are more complex.
[0143] Classification is performed based on the biological characteristics corresponding to the cell type annotation information. Specifically, the cell type annotation information of each sample in the multi-channel pathology sample image set is first analyzed to determine whether it belongs to epithelial tumor cells, mesenchymal tumor cells, or tumor cells of an undetermined degree of differentiation.
[0144] Step S420: In each round of training iteration, the verification operation is performed in the order of epithelial tumor cell subset, mesenchymal tumor cell subset, and tumor cell subset with undetermined differentiation degree. The verification operation of each subset includes extracting the classification error of the nuclear atypia index and the regression error of the immune cell aggregation density of the current network parameters in the corresponding subset.
[0145] The validation operation is performed sequentially in each round of training iteration in order to comprehensively evaluate the performance of the model on different types of cell subsets. The classification error of the nuclear atypia index is the difference between the classification result of the model for the nuclear atypia index and the true classification, and the regression error of the immune cell aggregation density is the difference between the predicted value of the model for the immune cell aggregation density and the true value. The process of extracting the classification error of the nuclear atypia index is, for example, inputting the samples of the corresponding subset into the current network, obtaining the classification result of the model for the nuclear atypia index, and then comparing it with the true classification to calculate the classification error. For example, the classification error can be calculated using a cross entropy loss function, which can measure the difference between the classification result of the model and the true label. The method of extracting the regression error of the immune cell aggregation density is to input the samples of the corresponding subset into the current network, obtain the predicted value of the model for the immune cell aggregation density, and then compare it with the true value to calculate the regression error. For example, the regression error can be calculated using the mean square error (MSE), which can measure the average error between the predicted value and the true value.
[0146] Step S430: The weighted sum of the classification error and the regression error of the epithelial tumor cell subset is used as the first verification indicator, the weighted sum of the classification error and the regression error of the mesenchymal tumor cell subset is used as the second verification indicator, and the weighted sum of the classification error and the regression error of the tumor cell subset with an undetermined degree of differentiation is used as the third verification indicator.
[0147] The weighted sum is to multiply the classification error and regression error by the corresponding weights respectively, and then add them together to get a comprehensive index. Setting different weights can adjust the importance of classification error and regression error in the validation index according to actual needs. For example, for the epithelial tumor cell subset, assuming that the weight of the classification error is 0.6 and the weight of the regression error is 0.4, multiply the classification error by 0.6 and the regression error by 0.4, and then add them together to get the first validation index. Similarly, calculate the second validation index and the third validation index.
[0148] Step S440: Dynamically adjust the learning rate of the fully connected layer and the loss function weight distribution ratio of the deep convolutional network based on the difference between the first verification indicator and the second verification indicator, and adjust the channel attention coefficient update step size of the convolutional layer based on the absolute value of the third verification indicator.
[0149] Dynamically adjusting the learning rate of the fully connected layer and the weight distribution ratio of the loss function is to enable the model to learn better on different types of cell subsets. If the difference between the first validation index and the second validation index is large, it means that the performance of the model on the epithelial tumor cell subset and the mesenchymal tumor cell subset is quite different. It is necessary to adjust the learning rate of the fully connected layer and the weight distribution ratio of the loss function to balance the performance of the model on these two subsets. For example, if the first validation index is greater than the second validation index, it means that the performance of the model on the epithelial tumor cell subset is poor. The learning rate of the fully connected layer on the epithelial tumor cell subset can be appropriately increased, and the weight distribution ratio of the loss function can be adjusted at the same time, so that the model pays more attention to the classification and regression errors of the epithelial tumor cell subset.
[0150] The purpose of adjusting the update step of the channel attention coefficient of the convolution layer based on the absolute value of the third verification indicator is to enable the model to better handle the subset of tumor cells with uncertain differentiation degrees. If the absolute value of the third verification indicator is large, it means that the performance of the model on the subset of tumor cells with uncertain differentiation degrees is poor, and it is necessary to increase the update step of the channel attention coefficient of the convolution layer so that the model pays more attention to the feature channels related to tumor cells with uncertain differentiation degrees. For example, an update step adjustment function h is set, the input of which is the absolute value of the third verification indicator, and the output is the update step of the channel attention coefficient of the convolution layer. When the absolute value of the third verification indicator is large, h outputs a larger update step; when the absolute value of the third verification indicator is small, h outputs a smaller update step.
[0151] Step S450: When the difference direction between the first verification indicator and the second verification indicator remains consistent in three consecutive training iterations and the third verification indicator does not exceed the preset fluctuation range, stop adjusting the network parameters and lock the current fully connected layer weights and convolutional layer channel attention coefficients as the final verification parameters.
[0152] The difference between the first and second validation indicators in three consecutive training iterations remained consistent, indicating that the performance difference trend of the model on epithelial tumor cell subsets and mesenchymal tumor cell subsets was stable. The third validation indicator did not exceed the preset fluctuation range, indicating that the performance of the model on tumor cell subsets with an uncertain degree of differentiation was relatively stable. When these two conditions are met, it means that the performance of the model has stabilized. At this time, stop adjusting the network parameters and lock the current fully connected layer weights and convolutional layer channel attention coefficients as the final validation parameters.
[0153] The preset fluctuation range is a preset fluctuation threshold of the third verification indicator, for example, it can be set to ±10% of the average value of the third verification indicator. In three consecutive training iterations, the difference direction between the first verification indicator and the second verification indicator and the value of the third verification indicator are monitored. When the difference direction between the first verification indicator and the second verification indicator is consistent for three consecutive times and the value of the third verification indicator is within the preset fluctuation range, the network parameters are stopped from being adjusted.
[0154] Step S500: generating a final recognition model based on the verified and adjusted network parameters, wherein the final recognition model is configured to receive a clinical pathology image stream and output the location coordinates and classification confidence of circulating tumor cells.
[0155] The final recognition model is the model obtained after the aforementioned training and verification adjustment, which has good performance and stability. The model can continuously process multiple clinical pathology images by receiving the clinical pathology image stream, which can come from the pathology examination equipment of the hospital. The model can output the location coordinates and classification confidence of circulating tumor cells, that is, the model can determine the location of circulating tumor cells in the image and give the confidence that each cell belongs to a different type.
[0156] As an implementation manner, step S500, generating a final recognition model based on the verified adjusted network parameters, may specifically include the following steps S510-S560:
[0157] Step S510: Load the verified and adjusted network parameters into the convolutional layer and fully connected layer of the deep convolutional network, and connect the spatial positioning head and the classification head in parallel after the last convolutional feature map of the deep convolutional network. The spatial positioning head uses three sets of deconvolution layers to generate a pixel-level coordinate offset map and a bounding box size map, and the classification head uses a channel attention mechanism to generate a multi-category probability distribution map.
[0158] The network parameters adjusted after verification are loaded into the convolutional layer and fully connected layer of the deep convolutional network in order to apply the optimal parameters obtained from the previous training and verification adjustment to the model so that the model has better performance. The spatial localization head and the classification head are connected in parallel after the last convolutional feature map of the deep convolutional network in order to simultaneously locate and classify circulating tumor cells. The spatial localization head uses three sets of deconvolution layers to generate pixel-level coordinate offset maps and bounding box size maps. The deconvolution layer is a convolution layer used for upsampling, which can enlarge the size of the feature map. The three sets of deconvolution layers can gradually enlarge the size of the last layer of convolutional feature maps to the same size as the original image, and generate pixel-level coordinate offset maps and bounding box size maps at the same time. The pixel-level coordinate offset map represents the coordinate offset of each pixel relative to the center point of the circulating tumor cell, and the bounding box size map represents the width and height of the bounding box of each circulating tumor cell. The classification head uses the channel attention mechanism to generate multi-category probability distribution maps. The channel attention mechanism can automatically focus on the important parts of the feature channel and enhance the response strength of the feature channel related to classification. Through the channel attention mechanism, the classification head can generate the probability distribution of each pixel belonging to different cell types based on the input feature map, forming a multi-category probability distribution map. For example, for a classification task that includes three types of epithelial tumor cells, mesenchymal tumor cells, and normal cells, the classification head can generate the probability distribution of each pixel belonging to these three types.
[0159] The specific implementation process of step S510 is, for example, loading the verified and adjusted network parameters into the convolutional layer and the fully connected layer of the deep convolutional network to ensure that the parameters of the model are optimized. Then, add the spatial positioning head and the classification head after the last layer of convolutional feature map. For the spatial positioning head, three groups of deconvolution layers are constructed, and the parameters such as the convolution kernel size, step size and padding of each group of deconvolution layers can be adjusted according to the specific situation. For example, the convolution kernel size of the first group of deconvolution layers can be 4×4, the step size is 2, and the padding is 1; the parameters of the second and third groups of deconvolution layers can be adjusted as needed to gradually enlarge the size of the feature map to the same as the original image. During the training process, the parameters of the deconvolution layer are updated by the back propagation algorithm so that it can accurately generate the pixel-level coordinate offset map and the bounding box size map. For the classification head, a channel attention mechanism is introduced, for example, a Squeeze-and-Excitation (SE) module can be used. The SE module can generate channel attention coefficients through global average pooling, fully connected layers, and Sigmoid activation functions, and then multiply the channel attention coefficients with the input feature maps to enhance the response strength of the feature channels related to classification. During the training process, the parameters of the classification head are updated through the back-propagation algorithm so that it can accurately generate multi-category probability distribution maps.
[0160] Step S520: Input the fused feature vector into the spatial positioning head, upsample the feature map through the deconvolution layer and perform weighted fusion with the immune cell aggregation density in the microenvironment feature, and output the horizontal and vertical coordinate offsets and the width and height of the bounding box of each pixel position.
[0161] After the fused feature vector is input into the spatial localization head, the deconvolution layer upsamples the feature map to enlarge the size of the feature map to the same size as the original image, so that the coordinate offset and bounding box size can be generated for each pixel position. The weighted fusion with the immune cell aggregation density in the microenvironment feature is to utilize the microenvironment information of the immune cell aggregation density to improve the accuracy of circulating tumor cell localization.
[0162] The upsampling process of the deconvolution layer is achieved through convolution operations, which can convert low-resolution feature maps into high-resolution feature maps. In the deconvolution process, the convolution kernel slides on the input feature map, and the output feature map is generated by calculating the convolution result of the convolution kernel and the input feature map. For example, for an input feature map with a size of 32×32, after a set of deconvolution layers, the size of the output feature map can be changed to 64×64.
[0163] The process of weighted fusion of immune cell aggregation density is, for example, to assign a weight coefficient to the immune cell aggregation density, multiply it element-by-element with the feature map output by the deconvolution layer, and then add the multiplication result to the feature map output by the deconvolution layer to obtain the weighted fused feature map. Through the weighted fused feature map, the spatial positioning head outputs the horizontal and vertical coordinate offsets and the bounding box width and height of each pixel position. The specific implementation is that on the weighted fused feature map, each channel corresponds to an output information, for example, one channel corresponds to the horizontal coordinate offset, one channel corresponds to the vertical coordinate offset, one channel corresponds to the bounding box width, and one channel corresponds to the bounding box height. By decoding the values of these channels, the horizontal and vertical coordinate offsets and the bounding box width and height of each pixel position can be obtained.
[0164] Step S530: Input the fused feature vector into the classification head, enhance the feature channel response intensity related to the nuclear atypia index through the channel attention mechanism, and output the classification probability of different cell types corresponding to each pixel position.
[0165] After the fused feature vector is input into the classification head, the channel attention mechanism can automatically focus on the feature channels related to the nuclear atypia index, enhance the response strength of these channels, and thus improve the accuracy of circulating tumor cell classification. The nuclear atypia index is an indicator to measure the degree of abnormal nuclear morphology and is closely related to the type of circulating tumor cells.
[0166] The channel attention mechanism can adopt the aforementioned SE module. Specifically, the input fused feature vector is first globally averaged pooled to compress the spatial dimension of the feature map into a scalar to obtain the global feature information of each channel. Then, the global feature information is input into a fully connected layer, and the channel attention coefficient of each channel is obtained through the Sigmoid activation function. Finally, the channel attention coefficient is element-wise multiplied with the input fused feature vector to enhance the response strength of the feature channel related to the nuclear atypia index.
[0167] Through the enhanced feature vector, the classification head outputs the classification probability of each pixel position corresponding to different cell types. In specific implementation, a fully connected layer is added to the classification head, the enhanced feature vector is input into the fully connected layer, and the output is converted into a probability distribution through the Softmax activation function. For example, for a classification task involving three cell types, the output of the fully connected layer is a three-dimensional vector. After the Softmax activation function, the classification probability of each pixel position corresponding to the three cell types is obtained.
[0168] Step S540: Perform peak point detection on the coordinate offset map and the bounding box size map, extract the pixel point with the local response maximum value as the candidate cell center point, and generate an initial detection box set according to the coordinate offset of the candidate cell center point and the bounding box size.
[0169] Peak point detection is used to find local maxima in an image, which can help determine the center of circulating tumor cells. Peak point detection is performed on the coordinate offset map and the bounding box size map to find the pixels with the largest changes in coordinate offset and bounding box size. These pixels are more likely to be the center of circulating tumor cells.
[0170] Peak point detection can use the non-maximum suppression (NMS) algorithm. When the NMS algorithm is implemented, the coordinate offset map and the bounding box size map are searched for local maximum values to find the maximum value pixel in each local area. Then, based on the coordinate offset and bounding box size of these maximum value pixels, the bounding box corresponding to each pixel is calculated. Next, these bounding boxes are sorted and the bounding box with the highest score is selected as the candidate bounding box. Finally, the remaining bounding boxes are traversed and the bounding boxes whose overlap with the candidate bounding box exceeds the preset threshold are deleted to obtain the final candidate cell center point.
[0171] The method of generating the initial detection frame set according to the coordinate offset and bounding box size of the candidate cell center point is to calculate the real coordinates of each candidate cell center point in the original image according to its coordinate offset, and then generate the corresponding bounding box according to the bounding box size. The bounding boxes corresponding to all candidate cell center points are combined to form the initial detection frame set.
[0172] Step S550: Match the initial detection frame set with the probability values of the corresponding areas in the multi-category probability distribution map, and select the detection frames whose probability values exceed the dynamic threshold as valid detection frames. The dynamic threshold is dynamically adjusted according to the probability distribution variance output by the classification head.
[0173] The purpose of matching the initial detection frame set with the probability values of the corresponding areas in the multi-class probability distribution map is to determine the probability that the cells in each detection frame belong to different types. In the multi-class probability distribution map, each pixel position corresponds to the classification probability of different cell types. By overlapping the initial detection frame with the multi-class probability distribution map, extracting the classification probability of each pixel position in the detection frame, and then performing statistical analysis on these probabilities, the probability of cells in the detection frame belonging to different types is obtained.
[0174] The detection frames with probability values exceeding the dynamic threshold are selected as valid detection frames in order to remove those detection frames with low classification probabilities and improve the accuracy of detection. The dynamic threshold is dynamically adjusted according to the variance of the probability distribution output by the classification head in order to adapt the threshold to different images and classification situations. If the variance of the probability distribution is large, it means that the uncertainty of the classification result is large, and the dynamic threshold can be appropriately lowered; if the variance of the probability distribution is small, it means that the uncertainty of the classification result is small, and the dynamic threshold can be appropriately increased.
[0175] The method of dynamically adjusting the threshold is, for example, to set a basic threshold T 0 and an adjustment factor k, based on the variance of the probability distribution output by the classification head Calculating dynamic thresholds The maximum probability that cells in each detection frame in the initial detection frame set belong to different types is compared with the dynamic threshold, and the detection frame whose probability value exceeds the dynamic threshold is regarded as a valid detection frame.
[0176] Step S560: Confidence calibration is performed on the effective detection frame based on the channel attention coefficient in the verified and adjusted network parameters, and the immune cell aggregation density and intercellular connection density index are used as spatial weight factors to calculate the final classification confidence of the effective detection frame and output the positioning coordinates.
[0177] The confidence calibration of the effective detection frame is performed based on the channel attention coefficient in the network parameters adjusted after verification, in order to use the channel attention coefficient to adjust the classification confidence of the effective detection frame, so that the confidence can more accurately reflect the reliability of the detection result. The channel attention coefficient can represent the importance of each feature channel. By weighted fusion of the channel attention coefficient and the original classification confidence of the effective detection frame, the confidence of the detection frame related to the important feature channel can be enhanced. The immune cell aggregation density and the cell-to-cell connection density index are used as spatial weight factors to consider the influence of the cell's microenvironment information on the classification confidence. The immune cell aggregation density and the cell-to-cell connection density index can reflect the microenvironment state of the cell, and different microenvironment states may affect the classification results of the cell. By using these indicators as spatial weight factors and multiplying them with the classification confidence of the effective detection frame, the classification confidence can more accurately reflect the true situation of the cell.
[0178] As an implementation mode, step S560 performs confidence calibration on the effective detection frame based on the channel attention coefficient in the verified adjusted network parameters, uses the immune cell aggregation density and the cell-to-cell connection density index as the spatial weight factor, calculates the final classification confidence of the effective detection frame and outputs the positioning coordinates, which may specifically include the following steps S561-S566:
[0179] Step S561: extract the channel weighted value of the channel attention coefficient in the adjusted network parameters, and perform channel-by-channel weighted summation of the immune cell aggregation density distribution map and the cell-to-cell connection density distribution map in the area covered by the effective detection frame according to the channel weighted value to generate a spatial weight factor map.
[0180] Verify that the channel weighted values of the channel attention coefficients in the adjusted network parameters represent the importance of each feature channel. Apply these channel weighted values to the immune cell aggregation density distribution map and the intercellular connection density distribution map in the effective detection frame coverage area, and perform channel-by-channel weighted summation to obtain a spatial weight factor map that comprehensively considers channel importance and microenvironment information. When performing channel-by-channel weighted summation, for each channel of the immune cell aggregation density distribution map and the intercellular connection density distribution map, multiply the channel weighted value by the pixel value of the corresponding channel, and then add the results of all channels to obtain the spatial weight factor map. For example, assuming that the immune cell aggregation density distribution map has C i channels, the distribution of cell-cell junction tightness is C c channels, and the channel weighted values of the channel attention coefficients are , the pixel value of the jth channel of the immune cell aggregation density distribution map is , the pixel value of the kth channel of the cell-cell connection density distribution map is , then the pixel value of the spatial weight factor map .
[0181] Step S562: Multiply the spatial weight factor map by the original classification probability distribution map of the effective detection box pixel by pixel to generate a weighted probability distribution map.
[0182] The purpose of pixel-by-pixel multiplication of the spatial weight factor map and the original classification probability distribution map of the effective detection box is to incorporate the microenvironment information into the classification probability so that the classification probability more accurately reflects the true situation of the cell. The pixel-by-pixel multiplication method is to multiply the pixel values of the spatial weight factor map and the original classification probability distribution map of the effective detection box at each pixel position of the two maps to obtain a weighted probability distribution map. For example, assuming that the pixel value of the spatial weight factor map at the pixel position (x, y) is S(x, y), and the pixel value of the original classification probability distribution map of the effective detection box at this position is P(x, y), then the pixel value of the weighted probability distribution map at this position is P w(x, y) =S(x, y)×P(x, y).
[0183] Step S563: Perform a regional integration operation on the weighted probability distribution map, count the maximum and mean values of the weighted probabilities within the coverage area of each valid detection box, and calculate the geometric mean of the maximum and mean values as the preliminary confidence level.
[0184] The regional integration operation is to integrate the weighted probability distribution map within the area covered by the valid detection frame, that is, to sum up all the pixel values within the area. By counting the maximum and mean values of the weighted probabilities within each valid detection frame coverage area, the highest and average values of the classification probability within the area can be obtained. Calculating the geometric mean of the maximum and mean values as the preliminary confidence level can comprehensively consider the maximum and mean values of the classification probability, so that the preliminary confidence level can more accurately reflect the reliability of the detection results. The calculation formula for the geometric mean is, for example, , where M is the maximum value and A is the mean value.
[0185] Step S564: According to the density gradient direction of the corresponding detection frame area in the immune cell aggregation density distribution map, the spatial consistency weight of the preliminary confidence is adjusted to generate an intermediate correction confidence.
[0186] The density gradient direction of the corresponding detection frame area in the immune cell aggregation density distribution map can reflect the trend and distribution of immune cell aggregation. The spatial consistency weight of the preliminary confidence is adjusted according to the density gradient direction, so that the preliminary confidence is more consistent with the spatial distribution characteristics of the cells. The method for adjusting the spatial consistency weight can be to calculate the density gradient direction of the corresponding detection frame area in the immune cell aggregation density distribution map, and adjust the preliminary confidence according to the consistency of the density gradient direction. If the density gradient direction is relatively consistent in the detection frame area, it means that the aggregation of immune cells has good spatial consistency. At this time, the spatial consistency weight of the preliminary confidence can be appropriately increased; if the density gradient direction is relatively scattered in the detection frame area, it means that the aggregation space consistency of immune cells is poor. At this time, the spatial consistency weight of the preliminary confidence can be appropriately reduced. Multiply the preliminary confidence by the spatial consistency weight to obtain the intermediate correction confidence.
[0187] Step S565: multiply the density difference ratio between inside and outside the detection box in the cell-to-cell connection density distribution map by the intermediate correction confidence to obtain the final classification confidence.
[0188] The density difference ratio between the inside and outside of the detection frame in the cell-to-cell connection density distribution map can reflect the difference in the density of the connection between the cells in the detection frame and the surrounding cells. The ratio is multiplied by the intermediate correction confidence to further consider the impact of the density of the cell-to-cell connection on the classification confidence.
[0189] The process of calculating the density difference ratio inside and outside the detection frame in the cell-to-cell connection density distribution map is, for example, to calculate the average value of the cell-to-cell connection density inside the detection frame and the average value of the cell-to-cell connection density outside the detection frame respectively, and then divide the average value inside the detection frame by the average value outside the detection frame to obtain the density difference ratio.
[0190] Step S566: Sort the valid detection frames according to the final classification confidence, remove the detection frames whose confidence is lower than the preset stability threshold, and retain the positioning coordinates of the remaining detection frames and the corresponding final classification confidence.
[0191] The purpose of sorting valid detection frames according to the final classification confidence is to put the detection frames with higher confidence in front for easy subsequent analysis and processing. The preset stability threshold is a pre-set confidence threshold used to remove detection frames with lower confidence and improve the reliability of the detection results. Detection frames with final classification confidence lower than the preset stability threshold are removed, and the positioning coordinates of the remaining detection frames and the corresponding final classification confidence are retained.
[0192] See also Figure 2 , Figure 2A schematic diagram of the structure of a computer system provided in an embodiment of the present invention, the computer system can be a stand-alone system, such as a high-performance desktop, a workstation, or a cluster system, such as a computing cluster, a cloud computing platform, etc. In the computer system, at least a processor 101, a communication interface 102 and a memory 103 are included. Among them, the processor 101, the communication interface 102 and the memory 103 can be connected via a bus or other means. Among them, the processor 101 (or the central processing unit (CPU)) is the computing core and control core of the computer system, which can parse various instructions in the computer system and process various data of the computer system. The communication interface 102 can optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.), which can be used to send and receive data under the control of the processor 101; the communication interface 102 can also be used for the transmission and interaction of data within the computer system. The memory 103 (Memory) is a memory device in the computer system for storing programs and data. It can be understood that the memory 103 here can include both the built-in memory of the computer system and the extended memory supported by the computer system. The memory 103 provides a storage space, which stores the operating system of the computer system, and the present invention does not limit this. In one embodiment, the processor 101 executes the deep learning-based circulating tumor cell recognition model training method provided in the above embodiment of the present invention by running the computer program in the memory 103.
Claims
1. A circulating tumor cell recognition model training method based on deep learning, characterized in that: The method comprises: Acquire a multi-channel pathological sample image set of circulating tumor cells, wherein the multi-channel pathological sample image set includes cell microscopic images processed with different stains, and each sample image carries corresponding cell type annotation information; Performing cross-modal feature extraction on the multi-channel pathological sample image to generate a fusion feature vector including cell morphology features and microenvironment features, wherein the microenvironment features include matrix component distribution features and immune cell infiltration features; Construct a deep convolutional network with a multi-stage training strategy. In the first stage, a cross-channel feature attention mechanism is used to optimize the convolution kernel parameters. In the second stage, adversarial sample generation is used to enhance feature discriminability. In the third stage, knowledge distillation is used to compress the network size. A stratified cross-validation mechanism is used to verify and adjust the network parameters of the deep convolutional network during the training process, wherein the stratified cross-validation mechanism is divided into multiple validation subsets according to the biological characteristics corresponding to the cell type annotation information, and the network parameters are adjusted based on the classification error of each validation subset; A final recognition model is generated based on the verified and adjusted network parameters, and the final recognition model is configured to receive a clinical pathology image stream and output the location coordinates and classification confidence of circulating tumor cells.
2. The method according to claim 1, characterized in that The cross-modal feature extraction of the multi-channel pathological sample image to generate a fusion feature vector including cell morphology features and microenvironment features includes: Performing a superpixel segmentation operation on the multi-channel pathological sample image to generate structural data covering the cell nucleus region, the cytoplasm region and the extracellular matrix, wherein the structural data includes cell nucleus morphology parameters, cell membrane curvature parameters and extracellular matrix distribution parameters of each region; Based on the structural data, a dual-branch feature extraction is performed, wherein the first branch extracts the cell nuclear atypia index and the cell-to-cell connection tightness index as cell morphological features through geometric topological analysis, and the second branch extracts the matrix fiber orientation features and the immune cell aggregation density as microenvironment features through spatial frequency transformation; The cell morphology features and microenvironment features are input into a cross-modal alignment module, which generates a feature alignment matrix by calculating the position correlation between the two types of features in a spatial coordinate system, wherein each element in the matrix represents the degree of matching between the cell morphology features and the microenvironment features at the corresponding spatial position; A multi-level feature interaction mechanism is used to process the feature alignment matrix, wherein the primary interaction layer captures the morphology-microenvironment association pattern in the local area through a convolution operation, and the high-level interaction layer establishes cross-regional feature dependencies through a self-attention mechanism; Dynamic gating fusion is performed on the interactively processed features to generate a four-dimensional feature tensor containing the morphological features of the cell nucleus level, the morphological features of the cell-to-cell connection, the microenvironmental features of the matrix components, and the microenvironmental features of the immune infiltration. The four-dimensional feature tensor is input into the compression excitation network, and the morphological and microenvironmental feature channels that are sensitive to diagnosis are strengthened through the channel attention mechanism, redundant feature channels are suppressed, and finally the fusion feature vector with dimension compression is output.
3. The method according to claim 2, characterized in that The first branch extracts the nuclear atypia index and the cell-to-cell connection density index as cell morphological features through geometric topological analysis, including: Performing edge detection operation on the structural data to generate cell nucleus contour boundary data and intercellular connection area boundary data, wherein the boundary data includes pixel-level coordinate sequence and curvature change information; Performing a polygon fitting operation on the cell nucleus contour boundary data to generate a polygon vertex distribution describing the cell nucleus morphology, calculating the nucleus-cytoplasm ratio deviation, contour concavity index and main axis deflection angle based on the vertex distribution, and fusing the parameters to generate a cell nucleus atypia index; Performing topological structure analysis on the boundary data of the intercellular connection area, extracting the statistical distribution of the overlapping length and contact angle of the adjacent cell contact surfaces, and generating an intercellular connection tightness index based on the area ratio of the connection area; The second branch extracts matrix fiber orientation characteristics and immune cell aggregation density as microenvironment characteristics through spatial frequency transformation, including: Performing a spatial frequency transformation operation on the extracellular matrix region in the structural data, converting the pixel distribution into a frequency domain energy distribution, and extracting the main frequency direction as a matrix fiber orientation feature; Performing regional clustering analysis on the frequency domain energy distribution to identify the spatial density distribution of immune cell aggregation regions, and generating immune cell aggregation density by combining the area of the cluster region with the contact boundary length of the adjacent matrix region; The nuclear atypia index and the cell-to-cell connection tightness index are combined into a cell morphology feature, and the matrix fiber orientation feature and the immune cell aggregation density are combined into a microenvironment feature.
4. The method according to claim 1, characterized in that The multi-stage training strategy includes: Inputting the initial parameters of the deep convolutional network into the first-stage channel optimization module, generating a channel enhancement coefficient matrix by calculating the spatial correlation between each feature channel activation map and the cell nuclear atypia index, performing channel-by-channel weighting on the convolution kernel based on the channel enhancement coefficient matrix, and outputting the first-stage optimization parameters; Inputting the first-stage optimization parameters into the second-stage adversarial training module, generating a perturbation sample set based on the optimized convolution kernel, wherein the perturbation samples in the perturbation sample set are constructed by adjusting the spatial co-occurrence pattern of the intercellular connection density index and the immune cell aggregation density, and maintaining the continuity constraint of the matrix fiber orientation feature; Alternately inputting the perturbed sample set and the original sample into the deep convolutional network, updating the fully connected layer parameters by comparing the distribution offsets of the original features and the perturbed features in the microenvironment feature space, and outputting the second-stage optimization parameters; The second-stage optimization parameters are input into the third-stage compression module. By matching the feature activation path differences between the teacher network and the student network in the process of extracting the nuclear atypia index, a channel importance ranking algorithm is used to remove redundant channels, and the lightweight third-stage optimization parameters are output.
5. The method according to claim 1, characterized in that The layered cross-validation mechanism is used to verify and adjust the network parameters of the deep convolutional network during the training process, including: Dividing the multi-channel pathology sample image set into an epithelial tumor cell subset, a mesenchymal tumor cell subset, and a tumor cell subset with an undetermined degree of differentiation according to the biological characteristics corresponding to the cell type annotation information; In each round of training iteration, the validation operation is performed in the order of the epithelial tumor cell subset, the mesenchymal tumor cell subset, and the tumor cell subset with an undetermined degree of differentiation, wherein the validation operation of each subset includes extracting the classification error of the cell nuclear atypia index and the regression error of the immune cell aggregation density of the current network parameters on the corresponding subset; The weighted sum of the classification error and the regression error of the epithelial tumor cell subset is used as the first verification indicator, the weighted sum of the classification error and the regression error of the mesenchymal tumor cell subset is used as the second verification indicator, and the weighted sum of the classification error and the regression error of the tumor cell subset with an undetermined degree of differentiation is used as the third verification indicator; Dynamically adjust the learning rate of the fully connected layer and the weight distribution ratio of the loss function of the deep convolutional network based on the difference between the first verification indicator and the second verification indicator, and adjust the channel attention coefficient update step size of the convolutional layer based on the absolute value of the third verification indicator; When the difference direction between the first verification indicator and the second verification indicator remains consistent in three consecutive training iterations and the third verification indicator does not exceed the preset fluctuation range, stop adjusting the network parameters and lock the current fully connected layer weights and convolutional layer channel attention coefficients as the final verification parameters.
6. The method according to claim 1, characterized in that The generating of the final recognition model based on the verified and adjusted network parameters comprises: The verified and adjusted network parameters are loaded into the convolutional layer and the fully connected layer of the deep convolutional network, and the spatial positioning head and the classification head are connected in parallel after the last convolutional feature map of the deep convolutional network, wherein the spatial positioning head uses three sets of deconvolution layers to generate a pixel-level coordinate offset map and a bounding box size map, and the classification head uses a channel attention mechanism to generate a multi-category probability distribution map; The fused feature vector is input into the spatial positioning head, the feature map is upsampled through the deconvolution layer and weightedly fused with the immune cell aggregation density in the microenvironment feature, and the horizontal and vertical coordinate offsets and the width and height of the bounding box of each pixel position are output; Inputting the fused feature vector into the classification head, enhancing the feature channel response intensity related to the nuclear atypia index through the channel attention mechanism, and outputting the classification probability of different cell types corresponding to each pixel position; Perform peak point detection on the coordinate offset map and the bounding box size map, extract the pixel point with the maximum local response as the candidate cell center point, and generate an initial detection box set according to the coordinate offset and bounding box size of the candidate cell center point; Matching the initial detection frame set with the probability values of the corresponding areas in the multi-category probability distribution map, and screening the detection frames whose probability values exceed a dynamic threshold as valid detection frames, wherein the dynamic threshold is dynamically adjusted according to the probability distribution variance output by the classification head; Based on the channel attention coefficient in the verified and adjusted network parameters, the confidence calibration of the effective detection frame is performed, the immune cell aggregation density and the cell-to-cell connection density index are used as spatial weight factors, the final classification confidence of the effective detection frame is calculated and the positioning coordinates are output.
7. The method according to claim 6, characterized in that The confidence calibration of the effective detection frame is performed based on the channel attention coefficient in the network parameters adjusted by the verification, the immune cell aggregation density and the cell-to-cell connection density index are used as spatial weight factors, the final classification confidence of the effective detection frame is calculated, and the positioning coordinates are output, including: Extracting the channel weighted value of the channel attention coefficient in the verified and adjusted network parameters, and performing channel-by-channel weighted summation of the immune cell aggregation density distribution map and the intercellular connection density distribution map in the area covered by the effective detection frame according to the channel weighted value to generate a spatial weight factor map; Multiplying the spatial weight factor map by the original classification probability distribution map of the effective detection box pixel by pixel to generate a weighted probability distribution map; Performing a regional integration operation on the weighted probability distribution map, counting the maximum and mean values of the weighted probabilities in the coverage area of each valid detection box, and calculating the geometric mean of the maximum and mean values as a preliminary confidence level; According to the density gradient direction of the corresponding detection frame area in the immune cell aggregation density distribution map, the spatial consistency weight of the preliminary confidence is adjusted to generate an intermediate correction confidence; The final classification confidence is obtained by multiplying the density difference ratio between inside and outside the detection frame in the intercellular connection density distribution map by the intermediate correction confidence; The valid detection frames are sorted according to the final classification confidence, the detection frames with confidence lower than a preset stability threshold are eliminated, and the positioning coordinates of the remaining detection frames and the corresponding final classification confidences are retained.
8. The method according to claim 1, characterized in that The generation process of the adversarial sample includes: Extracting the characteristic gradient distribution of the deep convolutional network on the multi-channel pathological sample image set, and identifying the cross-modal gradient association pattern of the cell nuclear atypia index and the immune cell aggregation density based on the characteristic gradient distribution; generating a disturbance direction template according to the cross-modal gradient association pattern, wherein the disturbance direction template includes a pixel offset direction that produces a coordinated gradient change in the cell nucleus morphological parameters and the matrix fiber orientation characteristics; Inputting the multi-channel pathological sample image into the deep convolutional network, extracting the spatial gradient response map of the fused feature vector, and generating an initial perturbation mask based on the distribution similarity between the spatial gradient response map and the perturbation direction template; Performing a morphological boundary constraint operation on the initial disturbance mask, eliminating isolated disturbance points that have no intersection with the cytoplasm region boundary, and retaining the disturbance region that overlaps with the cell nucleus contour boundary data space; The retained disturbance region is superimposed on the original multi-channel pathological sample image to generate a disturbance image and input it into the deep convolutional network, and the consistency index of the change direction of the nuclear atypia index and the change direction of the immune cell aggregation density before and after the disturbance is calculated; The intensity distribution of the perturbation direction template is dynamically adjusted based on the consistency index to generate an adversarial sample set, wherein the perturbation intensity of each sample in the adversarial sample set is inversely correlated with the local change rate of the cell membrane curvature parameter.
9. The method according to claim 1, characterized in that: The process of knowledge distillation and compression includes: Using the convolution layer parameters of the deep convolutional network as the teacher network parameters, initializing the convolution layer parameters of the student network so that its channel dimension matches the teacher network parameters but with a reduced number of channels; Extracting a set of intermediate feature activation maps of the teacher network on the multi-channel pathological sample image, wherein the set of intermediate feature activation maps includes the spatial response distribution of the cell morphological features and microenvironmental features in different convolutional layers; Performing cross-layer spatial alignment on the intermediate feature activation map of the student network and the intermediate feature activation map of the teacher network, compensating for the resolution difference between the two through dilated convolution and generating a feature similarity distribution map; Dynamically adjusting the channel attention coefficient of the student network based on the feature similarity distribution map, suppressing channels with significant differences in response mode from the teacher network and enhancing the weights of spatially consistent channels; Inputting the compressed feature vector of the student network into the classifier, calculating the difference loss between the student network and the teacher network classification probability distribution, and generating an adaptive distillation loss coefficient in combination with the cell type annotation information; The fully connected layer parameters of the student network are iteratively updated according to the adaptive distillation loss coefficient until the compression process is terminated when the difference between the comprehensive performance index of the student network under the stratified cross-validation mechanism and the performance of the teacher network is stabilized within a preset range.
10. A computer system, characterized in that: include: a memory, wherein a computer program is stored in the memory; A processor, used to load the computer program to implement the deep learning-based circulating tumor cell recognition model training method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Deep learning-based child posterior fossa tumor diagnosis method
CN117198511A
Multi-modal melanoma immunotherapy prediction method based on deep learning network
CN117936105A
Multi-center, multi-tumor-species and multi-prescription dosage prediction method for pelvic tumor and related equipment
CN119400348A
Disease monitoring method and system based on medical image data recognition
CN119811650A
Whole-slide pathology image classification system and construction method considering tumor microenvironment
JP7312510B1
Cited By
Deep learning-based breast lump ultrasonic image classification method and system
CN120451685A
IMC image segmentation method based on multi-channel cell collaborative segmentation architecture
CN120747521A
Imc image segmentation method based on multi-channel cell synergistic segmentation architecture
CN120747521B
Detection flushing method and system based on deep infectious wound sinus tract
CN120789380A
Training method and device of image recognition model
CN120913008A