Pathological section diagnosis screening analysis method based on deep neural network
The three-dimensional spatiotemporal Transformer model constructed through deep neural networks, combined with multi-spectral imaging and dynamic weight adjustment, solves the efficiency and accuracy of traditional pathological section diagnosis, and realizes efficient and accurate pathological section diagnosis, improving the comprehensiveness and reliability of pathological section diagnosis.
Patent Information
- Application Number
- CN202510356398.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional pathological section diagnosis is low, with poor accuracy and consistency, making it difficult to meet the precise analysis needs of complex pathological characteristics. The existing computer-assisted diagnostic technology has limitations in lesion feature extraction and data labeling, and is incompatible with clinical applications.
The pathological slice diagnostic screening method based on deep neural network is adopted, and digital images are obtained through multispectral imaging technology, a three-dimensional spatiotemporal Transformer network model is constructed, combined with super-resolution reconstruction and image registration technology, and the model is trained and optimized using transfer learning and dynamic weight adjustment mechanisms to achieve efficient and accurate diagnosis of pathological slices.
It improves the efficiency and accuracy of pathological diagnosis, reduces the rate of misdiagnosis and missed diagnosis, provides more comprehensive lesion information, enhances the reliability and adaptability of diagnosis, supports the detection of early micro lesions, and improves the ability to analyze complex pathological characteristics.
Smart Images

Figure CN120299672A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pathological section diagnosis and screening, and specifically to a pathological section diagnosis and screening analysis method based on a deep neural network. Background Art
[0002] In the field of modern medicine, pathological section diagnosis, as the "gold standard" for disease diagnosis, plays a crucial role in the accurate judgment of diseases, the formulation of treatment plans, and the prognosis evaluation of patients. However, traditional pathological section diagnosis mainly relies on pathologists to conduct manual observation and analysis through optical microscopes, and there are many limitations in this process. On the one hand, the efficiency of manual diagnosis is relatively low. Facing the increasing number of cases, the workload of pathologists is heavy, which easily leads to an extended diagnosis cycle and affects the timely treatment of patients. On the other hand, the accuracy of manual diagnosis largely depends on the experience and professional level of pathologists. There are differences in the diagnosis results among different pathologists, and it is difficult to ensure the consistency and reliability of diagnosis. In addition, with the continuous increase in the types of diseases and the increasing complexity of pathological features, higher requirements are put forward for the accuracy and comprehensiveness of pathological section diagnosis, and traditional manual diagnosis methods are difficult to meet these needs.
[0003] Traditional pathological section diagnosis technology mainly involves pathologists conducting manual observation and analysis of pathological sections under an optical microscope. Through years of professional study and practical accumulation, pathologists have mastered the morphological characteristics of various lesions and can judge the nature of lesions, such as benign or malignant, the type and severity of lesions, etc., based on characteristics such as the morphology, structure, arrangement pattern of cells, and changes in cell nuclei. Although this traditional technical solution has relatively high flexibility, pathologists can make comprehensive judgments on complex pathological conditions according to their experience and professional knowledge and can discover some atypical lesion characteristics. However, the efficiency of manual diagnosis is low. Pathologists need to carefully observe each pathological section one by one, which takes a lot of time. Especially when facing a large number of cases, the diagnosis cycle will be significantly extended, resulting in patients not being able to obtain the diagnosis results in a timely manner and affecting the timeliness of subsequent treatment; the diagnosis results are greatly affected by subjective factors of pathologists. There are differences in the experience, professional level, and diagnosis habits of different pathologists, which lead to inconsistent diagnosis results for the same pathological section, reducing the reliability and accuracy of diagnosis; it is difficult for manual diagnosis to conduct a comprehensive and in-depth analysis of the subtle features and complex data in pathological sections, and it is easy to miss some important information, thus affecting the accuracy of diagnosis.
[0004] With the development of computer technology and image processing technology, some computer-aided pathological section diagnosis technologies have gradually emerged. These existing technologies mainly use image processing algorithms to analyze pathological section images and classify and diagnose the extracted features through machine learning algorithms. Compared with traditional manual diagnosis, the existing technologies have certain advantages. First, the computer-aided diagnosis technology can quickly process a large number of pathological section images, greatly improving the diagnosis efficiency. Through an automated image analysis process, the preliminary screening of multiple sections can be completed in a short time, reducing the workload of pathologists. Second, the use of machine learning algorithms can quantitatively analyze the features of pathological section images, reducing the influence of subjective factors and improving the accuracy and consistency of diagnosis. In addition, some advanced image processing algorithms can enhance the details of the images, helping pathologists better observe and analyze the lesion features. However, the current machine learning algorithms still have certain limitations when dealing with complex pathological section images. The lesion features in pathological section images are complex and diverse, and the pathological features of different diseases overlap. The existing algorithms are difficult to accurately extract and distinguish these features, resulting in the need to further improve the diagnostic accuracy. On the other hand, the existing computer-aided diagnosis technologies often require a large amount of labeled data for training, and obtaining high-quality labeled data requires a lot of manpower, material resources and time. The quality of data labeling will also affect the performance of the algorithm. In addition, the combination of existing technologies and clinical practical applications is not close enough. In the actual use process, there are problems such as complex operations and incompatibility with the existing hospital diagnosis processes, which limit their wide application in clinical practice.
[0005] Therefore, in view of the above problems, the present application proposes a pathological section diagnosis screening analysis method based on a deep neural network, which solves the above-mentioned technical problems. Summary of the Invention
[0006] Based on the above content, the pathological section diagnosis screening analysis method based on a deep neural network of the present application includes the following steps:
[0007] S1. Obtain the digital image data of the pathological section, preprocess the image data to form an image data set;
[0008] S2. Construct a three-dimensional spatio-temporal Transformer network model, obtain the feature image of the pathological section in the spatio-temporal dimension, and form an image feature set;
[0009] S3. Transfer the model parameters pre-trained on a large-scale image data set to the constructed three-dimensional spatio-temporal Transformer network model through transfer learning, and fine-tune the model according to the pathological section image structure to optimize the model parameters;
[0010] S4. Use the labeled pathological slice image dataset to train the fine-tuned three-dimensional spatio-temporal Transformer network model. Use the cross-entropy loss function as the optimization objective and adopt the stochastic gradient descent algorithm to iteratively optimize the model;
[0011] S5. Input the pathological slice image to be diagnosed and screened into the trained three-dimensional spatio-temporal Transformer network model, output the diagnosis result, visually display the diagnosis result, and overlay and display the diagnosis result on the pathological slice image through color marking.
[0012] Preferably, in S1, the digital image data of the pathological slice is obtained by multi-spectral imaging technology. The pathological slice is photographed under different spectral bands to obtain images of multiple spectral channels. The super-resolution reconstruction algorithm is used to process the images of each spectral channel in combination with the information of adjacent slices to enhance the image resolution; through the image registration technology, the images of different spectral channels are aligned and fused to generate digital image data.
[0013] Preferably, the use of the super-resolution reconstruction algorithm to process the images of each spectral channel in combination with the information of adjacent slices, and the use of the image registration technology to align and fuse the images of different spectral channels to generate digital image data specifically includes:
[0014] Obtain the image I i (x, y) of the spectral channel, where (x, y) are the pixel coordinates of the image. Extract the features of the adjacent slice image I a (x, y) to obtain the high-frequency feature F a (x, y), and adopt the feature fusion formula where w1 and w2 are the corresponding weight coefficients, is the gradient feature of the i-th spectral channel image. Reconstruct I i (x, y) through the fused feature to obtain the image I i-q (x, y) with enhanced resolution; perform image registration through the mutual information function, and the formula is: where I m and I n are images of different spectral channels, p mn (x, y) is the joint probability distribution of I m and I n at (x, y), and p m (x, y), p n (x, y) are the probability distributions of I m and I n at (x, y) respectively. Iteratively find the optimal transformation parameter T, and transform the resolution-enhanced image I i-q (x, y) of each spectral channel, and the formula is: Ii-e (x, y) = T(I i-q (x, y)) realizes image alignment and adopts a weighted fusion formula α i is the fusion weight value, N is the number of spectral channels, and digital image data is generated.
[0015] Preferably, the digital image data of the pathological section in S1 is preprocessed to form an image dataset. Through the image normalization network of deep learning, the brightness and contrast feature distributions of the pathological section images are adaptively learned, the images are normalized, the pathological regions in the images are initially segmented, individual pathologies or pathological masses are separated, and the processed images are classified and sorted according to the lesion type and severity to form an image dataset.
[0016] Preferably, the construction of the three-dimensional spatio-temporal Transformer network model in S2 specifically includes:
[0017] Output three-dimensional pathological section images according to the image dataset where T represents the time dimension, H and W are the height and width of the image, C is the number of channels, and the model is stacked by multiple spatio-temporal Transformer blocks; within the spatio-temporal Transformer block, spatio-temporal position encoding is performed on the input feature map X to obtain the position encoding vector P, and the feature X after adding the position encoding is obtained pe , and the formula is: X pe = X + P. Project X pe into multiple low-dimensional spaces respectively to obtain Q j , K j , V j , where j = 1 ··· M, M is the number of heads, and the attention output of each head is calculated through the formula , where d k is the projection dimension. After splicing the multi-head outputs and then passing through a linear transformation, the self-attention result Z a is obtained. Process Z a through a feed-forward neural network. The feed-forward neural network contains two fully connected layers, and the ReLU activation function is used in the middle. The formula is: FFN(Z a ) = max(0, W1Z a + b1)W2 + b2, where W1, W2 are the corresponding weight coefficients, and b1, b2 are the corresponding bias vectors. After being processed layer by layer by multiple spatio-temporal Transformer blocks, the output feature image of the last block is obtained, and the feature images are sorted in order to form an image feature set.
[0018] Preferably, in the three-dimensional spatio-temporal Transformer network model, a spatio-temporal convolutional layer is interspersed between adjacent Transformer blocks. The spatio-temporal convolutional layer performs convolution operations using multi-scale convolutional kernels. Convolutional kernels of different scales can capture different features in the temporal and spatial dimensions. During convolution, the input feature map is processed using convolutional kernels of each scale respectively to obtain convolutional features of different scales. Global average pooling is performed on each convolutional feature, and the convolutional features are weighted and summed to form a fused feature map. After normalization and activation function processing, the output feature map of the spatio-temporal convolutional layer is obtained as the input of the next Transformer block.
[0019] Preferably, when transferring the model parameters pre-trained on a large-scale image dataset to the three-dimensional spatio-temporal Transformer network model through transfer learning in S3, by hierarchically matching the structures of the pre-trained model and the target model, the parameters corresponding to the spatio-temporal feature learning part in the pre-trained model are directly mapped to the target model; the three-dimensional spatio-temporal Transformer network model is fine-tuned through a progressive freezing strategy. First, the bottom convolutional layer and part of the attention layer of the model are frozen, and only a small number of iterative trainings are performed on the fully connected layer near the output layer and the upper spatio-temporal convolutional layer, so that the model initially adapts to the feature distribution of the pathological slice data. According to the loss change and feature activation situation during training, gradually unfreeze part of the bottom layer, and use the adaptive learning rate adjustment algorithm to dynamically adjust the learning rate according to the sensitivity of different layers to the pathological slice features, and optimize the parameters of the three-dimensional spatio-temporal Transformer network model.
[0020] Preferably, training the fine-tuned three-dimensional spatio-temporal Transformer network model using the labeled pathological slice image dataset in S4 specifically includes: stratifying and dividing the labeled data according to the dimensions of lesion type and image complexity to construct a stratified training dataset. During training iteration, randomly extract different proportions of samples from each layer of the dataset to form a mixed training batch for training; calculate the loss through the cross-entropy loss function, obtain the importance difference of different lesion regions in the pathological slice, introduce a dynamic weight adjustment mechanism to adjust the weights of different samples in the cross-entropy loss function, and perform iterative optimization using the stochastic gradient descent algorithm, and adjust through the momentum factor; set a high momentum factor at the beginning of training to quickly move towards the optimal solution direction. As training progresses, dynamically adjust the momentum factor according to the performance of the model on the validation set. If the accuracy of the model on the validation set does not improve for multiple consecutive rounds, reduce the momentum factor and switch to the mini-batch gradient descent mode to finely optimize the model, forming a trained three-dimensional spatio-temporal Transformer network model.
[0021] Preferably, a dynamic weight adjustment mechanism is introduced into the cross-entropy loss function, and the weights of different samples in the cross-entropy loss function are adjusted according to the area ratio of the lesion area in the image and the severity of the lesion. Specifically:
[0022] Obtain the lesion area and lesion severity of the pathological section image sample, and calculate the weight adjustment coefficient w of the sample n , and the formula is: where N is the total number of samples in a training batch, s n is the lesion area of the nth image sample, and r n is the severity of the nth image sample. The weighted cross-entropy loss function is obtained through the cross-entropy loss function, and the formula is: where y n is the true label of sample n, and p n is the probability that the model predicts sample n as the positive class. In this way, when training the model, samples corresponding to small but severely diseased areas can be given higher weights, and samples corresponding to large but mildly diseased areas can be given relatively lower weights, forming differences in the importance of different samples for model training.
[0023] Preferably, the diagnostic results in S5 are visually displayed. A corresponding color mapping table is established according to different diagnostic result types. Benign lesions are marked in green, malignant lesions are marked in red, and suspected lesions are marked in yellow. The diagnostic results output by the trained three-dimensional spatio-temporal Transformer network model are parsed to determine the location and type of each lesion area. For the lesion areas in the pathological section image, according to the color mapping table, the corresponding color is selected for each lesion area, and the selected color is superimposed on the lesion area through a semi-transparent color filling method to display the location of the lesion area without obscuring the details of the original pathological section image.
[0024] Compared with the prior art, the technical solution of the present application has the following technical effects:
[0025] By using the multi-spectral imaging technology to obtain the digital image data of the pathological section, and combining the super-resolution reconstruction algorithm and the image registration technology to enhance the image resolution and fuse to generate the digital image data, the present invention solves the technical problems of insufficient resolution in the process of obtaining traditional pathological section images and the difficulty in effectively fusing images of different spectral channels. The images of different spectral channels can be accurately aligned and fused, and the super-resolution reconstruction algorithm can also improve the resolution by combining the information of adjacent sections, enabling more accurate observation of details such as the morphology and distribution of diseased cells, improving the detection ability of early micro-lesions, providing strong support for the early diagnosis of diseases, avoiding missed diagnoses and misdiagnoses caused by image quality problems, and improving the accuracy and reliability of pathological diagnosis.
[0026] The present invention solves the technical problem that traditional models are difficult to effectively extract the complex features of pathological sections in the spatio-temporal dimension by constructing a three-dimensional spatio-temporal Transformer network model and interspersing spatio-temporal convolutional layers therein, and using multi-scale convolutional kernels to capture different features. The spatio-temporal Transformer block performs spatio-temporal position encoding and multi-head attention mechanism processing on the input feature map, and can mine features from multiple spatio-temporal dimensions. The interspersed spatio-temporal convolutional layers capture features at different levels such as cell morphological changes and lesion development trends through multi-scale convolutional kernels. Taking tumor pathological section analysis as an example, it is possible to clearly observe the dynamic processes such as the proliferation and invasion of tumor cells at different time points, as well as the microscopic feature differences of cells in different regions, providing richer and more comprehensive information for judging the development stage and malignancy of tumors, helping doctors formulate more targeted treatment plans, and improving the comprehensiveness and accuracy of pathological diagnosis.
[0027] The present invention solves the technical problems of insufficient data, low training efficiency and easy overfitting of the model when training a pathological section diagnosis model by using a transfer learning method to transfer the model parameters pre-trained on a large-scale image dataset to the three-dimensional spatio-temporal Transformer network model, and adopting a progressive freezing strategy and an adaptive learning rate adjustment algorithm for fine-tuning and optimization. When training a pathological section diagnosis model, it is costly and time-consuming to obtain a large amount of high-quality labeled data. Training the model only relying on limited pathological section data is likely to lead to insufficient model learning, poor generalization ability and overfitting. By using transfer learning, the general feature parameters learned on a large-scale image dataset are transferred, providing rich prior knowledge for the model and enabling the model to adapt to pathological section data faster. The progressive freezing strategy and the adaptive learning rate adjustment algorithm avoid blindly training all parameters, reduce the training time, improve the training efficiency, and effectively prevent the model from overfitting. In practical applications, the model can quickly learn the key features of pathological sections with less training data, has good generalization ability on new pathological section samples, improves the stability and reliability of the model, and reduces the cost and difficulty of model training.
[0028] The present invention solves the technical problem that the traditional training method cannot fully consider the importance differences of different lesion regions in pathological sections, resulting in poor model training effects, by using a labeled pathological section image dataset to construct a hierarchical training dataset through hierarchical division according to lesion types and image complexities, introducing a dynamic weight adjustment mechanism, and combining the stochastic gradient descent algorithm and momentum factor adjustment for model training. The hierarchical training dataset enables the model to conduct targeted learning for lesions of different difficulties and types. The dynamic weight adjustment mechanism adjusts the sample weights according to the proportion of the lesion area and the severity of the lesion, assigning higher weights to regions with small areas but severe lesions, enabling the model to pay more attention to these key regions. During the training process, the training speed is adjusted through the momentum factor, approaching the optimal solution quickly in the initial stage and being finely optimized according to the performance of the validation set in the later stage. In the training of breast cancer pathological sections, the model can more accurately identify tiny malignant tumor cell regions, improving the ability to identify malignant lesions, enhancing the adaptability and robustness of the model to different lesion situations, making the trained model perform better in actual diagnosis, and improving the accuracy and reliability of pathological diagnosis.
[0029] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, so as to be implemented in accordance with the content of the specification, and in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following takes the preferred embodiments of the present application and describes them in detail in conjunction with the drawings as follows.
[0030] Those skilled in the art will understand the above and other purposes, advantages and features of the present application more clearly according to the following detailed description of the specific embodiments of the present application in conjunction with the drawings. Brief Description of the Drawings
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to actual scale.
[0032] Figure 1 It is a flowchart of a pathological section diagnosis and screening analysis method based on a deep neural network according to the present invention;
[0033] Figure 2 It is a graph of the change of image features before and after processing by the super-resolution reconstruction algorithm according to the present invention;
[0034] Figure 3This is a diagram of image processing for deep learning in the present invention;
[0035] Figure 4 This is a structural diagram of a three-dimensional spatio-temporal Transformer network model in the present invention;
[0036] Figure 5 This is a color mapping marker diagram of a pathological section in the present invention;
[0037] Figure 6 This is a comparison diagram of the diagnostic accuracy of pathological sections between the present application and the prior art in the present invention. Detailed implementation manners
[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. In the following description, specific details such as specific configurations and components are provided only to assist in a comprehensive understanding of the embodiments of the present application. Therefore, those skilled in the art should clearly understand that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Additionally, descriptions of known functions and structures are omitted for clarity and conciseness in the embodiments.
[0039] It should be understood that the "one embodiment" or "the present embodiment" mentioned throughout the specification means that the specific features, structures, or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the "one embodiment" or "the present embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner.
[0040] In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplicity and clarity, and does not itself indicate the relationship between the various embodiments and / or arrangements discussed.
[0041] The term "and / or" in this document is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, B exists alone, and A and B exist simultaneously. The term " / and" in this document describes another association object relationship, indicating that two relationships may exist. For example, A / and B may represent: A exists alone, and A and B exist alone. Additionally, the character " / " in this document generally indicates that the associated objects before and after are in an "or" relationship.
[0042] As used herein, the term "at least one" is merely a description of the relationship between associated objects, indicating that there can be three relationships. For example, at least one of A and B can represent: A exists alone, A and B exist simultaneously, and B exists alone.
[0043] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion.
[0044] Embodiment 1
[0045] This embodiment mainly describes a pathological section diagnosis and screening analysis method based on a deep neural network, as Figure 1 shown, including the following steps:
[0046] S1. Obtain the digital image data of the pathological section, preprocess the image data to form an image data set;
[0047] S2. Construct a three-dimensional spatio-temporal Transformer network model, obtain the feature image of the pathological section in the spatio-temporal dimension, and form an image feature set;
[0048] S3. Transfer the model parameters pre-trained on a large-scale image data set to the constructed three-dimensional spatio-temporal Transformer network model through transfer learning, and fine-tune the model according to the pathological section image structure to optimize the model parameters;
[0049] S4. Use the labeled pathological section image data set to train the fine-tuned three-dimensional spatio-temporal Transformer network model, use the cross-entropy loss function as the optimization target, and adopt the stochastic gradient descent algorithm to iteratively optimize the model;
[0050] S5. Input the pathological section image to be diagnosed and screened into the trained three-dimensional spatio-temporal Transformer network model, output the diagnosis result, visually display the diagnosis result, and superimpose and display the diagnosis result on the pathological section image through color marking.
[0051] Furthermore, the digital image data of the pathological section in S1 is obtained through multi-spectral imaging technology. The pathological section is photographed in different spectral bands to obtain images of multiple spectral channels. The super-resolution reconstruction algorithm is used to process the images of each spectral channel in combination with the information of adjacent sections to enhance the image resolution; through image registration technology, the images of different spectral channels are aligned and fused to generate digital image data.
[0052] Further, the super-resolution reconstruction algorithm is used to process the images of each spectral channel by combining the information of adjacent slices. Through image registration technology, the images of different spectral channels are aligned and fused to generate digital image data, specifically including:
[0053] Obtain the image I of the spectral channel i (x, y), where (x, y) are the pixel coordinates of the image, and perform feature extraction on the adjacent slice image I a (x, y) to obtain the high-frequency feature F a (x, y), and adopt the feature fusion formula where w1 and w2 are the corresponding weight coefficients, is the gradient feature of the image of the i-th spectral channel. Reconstruct I i (x, y) through the fused feature to obtain the image I i-q (x, y) with enhanced resolution; perform image registration through the mutual information function, and the formula is: where I m and I n are images of different spectral channels, and p mn (x, y) is the joint probability distribution of I m and I n at (x, y). p m (x, y) and p n (x, y) are the probability distributions of I m and I n at (x, y) respectively. Iteratively find the optimal transformation parameter T to transform the resolution-enhanced image I i-q (x, y) of each spectral channel, and the formula is: I i-e (x, y) = T(I i-q (x, y)) to achieve image alignment, and adopt the weighted fusion formula α i is the fusion weight value, and N is the number of spectral channels to generate digital image data.
[0054] Further, the digital image data of the pathological section in S1 is preprocessed to form an image dataset. Through the image normalization network of deep learning, the brightness and contrast feature distributions of the pathological section images are adaptively learned, the images are normalized, the pathological regions in the images are preliminarily segmented, and individual pathologies or pathological masses are separated. The processed images are classified and sorted according to the lesion type and severity to form an image dataset.
[0055] Further, the construction of the three-dimensional spatio-temporal Transformer network model in S2 specifically includes:
[0056] Output three-dimensional pathological section images according to the image dataset where T represents the time dimension, H and W are the height and width of the image, C is the number of channels, and the model is stacked by multiple spatio-temporal Transformer blocks; within the spatio-temporal Transformer block, spatio-temporal position encoding is performed on the input feature map X to obtain the position encoding vector P, and the feature X after adding the position encoding is obtained pe , the formula is: X pe = X + P, project X pe onto multiple low-dimensional spaces respectively to obtain Q j , K j , V j , where j = 1···M, M is the number of heads, and the attention output of each head is calculated through the formula , where d k is the projection dimension. After splicing the multi-head outputs and then passing through a linear transformation, the self-attention result Z a is obtained. Process Z a through a feed-forward neural network. The feed-forward neural network contains two fully connected layers, and the ReLU activation function is used in the middle. The formula is: FFN(Z a ) = max(0, W1Z a + b1)W2 + b2, where W1 and W2 are the corresponding weight coefficients, and b1 and b2 are the corresponding bias vectors. After being processed layer by layer by multiple spatio-temporal Transformer blocks, the output feature image of the last block is obtained. The feature images are sorted in order to form an image feature set
[0057] Furthermore, in the three-dimensional spatio-temporal Transformer network model, spatio-temporal convolutional layers are inserted between adjacent Transformer blocks. The spatio-temporal convolutional layers perform convolution operations using multi-scale convolutional kernels. Different-scale convolutional kernels can capture different features in the time and space dimensions. During the convolution process, the input feature map is processed using convolutional kernels of each scale respectively to obtain convolutional features of different scales. Global average pooling is performed on each convolutional feature, and the convolutional features are weighted and summed to form a fused feature map. After being processed by normalization and activation functions, the output feature map of the spatio-temporal convolutional layer is obtained as the input of the next Transformer block
[0058] Further, when transferring the model parameters pre-trained on a large-scale image dataset to the three-dimensional spatio-temporal Transformer network model through transfer learning in S3, by hierarchically matching the structures of the pre-trained model and the target model, the parameters corresponding to the spatio-temporal feature learning part in the pre-trained model are directly mapped to the target model; the three-dimensional spatio-temporal Transformer network model is fine-tuned through a progressive freezing strategy. First, the underlying convolutional layer and part of the attention layer of the model are frozen, and only a small number of iterative trainings are performed on the fully connected layer near the output layer and the upper spatio-temporal convolutional layer, so that the model can initially adapt to the feature distribution of the pathological slice data. According to the loss change and feature activation situation during the training process, gradually unfreeze some of the underlying layers, and use the adaptive learning rate adjustment algorithm to dynamically adjust the learning rate according to the sensitivity of different layers to the pathological slice features, and optimize the parameters of the three-dimensional spatio-temporal Transformer network model.
[0059] Further, in S4, the fine-tuned three-dimensional spatio-temporal Transformer network model is trained using the annotated pathological slice image dataset, which specifically includes: stratifying and dividing the annotated data according to the dimensions of lesion type and image complexity to construct a stratified training dataset. During the training iteration, randomly extract different proportions of samples from each layer of the dataset to form a mixed training batch for training; calculate the loss through the cross-entropy loss function to obtain the importance difference of different lesion regions in the pathological slice, introduce a dynamic weight adjustment mechanism to adjust the weights of different samples in the cross-entropy loss function, and use the stochastic gradient descent algorithm for iterative optimization, and adjust through the momentum factor; set a high momentum factor at the beginning of the training to quickly move towards the optimal solution direction. As the training progresses, dynamically adjust the momentum factor according to the performance of the model on the validation set. If the accuracy of the model on the validation set does not improve for multiple consecutive rounds, reduce the momentum factor and switch to the mini-batch gradient descent mode to finely optimize the model, forming a trained three-dimensional spatio-temporal Transformer network model.
[0060] Further, by introducing a dynamic weight adjustment mechanism into the cross-entropy loss function, adjust the weights of different samples in the cross-entropy loss function according to the area ratio and lesion severity of the lesion region in the image. Specifically:
[0061] Obtain the lesion area and lesion severity of the pathological slice image sample, and calculate the weight adjustment coefficient w of the sample n , the formula is: where N is the total number of samples in a training batch, s n is the lesion area of the nth image sample, r n is the lesion severity of the nth image sample, and the weighted cross-entropy loss function is obtained through the cross-entropy loss function. The formula is: where y n is the true label of sample n, pn It is the probability that the model predicts the sample n as a positive class. In this way, when training the model, samples corresponding to small but severely diseased areas can be given higher weights, and samples corresponding to large but slightly diseased areas can be given relatively lower weights, forming differences in the importance of different samples for model training.
[0062] Furthermore, in S5, the diagnostic results are visually displayed. A corresponding color mapping table is established according to different diagnostic result types. Benign lesions are marked as green, malignant lesions are marked as red, and suspected lesions are marked as yellow. The diagnostic results output by the trained three-dimensional spatio-temporal Transformer network model are analyzed to determine the location and type of each lesion area. For the lesion areas in the pathological section images, according to the color mapping table, the corresponding color is selected for each lesion area, and the selected color is superimposed on the lesion area through a semi-transparent color filling method to display the location of the lesion area without obscuring the details of the original pathological section image.
[0063] This embodiment details the acquisition of digital image data of pathological sections in this application and preprocessing to form a dataset, then constructs a three-dimensional spatio-temporal Transformer network model to obtain an image feature set, subsequently fine-tunes the model using transfer learning, then trains the model using the labeled dataset, and finally inputs the slice to be diagnosed into the trained model and visually displays the results. This solves the problems of low efficiency and large subjective influence on accuracy in traditional pathological section diagnosis, improves the diagnosis efficiency, reduces the burden of manual diagnosis; improves the diagnosis accuracy, reduces the misdiagnosis and missed diagnosis rates through accurate feature extraction and dynamic weight adjustment; the visual display is clear and intuitive, assisting doctors to quickly judge the condition and providing a reliable basis for subsequent treatment.
[0064] Based on Embodiment 1, this embodiment details the experimental study of the pathological section diagnosis and screening analysis method of this application, specifically:
[0065] In this experiment, the performance of the pathological section diagnosis and screening technology based on deep neural network (hereinafter referred to as "this technology") and the prior art in pathological section diagnosis is compared to verify the advantages of this technology:
[0066] 2000 pathological section samples from multiple hospitals are selected, covering various common cancer types such as lung cancer, breast cancer, and gastric cancer. After preliminary diagnosis by professional pathologists, the lesion type and severity information are detailedly recorded; the samples are randomly divided into a training set (1400 cases), a validation set (300 cases), and a test set (300 cases).
[0067] The hardware environment uses a high-performance server equipped with NVIDIA Tesla V100 GPUs to accelerate the training process of deep learning models. The software environment is based on the Python programming language, using the PyTorch deep learning framework for model construction and training, and at the same time using the OpenCV library for image processing-related operations.
[0068] Using multi-spectral imaging technology, within the spectral range of 380nm - 780nm, 21 spectral channels are set at intervals of 20nm to photograph pathological sections. After obtaining the image data, a super-resolution reconstruction algorithm is used to process the images of each spectral channel in combination with adjacent section information; as Figure 2 shown, taking a lung cancer pathological section as an example, in the 400nm channel, the average pixel value of the pre-processed image is 118.4327, and the standard deviation is 14.2763; after processing, the average pixel value is increased to 135.6754, and the standard deviation becomes 12.1345. The image resolution is increased from 512×512 to 1024×1024, the detail clarity is significantly enhanced, and the cell texture and structure are more clearly distinguishable.
[0069] Through image registration technology, the images of different spectral channels are aligned and fused according to the mutual information function to generate digital image data. The registration errors of 100 lung cancer pathological sections are statistically analyzed, and the average registration error is 0.135 pixels, ensuring the accurate fusion of multi-spectral information. The image normalization network of deep learning is used to normalize the images, adaptively learning the brightness and contrast feature distributions of pathological section images; as Figure 3 shown, when processing breast cancer pathological sections, the average brightness value of the pre-processed image is 82.345, and the average contrast value is 0.286; after processing, the average brightness value is adjusted to 98.765, and the average contrast value is increased to 0.423, and the image quality is significantly improved, and the distinction between the lesion area and normal tissue is higher.
[0070] The pathological areas in the images are initially segmented, and a deep learning-based semantic segmentation model is used to separate individual pathologies or pathological masses. The segmentation tests are carried out on 300 pathological sections of different cancer types, and the average segmentation accuracy reaches 93.247%, which can effectively separate key pathological structures such as tumor cell masses and inflammatory areas. The processed images are classified and sorted according to the lesion type and severity to form an image dataset.
[0071] As Figure 4As shown in the figure, the size of the three-dimensional pathological section image input through the three-dimensional spatio-temporal Transformer network model is T×H×W×C, where T represents the time dimension (set to 8 in this experiment to simulate the pathological change characteristics at different time points), H = 1024, W = 1024, and C is the number of spectral channels (21 in this experiment); the model is stacked by 10 spatio-temporal Transformer blocks. Inside the spatio-temporal Transformer block, spatio-temporal position encoding is performed on the input feature map. Taking the spatio-temporal position (3, 512, 512, 10) as an example, the element values of the position encoding vector P are calculated as: [0.1234, 0.2345, -0.0987, 0.4567, -0.3456, 0.0876, 0.1567, 0.2890, -0.1456, 0.3789, 0.0654, -0.2345, 0.1789, 0.3210, -0.1123, 0.2678, -0.0789, 0.1901, 0.2456, -0.1678, 0.0567].
[0072] The features after adding position encoding are projected into the low-dimensional spaces of 12 heads (M = 12) respectively, and the projection dimension d k = 96. Through transfer learning, the model parameters pre-trained on the large-scale natural image dataset are transferred to the constructed three-dimensional spatio-temporal Transformer network model, and the structures of the pre-trained model and the target model are hierarchically matched. The parameters corresponding to the spatio-temporal feature learning part in the pre-trained model are directly mapped to the target model; the model is fine-tuned using a progressive freezing strategy. First, the bottom 4 convolutional layers and 3 attention layers are frozen, and only the fully connected layer near the output layer and the upper spatio-temporal convolutional layer are iteratively trained. At the beginning of training, the learning rate is set to 0.0008. As training progresses, an adaptive learning rate adjustment algorithm is used to dynamically adjust the learning rate according to the sensitivity of different layers to the pathological section features.
[0073] The labeled pathological section image dataset is used for training. The labeled data is stratified and divided according to the dimensions of lesion type and image complexity to construct a stratified training dataset. During training iteration, different proportions of samples are randomly selected from each layer of the dataset to form a mixed training batch for training. A dynamic weight adjustment mechanism is introduced to adjust the weights of different samples in the cross-entropy loss function according to the area ratio of the lesion area in the image and the severity of the lesion. During the training of 50 gastric cancer pathological sections, the area ratio of the lesion area of a certain sample is 15.673%, and the lesion severity score is 3.2. The calculated weight adjustment coefficient w n = 0.9124. During the training process, the stochastic gradient descent algorithm is used for iterative optimization. The momentum factor is set to 0.9 at the beginning of training. As training progresses, if the accuracy of the model on the validation set does not improve for 5 consecutive rounds, the momentum factor is reduced.
[0074] Input 300 pathological section images of the test set into the trained 3D spatio-temporal Transformer network model for diagnostic screening. The model outputs diagnostic results, including lesion type and severity information, and marks them through the constructed color mapping table, such as Figure 5 shown, and visually display the diagnostic results; conduct a detailed statistics on the diagnostic results of the test set, and the results are shown in the following table:
[0075]
[0076] For the experimental results of the prior art, a traditional method based on support vector machine (SVM) combined with simple image gray feature extraction was used as the prior art for comparative experiments. Diagnose the same 300 test samples, and the results are shown in the following table:
[0077]
[0078] Comparative analysis of experimental results. It can be clearly seen from the data in the above two tables that the proposed technology has significantly higher diagnostic accuracy in the diagnosis of various cancer lesions than the prior art. Taking the diagnosis of malignant lesions as an example, for the diagnosis accuracy of malignant lesions of lung cancer, breast cancer, and gastric cancer, the proposed technology is 20%, 20%, and 22% higher than the prior art respectively; the overall diagnosis accuracy of malignant lesions is 20.66% higher. In the diagnosis of benign lesions, the proposed technology also has obvious advantages, and the average diagnosis accuracy is 36.15% higher than the prior art. For the diagnosis of suspected lesions, the accuracy of the proposed technology is 5 times that of the prior art. The overall diagnosis accuracy of the proposed technology reaches 83.00%, while that of the prior art is only 60.00%.
[0079] As Figure 6 shown, it can be seen that regardless of lung cancer, breast cancer or gastric cancer, the diagnostic accuracy of the proposed technology is higher than that of the prior art, and the gap is relatively obvious. This indicates that the proposed technology has stronger adaptability and accuracy in dealing with the pathological section diagnosis of different types of cancers.
[0080] Further subdivide the lesion types and compare the performance of the two technologies in the diagnosis of different subtypes such as adenocarcinoma and squamous cell carcinoma. Taking lung cancer as an example, the diagnostic accuracy of the proposed technology for adenocarcinoma is 88% (44 cases were correctly diagnosed out of 50 adenocarcinoma samples), and the diagnostic accuracy for squamous cell carcinoma is 82% (25 cases were correctly diagnosed out of 30 squamous cell carcinoma samples); while the diagnostic accuracy of the prior art for adenocarcinoma is only 60% (30 cases were correctly diagnosed out of 50 adenocarcinoma samples), and the diagnostic accuracy for squamous cell carcinoma is 56.67% (17 cases were correctly diagnosed out of 30 squamous cell carcinoma samples). The detailed data are shown in the following table:
[0081]
[0082] As can be seen from the data in the table, in the diagnosis of cancer lesion subtypes, this technology also has significant advantages, being able to more accurately identify lesions of different subtypes and providing a more accurate basis for clinical treatment.
[0083] Through this experimental comparison, the pathological section diagnosis and screening technology based on deep neural networks has demonstrated excellent performance in pathological section diagnosis. Compared with the existing technology based on support vector machines combined with simple image feature extraction, in the image data acquisition and preprocessing stage of this technology, through technologies such as multispectral imaging, super-resolution reconstruction, and image normalization, the image quality and the recognition rate of lesion areas have been significantly improved; in the process of model construction and training, the application of three-dimensional spatio-temporal Transformer network models, transfer learning, progressive freezing strategies, and dynamic weight adjustment mechanisms enables the model to more effectively learn the features of pathological sections, improving the accuracy and adaptability of diagnosis.
[0084] This embodiment details that in the diagnosis of different cancer types, lesion subtypes, and different lesion natures (benign, malignant, suspected), the diagnostic accuracy of this technology is much higher than that of the existing technology. This result indicates that this technology has great application potential in the field of medical pathological diagnosis, and is expected to provide more efficient and accurate technical support for clinical diagnosis, promoting the development and progress of pathological diagnosis technology.
[0085] The above are only the preferred embodiments of the present invention, and it does not limit the protection scope of the present invention accordingly. For those skilled in the art, the present invention can have various changes and modifications; all changes, modifications, substitutions, integrations, and parameter changes made to these embodiments by means of conventional substitutions or capable of achieving the same functions without departing from the principle and spirit of the present invention fall within the protection scope of the present invention.
Claims
1. A pathological section diagnosis screening and analysis method based on a deep neural network, characterized in that, Including: S1. Obtain the digital image data of the pathological section, preprocess the image data to form an image data set; S2. Construct a three-dimensional spatio-temporal Transformer network model, obtain the feature images of the pathological section in the spatio-temporal dimension to form an image feature set; S3. Transfer the model parameters pre-trained on a large-scale image data set to the constructed three-dimensional spatio-temporal Transformer network model through transfer learning, and fine-tune the model according to the pathological section image structure to optimize the model parameters; S4. Use the labeled pathological section image data set to train the fine-tuned three-dimensional spatio-temporal Transformer network model, use the cross-entropy loss function as the optimization target, and use the stochastic gradient descent algorithm to iteratively optimize the model; S5. Input the pathological section image to be diagnosed and screened into the trained three-dimensional spatio-temporal Transformer network model, output the diagnosis result, visually display the diagnosis result, and overlay and display the diagnosis result on the pathological section image through color marking.
2. The pathological section diagnosis screening and analysis method based on a deep neural network according to claim 1, characterized in that In S1, the digital image data of the pathological section is obtained through multi-spectral imaging technology. The pathological section is photographed in different spectral bands to obtain images of multiple spectral channels. The super-resolution reconstruction algorithm is used to process the images of each spectral channel in combination with the information of adjacent sections to enhance the image resolution. Through image registration technology, the images of different spectral channels are aligned and fused to generate digital image data.
3. The pathological section diagnosis screening and analysis method based on a deep neural network according to claim 2, wherein, The use of the super-resolution reconstruction algorithm to process the images of each spectral channel in combination with the information of adjacent sections, and the alignment and fusion of the images of different spectral channels through image registration technology to generate digital image data specifically includes: Obtain the image I of the spectral channel i (x, y), where (x, y) are the pixel coordinates of the image, perform feature extraction on adjacent slice images I a (x, y) to obtain the high-frequency feature F a (x, y), and adopt the feature fusion formula where w1 and w2 are the corresponding weight coefficients, is the gradient feature of the i-th spectral channel image. Reconstruct I i (x, y) to obtain the image I i-q (x, y) with enhanced resolution; perform image registration through the mutual information function. The formula is: where I m and I n are images of different spectral channels, and p mn (x, y) is the joint probability distribution of I m and I n at (x, y). p m (x, y) and p n (x, y) are the probability distributions of I m and I n at (x, y) respectively. Iteratively find the optimal transformation parameter T, and transform the resolution-enhanced image I i-q (x, y) of each spectral channel. The formula is: I i-e (x, y) = T(I i-q (x, y)) to achieve image alignment. Adopt the weighted fusion formula α i is the fusion weight value, and N is the number of spectral channels to generate digital image data.
4. The pathological section diagnosis screening and analysis method based on a deep neural network according to claim 1, characterized in that The digital image data of the pathological section in S1 is preprocessed to form an image data set. Through the image normalization network of deep learning, the brightness and contrast feature distributions of the pathological section images are adaptively learned, the images are normalized, the pathological regions in the images are initially segmented, individual pathologies or pathological masses are separated, and the processed images are classified and sorted according to the lesion type and severity to form an image data set.
5. The pathological section diagnosis screening and analysis method based on a deep neural network according to claim 1, wherein The construction of the three-dimensional spatio-temporal Transformer network model in S2 specifically includes: Output three-dimensional pathological section images according to the image dataset where T represents the time dimension, H and W are the height and width of the image, C is the number of channels, and the model is stacked by multiple spatio-temporal Transformer blocks; within the spatio-temporal Transformer block, spatio-temporal position encoding is performed on the input feature map X to obtain the position encoding vector P, and the feature X after adding the position encoding is obtained pe , the formula is: X pe = X + P, project X pe into multiple low-dimensional spaces respectively to obtain Q j , K j , V j , where j = 1···M, M is the number of heads, and calculate the attention output of each head through the formula , where d k is the projection dimension, splice the multi-head outputs and then perform a linear transformation to obtain the self-attention result Z a , process Z a through a feed-forward neural network. The feed-forward neural network contains two fully connected layers, and the ReLU activation function is used in the middle. The formula is: FFN(Z a ) = max(0, W1Z a + b1)W2 + b2, where W1 and W2 are the corresponding weight coefficients, and b1 and b2 are the corresponding bias vectors. After being processed layer by layer by multiple spatio-temporal Transformer blocks, the output feature image of the last block is sorted in order to form an image feature set.
6. The pathological section diagnosis and screening analysis method based on a deep neural network according to claim 5, wherein In the three-dimensional spatio-temporal Transformer network model, spatio-temporal convolutional layers are interspersed between adjacent Transformer blocks. The spatio-temporal convolutional layers perform convolutional operations using multi-scale convolutional kernels. Different-scale convolutional kernels can capture different features in the time and space dimensions. During the convolution process, the input feature maps are processed using convolutional kernels of each scale respectively to obtain convolutional features of different scales. Global average pooling is performed on each convolutional feature, and the convolutional features are weighted and summed to form a fused feature map. After normalization and activation function processing, the output feature map of the spatio-temporal convolutional layer is obtained as the input of the next Transformer block.
7. The pathological section diagnosis screening and analysis method based on a deep neural network according to claim 1, wherein When migrating the model parameters pre-trained on a large-scale image dataset to the 3D spatio-temporal Transformer network model through transfer learning in S3, by hierarchically matching the structures of the pre-trained model and the target model, the parameters corresponding to the spatio-temporal feature learning part in the pre-trained model are directly mapped to the target model; the 3D spatio-temporal Transformer network model is fine-tuned through a progressive freezing strategy. First, the bottom convolutional layer and some attention layers of the model are frozen, and only a small number of iterations are performed on the fully connected layer near the output layer and the upper spatio-temporal convolutional layer to make the model initially adapt to the feature distribution of the pathological slice data. According to the loss change and feature activation situation during the training process, gradually unfreeze some of the bottom layers, and use the adaptive learning rate adjustment algorithm to dynamically adjust the learning rate according to the sensitivity of different layers to the pathological slice features, and optimize the parameters of the 3D spatio-temporal Transformer network model.
8. The pathological section diagnosis screening and analysis method based on a deep neural network according to claim 1, characterized in that, In S4, the fine-tuned 3D spatio-temporal Transformer network model is trained using the labeled pathological slice image dataset, specifically including: dividing the labeled data into layers according to the dimensions of lesion type and image complexity to construct a hierarchical training dataset. During the training iteration, randomly extract different proportions of samples from each layer of the dataset to form a mixed training batch for training; calculate the loss through the cross-entropy loss function to obtain the importance difference of different lesion regions in the pathological slice, introduce a dynamic weight adjustment mechanism to adjust the weights of different samples in the cross-entropy loss function, and use the stochastic gradient descent algorithm for iterative optimization, which is adjusted by the momentum factor; set a high momentum factor at the beginning of training to quickly move towards the optimal solution direction. As the training progresses, dynamically adjust the momentum factor according to the performance of the model on the validation set. If the accuracy of the model on the validation set does not improve for several consecutive rounds, reduce the momentum factor and switch to the mini-batch gradient descent mode to finely optimize the model to form a trained 3D spatio-temporal Transformer network model.
9. The pathological section diagnosis screening and analysis method based on a deep neural network according to claim 8, wherein By introducing a dynamic weight adjustment mechanism in the cross-entropy loss function, the weights of different samples in the cross-entropy loss function are adjusted according to the area ratio and severity of the lesion regions in the image, specifically as follows: Obtain the lesion area and lesion severity of the pathological section image sample, and calculate the weight adjustment coefficient w of the sample n , the formula is: where N is the total number of samples in a training batch, s n is the lesion area of the nth image sample, r n is the lesion severity of the nth image sample. Obtain the weighted cross-entropy loss function through the cross-entropy loss function. The formula is: where y n is the true label of sample n, p n is the probability that the model predicts sample n as the positive class. In this way, when training the model, samples corresponding to small areas but severe lesions can be given higher weights, and samples with large areas but mild lesions can be given relatively lower weights, forming differences in the importance of different samples for model training.
10. The method for pathological section diagnosis screening and analysis based on a deep neural network according to claim 1, wherein In S5, the diagnostic results are visually displayed. A corresponding color mapping table is established according to different diagnostic result types, with benign lesions marked as green, malignant lesions marked as red, and suspected lesions marked as yellow. The diagnostic results output by the trained 3D spatio-temporal Transformer network model are parsed to determine the location and type of each lesion region. For the lesion regions in the pathological slice image, according to the color mapping table, select the corresponding color for each lesion region, and overlay the selected color on the lesion region through a semi-transparent color filling method to display the location of the lesion region without obscuring the details of the original pathological slice image.