Intelligent analysis method and system for automatic classification of pathological images
By combining hyperspectral imaging technology and deep learning technology, intelligent analysis of pathological images is solved, and the problem that traditional pathological image analysis relies on manual observation is achieved, efficient and accurate automatic classification and analysis is achieved, and the quality and efficiency of pathological diagnosis is improved.
Patent Information
- Application Number
- CN202510058478.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional pathological image analysis relies on manual observation and is susceptible to subjective factors, and accuracy and consistency are difficult to guarantee.
An intelligent analysis method is adopted, combining hyperspectral imaging technology and deep learning technology to pre-process hyperspectral images and pathological slice images, deep analysis, feature fusion, pseudo-label generation and loss function optimization, dynamic optimization and model construction, and multi-level interpretive analysis.
It realizes efficient, accurate and automatic classification and analysis of pathological images, improves the quality and efficiency of pathological diagnosis, and enhances the robustness and credibility of the model.
Smart Images

Figure CN119992181A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic classification of medical record images, and in particular to an intelligent analysis method and system for automatic classification of pathological images. Background Art
[0002] In the field of pathology, accurate classification and analysis of pathological images are of vital importance for disease diagnosis, treatment plan formulation, and prognosis assessment. Traditional pathological image analysis mainly relies on manual observation and empirical judgment by pathologists. However, this method has many limitations. On the one hand, manual analysis is time-consuming and laborious. Faced with massive amounts of pathological image data, it is inefficient and difficult to meet the needs of rapid clinical diagnosis. On the other hand, subjective judgment is easily affected by factors such as observer fatigue and experience differences, making it difficult to ensure the accuracy and consistency of diagnostic results.
[0003] With the rapid development of computer technology and imaging technology, hyperspectral imaging technology has gradually emerged and has shown great potential in the field of pathological image analysis. Hyperspectral images can simultaneously obtain information about pathological tissues in multiple spectral bands, and contain richer details and features than traditional pathological images. However, hyperspectral data has the characteristics of high dimensionality and large data volume, which brings challenges to image processing and analysis. Deep learning technology has achieved remarkable results in image recognition, classification and other fields, and has provided new ideas and methods for the automatic classification of pathological images. Combining deep learning with hyperspectral imaging technology is expected to overcome the shortcomings of traditional pathological image analysis, realize efficient, accurate and automatic classification and analysis of pathological images, improve the quality and efficiency of pathological diagnosis, and promote the development of the field of pathology towards intelligence and precision. Summary of the invention
[0004] The purpose of the present invention is to provide an intelligent analysis method and system for automatic classification of pathological images, which solves the problem that traditional pathological image analysis relies on manual observation, is easily affected by subjective factors, and is difficult to ensure accuracy and consistency.
[0005] To achieve the above object, the present invention provides an intelligent analysis method for automatic classification of pathological images, comprising the following steps:
[0006] S1, preprocessing the hyperspectral image and pathological section image;
[0007] S2, performing in-depth analysis on the pre-processed hyperspectral images and pathological slice images;
[0008] S3, global feature fusion in the channel dimension through a three-stage fusion strategy;
[0009] S4, pseudo label generation and loss function optimization;
[0010] S5, dynamic optimization and model building;
[0011] S6. Use attention mechanism to perform multi-level interpretive analysis on hyperspectral data.
[0012] Preferably, in S1, the preprocessing operation on the hyperspectral image in the data processing module is specifically:
[0013] De-noising is used to eliminate random noise introduced by sensor noise or environmental interference during the acquisition process, and the low-rank matrix decomposition method is used to optimize the hyperspectral image. Assuming that the original hyperspectral image data is recorded as X, the goal is to minimize the error function:
[0014]
[0015] Among them, ||·|| F represents the Frobenius norm, the matrix L is the low-rank component, and S represents sparse noise;
[0016] Add regularization constraints:
[0017] ||L|| * +λ||S||1;
[0018] Among them, ||L|| * represents the nuclear norm of L, ||S||1 represents the l1 norm of sparse noise, and λ is the regularization parameter;
[0019] The spectral matching transformation technology is used to geometrically align the spectral data of different bands, specifically:
[0020] The improved chroma-brightness enhancement method is used to linearly adjust the brightness and contrast of the hyperspectral image. Its mathematical expression is:
[0021] I'=α·(I-μ)+β;
[0022] Among them, I' represents the new image after contrast and brightness adjustment, I represents the original image pixel value matrix, μ is the global mean of pixel values, α is the contrast adjustment coefficient, and β is the brightness compensation term.
[0023] Preferably, in S1, the preprocessing operation of the pathological slice image in the data processing module is:
[0024] The Laplace operator is used to enhance the morphological features, and the edge features in the image are reflected by calculating the gradient change of the pixel gray value;
[0025] A saliency detection algorithm is used to annotate the salient areas in the case slice images, and the areas containing key pathological information are located by analyzing the texture, color and brightness features of the images.
[0026] Preferably, in S2, in the feature extraction module, the preprocessed hyperspectral image and pathological slice image are deeply analyzed, which specifically includes the following steps:
[0027] S21. In the spectral feature extraction stage, a one-dimensional convolutional network is used to mine the spectral dimension characteristics of the hyperspectral image data. The one-dimensional convolution extracts the spectral features through the following formula:
[0028] F spectrum =Conv1D(X;W s ,b s );
[0029] Among them, F spectrum represents the spectral features after one-dimensional convolution operation, W s represents the weight matrix of the one-dimensional convolution kernel, b s Represents the bias term of the convolutional layer;
[0030] S22. In the spatial feature extraction stage, a two-dimensional convolutional network is used to capture the spatial texture information of the image. The two-dimensional convolution extracts the spatial features through the following formula:
[0031] F spatial =Conv2D(X;W sp ,b sp );
[0032] Among them, F spatial represents the extracted spatial texture features, W sp Represents the weight matrix of the two-dimensional convolution kernel, b sp The offset term corresponding to the two-dimensional convolution layer is used to adjust the offset of the convolution result.
[0033] S23, introduce the multi-head attention mechanism to fuse local and global information by calculating the global correlation of input features. The calculation formula is:
[0034]
[0035] In the formula, Q, K, and V represent the query, key, and value of the feature matrix respectively, and d k is the number of columns for which the scaling factor is equal to K.
[0036] Preferably, in S3, a three-stage fusion strategy is used in the feature fusion module to perform global fusion in the channel dimension, and the specific process is as follows:
[0037] In the shallow fusion stage, the linear weighted fusion of local features and global features is performed using the following formula:
[0038]
[0039] Among them, F CNN represents the local features extracted by the convolutional neural network, F Transformer represents the global features extracted by the Transformer module, ω1 and ω2 are dynamic weights, which are used to regulate the contribution of local and global features in the fusion results respectively;
[0040] In the deep fusion stage, the features generated by the shallow fusion Input to the MBConv module for processing to obtain the deep fusion features:
[0041]
[0042] In the final global feature fusion stage, the shallow fusion features are fused by the following formula Deep fusion features Concatenate in the channel dimension:
[0043]
[0044] Preferably, in S4, pseudo label generation and loss function optimization are performed in the automatic classification module, and the specific process is as follows:
[0045] In the pseudo-label generation stage, the model dynamically generates pseudo-labels based on the prediction results of unlabeled data and the confidence threshold. For each sample, the predicted probability distribution is expressed as P c Indicates that, P c is the predicted probability that the sample belongs to category c;
[0046] When the highest predicted probability of the sample max(P c ) exceeds the confidence threshold τ, a label is generated And its category is determined as the classification result with the highest probability through the following formula:
[0047]
[0048] Consider the supervised learning loss of labeled data and the semi-supervised learning loss of unlabeled data. In the labeled data part, the cross entropy loss function is used:
[0049] L sup =CE(y,P model );
[0050] Among them, y is the true label, P model is the predicted probability distribution of the model output;
[0051] For unlabeled data, the consistency regularization loss L is introduced unsup , when the highest prediction probability max(P) of unlabeled data exceeds the confidence threshold τ, consistency regularization is enabled;
[0052] The original input is transformed through strong data enhancement technology, and the predicted probability distribution P of the enhanced sample is calculated strong , the consistency regularization loss is defined by the cross entropy formula:
[0053]
[0054] Among them, 1 {max(P)>β} represents the indicator function, which is activated only when the prediction confidence satisfies the threshold β; P strong is the predicted distribution after strong enhancement;
[0055] The total loss function combines the supervised loss and the unsupervised loss and is defined as:
[0056] L=L sup +μL unsup ;
[0057] Among them, μ is the weight coefficient of the unsupervised loss.
[0058] Preferably, in S5, dynamic optimization and model building are performed in the dynamic adaptive optimization module and the model building module, and the specific process is:
[0059] First, we introduce consistency regularization as the optimization objective and assume that the prediction distributions generated by the model for the two enhancement methods are P aug1 and P aug2 , then the consistency regularization term is expressed as:
[0060] R consistency =||P aug1 -P aug2 ||2;
[0061] Among them, ||·||2 represents the Euclidean norm, which is used to quantify the difference between the two prediction distributions;
[0062] By dynamically adjusting the loss weight ω t To balance the contribution between different objectives, the loss weight is calculated as:
[0063]
[0064] Among them, t represents the current number of training rounds, and γ is an adjustment rate parameter used to control the speed of weight change.
[0065] Preferably, in S6, in the result diagnosis module, the attention mechanism is used to perform multi-level explanatory analysis on the hyperspectral data, and the specific process is:
[0066] Generate a saliency score S based on the attention weight attention , and its calculation formula is:
[0067] S attention =Softmax(W a ·F final );
[0068] Among them, S attention represents the saliency score of each pixel or region, W a is the attention weight matrix, F final Represents the final feature matrix extracted by the neural network, which contains the comprehensive information of the image in spatial and spectral dimensions.
[0069] The present invention also provides an intelligent analysis system for automatic classification of pathological images, comprising a data processing module, a feature extraction module, a feature fusion module, an automatic classification module, a model building module, a result diagnosis module and a dynamic adaptive optimization module;
[0070] The data processing module combines the spatial-spectral feature dimensionality reduction technology of hyperspectral images and the morphological feature enhancement technology of pathological images, and adopts a multi-scale feature compression strategy based on a deep convolutional encoder to map high-dimensional data into low-dimensional and representative feature representations. At the same time, a significant region detection algorithm is introduced to automatically mark high-correlation regions.
[0071] The feature extraction module uses a method that combines convolutional neural networks with self-supervised learning to extract features from the region of interest (ROI) in hyperspectral pathology images, including using one-dimensional convolution to extract spectral features and two-dimensional convolution to capture spatial texture features, and combines a multi-head attention mechanism to improve local and global feature representation capabilities;
[0072] The feature fusion module adopts a multi-level feature fusion architecture. Through a three-stage fusion strategy, it fuses the features of images from different sources at the shallow layer, deep layer, and interactive layer, and dynamically weighs the weights of different features to maximize the complementarity of multimodal information.
[0073] The automatic classification module uses a deep learning model based on contrastive learning and semi-supervision, generates pseudo labels using a small amount of labeled data, and iteratively optimizes by comparing the feature similarities of positive and negative samples;
[0074] The model building module enhances the robustness and accuracy of the deep learning network by introducing adaptive consistency regularization and dynamic loss weighting strategy. The deep learning network adopts dynamic data enhancement technology based on hybrid augmentation, and combines random cropping, flipping, spectral perturbation and regional hybrid enhancement to improve the model's adaptability to data diversity.
[0075] The result diagnosis module is used for multi-level explanatory analysis of pathological image label prediction. The specific numerical values of the labels generated by it are not only consistent with the training labels, but also come with the significance scores of the pathological areas.
[0076] Preferably, the data processing module retains the staining consistency of the pathological image by introducing an improved chromaticity-brightness enhancement technology, and also automatically and evenly processes the brightness distribution of the hyperspectral image by spectral normalization and dynamic contrast adjustment;
[0077] The deep learning network includes consistency regularization, pseudo-label generation module, dynamic data enhancement module and composite loss function module. It adopts a progressive learning strategy and divides the complex network training process into multiple stages.
[0078] Therefore, the present invention adopts the above-mentioned intelligent analysis method and system for automatic classification of pathological images, and the beneficial effects are as follows:
[0079] (1) The data processing module provided in the present invention adopts a variety of advanced technologies, such as spatial-spectral feature dimensionality reduction technology, morphological feature enhancement technology, chromaticity-brightness enhancement technology, etc., which effectively improves the data quality and enhances the representativeness and robustness of the features; the feature extraction module combines convolutional neural network with self-supervised learning and multi-head attention mechanism, which can comprehensively and accurately extract the spectral and spatial features of hyperspectral pathological images, provide better feature input for subsequent classification, and thus improve classification accuracy.
[0080] (2) The multi-level feature fusion architecture of the feature fusion module set up in the present invention fully combines the advantages of convolutional neural networks and visual Transformers. Through the phased fusion strategy and dynamic weighing of feature weights, it maximizes the complementarity of multimodal information and improves the richness and discrimination ability of feature expression. The automatic classification module adopts a deep learning model that combines contrastive learning and semi-supervision, and introduces a pseudo-label generation strategy with a dynamic confidence adjustment mechanism, which can effectively utilize limited labeled data and a large amount of unlabeled data, improve the generalization ability of the model on diversified data sets, reduce the risk of error propagation, and achieve more accurate classification.
[0081] (3) The model building module in the present invention significantly enhances the robustness and accuracy of the deep learning network through adaptive consistency regularization and dynamic loss weighting strategy and dynamic data enhancement technology of hybrid augmentation, so that it can better adapt to data diversity; the result diagnosis module introduces an attention mechanism based on multimodal features, and the generated labels are accompanied by the significance scores of the pathological areas, which not only makes the label values consistent with the training labels, but also provides a reliable quantitative basis, realizes multi-level explanatory analysis of pathological image label predictions, enhances the transparency and credibility of the model, helps to verify the consistency of model predictions with pathological features, and assists in pathological diagnosis decisions.
[0082] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 It is an overall flow chart of an intelligent analysis method and system embodiment for automatic classification of pathological images of the present invention;
[0084] Figure 2 It is a feature fusion schematic diagram of an intelligent analysis method and system embodiment for automatic classification of pathological images of the present invention. DETAILED DESCRIPTION
[0085] The technical solution of the present invention is further described below through the accompanying drawings and embodiments.
[0086] Unless otherwise defined, technical or scientific terms used in the present invention shall have the common meanings understood by one having ordinary skills in the field to which the present invention belongs.
[0087] like Figure 1 As shown, an intelligent analysis system for automatic classification of pathological images includes a data processing module, a feature extraction module, a feature fusion module, an automatic classification module, a model building module, a result diagnosis module and a dynamic adaptive optimization module.
[0088] The data processing module combines the spatial-spectral feature dimensionality reduction technology of hyperspectral images and the morphological feature enhancement technology of pathological images, and adopts a multi-scale feature compression strategy based on a deep convolutional encoder to map high-dimensional data into low-dimensional and representative feature representations. At the same time, a significant region detection algorithm is introduced to automatically mark high-correlation areas.
[0089] At the same time, the data processing module not only retains the staining consistency of pathological images by introducing improved chroma-brightness enhancement technology, but also automatically balances the brightness distribution of hyperspectral images through spectral normalization and dynamic contrast adjustment, thereby improving the robustness of subsequent feature extraction.
[0090] The feature extraction module extracts features of the region of interest (ROI) in the hyperspectral pathology image based on a method that combines convolutional neural networks with self-supervised learning. The feature extraction method includes using one-dimensional convolution to extract spectral features and using two-dimensional convolution to capture spatial texture features, and combining it with a multi-head attention mechanism to enhance the local and global feature representation capabilities.
[0091] At the same time, the feature fusion module adopts a multi-level feature fusion architecture, which combines the advantages of convolutional neural network (CNN) and visual transformer (ViT). Through a three-stage fusion strategy, it fuses the features of images from different sources at the shallow, deep and interactive layers, and dynamically weighs the weights of different features to maximize the complementarity of multimodal information.
[0092] The automatic classification module adopts a deep learning model based on contrastive learning and semi-supervision, generates pseudo labels using a small amount of labeled data, and iteratively optimizes by comparing the feature similarities of positive and negative samples to ensure that the classification model can achieve high generalization capabilities on diverse data sets.
[0093] The pseudo-label generation strategy further introduces a dynamic confidence adjustment mechanism, which updates the credibility of the pseudo-label in real time according to the current learning state of the model, thereby improving the quality of the pseudo-label and effectively reducing the risk of error propagation.
[0094] The model building module enhances the robustness and accuracy of the deep learning network by introducing adaptive consistency regularization and dynamic loss weighting strategies. The deep learning network adopts a dynamic data enhancement technology based on hybrid augmentation, combined with random cropping, flipping, spectral perturbation and regional hybrid enhancement, which improves the model's adaptability to data diversity.
[0095] Through modular design in the deep learning network, including consistency regularization, pseudo-label generation module, dynamic data enhancement module and composite loss function module, a progressive learning strategy is adopted to divide the complex network training process into multiple stages.
[0096] The result diagnosis module introduces an attention mechanism based on multimodal features to achieve multi-level explanatory analysis of pathological image label prediction. The specific numerical values of the labels generated are not only consistent with the training labels, but also come with the significance scores of the pathological areas, providing a reliable quantitative basis for diagnosis.
[0097] like Figure 2 As shown, an intelligent analysis method for automatic classification of pathological images includes the following steps:
[0098] S1. Preprocessing of hyperspectral images and pathological slice images in the data processing module is the basis for accurate analysis and classification. Through targeted processing methods, the quality of data and the effectiveness of information extraction can be significantly improved, specifically:
[0099] In the preprocessing operation of the hyperspectral image, the denoising step is first used to eliminate the random noise introduced by sensor noise or environmental interference during the acquisition process. To this end, the low-rank matrix decomposition method is used to optimize the hyperspectral image. Assuming that the original hyperspectral image data is recorded as X, the goal is to minimize the error function:
[0100]
[0101] Among them, ||·|| F represents the Frobenius norm, which is used to measure the overall error of the matrix; the matrix L is a low-rank component, which reflects the main structural information of the image; and S represents sparse noise, which is used to characterize those relatively discrete interference data.
[0102] At the same time, in order to ensure that the decomposition results are more realistic, regularization constraints are added:
[0103] ||L|| * +λ||S||1;
[0104] Among them, ||L|| * represents the nuclear norm of L, that is, the sum of its singular values, which is used to limit its low rank; ||S||1 represents the l1 norm of sparse noise, which measures the absolute value and non-zero elements in the sparse matrix; λ is the regularization parameter, which is used to balance the weight relationship between low-rank components and sparse noise.
[0105] This method can effectively separate the main information and random noise in the hyperspectral image, significantly improving the reliability of subsequent processing. The denoised hyperspectral image usually has a certain geometric distortion, which is caused by the optical characteristics of the imaging device or errors in operation.
[0106] On this basis, geometric correction is required to ensure that the spatial structure of the image is consistent with the actual physical space. The spectral matching transformation technology is used to geometrically align the spectral data of different bands to eliminate the distortion introduced by the equipment. This operation provides an accurate geometric benchmark for subsequent feature extraction and analysis. In terms of color and contrast processing of hyperspectral images, in order to further improve the visual effect and analysis stability, an improved chroma-brightness enhancement method is used, specifically:
[0107] The improved chroma-brightness enhancement method is used to optimize the dynamic range of the spectral image by linearly adjusting the brightness and contrast of the hyperspectral image. Its mathematical expression is:
[0108] I'=α·(I-μ)+β;
[0109] Among them, I' represents the new image after contrast and brightness adjustment, I represents the original image pixel value matrix, μ is the global mean of pixel values, which is used to eliminate the overall brightness offset; α is the contrast adjustment coefficient, which controls the dynamic range expansion of pixel values; β is the brightness compensation term, which adjusts the overall brightness level of the image. In this way, the spectral distribution characteristics of the image can be effectively improved, the visibility of the target features can be enhanced, and the recognition ability of the subsequent algorithm for key areas can also be improved.
[0110] The preprocessing operation of pathological slice images in the data processing module is:
[0111] First, in order to highlight the structural features of pathological tissue, the Laplace operator is used to enhance the morphological features. The Laplace operator is an image processing method based on the second-order derivative. By calculating the gradient change of pixel grayscale values, it can reflect the edge features in the image and make the tissue boundaries in the pathological sections clearer, thus providing a more intuitive basis for disease diagnosis.
[0112] On the basis of morphological enhancement, the saliency detection algorithm is further used to annotate the salient areas in the case slice images. Saliency detection is a technology that can automatically identify the region of interest in an image. Its core idea is to locate the area that is most likely to contain key pathological information by analyzing the image's texture, color, and brightness.
[0113] This method can not only reduce the intervention of manual operation, but also significantly improve the efficiency and accuracy of subsequent analysis. Through the automatic extraction of significant regions, the focus of analysis can be focused on relevant target areas, thereby optimizing the image processing process and ultimately providing high-quality input data for pathological research and diagnosis.
[0114] S2. In the feature extraction module, the pre-processed hyperspectral images and pathological slice images are deeply analyzed. By combining convolutional neural networks and self-supervised learning technology, the spectral characteristics and spatial information of the data can be fully captured. At the same time, the multi-head attention mechanism is used to improve the feature expression ability, thereby providing more accurate input for downstream tasks. The specific steps include:
[0115] S21. In the spectral feature extraction stage, a one-dimensional convolutional network is used to fully exploit the spectral dimension characteristics of hyperspectral image data.
[0116] A hyperspectral image data X contains a multi-dimensional tensor of multiple band information, where each dimension corresponds to the spectral intensity of a specific wavelength. One-dimensional convolution extracts spectral features through the following formula:
[0117] F spectrum =Conv1D(X;W s ,b s );
[0118] Among them, F spectrum represents the spectral features after one-dimensional convolution operation, W s Represents the weight matrix of the one-dimensional convolution kernel, controls the filtering mode of the convolution, and captures the local characteristics in the spectral dimension; b s Represents the bias term of the convolution layer, which is used to balance the output. By performing a convolution operation on the input data X, its local correlation can be extracted, revealing the deep connection between different bands, thereby better characterizing the spectral characteristics of the hyperspectral data.
[0119] S22. In the spatial feature extraction stage, in order to capture the spatial texture information of the image, a two-dimensional convolutional network is used to capture the spatial texture information of the image. Hyperspectral data is often presented as a multi-dimensional image in space, in which the value of each pixel corresponds to the reflection intensity of a specific wavelength. The two-dimensional convolution extracts spatial features through the following formula:
[0120] F spatial =Conv2D(X;W sp ,b sp );
[0121] Among them, F spatial represents the extracted spatial texture features, W sp b represents the weight matrix of a two-dimensional convolution kernel, which is used to detect local patterns in an image, such as edges, textures, and shapes; sp Represents the bias term corresponding to the two-dimensional convolution layer, which is used to adjust the output value, that is, the parameter used to offset the convolution result. The role of two-dimensional convolution is to capture the spatial relationship between pixels, thereby effectively representing the spatial pattern in the hyperspectral image.
[0122] S23. In order to further integrate spectral and spatial features and enhance the representation ability of the model, a multi-head attention mechanism is introduced. By calculating the global correlation of input features, it can effectively fuse local and global information, thereby improving the expressiveness and semantic consistency of features. The core calculation formula is:
[0123]
[0124] In the formula, Q, K, and V represent the query, key, and value of the feature matrix, respectively. They are generated from the input features through linear transformation and are used to characterize the correlation between different features. k is the number of columns with a scaling factor equal to K, which is used to prevent the attention score from becoming too large when the feature dimension is high, affecting the training stability.
[0125] Through QK TNormalization (Softmax operation) can generate an attention weight matrix, which is used to measure the importance of each feature point, thereby achieving the fusion of global context information. Finally, under the guidance of the attention weight, the feature matrix V is weighted and summed to form the output feature F attention , this process can significantly enhance the global expressiveness of features.
[0126] S3. In the feature fusion module, a three-stage fusion strategy is used to fully combine the characteristics of convolutional neural networks and Transformers to improve the richness and discrimination ability of multimodal data feature expression. The overall process starts with the initial fusion of shallow features, gradually goes deep into the extraction and fusion of deep semantic features, and finally performs global fusion in the channel dimension to generate more expressive final features. The specific process is as follows:
[0127] In the first stage, i.e. the shallow fusion stage, the linear weighted fusion of local features and global features is performed using the following formula:
[0128]
[0129] Among them, F CNN represents the local features extracted by the convolutional neural network, which is used to capture the fine-grained spatial texture information in the image; F Transformer represents the global features extracted by the Transformer module, which is used to capture long-distance dependencies and overall semantic information; ω1 and ω2 are dynamic weights, which are used to regulate the contribution of local and global features in the fusion results respectively; this linear weighting method can optimize feature expression by dynamically adapting to different data distributions by adjusting weights while retaining the original advantages of the two types of features.
[0130] The second stage is the deep fusion stage, which further performs deep semantic extraction and nonlinear enhancement on the features, and combines the features generated by the shallow fusion Input to the MBConv module for processing to obtain the deep fusion features:
[0131]
[0132] MBConv is an efficient convolutional structure that uses depthwise separable convolution and expansion factors to enhance computational efficiency and feature extraction capabilities. Through this process, the fused features not only obtain deeper semantic information, but also effectively suppress the noise and redundancy that may exist in shallow features.
[0133] Finally, in the final global feature fusion stage, the shallow fusion features are fused by the following formula and deep fusion features Concatenate in the channel dimension:
[0134]
[0135] This concatenation operation not only retains the local details and global relationships of the shallow features, but also combines the high-level semantic information of the deep features, thereby generating the final feature F with multi-dimensional information expression capabilities. final This design enables the fused features to represent both low-level texture details and high-level semantic information, greatly improving the feature discrimination and generalization performance.
[0136] S4. In the automatic classification module, the design of contrastive learning and semi-supervised strategy is combined to make full use of limited labeled data and a large amount of unlabeled data to improve the performance and generalization ability of the classification model. The core process of this module is to generate pseudo labels and optimize the loss function. By dynamically adjusting pseudo labels and introducing consistency regularization constraints, the coordinated training of supervised and unsupervised data is achieved. The specific process is as follows:
[0137] In the pseudo-label generation stage, the model dynamically generates pseudo-labels based on the prediction results of unlabeled data and the confidence threshold. Specifically, for each sample, the predicted probability distribution is expressed as P c Indicates that, P c is the predicted probability that the sample belongs to category c;
[0138] When the highest predicted probability of the sample max(P c ) exceeds the confidence threshold τ, the sample is considered to have a high classification credibility, and the label is generated at this time And its category is determined as the classification result with the highest probability through the following formula:
[0139]
[0140] Here, τ is a key parameter used to control the credibility of pseudo labels, and its value affects the quality and quantity of pseudo label generation. Through this mechanism, the noise of pseudo labels can be effectively reduced and the efficiency of the model's use of unlabeled data can be improved.
[0141] In the design of the loss function, both the supervised learning loss of labeled data and the semi-supervised learning loss of unlabeled data are considered. In the labeled data part, the classic cross entropy loss function is used:
[0142] L sup =CE(y,P model );
[0143] Among them, y is the true label, P model is the predicted probability distribution of the model output;
[0144] The cross entropy loss is used to measure the difference between the predicted distribution and the true label. The goal is to minimize this difference, thereby improving the classification accuracy of the model for labeled data.
[0145] For unlabeled data, the consistency regularization loss L is introduced unsup , by constraining the prediction consistency of the model under different data enhancement conditions, the robustness and generalization ability of the model can be improved. Specifically, when the highest prediction probability max(P) of the unlabeled data exceeds the confidence threshold τ, consistency regularization is enabled; on this basis, the original input is transformed through strong data enhancement technology, and the prediction probability distribution P of the enhanced sample is calculated strong , the consistency regularization loss is defined by the cross entropy formula:
[0146]
[0147] Among them, 1 {max(P)>β} represents the indicator function, which is activated only when the prediction confidence satisfies the threshold β; P strong It is the predicted distribution after strong enhancement, which aims to examine whether the model can maintain the pseudo label under different data enhancement conditions. prediction consistency.
[0148] The total loss function combines supervised loss and unsupervised loss to achieve coordinated optimization of the two. The total loss function is defined as:
[0149] L=L sup +μL unsup ;
[0150] Among them, μ is the weight coefficient of the unsupervised loss, which is used to balance the importance of supervised and unsupervised data in training. By adjusting λ, the learning intensity of the model for unlabeled data can be flexibly controlled.
[0151] S5. Dynamic optimization and model construction are performed in the dynamic adaptive optimization module and the model construction module. In order to improve the robustness and generalization ability of the model, a method combining adaptive consistency regularization and dynamic loss weighting strategy is designed. This method optimizes the model performance by constraining the prediction consistency of the model under different data enhancements and dynamically adjusting the loss weights during the training process. The specific process is as follows:
[0152] First, consistency regularization is introduced as one of the optimization objectives to improve the robustness of the model to perturbations. The core idea of consistency regularization is that after applying different data augmentation operations to the same input image, its prediction results should remain consistent.
[0153] Specifically, let the prediction distributions generated by the model for the two enhancement methods be P aug1 and Paug2 , then the consistency regularization term can be expressed as:
[0154] R consistency =||P aug1 -P aug2 ||2;
[0155] Among them, ||·||2 represents the Euclidean norm, which is used to quantify the difference between the two predicted distributions. By minimizing this regularization term, the model is guided to learn consistent predictions for different enhancement methods, thereby improving robustness to noise and input perturbations.
[0156] On this basis, in order to further adapt to the optimization requirements of different training stages, a dynamic loss weighting strategy is designed to dynamically adjust the loss weight ω t To balance the contribution between different objectives, the loss weight is calculated as:
[0157]
[0158] Among them, t represents the current number of training rounds, and γ is an adjustment rate parameter used to control the speed of weight change.
[0159] At the beginning of training, the loss weight ω t Close to 0, which means that the model will pay more attention to the main loss terms (such as classification loss or regression loss) in the initial stage to ensure the convergence of basic tasks. As the training progresses, the loss weight gradually increases, which gradually strengthens the influence of consistency regularization, thereby improving its robustness and adaptability to data perturbations as the model performance gradually converges.
[0160] In each round of training, two different data augmentation methods are applied to the input image to generate enhanced sample input models and obtain the predicted distribution P aug1 and P aug2 At the same time, the loss weight ω is calculated according to the current training round number t t The final optimization objective function consists of the main loss term and the consistency regularization term, where the weight of the consistency regularization term is ω t Dynamic adjustment. This dynamic optimization strategy can achieve a balance of objectives at different training stages, thereby effectively improving the model’s robustness, generalization ability, and adaptability to the diversity of input data.
[0161] S6. In the result diagnosis module, the attention mechanism is used to perform multi-level explanatory analysis on the hyperspectral data, thereby providing an intuitive understanding of the pathological or classification results. This module generates a significance score to evaluate the importance of specific areas in the input image, helping users understand the decision basis of the model. The specific process is as follows:
[0162] The result diagnosis module generates a significance score S based on the attention weights. attention , and its calculation formula is:
[0163] S attention =Softmax(W a ·F final );
[0164] Among them, S attention W represents the saliency score of each pixel or region, reflecting its role in the final decision. a is the attention weight matrix, which is used to measure the importance of input features at different positions. final Represents the final feature matrix extracted by the neural network, which contains the comprehensive information of the image in spatial and spectral dimensions.
[0165] In this formula, W a The specific calculation of F is automatically optimized by the model's learning process, usually by back-propagating the parameters adjusted according to the objective function. It reflects the model's attention distribution in the input feature space. final It is the output of the last layer in the feature extraction process, which aggregates the deep semantic information after multiple layers of convolution or transformation operations and can describe the potential important patterns in the image. a With F final Multiplying together, the attention mechanism can highlight important features related to the current task and weaken background or irrelevant information.
[0166] The scores are then normalized using the Softmax function to ensure that the sum of all saliency scores is 1. This normalization allows the saliency score to intuitively represent the relative contribution of each pixel or region to the overall decision. The results of the saliency score can be mapped back to the spatial location of the input image to generate a saliency heat map, which can intuitively show the areas that the model pays the most attention to. This explanatory analysis not only enhances the transparency of the model, but also helps users verify whether the model's predictions are consistent with pathological features.
[0167] Therefore, the present invention adopts the above-mentioned intelligent analysis method and system for automatic classification of pathological images, which improves data quality and feature expression while enhancing model accuracy and credibility through the collaborative operation of multiple modules and the integration of multiple technologies and strategies.
[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. An intelligent analysis method for automatic classification of pathological images, characterized in that: The following steps are involved: S1, preprocessing the hyperspectral image and pathological section image; S2, performing in-depth analysis on the pre-processed hyperspectral images and pathological slice images; S3, global feature fusion in the channel dimension through a three-stage fusion strategy; S4, pseudo label generation and loss function optimization; S5, dynamic optimization and model building; S6. Use attention mechanism to perform multi-level interpretive analysis on hyperspectral data.
2. The intelligent analysis method for automatic classification of pathological images according to claim 1, characterized in that: In S1, the preprocessing operation of the hyperspectral image in the data processing module is as follows: De-noising is used to eliminate random noise introduced by sensor noise or environmental interference during the acquisition process, and the low-rank matrix decomposition method is used to optimize the hyperspectral image. Assuming that the original hyperspectral image data is recorded as X, the goal is to minimize the error function: Among them, ||·|| F represents the Frobenius norm, the matrix L is the low-rank component, and S represents sparse noise; Add regularization constraints: ||L|| * +λ||S||1; Among them, ||L|| * represents the nuclear norm of L, ||S||1 represents the l1 norm of sparse noise, and λ is the regularization parameter; The spectral matching transformation technology is used to geometrically align the spectral data of different bands, specifically: The improved chroma-brightness enhancement method is used to linearly adjust the brightness and contrast of the hyperspectral image. Its mathematical expression is: I'=α·(I-μ)+β; Among them, I' represents the new image after contrast and brightness adjustment, I represents the original image pixel value matrix, μ is the global mean of pixel values, α is the contrast adjustment coefficient, and β is the brightness compensation term.
3. The intelligent analysis method for automatic classification of pathological images according to claim 2, characterized in that: In S1, the preprocessing operation of the pathological slice image in the data processing module is: The Laplace operator is used to enhance the morphological features, and the edge features in the image are reflected by calculating the gradient change of the pixel gray value; A saliency detection algorithm is used to annotate the salient areas in the case slice images, and the areas containing key pathological information are located by analyzing the texture, color and brightness features of the images.
4. The intelligent analysis method for automatic classification of pathological images according to claim 3, characterized in that: In S2, the pre-processed hyperspectral image and pathological slice image are deeply analyzed in the feature extraction module, which specifically includes the following steps: S21. In the spectral feature extraction stage, a one-dimensional convolutional network is used to mine the spectral dimension characteristics of the hyperspectral image data. The one-dimensional convolution extracts the spectral features through the following formula: F spectrum =Conv1D(X;W s ,b s ); Among them, F spectrum represents the spectral features after one-dimensional convolution operation, W s represents the weight matrix of the one-dimensional convolution kernel, b s Represents the bias term of the convolutional layer; S22. In the spatial feature extraction stage, a two-dimensional convolutional network is used to capture the spatial texture information of the image. The two-dimensional convolution extracts the spatial features through the following formula: F spatial =Conv2D(X;W sp ,b sp ); Among them, F spatial represents the extracted spatial texture features, W sp Represents the weight matrix of the two-dimensional convolution kernel, b sp The offset term corresponding to the two-dimensional convolution layer is used to adjust the offset of the convolution result. S23, introduce the multi-head attention mechanism to fuse local and global information by calculating the global correlation of input features. The calculation formula is: In the formula, Q, K, and V represent the query, key, and value of the feature matrix respectively, and d k is the number of columns for which the scaling factor is equal to K.
5. The intelligent analysis method for automatic classification of pathological images according to claim 4, characterized in that: In S3, a three-stage fusion strategy is used in the feature fusion module to perform global fusion in the channel dimension. The specific process is as follows: In the shallow fusion stage, the linear weighted fusion of local features and global features is performed using the following formula: Among them, F CNN represents the local features extracted by the convolutional neural network, F Transformer represents the global features extracted by the Transformer module, ω1 and ω2 are dynamic weights, which are used to regulate the contribution of local and global features in the fusion results respectively; In the deep fusion stage, the features generated by the shallow fusion Input to the MBConv module for processing to obtain the deep fusion features: In the final global feature fusion stage, the shallow fusion features are fused by the following formula and deep fusion features Concatenate in the channel dimension:
6. The intelligent analysis method for automatic classification of pathological images according to claim 5, characterized in that: In S4, pseudo-label generation and loss function optimization are performed in the automatic classification module. The specific process is as follows: In the pseudo-label generation stage, the model dynamically generates pseudo-labels based on the prediction results of unlabeled data and the confidence threshold. For each sample, the predicted probability distribution is expressed as P c Indicates that, P c is the predicted probability that the sample belongs to category c; When the highest predicted probability of the sample max(P c ) exceeds the confidence threshold τ, a label is generated And its category is determined as the classification result with the highest probability through the following formula: Consider the supervised learning loss of labeled data and the semi-supervised learning loss of unlabeled data. In the labeled data part, the cross entropy loss function is used: IT sup =CE(y,P model ); Among them, y is the true label, P model is the predicted probability distribution of the model output; For unlabeled data, the consistency regularization loss L is introduced unsup , when the highest prediction probability max(P) of unlabeled data exceeds the confidence threshold τ, consistency regularization is enabled; The original input is transformed through strong data enhancement technology, and the predicted probability distribution P of the enhanced sample is calculated strong , the consistency regularization loss is defined by the cross entropy formula: Among them, 1 {max(P)>β} represents the indicator function, which is activated only when the prediction confidence satisfies the threshold β; P strong is the predicted distribution after strong enhancement; The total loss function combines the supervised loss and the unsupervised loss and is defined as: L=L sup +μL unsup ; Among them, μ is the weight coefficient of the unsupervised loss.
7. The intelligent analysis method for automatic classification of pathological images according to claim 6, characterized in that: In S5, dynamic optimization and model building are performed in the dynamic adaptive optimization module and the model building module. The specific process is as follows: First, we introduce consistency regularization as the optimization objective and assume that the prediction distributions generated by the model for the two enhancement methods are P aug1 and P aug2 , then the consistency regularization term is expressed as: R consistency =||P aug1 -P aug2 ‖2; Among them, ||·||2 represents the Euclidean norm, which is used to quantify the difference between the two prediction distributions; By dynamically adjusting the loss weight ω t To balance the contribution between different objectives, the loss weight is calculated as: Among them, t represents the current number of training rounds, and γ is an adjustment rate parameter used to control the speed of weight change.
8. The intelligent analysis method for automatic classification of pathological images according to claim 7, characterized in that: In S6, in the result diagnosis module, the attention mechanism is used to perform multi-level interpretive analysis on the hyperspectral data. The specific process is as follows: Generate a saliency score S based on the attention weight attention , and its calculation formula is: S attention =Softmax(W a ·F final ); Among them, S attention represents the saliency score of each pixel or region, W a is the attention weight matrix, F final Represents the final feature matrix extracted by the neural network, which contains the comprehensive information of the image in spatial and spectral dimensions.
9. An intelligent analysis system for automatic classification of pathological images, characterized in that: It includes data processing module, feature extraction module, feature fusion module, automatic classification module, model building module, result diagnosis module and dynamic adaptive optimization module; The data processing module combines the spatial-spectral feature dimensionality reduction technology of hyperspectral images and the morphological feature enhancement technology of pathological images, and adopts a multi-scale feature compression strategy based on a deep convolutional encoder to map high-dimensional data into low-dimensional and representative feature representations. At the same time, a significant region detection algorithm is introduced to automatically mark high-correlation regions. The feature extraction module uses a method that combines convolutional neural networks with self-supervised learning to extract features from the region of interest (ROI) in hyperspectral pathology images, including using one-dimensional convolution to extract spectral features and two-dimensional convolution to capture spatial texture features, and combines a multi-head attention mechanism to improve local and global feature representation capabilities; The feature fusion module adopts a multi-level feature fusion architecture. Through a three-stage fusion strategy, it fuses the features of images from different sources at the shallow layer, deep layer, and interactive layer, and dynamically weighs the weights of different features to maximize the complementarity of multimodal information. The automatic classification module uses a deep learning model based on contrastive learning and semi-supervision, generates pseudo labels using a small amount of labeled data, and iteratively optimizes by comparing the feature similarities of positive and negative samples; The model building module enhances the robustness and accuracy of the deep learning network by introducing adaptive consistency regularization and dynamic loss weighting strategy. The deep learning network adopts dynamic data enhancement technology based on hybrid augmentation, and combines random cropping, flipping, spectral perturbation and regional hybrid enhancement to improve the model's adaptability to data diversity. The result diagnosis module is used for multi-level explanatory analysis of pathological image label prediction. The specific numerical values of the labels generated by it are not only consistent with the training labels, but also come with the significance scores of the pathological areas.
10. The intelligent analysis system for automatic classification of pathological images according to claim 9, characterized in that: The data processing module retains the staining consistency of pathological images by introducing improved chroma-brightness enhancement technology, and automatically balances the brightness distribution of hyperspectral images through spectral normalization and dynamic contrast adjustment; The deep learning network includes consistency regularization, pseudo-label generation module, dynamic data enhancement module and composite loss function module. It adopts a progressive learning strategy and divides the complex network training process into multiple stages.
Citation Information
Cited By
Hyperspectral imaging-based cervical cancer detection system, method, equipment and medium
CN120543545A
Intelligent image recognition and classification system based on deep learning
CN120783091A
Neurological pathological feature analysis method and system
CN120809177A
Cinnamomum camphora yellows control system and method based on machine vision
CN120876144A
Digestive enteritis pathological section image adaptive feature processing method
CN120953150A