Method, apparatus and readable storage medium for multi-label recognition of pathological sections
A multi-label recognition model optimizes focus settings for pathology slides by identifying types and surface features, improving scanning efficiency and accuracy without additional hardware, addressing the inefficiencies of fixed focus and high-cost real-time adjustments.
Patent Information
- Application Number
- CN202510638057.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-19
AI Technical Summary
Existing pathological slice scanners are inefficient and costly when identifying slice types and setting focus density, cannot be dynamically optimized, and it is difficult to balance scanning efficiency and image clarity.
A multi-label recognition model based on cubic spline interpolation verification is adopted. By constructing a multi-label recognition model combining pathological slice type and focus density type, the minimum effective focus density is dynamically calculated using the focus point three-dimensional coordinate data recorded by the scanning device, and combining the residual convolution module and dynamic threshold post-processing algorithm to achieve scanning efficiency optimization without hardware upgrades.
It significantly improves scanning efficiency (30%-50% shortens time), reduces hardware costs, improves classification accuracy (99.1%) and segmentation accuracy (98.3%), enhances the robustness and versatility of the model, and adapts to different production methods and sample shapes.
Smart Images

Figure CN120164048B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing and medical equipment control technology, and in particular to a method and device for multi-label recognition of pathological sections and a readable storage medium thereof. Background Art
[0002] Pathology section scanning is a key step in pathology diagnosis. When converting pathology sections into digital images through a scanner, the focus density setting directly affects the scanning time and image clarity. Existing scanners usually use fixed high-density focus points or real-time focus technology:
[0003] 1. Although fixed high-density focus points can ensure clear images, the scanning time is long and the efficiency is low;
[0004] 2. Real-time focus technology dynamically adjusts the focus height through a small number of focus points. Although it reduces time, it relies on high-precision hardware and significantly increases equipment costs.
[0005] In addition, the existing pathological tissue recognition model can only perform single-label classification (such as staining method classification) and cannot simultaneously deal with the focus point density setting problem that is closely related to tissue morphology. The surface undulation differences caused by different preparation methods (such as TCT, biopsy puncture) and sample shapes (regular tissue, adipose tissue) require matching different optimal focus point densities, and traditional methods lack the ability to dynamically optimize this key parameter.
[0006] Therefore, how to synchronously identify the pathological section type and dynamically determine the minimum effective focus point density through algorithms without increasing hardware costs has become a technical problem that needs to be urgently solved in this field. Summary of the invention
[0007] The embodiments of the present invention provide a method and device for multi-label identification of pathological sections and a readable storage medium thereof, which address the problems existing in current technologies, such as the inability to simultaneously identify the type of pathological sections (such as staining method, preparation method) and surface undulation characteristics (sample shape), resulting in the focus density having to be fixed or relying on high-cost hardware adjustment, making it difficult to achieve an optimal balance between scanning efficiency and image clarity.
[0008] The core technology of the present invention is mainly a multi-label recognition model based on cubic spline interpolation verification. By constructing a multi-label recognition model that combines pathological section types (6 categories) and focus density types (4 categories), the model is trained using the three-dimensional coordinate data of the focus points recorded by the scanning equipment. The minimum effective focus point density is dynamically calculated based on the undulating characteristics of the tissue surface, thereby achieving scanning efficiency optimization without the need for hardware upgrades.
[0009] In a first aspect, the present invention provides a method for multi-label recognition of pathological sections, the method comprising the following steps:
[0010] S1. Obtain the training dataset:
[0011] Based on the focus height data of the pathological slice scan images, fit the reference focus plane through cubic spline interpolation, and determine the fitting error threshold by combining the objective lens depth of field parameter and the device motion error. Divide the focus density type labels according to the fitting result;
[0012] S2. Construct a multi-label recognition model:
[0013] Adopt a neural network architecture containing a tandem residual convolution module to perform multi-scale feature extraction on the input pathological slice scan images;
[0014] Generate shallow features and deep features through upsampling and feature splicing. The shallow features are used to identify the focus density type, and the deep features are combined with the processing results of the shallow features to identify the slice type; the multi-label recognition model outputs the slice type classification probability vector, the focus density type classification probability vector, the slice type segmentation probability map, and the focus density type segmentation probability map;
[0015] S3. Dynamic threshold post-processing:
[0016] Dynamically adjust the segmentation threshold according to the confidence of the slice type classification probability vector, process the slice type segmentation probability map and the focus density type segmentation probability map, and determine the final slice type and focus density type in combination with the processed segmentation results.
[0017] Furthermore, the specific steps of S1 include:
[0018] Obtain the set of three-dimensional coordinates of the focus points of each tile in the pathological slice scan image through a pathological slice scanner. The three-dimensional coordinates of the focus points include position coordinates and focus height;
[0019] Uniformly sample the focus points at a preset interval, and use the cubic spline interpolation method to fit and generate the focus surface;
[0020] Calculate the vertical distance between the unsampled focus points and the focus surface at the corresponding horizontal and vertical coordinates, determine the distance threshold according to the objective lens depth of field and the device motion error, and count the proportion of the focus points with a vertical distance less than the distance threshold. When the proportion is not lower than the preset fitting error threshold, determine the current interval as the best focus interval for this pathological slice;
[0021] Based on the best focus interval, divide the focus density into multiple categories and label the density labels.
[0022] Furthermore, in the specific steps of S1, the method for classifying the focus density categories includes:
[0023] The optimal focus interval is divided into 4 intervals greater than 5 mm, 5 - 4 mm, 4 - 3 mm, and 3 - 2 mm at 1 - mm intervals, corresponding to 4 types of focus density type labels respectively.
[0024] Further, in step S2, the residual convolution module includes multiple convolution modules. The convolution modules sequentially perform convolution, batch normalization, and ReLU activation processing on the input features; the residual convolution module splices the features after different convolution processes and outputs the fused features.
[0025] Further, in step S3, the dynamic threshold post - processing includes:
[0026] When the classification confidence of the slice type is greater than or equal to the first threshold, a higher segmentation threshold is used to process the segmentation probability maps of the slice type and the focus density type, and the union of the segmentation masks of the two is used as the recognition result;
[0027] When the classification confidence of the slice type is between the first threshold and the second threshold, a lower segmentation threshold is used to process the segmentation probability map, and the union of the segmentation masks of the two is used as the recognition result;
[0028] When the classification confidence of the slice type is less than or equal to the second threshold, the segmentation threshold of the focus density type is dynamically adjusted according to the classification confidence of the focus density type, and the segmentation result of the focus density type is used as the recognition result.
[0029] Further, the first threshold is set to 0.8, the second threshold is set to 0.1; the higher segmentation threshold is 0.5, and the lower segmentation threshold is 0.2.
[0030] Further, in the training process of the multi - label recognition model in step S2, the detection loss function comprehensively considers the differences in the position, size, and overlapping area between the predicted box and the ground - truth box. The classification loss function uses the binary cross - entropy loss function, and the total loss function balances the training effect by cyclically adjusting the weights of the detection loss and the classification loss.
[0031] In the second aspect, the present invention provides a multi - label recognition device for pathological slices, including:
[0032] A dataset construction module, based on the focus height data of the pathological slice scan image, fits the reference focus plane through cubic spline interpolation, combines the objective lens depth - of - field parameter and the device motion error to determine the fitting error threshold, and divides the focus density type labels according to the fitting result;
[0033] A multi - label recognition model construction module, which uses a neural network architecture including a series - connected residual convolution module to perform multi - scale feature extraction on the input pathological slice scan image;
[0034] Generate shallow features and deep features through upsampling and feature splicing. The shallow features are used to identify the focus density type, and the deep features are combined with the processing results of the shallow features to identify the slice type; the multi-label recognition model outputs a slice type classification probability vector, a focus density type classification probability vector, a slice type segmentation probability map, and a focus density type segmentation probability map;
[0035] A dynamic threshold post-processing module dynamically adjusts the segmentation threshold according to the confidence of the slice type classification probability vector, processes the slice type segmentation probability map and the focus density type segmentation probability map, and determines the final slice type and focus density type in combination with the processed segmentation results.
[0036] In a third aspect, the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the above-mentioned pathological slice multi-label recognition method.
[0037] In a fourth aspect, the present invention provides a readable storage medium. A computer program is stored in the readable storage medium, and the computer program includes program code for controlling a process to execute the process, and the process includes the above-mentioned pathological slice multi-label recognition method.
[0038] The main contributions and innovations of the present invention are as follows:
[0039] 1. Efficiency improvement: By dynamically setting the focus density, the number of foci is significantly reduced while ensuring image clarity, shortening the scanning time (for example, for TCT samples with less surface undulation, a larger focus interval can be used, and the scanning time is reduced by 30%-50%).
[0040] 2. Cost advantage: The model can be deployed on general-purpose hardware such as CPU / GPU, without the need for dedicated high-precision focusing equipment, reducing the hardware cost of the scanner.
[0041] 3. Breakthrough in recognition accuracy: Compared with existing single-label models (such as MaskRCNN, Unet), the multi-label segmentation model of the present invention has improved the classification accuracy to 99.1% and the segmentation accuracy (MPA50) to 98.3% (see Table 1), especially the recognition effect on complex morphologies such as adipose tissue and puncture samples has been significantly enhanced.
[0042] 4. Enhanced robustness: The dynamic threshold post-processing algorithm can adaptively process new sample types outside the training data, and dynamically adjusts the segmentation threshold by combining the confidence of the slice type and the focus density, avoiding missed detection or misjudgment.
[0043] 5. General-purpose design: Covers various slice types such as TCT, HE, and IHC, is compatible with different specimen preparation methods and sample shapes, and forms a standardized focus strategy optimization scheme.
[0044] The details of one or more embodiments of the present invention are set forth in the following drawings and description to make other features, objects, and advantages of the present invention more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0046] Figure 1 is a flowchart of a method for multi-label recognition of pathological sections according to an embodiment of the present invention;
[0047] Figure 2 is a distribution diagram of the optimal focus interval according to an embodiment of the present invention;
[0048] Figure 3 is a model framework diagram according to an embodiment of the present invention;
[0049] Figure 4 is a convolution module diagram according to an embodiment of the present invention;
[0050] Figure 5 is a residual convolution module diagram according to an embodiment of the present invention;
[0051] Figure 6 is a flowchart of post-processing with a dynamic threshold according to an embodiment of the present invention;
[0052] Figure 7 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0054] It should be noted that: In other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0055] In the prior art, it is impossible to dynamically determine the optimal focus density based on the morphological features of pathological sections, resulting in difficulty in balancing scanning efficiency and image quality.
[0056] Based on this, the present invention is based on a multi-label recognition model verified by cubic spline interpolation. By synchronously identifying the sample type based on the shape features of pathological sections and predicting the optimal focus density category (0-3 classes), the problems existing in the prior art are solved.
[0057] Embodiment 1
[0058] The present invention aims to propose a multi-label recognition method for pathological sections. Specifically, referring to Figure 1 , the method includes:
[0059] S1. Dataset construction (obtaining a training dataset):
[0060] Based on the focus height data of pathological section scanning images, a reference focus plane is fitted by cubic spline interpolation, and a fitting error threshold is determined in combination with the objective lens depth of field parameter and the device movement error. The focus density type label is divided according to the fitting result;
[0061] In this embodiment, the essence of setting the focus density is to set the sample sampling rate, and the optimal sampling rate is related to the data situation. The surface undulation of pathological tissues and the section type are related to the tissue shape. For example, the TCT preparation method makes the surface of such samples present as a plane; the tissue obtained by biopsy puncture is strip-shaped with large surface undulation; the surface of tissues with more fat or cavities usually has large undulation, while the surface of complete and regular tissues is closer to a plane. Pathological section scanners usually have the function of saving the focus height and position of each tile. The focus height of clear tiles can be considered accurate. Therefore, the position and height data of tiles can be obtained from existing pathological section scanning images, and then a training dataset is formed. Each tile in each pathological image can be regarded as a focus, and the three-dimensional coordinates of the focus are recorded as the set . To ensure the generalization of the algorithm, data needs to be extracted from different types of slice images.
[0062] For quantitative analysis of the focus density setting, uniform sampling is performed at a certain distance interval G, and the selected foci are used for cubic spline interpolation fitting to obtain a fitting surface, where G ranges from 200 at an interval of 200 to 5000 Take values. Then calculate the vertical distance E between the unselected focus points and the corresponding horizontal and vertical coordinates of the fitted surface. Consider the focus points with smaller gaps as correctly fitted, and vice versa, consider the fitting height as incorrect. The threshold T for correct fitting is determined by the depth of field D of the objective lens and the device movement error V. The imaging within the depth of field is clear, but there are errors and vibrations in the platform movement. Therefore, the final threshold is defined as T = D / 2 - V. (In the present invention, D = 1 , V = 0.2 ). Statistically, calculate the proportion of focus points with a gap E less than T among all focus points. When the proportion is less than the fitting error threshold α, it is considered that the scan is clear, and its value range is recommended to be between 0% and 10%. The value of α is related to the clarity requirement and the preparation quality of the specimen.
[0063] Specifically, if a more accurate virtual surface to be fitted is required, and thus the scanned image is clearer, then a smaller value of α is set. However, the smaller the α is set, the more focus points will be generated, resulting in an increase in time consumption. In particular, because there is a situation where the preparation quality is low, resulting in large fluctuations in the sliced tissue, it is difficult for the fitted virtual surface to fully conform to the actual focal plane. Therefore, it is not recommended to set α too small. In the present invention, α is set to 5%, and the virtual surface fitting method uses the cubic spline surface interpolation method.
[0064] Each slice will be judged as clear or blurred under different intervals. The distance when it is judged as clear and the interval is the largest is used as the best focus interval distance for this slice. Then, based on the best focus interval distance of the slice, it is used as the basis for density label division. Figure 2 is the distribution probability of the best focus interval distance, where the smallest distance is close to 2 mm and the largest distance exceeds 5 mm. Therefore, start from 2 mm and divide the density types in units of 1 mm. Since the condition of clear scanning is satisfied when the interval exceeds 5 mm, and it is also satisfied when set to 5 mm, distances exceeding 5 mm are regarded as one category. Finally, the focus interval distances are set to greater than 5 mm, 5 mm - 4 mm, 4 mm - 3 mm, 3 mm - 2 mm, and the corresponding labels are set to 0 - 3 respectively. In particular, in the present invention, the categories are divided at intervals of 1 mm, and they can also be divided at smaller intervals. If divided at smaller intervals (such as 0.2 mm), since the interval becomes smaller, the categories of adjacent intervals will be relatively similar, such as 2.2 mm - 2.4 mm and 2.4 mm - 2.6 mm. This will cause the model to be difficult to distinguish adjacent categories. Eventually, it will be misclassified into adjacent categories, such as the category of 2.2 mm - 2.4 mm being recognized as the category of 2.4 mm - 2.6 mm or the category of 2.6 mm - 2.8 mm.
[0065] S2. Construct a multi-label recognition model:
[0066] Adopt a neural network architecture containing a series-connected residual convolution module to perform multi-scale feature extraction on the input scanned pathological section image;
[0067] Generate shallow features and deep features through upsampling and feature concatenation. The shallow features are used to identify the focus density type, and the deep features combine the processing results of the shallow features to identify the section type; The multi-label recognition model outputs a section type classification probability vector, a focus density type classification probability vector, a section type segmentation probability map, and a focus density type segmentation probability map.
[0068] In this embodiment, existing scanner algorithms all use neural networks for single-label multi-class segmentation tasks and can only identify the type of the section. In order to simultaneously identify the section type and the focus setting type, the present invention designs a multi-label recognition model that can identify 6 section types (TCT, HE, IHC, PAS, oil red, and other special stains) and 4 focus density types.
[0069] The architecture of the multi-label recognition model is as Figure 3 shown. In the figure, [w, h, c] is used to represent the feature dimension, w represents the width, h represents the height, and c represents the number of channels. The multi-label recognition model uses five series-connected residual convolution modules to extract features of different scales, and then obtains shallow features [80, 80, 768] and deep features [40, 40, 1024] through upsampling and concatenating features of different scales. The shallow features will pass through the residual convolution module to obtain density type features [80, 80, 256], and finally pass through the convolution module and convolution to obtain the detection and classification results.
[0070] The basis for density type division is mainly shape, and shape is also one of the bases for section type division, and the segmentation results of the two types are the same. Therefore, the density type features can be used for the classification and recognition of section types. The model further processes the density type features through the convolution module, concatenates the obtained features with the deep features, then the concatenated features pass through the residual convolution module to extract features and obtain section type features, and finally pass through the convolution module and convolution to obtain the detection and classification results.
[0071] As Figure 4As shown, the convolutional module consists of convolution, batch normalization, and the activation function Relu. k, s, p, and c are several important hyperparameters. k represents the size of the convolutional kernel. s is the stride, indicating the interval at which the convolutional kernel slides on the input data. p represents the number of additional pixels added around the boundary of the input data. c_in represents the number of channels of the input data or the convolutional kernel. Since the number of channels of the input data is the same as that of the convolutional kernel, the size of the convolutional kernel c_in will not be specifically described in the figure or subsequent text. At the same time, the dimension c_out of the output data is the number of convolutional kernels and will not be specifically described either. Batch normalization normalizes each feature dimension separately to make its mean 0 and variance 1.
[0072] Among them, bilinear interpolation is used for upsampling, which can enlarge the width and height of the feature map. For example, the feature [20, 20, 512] is upsampled by a factor of two to become [40, 40, 512]. Concatenation requires the input features to have the same width and height and can merge the input in the channel dimension. For example, the features [40, 40, 512] and [40, 40, 256] become [40, 40, 768] after concatenation.
[0073] As Figure 5 shown, the residual convolutional module is composed of multiple convolutional modules combined. This module splices the features processed by different numbers of convolutional modules together, and then inputs the spliced features into the last convolutional module to obtain the output.
[0074] In this embodiment, the training loss function of the multi-label recognition model includes detection loss and classification loss. The model first identifies the region where the sample is located and outputs a prediction box, and then classifies the pixels within each prediction box. The position and size of the prediction box relative to the ground truth box (the minimum bounding rectangle of the sample) are the criteria for measuring the detection effect. The detection loss comprehensively considers the differences between the prediction box and the ground truth box to maximize the overlapping area between the two and minimize the differences in the center distance and aspect ratio. The specific formula is as follows:
[0075]
[0076] Among them, IoU represents the intersection over union of the prediction box and the ground truth box, d() calculates the Euclidean distance between two points, p represents the center point of the prediction box, p gt represents the center point of the ground truth box, c gt represents the diagonal length of the ground truth box, a is a hyperparameter, default set to 0.2, w and h represent the width and height of the prediction box, and wgt and hgt represent the width and height of the ground truth box.
[0077] Preferably, the multi-label recognition model uses the binary cross-entropy loss function as the classification loss, that is, the cross-entropy loss is calculated separately for each pixel within the prediction box, and then the mean value of all cross-entropy loss values is obtained. The mathematical formula is as follows:
[0078]
[0079] where N represents the number of samples, y i represents the true label, and p i represents the predicted probability value.
[0080] Since the two types of loss functions use different features and labels, it is necessary to balance the loss functions of both. Otherwise, there will be a situation where one type has a good effect while the other is poor. Therefore, the present invention uses a loss function hyperparameter that changes cyclically, enabling the model to focus on both types during training. The total loss function L is defined as follows:
[0081]
[0082] where, and represent the density type detection loss and the classification loss, and represent the slice type detection loss and the classification loss. iter is the total number of training iterations, which can be adjusted and is default set to 200. k is the k-th iteration of training, that is, after training all the data once, k is incremented by one. Therefore, the range of k is [1, iter].
[0083] S3, Dynamic Threshold Post-Processing:
[0084] Dynamically adjust the segmentation threshold according to the confidence of the slice type classification probability vector, process the slice type segmentation probability map and the focal density type segmentation probability map, and determine the final slice type and focal density type based on the processed segmentation results.
[0085] In this embodiment, when the input image is fed into the multi-label recognition model, it will output the classification probability vector and the segmentation probability map of the slice type and the density type. The maximum value in the classification probability vector is the confidence, and the position of the vector where the maximum value is located is the classification result. Define the slice type classification confidence as x1 and the density type classification confidence as x2. To obtain an accurate recognition result, the present invention proposes a dynamic threshold post-processing algorithm, which combines the classification and segmentation results of both types to calculate the final recognition result. The specific method is as Figure 6 shown.
[0086] The segmentation probability maps output by the model will have overlaps. The non-maximum suppression algorithm is respectively applied to the slice type and density type, where the confidence threshold is 0.1 and the intersection over union threshold is 0.75. After processing, most of the overlapping probability maps will be excluded, and each slice type segmentation probability map will correspond to a density type segmentation probability map. The two probability maps together represent the recognized target. Then, the threshold for post-processing is dynamically adjusted according to the classification confidence of the slice type.
[0087] When the classification confidence x1 of the slice type is greater than a1 = 0.8, it indicates that the model can accurately identify the category of the pathological slice sample, and the corresponding segmentation result has high accuracy. Therefore, a higher threshold t1 = 0.5 can be used to process the segmentation probability map of the slice type. Since the segmentation result of the slice type is relatively accurate, and the segmentation result of the sample shape is only used for minor supplementation, a higher threshold is set to process the segmentation probability map of the density type, with a2 = 0.8 and t2 = 0.5. After processing the segmentation probability maps, segmentation masks will be obtained, and the union of the two mask maps is used as the final recognition result.
[0088] When x1 is between a1 and b1, the accuracy of the recognized slice type decreases. Therefore, the results of the density type need to be combined. So, lower thresholds are set to process the segmentation probability maps. Set t1 = 0.1, a2 = 0.4, and t2 = 0.2. The union of the two mask maps is used as the recognition result.
[0089] When x1 is less than b1, it means that the sample of the slice type is not recognized. In addition, in actual use, there may be pathological tissue types not included in the training data. At this time, the confidence value of the slice type obtained by model inference will be relatively low. So, when the input image is a new pathological tissue slice type, it will also result in the failure to recognize the sample. To be able to recognize slice types not included in the training data, the segmentation result of the density type can be used as the final result. Because the higher the confidence x2 of the density type, the more accurate the segmentation result will be, the threshold of the segmentation probability map can be set lower. Therefore, set a2 = 0.1 and t1 = 0.6 - x2 / 2.
[0090] Since most methods do not support multi-label segmentation, only the classification accuracy of the slice type and the overall segmentation accuracy are compared. Table 1 shows the comparison results between the present invention and existing methods. In terms of the classification and segmentation accuracy of pathological slices, the accuracy of the present invention is higher than that of all comparison methods. The experimental results prove that the multi-label segmentation model and dynamic threshold processing designed by the present invention can effectively improve the recognition effect of pathological slices.
[0091] Table 1
[0092]
[0093] The following is a supplementary explanation of the technical terms involved in the present invention to help understand the core concepts of the technical solution:
[0094] 1. Cubic spline interpolation
[0095] Explanation: A mathematical interpolation method that fits discrete data points through piecewise cubic polynomial functions to generate a smooth surface (such as the focus plane). In the present invention, it is used to fit the focus point height data, calculate the vertical distance between the unsampled points and the fitted surface, and evaluate the accuracy of the focus plane.
[0096] 2. Non-maximum suppression algorithm (NMS)
[0097] Explanation: A post-processing algorithm used to remove overlapping prediction boxes in object detection or segmentation and retain the result with the highest confidence. In the present invention, NMS is applied to the segmentation probability maps of slice types and focus density types to reduce redundant predictions and improve the accuracy of the recognition results.
[0098] Embodiment 2
[0099] Based on the same concept, the present invention also proposes a multi-label recognition device for pathological slices, including:
[0100] A dataset construction module, based on the focus height data of the pathological slice scan image, fits a reference focus plane through cubic spline interpolation, determines the fitting error threshold in combination with the objective lens depth of field parameter and the device movement error, and divides the focus density type labels according to the fitting result;
[0101] A multi-label recognition model construction module, using a neural network architecture including a tandem residual convolution module, performs multi-scale feature extraction on the input pathological slice scan image;
[0102] Generate shallow features and deep features through upsampling and feature splicing. The shallow features are used to identify the focus density type, and the deep features are combined with the processing results of the shallow features to identify the slice type; the multi-label recognition model outputs a slice type classification probability vector, a focus density type classification probability vector, a slice type segmentation probability map, and a focus density type segmentation probability map;
[0103] A dynamic threshold post-processing module dynamically adjusts the segmentation threshold according to the confidence of the slice type classification probability vector, processes the slice type segmentation probability map and the focus density type segmentation probability map, and determines the final slice type and focus density type in combination with the processed segmentation results.
[0104] Embodiment 3
[0105] This embodiment also provides an electronic device, refer to Figure 7, including a memory 404 and a processor 402, wherein a computer program is stored in the memory 404, and the processor 402 is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0106] Specifically, the above-mentioned processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0107] Among them, the memory 404 may include a mass storage 404 for data or instructions. By way of example and not limitation, the memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 404 may include removable or non-removable (or fixed) media. Where appropriate, the memory 404 may be internal or external to the data processing device. In a particular embodiment, the memory 404 is non-volatile memory. In a particular embodiment, the memory 404 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory (FLASH), or a combination of two or more of these. Where appropriate, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0108] The memory 404 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402.
[0109] The processor 402 reads and executes the computer program instructions stored in the memory 404 to implement any one of the pathological section multi-label recognition methods in the above embodiments.
[0110] Optionally, the above electronic device may further include a transmission device 406 and an input / output device 408, wherein the transmission device 406 is connected to the above processor 402, and the input / output device 408 is connected to the above processor 402.
[0111] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above network may include wired or wireless networks provided by the communication provider of the electronic device. In one example, the transmission device includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 406 can be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0112] The input / output device 408 is used to input or output information.
[0113] Embodiment 4
[0114] This embodiment also provides a readable storage medium, in which a computer program is stored. The computer program includes program codes for controlling a process to execute the process, and the process includes the pathological section multi-label recognition method according to Embodiment 1.
[0115] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.
[0116] Generally, various embodiments can be implemented in hardware or special circuits, software, logic, or any combination thereof. Some aspects of the present invention can be implemented in hardware, while other aspects can be implemented by firmware or software executed by a controller, a microprocessor, or other computing devices, but the present invention is not limited thereto. Although various aspects of the present invention can be shown and described as block diagrams, flowcharts, or using some other graphical representations, it should be understood that, as a non-limiting example, the blocks, devices, systems, technologies, or methods described herein can be implemented in hardware, software, firmware, special circuits or logic, general hardware or a controller, or other computing devices, or some combination thereof.
[0117] Embodiments of the present invention can be implemented by computer software, which can be executed by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. A computer software or program (also referred to as a program product), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product can include one or more computer-executable components configured to execute the embodiments when the program runs. One or more computer-executable components can be at least one software code or a part thereof. Additionally, at this point, it should be noted that any box in the logical flow, as Figure 1 shown in [reference], can represent a program step, or interconnected logic circuits, boxes, and functions, or a combination of program steps and logic circuits, boxes, and functions. The software can be stored on physical media such as memory chips or storage blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs. The physical media are non-transitory media.
[0118] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.
[0119] The above embodiments only represent several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A multi-label recognition method for pathological sections, characterized in that, It includes the following steps: S1. Obtain the training data set: Based on the focus height data of the pathological slice scanning image, fit the reference focus plane through cubic spline interpolation, and determine the fitting error threshold by combining the objective lens depth of field parameter and the device motion error. Divide the focus density type labels according to the fitting result; S2. Construct a multi-label recognition model: Adopt a neural network architecture containing a series-connected residual convolution module to perform multi-scale feature extraction on the input pathological slice scanning image; Generate shallow features and deep features through upsampling and feature splicing. The shallow features are used to identify the focus density type, and the deep features are combined with the processing results of the shallow features to identify the slice type; the multi-label recognition model outputs a slice type classification probability vector, a focus density type classification probability vector, a slice type segmentation probability map, and a focus density type segmentation probability map; S3. Dynamic threshold post-processing: Dynamically adjust the segmentation threshold according to the confidence of the slice type classification probability vector, process the slice type segmentation probability map and the focus density type segmentation probability map, and determine the final slice type and focus density type in combination with the processed segmentation results; Among them, the dynamic threshold post-processing includes: When the slice type classification confidence is greater than or equal to the first threshold, use a higher segmentation threshold to process the segmentation probability maps of the slice type and the focus density type, and use the union of the segmentation masks of the two as the recognition result; When the slice type classification confidence is between the first threshold and the second threshold, use a lower segmentation threshold to process the segmentation probability map, and use the union of the segmentation masks of the two as the recognition result; When the slice type classification confidence is less than or equal to the second threshold, dynamically adjust the segmentation threshold of the focus density type according to the focus density type classification confidence, and use the segmentation result of the focus density type as the recognition result.
2. The method for multi-label recognition of pathological sections according to claim 1, wherein The specific steps of S1 include: Obtain the set of three-dimensional coordinates of the focus points of each tile in the pathological slice scanning image through a pathological slice scanner. The three-dimensional coordinates of the focus points include position coordinates and focus height; Uniformly sample the focus points at a preset interval, and use the cubic spline interpolation method to fit and generate a focus surface; Calculate the vertical distance between the unsampled focus points and the focus surface at the corresponding horizontal and vertical coordinates, determine the distance threshold according to the objective lens depth of field and the device motion error, and count the proportion of the focus points with a vertical distance less than the distance threshold. When the proportion is not less than the preset fitting error threshold, determine the current interval as the best focus interval of the pathological slice; Based on the best focus interval, divide the focus density into multiple categories and label the density labels.
3. The method for multi-label recognition of pathological sections according to claim 2, characterized in that In the specific steps of step S1, the method for classifying the focus density includes: Divide the best focus interval into 4 intervals of greater than 5 mm, 5 - 4 mm, 4 - 3 mm, and 3 - 2 mm at 1 mm intervals, corresponding to 4 focus density type labels respectively.
4. The method for multi-label recognition of pathological sections according to claim 1, wherein, In step S2, the residual convolution module includes a plurality of convolution modules. The convolution modules sequentially perform convolution, batch normalization, and ReLU activation processing on the input features; the residual convolution module splices the features after different convolution processes and outputs a fused feature.
5. The method for multi-label recognition of pathological sections according to claim 4, wherein The first threshold is set to 0.8, and the second threshold is set to 0.1; the higher segmentation threshold is 0.5, and the lower segmentation threshold is 0.
2.
6. A method for multi-label recognition of pathological sections according to any one of claims 1 to 5, characterized in that, In step S2, during the training process of the multi-label recognition model, the detection loss function comprehensively considers the differences in the position, size, and overlapping area between the predicted bounding box and the ground truth bounding box. The classification loss function uses the binary cross-entropy loss function, and the total loss function adjusts the weights of the detection loss and the classification loss through iteration to balance the training effect.
7. A multi-label recognition device for pathological sections, characterized in that, It includes: A dataset construction module, based on the focus height data of the pathological slice scan image, fits the reference focal plane through cubic spline interpolation, combines the objective lens depth of field parameter and the device motion error to determine the fitting error threshold, and divides the focus density type label according to the fitting result; A multi-label recognition model construction module, using a neural network architecture including a tandem residual convolution module, performs multi-scale feature extraction on the input pathological slice scan image; Generates shallow features and deep features through upsampling and feature splicing. The shallow features are used to identify the focus density type, and the deep features combine the processing results of the shallow features to identify the slice type; the multi-label recognition model outputs a slice type classification probability vector, a focus density type classification probability vector, a slice type segmentation probability map, and a focus density type segmentation probability map; A dynamic threshold post-processing module, dynamically adjusts the segmentation threshold according to the confidence of the slice type classification probability vector, processes the slice type segmentation probability map and the focus density type segmentation probability map, and determines the final slice type and focus density type in combination with the processed segmentation results; Among them, the dynamic threshold post-processing includes: When the slice type classification confidence is greater than or equal to the first threshold, the higher segmentation threshold is used to process the segmentation probability maps of the slice type and the focus density type, and the union of the segmentation masks of the two is used as the recognition result; When the slice type classification confidence is between the first threshold and the second threshold, the lower segmentation threshold is used to process the segmentation probability map, and the union of the segmentation masks of the two is used as the recognition result; When the slice type classification confidence is less than or equal to the second threshold, the segmentation threshold of the focus density type is dynamically adjusted according to the focus density type classification confidence, and the segmentation result of the focus density type is used as the recognition result.
8. An electronic device, comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to execute the pathological slice multi-label recognition method according to any one of claims 1 to 6.
9. A readable storage medium, characterized in that, The readable storage medium stores a computer program, and the computer program includes program codes for controlling a process to execute the process, and the process includes the pathological slice multi-label recognition method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Self-adaptive focusing method and device and readable storage medium
CN118134920A
Multistream fusion encoder for prostate lesion segmentation and classification
US20230162353A1