Pathological section multi-label identification method and device and readable storage medium thereof

By using a multi-label recognition model in a pathological section scanner, dynamically identifying the slice type and focus density, the problem of difficulty in taking into account scanning efficiency and image clarity in the prior art is solved, and efficient and low-cost pathological section scanning is achieved.

CN120164048AActive Publication Date: 2025-06-17SHENZHEN SHENGQIANG TECH

Patent Information

Application Number
CN202510638057.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-17
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

When existing pathological slice scanners identify slice types and focus density settings, it is difficult to achieve an optimal balance between scanning efficiency and image sharpness, and rely on high-cost hardware or fixed settings and cannot be dynamically optimized.

Method used

A multi-label recognition model based on cubic spline interpolation verification is adopted. By constructing a multi-label recognition model combining pathological slice type and focus density type, the three-dimensional coordinate data of focus recorded by the scanning device is used to dynamically calculate the lowest effective focus density, and the scanning efficiency optimization without hardware upgrade is achieved.

Benefits of technology

It significantly reduces the number of focus points, shortens the scanning time, and reduces hardware costs, improves the recognition accuracy and segmentation accuracy of pathological slices, and enhances the robustness and versatility of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164048A_ABST
    Figure CN120164048A_ABST
Patent Text Reader

Abstract

The invention provides a pathological section multi-label identification method and device and a readable storage medium thereof. The pathological section multi-label identification method specifically comprises the steps that focusing point three-dimensional coordinate data recorded by scanning equipment is utilized, a focusing curved surface is fitted through cubic spline interpolation, a fitting error threshold value is determined based on the depth of field of an objective lens and a motion error, and density labels are divided at the optimal focus interval; designing a neural network containing a residual convolution module, and fusing multi-scale features to realize double-task joint learning; and through a dynamic threshold post-processing algorithm, the segmentation threshold is adaptively adjusted according to the classification confidence, and the recognition robustness is improved. The method does not need special hardware, can be deployed in a general CPU / GPU, remarkably reduces the number of focusing points and scanning time, improves classification precision and segmentation precision, effectively balances scanning efficiency and image quality, and is suitable for various pathological section types such as TCT, HE and IHC.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing and medical device control, and particularly to a multi-label recognition method, device and readable storage medium for pathological sections. Background Art

[0002] Pathological section scanning is a key link in pathological diagnosis. When converting a pathological section into a digital image by a scanner, the setting of the focus density directly affects the scanning time and image clarity. Existing scanners usually adopt a fixed high-density focus point or real-time focusing technology: 1. Although the fixed high-density focus point can ensure clear images, the scanning time is long and the efficiency is low; 2. The real-time focusing technology dynamically adjusts the focusing height through a small number of focus points. Although the time is reduced, it relies on high-precision hardware, significantly increasing the equipment cost.

[0003] In addition, existing pathological tissue recognition models can only perform single-label classification (such as staining method classification), and cannot simultaneously handle the problem of focus density setting closely related to tissue morphology - the surface undulation differences caused by different specimen preparation methods (such as TCT, biopsy puncture) and sample shapes (regular tissue, adipose tissue) require matching different optimal focus densities, while traditional methods lack the ability to dynamically optimize this key parameter.

[0004] Therefore, how to synchronously identify the type of pathological sections and dynamically determine the lowest effective focus density through an algorithm without increasing the hardware cost has become a technical problem urgently to be solved in this field. Summary of the Invention

[0005] Embodiments of the present invention provide a multi-label recognition method, device and readable storage medium for pathological sections, aiming at the problems existing in the current technology that it is impossible to simultaneously identify the type of pathological sections (such as staining method, specimen preparation method) and surface undulation characteristics (sample shape), resulting in that the focus density can only be fixed or adjusted relying on high-cost hardware, and it is difficult to achieve the optimal balance between scanning efficiency and image clarity.

[0006] The core technology of the present invention is mainly a multi-label recognition model based on cubic spline interpolation verification. By constructing a multi-label recognition model combining the type of pathological sections (6 types) and the type of focus density (4 types), training the model with the three-dimensional coordinate data of the focus points recorded by the scanning device, and dynamically calculating the lowest effective focus density based on the surface undulation characteristics of the tissue, the scanning efficiency optimization without hardware upgrade is realized.

[0007] In the first aspect, the present invention provides a multi-label recognition method for pathological sections, and the method includes the following steps: S1. Obtain a training data set: Based on the focus height data of the pathological section scanning image, a reference focus plane is fitted by cubic spline interpolation, and the fitting error threshold is determined by combining the objective lens depth of field parameter and the device motion error. The focus density type label is divided according to the fitting result; S2. Construct a multi-label recognition model: Adopt a neural network architecture containing a series of residual convolution modules to perform multi-scale feature extraction on the input pathological section scanning image; Generate shallow features and deep features through upsampling and feature splicing. The shallow features are used to identify the focus density type, and the deep features are combined with the processing results of the shallow features to identify the section type; the multi-label recognition model outputs the section type classification probability vector, the focus density type classification probability vector, the section type segmentation probability map, and the focus density type segmentation probability map; S3. Dynamic threshold post-processing: Dynamically adjust the segmentation threshold according to the confidence of the section type classification probability vector, process the section type segmentation probability map and the focus density type segmentation probability map, and determine the final section type and focus density type in combination with the processed segmentation results.

[0008] Furthermore, the specific steps of S1 include: Obtain the set of three-dimensional coordinates of the focus points of each tile in the pathological section scanning image through a pathological section scanner. The three-dimensional coordinates of the focus points include position coordinates and focus height; Uniformly sample the focus points at a preset interval, and use the cubic spline interpolation method to fit and generate a focus surface; Calculate the vertical distance between the unsampled focus points and the focus surface at the corresponding horizontal and vertical coordinates, determine the distance threshold according to the objective lens depth of field and the device motion error, and count the proportion of the focus points with a vertical distance less than the distance threshold. When the proportion is not less than the preset fitting error threshold, determine the current interval as the best focus interval for this pathological section; Based on the best focus interval, divide the focus density into multiple categories and label the density labels.

[0009] Furthermore, in the specific steps of S1, the method for classifying the focus density includes: Divide the best focus interval into 4 intervals of more than 5 mm, 5 - 4 mm, 4 - 3 mm, and 3 - 2 mm at 1 mm intervals, corresponding to 4 focus density type labels respectively.

[0010] Furthermore, in S2, the residual convolution module includes multiple convolution modules. The convolution modules sequentially perform convolution, batch normalization, and ReLU activation processing on the input features; the residual convolution module splices the features after different convolution processes and outputs the fused features.

[0011] Furthermore, in S3, the dynamic threshold post-processing includes: When the classification confidence of the slice type is greater than or equal to the first threshold, a higher segmentation threshold is used to process the segmentation probability maps of the slice type and the focus density type, and the union of the segmentation masks of the two is used as the recognition result; When the classification confidence of the slice type is between the first threshold and the second threshold, a lower segmentation threshold is used to process the segmentation probability map, and the union of the segmentation masks of the two is used as the recognition result; When the classification confidence of the slice type is less than or equal to the second threshold, the segmentation threshold of the focus density type is dynamically adjusted according to the classification confidence of the focus density type, and the segmentation result of the focus density type is used as the recognition result.

[0012] Furthermore, the first threshold is set to 0.8, the second threshold is set to 0.1; the higher segmentation threshold is 0.5, and the lower segmentation threshold is 0.2.

[0013] Furthermore, in step S2, during the training process of the multi-label recognition model, the detection loss function comprehensively considers the differences in the position, size, and overlapping area between the predicted box and the ground truth box. The classification loss function uses the binary cross-entropy loss function, and the total loss function adjusts the weights of the detection loss and the classification loss in a loop to balance the training effect.

[0014] In a second aspect, the present invention provides a multi-label recognition device for pathological slices, including: A dataset construction module, based on the focus height data of the pathological slice scan image, fits a reference focus plane through cubic spline interpolation, combines the objective lens depth of field parameter and the device motion error to determine the fitting error threshold, and divides the focus density type labels according to the fitting result; A multi-label recognition model construction module, which uses a neural network architecture including a series of residual convolution modules to perform multi-scale feature extraction on the input pathological slice scan image; Generate shallow features and deep features through upsampling and feature splicing. The shallow features are used to identify the focus density type, and the deep features combine the processing results of the shallow features to identify the slice type; the multi-label recognition model outputs a slice type classification probability vector, a focus density type classification probability vector, a slice type segmentation probability map, and a focus density type segmentation probability map; A dynamic threshold post-processing module, which dynamically adjusts the segmentation threshold according to the confidence of the slice type classification probability vector, processes the slice type segmentation probability map and the focus density type segmentation probability map, and determines the final slice type and focus density type in combination with the processed segmentation results.

[0015] In a third aspect, the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the above-mentioned multi-label recognition method for pathological slices.

[0016] In a fourth aspect, the present invention provides a readable storage medium storing a computer program, the computer program including program code for controlling a process to execute the process, the process including the pathological section multi-label recognition method as described above.

[0017] The main contributions and innovations of the present invention are as follows: 1. Efficiency improvement: By dynamically setting the focus density, the number of foci is significantly reduced while ensuring image clarity, shortening the scanning time (for example, for TCT samples with less surface undulation, a larger focus interval can be used, and the scanning time is reduced by 30%-50%).

[0018] 2. Cost advantage: The model can be deployed on general-purpose hardware such as CPU / GPU, without the need for dedicated high-precision focusing equipment, reducing the hardware cost of the scanner.

[0019] 3. Breakthrough in recognition accuracy: Compared with existing single-label models (such as MaskRCNN, Unet), the multi-label segmentation model of the present invention improves the classification accuracy to 99.1% and the segmentation accuracy (MPA50) to 98.3% (see Table 1), and the recognition effect on complex morphologies such as adipose tissue and puncture samples is significantly enhanced.

[0020] 4. Enhanced robustness: The dynamic threshold post-processing algorithm can adaptively process new sample types outside the training data, and dynamically adjusts the segmentation threshold by combining the confidence of the slice type and the focus density, avoiding missed detection or misjudgment.

[0021] 5. General-purpose design: Covers various slice types such as TCT, HE, and IHC, is compatible with different specimen preparation methods and sample shapes, and forms a standardized optimization scheme for the focusing strategy.

[0022] Details of one or more embodiments of the present invention are set forth in the following drawings and description, so that other features, objects, and advantages of the present invention will become more clearly understood. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The illustrative embodiments and descriptions thereof of the present invention are used to explain the present invention, and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is a flowchart of the pathological section multi-label recognition method according to an embodiment of the present invention; Figure 2 is a distribution diagram of the optimal focus interval according to an embodiment of the present invention; Figure 3 is a model framework diagram according to an embodiment of the present invention; Figure 4It is a diagram of a convolution module according to an embodiment of the present invention; Figure 5 It is a diagram of a residual convolution module according to an embodiment of the present invention; Figure 6 It is a flowchart of post - processing with a dynamic threshold according to an embodiment of the present invention; Figure 7 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0024] Here, exemplary embodiments will be described in detail, and their examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0025] It should be noted that: In other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0026] The prior art cannot dynamically determine the optimal focus density based on the morphological features of pathological sections, resulting in difficulty in balancing scanning efficiency and image quality.

[0027] Based on this, the present invention is based on a multi - label recognition model verified by cubic spline interpolation, and solves the problems existing in the prior art by synchronously identifying the sample type based on the shape features of pathological sections and predicting the optimal focus density category (0 - 3 classes).

[0028] Embodiment 1 The present invention aims to propose a multi - label recognition method for pathological sections. Specifically, referring to Figure 1 , the method includes: S1. Dataset construction (obtaining a training dataset): Based on the focus height data of pathological section scanning images, a reference focus plane is fitted by cubic spline interpolation, and a fitting error threshold is determined by combining the objective lens depth - of - field parameter and the device motion error. The focus density type labels are divided according to the fitting result; In this embodiment, the essence of setting the focus density is to set the sample sampling rate, and the optimal sampling rate is related to the data situation. The surface undulation of pathological tissues and the section types are related to the tissue shape. For example, the TCT preparation method makes the surface of such samples present as a plane; the tissues obtained by biopsy puncture are strip-shaped with large surface undulations; the tissues with more fat or cavities usually have large surface undulations, while the surface of complete and regular tissues is closer to a plane. A pathological slice scanner usually has the function of saving the focus height and position of each tile, and the focus height of a clear tile can be considered accurate. Therefore, the position and height data of the tiles can be obtained from the existing scanned pathological slice images, and then a training data set can be formed. Each tile in each pathological image can be regarded as a focus, and the three-dimensional coordinates of the focus are recorded as a set . To ensure the generalization of the algorithm, data needs to be extracted from different types of slice images.

[0029] For quantitative analysis of the focus density setting, uniform sampling is performed at a certain distance interval G, and the selected foci are used for cubic spline interpolation fitting to obtain a fitting surface, where G takes values from 200 at intervals of 200 to 5000 . Then, calculate the vertical distance E between the unselected foci and the corresponding horizontal and vertical coordinates of the fitting surface. The foci with smaller differences are regarded as correctly fitted, otherwise the fitting height is considered incorrect. The threshold T for whether the fitting is correct is determined by the depth of field D of the objective lens and the device movement error V. The imaging within the depth of field is clear, but there are errors in the platform movement and vibrations. Therefore, the final threshold is defined as T = D / 2 - V. (In the present invention, D = 1 , V = 0.2 ). Calculate the proportion of foci with a gap E less than T in all foci. When the proportion is less than the fitting error threshold α, the scan is considered clear, and its value range is recommended to be 0% - 10%. The value of α is related to the clarity requirement and the preparation quality.

[0030] Specifically, if a more accurate fitting virtual surface is required, and thus the scanned image is clearer, a smaller value of α is set. However, the smaller the α is set, the more foci will be generated, resulting in an increase in time consumption. In particular, because there is a situation where the preparation quality is low, resulting in large undulations in the section tissues, it is difficult for the fitting virtual surface to fully conform to the actual focus surface. Therefore, α is not recommended to be set too small. In the present invention, α is set to 5%, and the virtual surface fitting method uses the cubic spline surface interpolation method.

[0031] Each slice will be judged as clear or blurred under different interval conditions, and the distance when it is judged as clear and the interval is the largest is used as the best focus interval distance for the slice. Then, based on the best focus interval distance of the slice, it is used as the basis for density label division.Figure 2 is the distribution probability of the optimal focus interval distance, where the minimum distance is close to 2 mm and the maximum distance exceeds 5 mm. Therefore, the density types are divided in units of 1 mm starting from 2 mm. Since the condition of clear scanning is satisfied when the interval exceeds 5 mm, and it is also satisfied when set to 5 mm, intervals exceeding 5 mm are regarded as one category. Finally, the focus interval distances are set to greater than 5 mm, 5 mm - 4 mm, 4 mm - 3 mm, 3 mm - 2 mm, and the corresponding labels are set to 0 - 3 respectively. In particular, the present invention divides the categories at intervals of 1 mm, and can also be divided at smaller intervals. If divided at smaller intervals (such as 0.2 mm), since the intervals become smaller, the categories of adjacent intervals will be relatively similar, such as 2.2 mm - 2.4 mm and 2.4 mm - 2.6 mm. This will cause the model to be difficult to distinguish adjacent categories. Eventually, it will be misclassified into adjacent categories, such as the category of 2.2 mm - 2.4 mm being recognized as the category of 2.4 mm - 2.6 mm or 2.6 mm - 2.8 mm.

[0032] S2. Construct a multi-label recognition model: Adopt a neural network architecture including a tandem residual convolution module to perform multi-scale feature extraction on the input pathological section scan image; Generate shallow features and deep features through upsampling and feature splicing. The shallow features are used to identify the focus density type, and the deep features are combined with the processing results of the shallow features to identify the section type; the multi-label recognition model outputs a section type classification probability vector, a focus density type classification probability vector, a section type segmentation probability map, and a focus density type segmentation probability map.

[0033] In this embodiment, existing scanner algorithms all use neural networks for single-label multi-category segmentation tasks and can only identify the type of the section. In order to simultaneously identify the section type and the focus setting type, the present invention designs a multi-label recognition model that can identify 6 section types (TCT, HE, IHC, PAS, oil red, and other special stains) and 4 focus density types.

[0034] The architecture of the multi-label recognition model is as Figure 3 shown. In the figure, [w, h, c] is used to represent the feature dimension, w represents the width, h represents the height, and c represents the number of channels. This multi-label recognition model uses five tandem residual convolution modules to extract features of different scales, and then through upsampling and splicing features of different scales, shallow features [80, 80, 768] and deep features [40, 40, 1024] are obtained respectively. The shallow features will pass through the residual convolution module to obtain density type features [80, 80, 256], and finally pass through the convolution module and convolution to obtain the detection and classification results.

[0035] The basis for density type division is mainly shape, which is also one of the bases for slice type division, and the segmentation results of the two types are the same. Therefore, the density type features can be used for the classification and recognition of slice types. The model further processes the density type features through a convolution module, concatenates the obtained features with deep features, then the concatenated features are used to extract features through a residual convolution module to obtain slice type features, and finally, detection and classification results are obtained through a convolution module and convolution.

[0036] As Figure 4 shown, the convolution module consists of convolution, batch normalization, and the activation function Relu. k, s, p, and c are several important hyperparameters. k represents the size of the convolution kernel. s is the stride, indicating the interval at which the convolution kernel slides on the input data. p represents the number of additional pixels added around the boundary of the input data. c_in represents the number of channels of the input data or the convolution kernel. Since the number of channels of the input data is the same as that of the convolution kernel, the size of the convolution kernel c_in will not be specifically described in the figure or subsequent text descriptions. At the same time, the dimension c_out of the output data is the number of convolution kernels, which will also not be specifically described. Batch normalization normalizes each feature dimension separately to make its mean 0 and variance 1.

[0037] Among them, bilinear interpolation is used for upsampling, which can enlarge the width and height of the feature map. For example, the feature [20, 20, 512] is upsampled by a factor of two to become [40, 40, 512]. Concatenation requires that the input features have the same width and height and can merge the input in the channel dimension. For example, the feature [40, 40, 512] and the feature [40, 40, 256] become [40, 40, 768] after concatenation.

[0038] As Figure 5 shown, the residual convolution module is composed of multiple convolution modules combined. This module concatenates the features processed by different numbers of convolution modules together, and then inputs the concatenated features into the last convolution module to obtain the output.

[0039] In this embodiment, the training loss function of the multi-label recognition model includes detection loss and classification loss. The model first identifies the region where the sample is located and outputs a prediction box, and then classifies the pixels within each prediction box. The position and size of the prediction box and the true box (the minimum bounding rectangle of the sample) are the criteria for measuring the detection effect. The detection loss comprehensively considers the differences between the prediction box and the true box to maximize the overlapping area between the two and minimize the differences in the center distance and aspect ratio. The specific formula is as follows:

[0040] Among them, IoU represents the intersection over union of the prediction box and the true box, d() calculates the Euclidean distance between two points, p represents the center point of the prediction box, pgt represents the center point of the ground truth box, c gt represents the diagonal length of the ground truth box, a is a hyperparameter, default set to 0.2, w and h represent the width and height of the predicted box, and wgt and hgt represent the width and height of the ground truth box.

[0041] Preferably, the multi-label recognition model uses the binary cross-entropy loss function as the classification loss, that is, calculates the cross-entropy loss for each pixel within the predicted box separately, and then finds the mean of all cross-entropy loss values. Its mathematical formula is as follows:

[0042] Among them, N represents the number of samples, y i represents the ground truth label, p i represents the predicted probability value.

[0043] Since the two types of loss functions use different features and labels, it is necessary to balance the loss functions of both, otherwise there will be a situation where one type has a good effect while the other is poor. Therefore, the present invention uses a loss function hyperparameter that changes cyclically, enabling the model to focus on both types during training. The total loss function L is defined as follows:

[0044] Among them, and represent the density type detection loss and classification loss, and represent the slice type detection loss and classification loss, iter is the total number of training iterations, which can be adjusted, default set to 200, k is the k-th iteration of training, that is, after training all data once, k is incremented by one, so the range of k is [1, iter].

[0045] S3, Dynamic Threshold Post-Processing: Dynamically adjust the segmentation threshold according to the confidence of the slice type classification probability vector, process the slice type segmentation probability map and the focus density type segmentation probability map, and determine the final slice type and focus density type in combination with the processed segmentation results.

[0046] In this embodiment, for the input image, the multi-label recognition model will output the classification probability vector and segmentation probability map of the slice type and density type. The maximum value in the classification probability vector is the confidence, and the position of the vector where the maximum value is located is the classification result. Define the slice type classification confidence as x1 and the density type classification confidence as x2. In order to obtain an accurate recognition result, the present invention proposes a dynamic threshold post-processing algorithm, which combines the classification and segmentation results of both types to calculate the final recognition result, and the specific method is as Figure 6 shown.

[0047] The segmentation probability maps output by the model will have overlaps. The non-maximum suppression algorithm is applied to the slice type and density type respectively, with a confidence threshold of 0.1 and an intersection over union threshold of 0.75. After processing, most of the overlapping probability maps will be excluded, and each slice type segmentation probability map will correspond to a density type segmentation probability map. The two probability maps together represent the recognized target. Then, the threshold for post-processing is dynamically adjusted according to the classification confidence of the slice type.

[0048] When the classification confidence x1 of the slice type is greater than a1 = 0.8, it indicates that the model can accurately identify the category of the pathological slice sample, and the corresponding segmentation result has high accuracy. Therefore, a higher threshold t1 = 0.5 can be used to process the segmentation probability map of the slice type. Since the segmentation result of the slice type is relatively accurate and the segmentation result of the sample shape is only used for fine supplementation, a higher threshold is set to process the segmentation probability map of the density type, with a2 = 0.8 and t2 = 0.5. After processing the segmentation probability maps, segmentation masks will be obtained, and the union of the two mask images is used as the final recognition result.

[0049] When x1 is between a1 and b1, the accuracy of the recognized slice type decreases, so the result of the density type needs to be combined. Therefore, a lower threshold is set to process the segmentation probability maps. Set t1 = 0.1, a2 = 0.4, and t2 = 0.2. The union of the two mask images is used as the recognition result.

[0050] When x1 is less than b1, it means that the sample of the slice type is not recognized. In addition, in actual use, there may be pathological tissue types not included in the training data. At this time, the confidence value of the slice type obtained by model inference will be relatively low. Therefore, when the input image is a new pathological tissue slice type, the sample may not be recognized. To recognize the slice type not included in the training data, the segmentation result of the density type can be used as the final result. Because the higher the confidence x2 of the density type, the more accurate the segmentation result, the threshold of the segmentation probability map can be set lower. Therefore, set a2 = 0.1 and t1 = 0.6 - x2 / 2.

[0051] Since most methods do not support multi-label segmentation, only the classification accuracy of the slice type and the overall segmentation accuracy are compared. Table 1 shows the comparison results between the present invention and existing methods. In terms of the classification and segmentation accuracy of pathological slices, the accuracy of the present invention is higher than that of all comparison methods. The experimental results prove that the multi-label segmentation model and dynamic threshold processing designed by the present invention can effectively improve the recognition effect of pathological slices.

[0052] Table 1

[0053] The following is a supplementary explanation of the technical terms involved in the present invention to help understand the core concepts of the technical solution: 1. Cubic spline interpolation Explanation: A mathematical interpolation method that fits discrete data points through piecewise cubic polynomial functions to generate a smooth surface (such as the focus plane). In the present invention, it is used to fit the focus height data, calculate the vertical distance between the unsampled points and the fitted surface, and evaluate the accuracy of the focus plane.

[0054] 2. Non-maximum suppression algorithm (NMS) Explanation: A post-processing algorithm used to remove overlapping prediction boxes in object detection or segmentation and retain the result with the highest confidence. In the present invention, NMS is applied to the segmentation probability maps of slice types and focus density types to reduce redundant predictions and improve the accuracy of the recognition results.

[0055] Embodiment 2 Based on the same concept, the present invention also proposes a multi-label recognition device for pathological slices, including: A dataset construction module that fits a reference focus plane through cubic spline interpolation based on the focus height data of the pathological slice scan image, determines the fitting error threshold in combination with the objective lens depth of field parameter and the device motion error, and divides the focus density type labels according to the fitting result; A multi-label recognition model construction module that uses a neural network architecture including a series of residual convolution modules to perform multi-scale feature extraction on the input pathological slice scan image; Generate shallow features and deep features through upsampling and feature splicing. The shallow features are used to identify the focus density type, and the deep features, combined with the processing results of the shallow features, are used to identify the slice type; the multi-label recognition model outputs a slice type classification probability vector, a focus density type classification probability vector, a slice type segmentation probability map, and a focus density type segmentation probability map; A dynamic threshold post-processing module that dynamically adjusts the segmentation threshold according to the confidence of the slice type classification probability vector, processes the slice type segmentation probability map and the focus density type segmentation probability map, and determines the final slice type and focus density type in combination with the processed segmentation results.

[0056] Embodiment 3 This embodiment also provides an electronic device, refer to Figure 7 , including a memory 404 and a processor 402. The memory 404 stores a computer program, and the processor 402 is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0057] Specifically, the above-mentioned processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present invention.

[0058] Among them, the memory 404 may include a mass memory 404 for data or instructions. By way of example and not limitation, the memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 404 may include removable or non-removable (or fixed) media. Where appropriate, the memory 404 may be internal or external to the data processing device. In a particular embodiment, the memory 404 is non-volatile memory. In a particular embodiment, the memory 404 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. Where appropriate, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.

[0059] The memory 404 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402.

[0060] The processor 402 reads and executes the computer program instructions stored in the memory 404 to implement any one of the pathological section multi-label recognition methods in the above embodiments.

[0061] Optionally, the above electronic device may further include a transmission device 406 and an input / output device 408. Among them, the transmission device 406 is connected to the above processor 402, and the input / output device 408 is connected to the above processor 402.

[0062] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above network may include wired or wireless networks provided by the communication provider of the electronic device. In one example, the transmission device includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the transmission device 406 can be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0063] The input / output device 408 is used to input or output information.

[0064] Embodiment 4 This embodiment also provides a readable storage medium, in which a computer program is stored. The computer program includes program codes for controlling a process to execute the process, and the process includes the pathological section multi-label recognition method according to Embodiment 1.

[0065] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be repeated here.

[0066] Generally, various embodiments can be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects of the present invention can be implemented in hardware, while other aspects can be implemented by firmware or software executed by a controller, a microprocessor, or other computing devices, but the present invention is not limited thereto. Although various aspects of the present invention can be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, as a non-limiting example, the blocks, devices, systems, technologies, or methods described herein can be implemented in hardware, software, firmware, dedicated circuits or logic, general hardware or a controller, or other computing devices, or some combination thereof.

[0067] Embodiments of the present invention can be implemented by computer software, which can be executed by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. A computer software or program (also called a program product), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product can include one or more computer-executable components configured to execute the embodiments when the program runs. One or more computer-executable components can be at least one software code or a part thereof. Additionally, in this regard, it should be noted that any box in the logical flow, as shown in Figure 1 can represent a program step, or interconnected logic circuits, boxes, and functions, or a combination of program steps and logic circuits, boxes, and functions. The software can be stored on physical media such as memory chips or storage blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs. The physical media is a non-transitory medium.

[0068] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0069] The above embodiments only represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.

Claims

1. A multi-label recognition method for pathological sections, characterized in that: The following steps are involved: S1. Get the training data set: Based on the focus height data of the pathological slice scanning image, the reference focus plane is fitted by cubic spline interpolation, and the fitting error threshold is determined by combining the objective lens depth of field parameter and the equipment motion error. The focus density type label is divided according to the fitting result; S2. Build a multi-label recognition model: A neural network architecture including a series of residual convolution modules is used to extract multi-scale features from the input pathological slice scan images; Shallow features and deep features are generated by upsampling and feature concatenation, wherein the shallow features are used to identify the focal density type, and the deep features are combined with the shallow feature processing results to identify the slice type; the multi-label recognition model outputs a slice type classification probability vector, a focal density type classification probability vector, a slice type segmentation probability map, and a focal density type segmentation probability map; S3, dynamic threshold post-processing: The segmentation threshold is dynamically adjusted according to the confidence of the slice type classification probability vector, the slice type segmentation probability map and the focal density type segmentation probability map are processed, and the final slice type and focal density type are determined in combination with the processed segmentation results.

2. A pathological section multi-label recognition method as claimed in claim 1, characterized in that: The specific steps of S1 include: Acquire a set of three-dimensional coordinates of the focus points of each image block in the pathology slice scan image by a pathology slice scanner, wherein the three-dimensional coordinates of the focus points include position coordinates and focus height; The focus points are uniformly sampled at preset intervals, and a focus surface is generated by fitting using a cubic spline interpolation method; Calculate the vertical distance between the unsampled focus point and the focus surface at the corresponding horizontal and vertical coordinates, determine the distance threshold according to the depth of field of the objective lens and the device motion error, count the proportion of focus points whose vertical distance is less than the distance threshold, and when the proportion is not less than a preset fitting error threshold, determine the current interval as the optimal focus interval of the pathological slice; Based on the optimal focus interval, the focus density is divided into multiple categories and density labels are marked.

3. A pathological section multi-label recognition method as claimed in claim 2, characterized in that: In the specific steps of step S1, the method for classifying the focus density includes: The optimal focus interval is divided into four intervals of greater than 5 mm, 5-4 mm, 4-3 mm, and 3-2 mm at 1 mm intervals, corresponding to four focus density type labels respectively.

4. A method for multi-label recognition of pathological sections as claimed in claim 1, characterized in that: In step S2, the residual convolution module includes multiple convolution modules, and the convolution modules sequentially perform convolution, batch normalization and ReLU activation processing on the input features; the residual convolution module splices the features after different convolution processing and outputs fused features.

5. A method for multi-label recognition of pathological sections as claimed in claim 1, characterized in that: In step S3, the dynamic threshold post-processing includes: When the slice type classification confidence is greater than or equal to the first threshold, a higher segmentation threshold is used to process the segmentation probability map of the slice type and the focal density type, and the union of the segmentation masks of the two is taken as the recognition result; When the slice type classification confidence is between the first threshold and the second threshold, the lower segmentation threshold is used to process the segmentation probability map, and the union of the segmentation masks of the two is taken as the recognition result; When the slice type classification confidence is less than or equal to the second threshold, the focus density type segmentation threshold is dynamically adjusted according to the focus density type classification confidence, and the focus density type segmentation result is used as the recognition result.

6. A method for multi-label recognition of pathological sections as claimed in claim 5, characterized in that: The first threshold is set to 0.8, the second threshold is set to 0.1; the upper segmentation threshold is 0.5, and the lower segmentation threshold is 0.

2.

7. A method for multi-label recognition of pathological sections according to any one of claims 1 to 6, characterized in that: In step S2, during the training of the multi-label recognition model, the detection loss function comprehensively considers the differences in position, size, and overlapping area between the predicted box and the true box, the classification loss function adopts a binary cross entropy loss function, and the total loss function balances the training effect by cyclically adjusting the weights of the detection loss and the classification loss.

8. A pathological section multi-label recognition device, characterized in that: include: The dataset construction module is based on the focus height data of the pathological slice scan image, fits the reference focus plane through cubic spline interpolation, and determines the fitting error threshold by combining the objective lens depth of field parameter and the device motion error, and divides the focus density type label according to the fitting result; The multi-label recognition model building module uses a neural network architecture containing a series of residual convolution modules to extract multi-scale features from the input pathological slice scan images; Shallow features and deep features are generated by upsampling and feature concatenation, wherein the shallow features are used to identify the focal density type, and the deep features are combined with the shallow feature processing results to identify the slice type; the multi-label recognition model outputs a slice type classification probability vector, a focal density type classification probability vector, a slice type segmentation probability map, and a focal density type segmentation probability map; The dynamic threshold post-processing module dynamically adjusts the segmentation threshold according to the confidence of the slice type classification probability vector, processes the slice type segmentation probability map and the focal density type segmentation probability map, and determines the final slice type and focal density type based on the processed segmentation results.

9. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor is configured to run the computer program to execute the pathological section multi-label identification method according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, wherein the computer program includes a program code for controlling a process to execute a process, wherein the process includes the pathological section multi-label identification method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Self-adaptive focusing method and device and readable storage medium

    CN118134920A

  • Multistream fusion encoder for prostate lesion segmentation and classification

    US20230162353A1

Cited By

  • Pathological section row-level discrete speed planning and scanning method and application thereof

    CN122120384A