Deep learning-based method and system for inspecting embedding box before batch dehydration
By combining a three-level cascade strategy with EGhost_YOLO and PP-OCRv4 models, the automated identification and verification of pathology specimen embedding cassettes was achieved, solving the problems of low identification accuracy and efficiency bottlenecks in existing technologies, and improving the automated management capabilities of the pathology department before dehydration.
Patent Information
- Application Number
- CN202511396206.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2025-10-31
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies for the intelligent management of pathological specimen embedding cassettes suffer from problems such as low recognition accuracy, efficiency bottlenecks, high computational burden, and slow reasoning speed, making it difficult to meet the rapid verification requirements of large batches of embedding cassettes, especially in the automated recognition and management of pre-dehydration stages.
A three-level cascaded strategy is adopted, combining the EGhost_YOLO object detection model and the PP-OCRv4 text recognition model. By rotating the bounding box to detect and locate the embedding box, cropping the numbering area, and comparing it with the preset list, the system automatically marks the missing, duplicate, and abnormal numbers, achieving accurate tracking throughout the entire process.
It improves the accuracy and efficiency of pathology sample management, reduces human error, ensures the reliability of diagnostic data, and meets the real-time verification needs of high-throughput scenarios in pathology departments.
Smart Images

Figure CN120877307A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology, and in particular relates to a method and system for pre-dehydration inspection of embedding cassettes based on deep learning. Background Technology
[0002] Currently, intelligent management of pathology specimen embedding cassettes mainly relies on label recognition technology. The closest existing technology (such as patent CN115019306A) proposes a batch identification method for embedding cassette labels based on deep learning and machine vision. This method automatically extracts the embedding cassette number and identification information through image acquisition, target detection, OCR recognition, and QR code parsing. The system is based on an improved YOLO network and integrates a spatial pyramid pooling module, combined with image enhancement strategies to improve model robustness, achieving automatic identification and information registration of embedding cassettes before dehydration, suitable for sample tracking and management in pathology procedures.
[0003] This method uses horizontal rectangular bounding boxes for target detection, which is ill-suited to situations where embedding cassettes are tilted, obscured, or stacked in the dehydration basket. This leads to the cropped area easily including information from adjacent cassettes, interfering with subsequent OCR text recognition and QR code parsing, significantly reducing recognition accuracy. Secondly, the system only performs format rule verification on the character number of a single label, lacking a global consistency comparison mechanism for the numbering. It cannot automatically detect duplicate numbers or missing cassettes, relying on manual secondary verification, thus failing to fundamentally address the efficiency bottleneck. Furthermore, the QR code detection process still relies on traditional image processing methods such as grayscale conversion, threshold segmentation, and morphological processing, resulting in a complex process and high computational burden. Especially in scenarios with a large number of embedding cassettes or degraded image quality, recognition performance becomes unstable. Finally, although the detection model structure has been optimized to some extent, the overall inference speed still cannot meet the timeliness requirements for rapid verification of large batches of embedding cassettes before dehydration, limiting the system's practical application in high-throughput clinical scenarios.
[0004] Therefore, this invention proposes to combine the EGhost_YOLO object detection model with the PP-OCRv4 text recognition model, employing a three-level cascaded strategy: First, the embedded cassettes in the overall image are rotated and located (coarse-grained); then, the embedded cassette area is cropped and its numbering area is precisely detected (fine-grained); finally, the numbering information is identified and compared with a preset list, automatically marking missing, duplicate, and abnormal numbers. This is applied to the pre-dehydration stage of the paraffin embedding process in pathology departments. By automatically identifying embedded cassette numbers and verifying omissions and duplicates, accurate tracking of pathology samples throughout the entire process is achieved, ensuring the reliability of diagnostic data. Summary of the Invention
[0005] The purpose of this invention is to provide a method and system for pre-dehydration inspection of embedded cassettes based on deep learning. It employs a three-level cascade strategy: first, rotating bounding boxes are used to detect and locate the embedded cassettes in the overall image (coarse-grained); then, the embedded cassette area is cropped and its numbering region is precisely detected (fine-grained); finally, the numbering information is identified and compared with a preset list, automatically marking missing, duplicate, and abnormal numbers. This method is applied to the pre-dehydration stage of the paraffin embedding process in pathology departments. By automatically identifying embedded cassette numbers and verifying for omissions and duplicates, it achieves precise tracking of pathological samples throughout the entire process, ensuring the reliability of diagnostic data.
[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: Firstly, a deep learning-based method for pre-dehydration testing of embedding boxes is provided, including the following steps: S1: Collect images of dehydrated blue embedding boxes, and construct an embedding box dataset after annotating the images; S2: Improve the YOLOv11s model by using the EGhostC3 module and depthwise separable convolution as the main modules to construct the Eghost-YOLO model. Train the Eghost-YOLO model using the aforementioned embedding box dataset. Train the two models to identify the embedding boxes in the image and their corresponding embedding box numbers, respectively. S3: Input the dehydrated blue image to be analyzed into the trained Eghost-YOLO embedding box detection model to detect and crop each embedding box; S4: Input the cropped embedding box image into the Eghost-YOLO number detection model to detect and crop the corresponding number region; S5: Based on the pre-trained PP-OCRv4 character recognition model, the detected embedding box number region is identified, and the identified number is compared with the real number to output the number verification status.
[0007] Preferably, the specific process of step S1 is as follows: S11: Use a high-speed document scanner to capture images of dehydrated blue pathological specimens from above, showing that the embedding cassettes are densely and regularly arranged in the images; S12: After acquiring the image, use the Labelme tool to annotate the outline of each embedding box in the image one by one. The annotation format is polygon to accurately represent the edge of the embedding box. S13: Calculate the minimum bounding rectangle of each labeled area, extract the center point coordinates, width and height, and tilt angle of the rectangle as the position and orientation information of each embedding box, which is used for subsequent area detection or verification positioning. S14: Use the LabelImg tool to individually label the number above or on the surface of each embedding box. The label type is horizontal box. Save the labeling results in YOLO format for training the number recognition model. To ensure the generalization ability of the data, divide all samples into training set, validation set and test set in a ratio of 7:2:1.
[0008] Preferably, the specific process of training the Eghost-YOLO model using the embedded box dataset in step S2 is as follows: S21: Before inputting the embedding box dataset into the Eghost-YOLO model, the image is filled with a gray background square to maintain the original image ratio and adapt to the model input size. Then it is uniformly adjusted to 1024×1024 for embedding box detection. S22: The detection result is located on the original image, and the corresponding embedding box area is then cropped from the original high-resolution image to obtain a clearer embedding box image for number detection and recognition. The number detection model uses the same filling method, and the input size is 256×256. S23: Using the preset total loss function Loss Calculate the total loss of the Eaglet-YOLO model, where the total loss function is derived from the bounding box regression loss function. L box Category classification loss function L cls and distribution-focused loss function L DFL The composition, specifically the formula, is as follows: ; in, , , The values represent the weights of each part of the loss, and are 0.5, 7.5, and 1.5 during the training process. ; in, C It is the number of categories; N It is the number of positive samples; Is the i-th sample in the i-th order? c The true label of the class; No. i The sample at the th c The predicted value of the class; This represents the Sigmoid activation function. ; ; ; ; in, B It is a prediction box; It is a real frame; w It is the predicted width; h It is the predicted bounding box height; It is the actual frame width; It is the actual frame height; It is the squared Euclidean distance between the predicted bounding box and the ground truth bounding box; It is the square of the diagonal length of the smallest closure rectangle containing the two boxes; This constitutes a center point distance penalty term; This constitutes a length-to-width ratio penalty term; ; in, and It is the softmax output predicted by the model, corresponding to the... i The and the first i +1 discrete point probability; y For true continuous numerical labels; and This represents two adjacent discretization points after the continuous value has been discretized.
[0009] Preferably, the specific process of step S3 is as follows: S31: Input the dehydrated blue image to be analyzed into the trained Eghost-YOLO embedding box detection model to obtain the detection results of the rotated rectangle positioning box of each embedding box; S32: Map the detection results back to the original high-resolution image proportionally, crop the embedding box region on the original high-resolution image according to the rotating frame, and correct its angle to obtain a standardized embedding box image, providing a more stable and accurate input basis for subsequent numbering detection and recognition.
[0010] Preferably, in step S4, the cropped embedding box image is input into the Eghost-YOLO number detection model, and the specific process of detecting and cropping the corresponding numbered regions is as follows: S41: Input the cropped embedding box image into the Eghost-YOLO number detection model, which accurately locates and crops the large number and multiple small number regions in the cropped embedding box image; S42: Resample the smaller numbered images cut out from the same embedding box proportionally to the height of the larger numbered images; S43: Connect the smaller number image and the larger number image using the "*" symbol to form a complete character image; S44: Input the stitched complete character image into the PP-OCRv4 character recognition model.
[0011] Preferably, the size of the stitched complete character image is padded before being input into the PP-OCRv4 text recognition model, and uniformly adjusted to 320×48 to adapt to the structure of the PP-OCRv4 text recognition model. The specific process of step S5 is as follows: S51: The PP-OCRv4 text recognition model extracts the large and small number information corresponding to each embedding box; S52: The large and small number information corresponding to each extracted embedding box is compared with the number list that should exist for this batch, and the number verification is completed automatically; S53: Outputs a visual annotation result containing anomaly number information and the original image.
[0012] The visualization annotation results in step S53 are as follows: The total number of detected embedding boxes is marked in the upper left corner of the image. Red dots are used to mark the locations of all identified embedding boxes. A large number of abnormal numbers is marked with a red box, a small number of redundant numbers is marked with a blue box, and a small number of duplicate numbers is marked with a green box.
[0013] Secondly, a deep learning-based pre-dehydration inspection system for embedding cassettes is provided to implement the aforementioned deep learning-based pre-dehydration inspection method for embedding cassettes, including... The data acquisition and construction module is used to acquire images of dehydrated blue embedding boxes, and to construct an embedding box dataset after annotating the images. The model training module is used to improve the YOLOv11s model. It uses the EGhostC3 module and depthwise separable convolution as the main modules to build the Eghost-YOLO model. The Eghost-YOLO model is trained using the embedded box dataset to obtain two models, one for recognizing embedded boxes in images and the other for recognizing the corresponding embedded box numbers. The embedding box detection and cropping module is used to receive the dehydrated blue image to be analyzed, input it into the trained Eagle-YOLO embedding box detection model, and detect and crop out each embedding box. The numbered region detection and cropping module is used to receive the cropped embedding box image, input it into the Eagle-YOLO number detection model, and detect and crop out the corresponding numbered regions. The number recognition and verification module is used to identify the detected embedding box number region based on the pre-trained PP-OCRv4 text recognition model, compare the identified number with the real number, and output the number verification status.
[0014] The beneficial effects of this invention include: This invention provides a deep learning-based method and system for pre-dehydration inspection of embedding cassettes. First, images of densely and regularly arranged embedding boxes are acquired using a high-speed scanner. The outlines and numbers of the embedding boxes are accurately labeled using the Labelme and LabelImg tools, respectively. The position and orientation information of the embedding boxes are also calculated and the dataset is divided reasonably. This ensures the accuracy and richness of the data and improves the generalization ability of the dataset, making the trained model more adaptable to actual detection scenarios.
[0015] Secondly, the YOLOv11s model was improved to obtain the Eghost-YOLO model, which adopts the EGhostC3 module and depthwise separable convolution, effectively reducing model complexity and computational cost while ensuring detection accuracy. Simultaneously, appropriate input sizes were set for different detection tasks, and a carefully designed total loss function was combined to further improve the model's localization and classification accuracy, providing reliable model support for subsequent detection pruning.
[0016] Furthermore, the trained model is used to detect the embedding boxes and numbered regions. The results are then mapped back to the original high-resolution image, cropped, and the angles corrected. The numbered regions are also processed and stitched together, standardizing and normalizing the image input to the character recognition model, reducing interference factors, and improving the accuracy of number recognition. The PP-OCRv4 character recognition model is used to extract numbering information, which is automatically compared with the required numbering list to complete the verification, significantly improving efficiency while avoiding subjective errors and fatigue-related mistakes caused by manual operation, ensuring more reliable verification results.
[0017] Finally, the visualization annotation results clearly show the total number of embedding boxes, the location and type of anomaly numbers, allowing staff to quickly locate problematic embedding boxes and take timely measures to improve the smoothness of the workflow and the efficiency of problem solving. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating the deep learning-based pre-dehydration inspection method for embedding cassettes according to the present invention.
[0019] Figure 2 This is a schematic diagram of the architecture of the Eaglehost-YOLO detection model of the present invention.
[0020] Figure 3 This is a schematic diagram of the architecture of the DWConv module and GhostC3 module of the present invention.
[0021] Figure 4 This is a schematic diagram of the architecture of the PP-OCRv4 text recognition model of the present invention. Detailed Implementation
[0022] The following is in conjunction with the appendix Figures 1-4 The present invention will be further described in detail below: Example 1 To address the problems of low efficiency, easy omissions or duplications, and difficulty in traceability of manual verification of embedding cassettes before dehydration in the tissue processing workflow of pathology departments, this invention aims to develop a batch pre-dehydration verification system for embedding cassettes based on image recognition and automated detection. By introducing deep learning technology, the system can automatically identify, accurately verify, and visually label the quantity and number of embedding cassettes, thereby improving the accuracy and efficiency of sample management, reducing labor costs and error risks, and ensuring the quality of pathological diagnosis.
[0023] See appendix Figure 1 As shown, the deep learning-based pre-dehydration inspection method for embedding cassettes of the present invention includes the following steps: S1: Collect images of dehydrated blue embedding boxes, and construct an embedding box dataset after annotating the images; S2: Improve the YOLOv11s model by using the EGhostC3 module and depthwise separable convolution as the main modules to construct the Eghost-YOLO model. Train the Eghost-YOLO model using the aforementioned embedding box dataset. Train the two models to identify the embedding boxes in the image and their corresponding embedding box numbers, respectively. S3: Input the dehydrated blue image to be analyzed into the trained Eghost-YOLO embedding box detection model to detect and crop each embedding box; S4: Input the cropped embedding box image into the Eghost-YOLO number detection model to detect and crop the corresponding number region; S5: Based on the pre-trained PP-OCRv4 character recognition model, the detected embedding box number region is identified, and the identified number is compared with the real number to output the number verification status.
[0024] In this embodiment, a high-speed document scanner (S1500A3AF) with a resolution of 3264×2448 is used to acquire dehydrated blue images. The server configuration is as follows: CPU (13th Gen Intel(R) Core(TM) i5-13500 2.50 GHz); GPU (Nvidia RTX 3060 12G); Memory: 16G DDR4.
[0025] First, a high-speed document scanner acquires dehydrated blue images (e.g., containing 200 embedding cassettes) and transmits the acquired images to the hospital's pathology information management system. Then, the corresponding identification number information for each image and the acquired image are transmitted to the algorithm interface. The Eghost-YOLO embedding cassette detection model rotates and locates all embedding cassettes in the image, then crops and corrects 200 sub-images of the embedding cassettes from the original image based on the coordinate boxes. These 200 cropped embedding cassette images are then fed into the Eghost-YOLO identification numbering model. The model locates the positions of the large and small identification numbers and crops them to obtain 200 large and 200 small identification number images. Images with the same identification number belonging to the same embedding cassette are then stitched together to obtain 200 identification number images. These identification number images are then uniformly sized to 320×48 and fed into the PP-OCRv4 recognition model to obtain the identification number information. Finally, the obtained identification number information is compared with the correct identification number information to identify abnormal identification numbers, which are then marked on the original image. The abnormal identification number information and the marked images are then fed back to the hospital's pathology information management system.
[0026] This invention employs a three-tiered cascaded high-resolution processing architecture: a three-level processing flow of "dehydrated blue whole → single embedded box → numbered region," effectively solving the problem of high-resolution image processing with multiple embedded boxes, significantly improving OCR accuracy and eliminating background interference. A lightweight and efficient detection model is introduced: the EGhostC3 module (feature redundancy utilization + channel attention focusing) and depthwise separable convolution (computational decoupling and cost reduction) are introduced into the core of the Eghost-YOLO system. While maintaining accuracy in detecting complex and blurred targets, this significantly reduces the computational load and parameter count of the model, achieving efficient real-time detection. Automated batch verification and visual output: the system automatically identifies the number of embedded boxes in the entire basket, compares them with a preset list, and detects anomalies (omissions, duplicates, redundancies, etc.), outputting visual annotation results (total number, location, and anomaly number boxes), replacing manual verification, significantly reducing the risk of sample confusion, and improving efficiency and operability.
[0027] This invention enables automated batch verification of high-resolution images of whole-basket embedding cassettes. It employs a step-by-step focusing mechanism to effectively suppress interference from complex backgrounds, improving the accuracy and stability of identification and verification. While ensuring detection and identification accuracy, it significantly increases processing speed, meeting the real-time requirements of pathology departments' pre-dehydration workflows.
[0028] The specific process of step S1 is as follows: S11: Use a high-speed document scanner to capture images of dehydrated blue pathological specimens from above, showing that the embedding cassettes are densely and regularly arranged in the images; S12: After acquiring the image, use the Labelme tool to annotate the outline of each embedding box in the image one by one. The annotation format is polygon to accurately represent the edge of the embedding box. S13: Calculate the minimum bounding rectangle of each labeled area, extract the center point coordinates, width and height, and tilt angle of the rectangle as the position and orientation information of each embedding box, which is used for subsequent area detection or verification positioning. S14: Use the LabelImg tool to individually label the number above or on the surface of each embedding box. The label type is horizontal box. Save the labeling results in YOLO format for training the number recognition model. To ensure the generalization ability of the data, divide all samples into training set, validation set and test set in a ratio of 7:2:1. When dividing, ensure that different image sources are evenly distributed to avoid data bias.
[0029] The Eghost-YOLO model includes the original modules of YOLOv11, consisting of three parts: Backbone, Neck, and Head. The Eghost-YOLO model architecture is as follows: Figure 2 As shown. The Backbone is used for lightweight feature extraction of images, the Neck is used for multi-scale feature fusion, and the Head is used for precise localization of the rotated bounding box. Compared to YOLOv11, this invention, Eaglet-YOLO, introduces the EGhostC3 module and depthwise separable convolution in the backbone and neck parts, effectively improving the detection accuracy and inference efficiency of the model in complex backgrounds.
[0030] Depthwise separable convolution is an efficient alternative to ordinary convolution. It decomposes the standard convolution operation into two sub-operations: channel-wise convolution and pointwise convolution. This can significantly reduce the number of parameters and computational cost while maintaining good feature representation capabilities.
[0031] The processing procedure of the EGhostC3 module is as follows: Figure 3 In the EGhostC3 module, input features are first fed into two parallel GhostConv branches. One branch uses standard convolution, while the other uses depthwise separable convolution to generate redundant features. Finally, the two sets of features are combined to maintain accuracy while reducing computational cost. The backbone of the EGhostC3 module consists of multiple stacked GhostBottleneck modules. Each GhostBottleneck contains a three-layer structure: GhostConv→DWConv→GhostConv, with residual connection capabilities. Its output is sequentially appended to the feature list and concatenated with the outputs of the initial two branches via channel dimension. The concatenated features are then upscaled to the target number of channels using a 1×1 GhostConv. If use_eca is enabled, the final feature will also pass through an ECA attention module for channel weighting to improve the responsiveness of key feature channels.
[0032] In practical applications, the background of the embedding box image is complex, and there are often problems such as occlusion, adhesion, and blurred edges between the boxes. The lightweight feature generation capability and channel attention mechanism of the EGhostC3 module can better focus on the target area and improve the discrimination and detection accuracy of the embedding box boundary.
[0033] Because the input image contains a large number of embedded boxes, and the embedded box numbers occupy a relatively small proportion of the entire image, to avoid a decrease in subsequent recognition accuracy due to blurry numbers, it is necessary to ensure image clarity while also considering the computational efficiency of the model. Directly scaling a high-resolution image to a lower size can easily lead to blurring and distortion of key details, weakening the detection model's ability to perceive boundaries and potentially causing unclear target contours and missing feature information, thus affecting the overall recognition performance. The original acquired image resolution is 3264×2448. Before being fed into the model, the image is first filled with a gray background using squares to maintain the original image proportions and adapt to the model's input size. Then, it is uniformly adjusted to 1024×1024 for embedded box detection. The detection results are located on the original image, and the corresponding embedded box regions are then cropped from the original high-resolution image to obtain clearer embedded box images for number detection and recognition. The number detection model uses the same filling method, with an input size of 256×256.
[0034] The specific process of training the Eaglehost-YOLO model using the aforementioned embedding box dataset in step S2 is as follows: S21: Before inputting the embedding box dataset into the Eghost-YOLO model, the image is filled with a gray background square to maintain the original image ratio and adapt to the model input size. Then it is uniformly adjusted to 1024×1024 for embedding box detection. S22: The detection result is located on the original image, and the corresponding embedding box area is then cropped from the original high-resolution image to obtain a clearer embedding box image for number detection and recognition. The number detection model uses the same filling method, and the input size is 256×256. S23: Using the preset total loss function Loss Calculate the total loss of the Eaglet-YOLO model, where the total loss function is derived from the bounding box regression loss function. L box Category classification loss function L cls and distribution-focused loss function L DFL The composition, specifically the formula, is as follows: ; in, , , The values represent the weights of each part of the loss, and are 0.5, 7.5, and 1.5 during the training process. ; in, C It is the number of categories; N It is the number of positive samples; Is the i-th sample in the i-th order? c The true label of the class; No. i The sample at the th c The predicted value of the class; This represents the Sigmoid activation function. ; ; ; ; in, B It is a prediction box; It is a real frame; w It is the predicted width; h It is the predicted bounding box height; It is the actual frame width; It is the actual frame height; It is the squared Euclidean distance between the predicted bounding box and the ground truth bounding box; It is the square of the diagonal length of the smallest closure rectangle containing the two boxes; This constitutes a center point distance penalty term; This constitutes a length-to-width ratio penalty term; ; in, and It is the softmax output predicted by the model, corresponding to the... i The and the first i +1 discrete point probability; y For true continuous numerical labels; and This represents two adjacent discretization points after the continuous value has been discretized.
[0035] Because the embedded boxes in the image are closely arranged and have a certain tilt angle, traditional horizontal rectangular detection boxes are difficult to accurately cover the target area. Therefore, in this invention, the head part of the Eagle-YOLO model uses rotated rectangular boxes as the output form to more accurately adapt to the actual geometry of the embedded boxes. In this step, the image is input into the trained embedded box detection model to obtain the rotated rectangular positioning box for each embedded box, and the detection results are mapped back to the original high-resolution image proportionally. A clear embedded box region is cropped from the original image based on the rotated boxes, and its angle is corrected to obtain a standardized embedded box image, providing a more stable and accurate input basis for subsequent numbering detection and recognition. This strategy effectively improves the detection robustness and image quality under complex arrangements and non-standard poses.
[0036] The specific process of step S3 is as follows: S31: Input the dehydrated blue image to be analyzed into the trained Eghost-YOLO embedding box detection model to obtain the detection results of the rotated rectangle positioning box of each embedding box; S32: The detection results are mapped back to the original high-resolution image proportionally. A clear embedding box region is then cropped from the original high-resolution image using a rotated bounding box, and its angle is corrected to obtain a standardized embedding box image. This provides a more stable and accurate input basis for subsequent numbering detection and recognition. This strategy effectively improves detection robustness and image quality under complex arrangements and non-standard poses.
[0037] The embedded box images obtained in the previous step, after preprocessing, are input into the numbering segment detection model. This model accurately locates and crops the large and multiple small number regions in the image. To unify the input format for subsequent recognition, the system resamples the small number images cropped from the same embedded box proportionally to the height of the large number image, and connects them with the large number image using an asterisk (*) to form a complete character image. This stitched image is then fed into the character recognition model as input. In this way, the numbering information of the same embedded box can be processed centrally, effectively reducing the number of recognition attempts and improving the overall recognition efficiency and system speed.
[0038] In step S4, the cropped embedding box image is input into the Eagle-YOLO number detection model. The specific process of detecting and cropping the corresponding numbered regions is as follows: S41: Input the cropped embedding box image into the Eghost-YOLO number detection model, which accurately locates and crops the large number and multiple small number regions in the cropped embedding box image; S42: To unify the input format for subsequent recognition, the system resamples the smaller numbered images cropped from the same embedding box proportionally according to the height of the larger numbered images; S43: Connect the smaller number image and the larger number image using the "*" symbol to form a complete character image; S44: Input the stitched complete character image into the PP-OCRv4 character recognition model.
[0039] This system uses the PP-OCRv4 model, and only uses its text recognition part. The model structure is as follows: Figure 4 As shown.
[0040] The first two stages of the PP-OCRv4 recognizer integrate local and global fusion modules. The former models short-range dependencies between characters, which helps in recognizing clustered or closely spaced numbered characters, while the latter models long-range relationships, enhancing the model's adaptability to long sequences and broken numbering. This collaborative modeling of local and global features enables the model to better handle complex interferences commonly found in embedding box numbering, such as unclear printing, blurred images, character tilt, and low contrast, thereby improving the accuracy and robustness of recognition.
[0041] Before inputting the stitched complete character image into the PP-OCRv4 text recognition model, its size is padded and uniformly adjusted to 320×48 to fit the structure of the PP-OCRv4 text recognition model. The specific process of step S5 is as follows: S51: The PP-OCRv4 text recognition model extracts the large and small number information corresponding to each embedding box; S52: The large and small number information corresponding to each extracted embedding box is compared with the number list that should exist for this batch, and the number verification is completed automatically; S53: Output a visual annotation result containing anomaly number information and the original image. The visual annotation result is as follows: the total number of detected embedding boxes is marked in the upper left corner of the image, all identified embedding box locations are marked with red dots, large anomaly numbers are marked with red boxes, small redundant numbers are marked with blue boxes, and duplicate small numbers are marked with green boxes.
[0042] A deep learning-based pre-dehydration inspection system for embedding cassettes is provided to implement the aforementioned deep learning-based pre-dehydration inspection method for embedding cassettes, including... The data acquisition and construction module is used to acquire images of dehydrated blue embedding boxes, and to construct an embedding box dataset after annotating the images. The model training module is used to improve the YOLOv11s model. It uses the EGhostC3 module and depthwise separable convolution as the main modules to build the Eghost-YOLO model. The Eghost-YOLO model is trained using the embedded box dataset to obtain two models, one for recognizing embedded boxes in images and the other for recognizing the corresponding embedded box numbers. The embedding box detection and cropping module is used to receive the dehydrated blue image to be analyzed, input it into the trained Eagle-YOLO embedding box detection model, and detect and crop out each embedding box. The numbered region detection and cropping module is used to receive the cropped embedding box image, input it into the Eagle-YOLO number detection model, and detect and crop out the corresponding numbered regions. The number recognition and verification module is used to identify the detected embedding box number region based on the pre-trained PP-OCRv4 text recognition model, compare the identified number with the real number, and output the number verification status.
Claims
1. A deep learning-based method for pre-dehydration inspection of embedding cassettes, characterized in that, Includes the following steps: S1: Collect images of dehydrated blue embedding boxes, and construct an embedding box dataset after annotating the images; S2: Improve the YOLOv11s model by using the EGhostC3 module and depthwise separable convolution as the main modules to construct the Eghost-YOLO model. Train the Eghost-YOLO model using the aforementioned embedding box dataset. Train the two models to identify the embedding boxes in the image and their corresponding embedding box numbers, respectively. S3: Input the dehydrated blue image to be analyzed into the trained Eghost-YOLO embedding box detection model to detect and crop each embedding box; S4: Input the cropped embedding box image into the Eghost-YOLO number detection model to detect and crop the corresponding number region; S5: Based on the pre-trained PP-OCRv4 character recognition model, the detected embedding box number region is identified, and the identified number is compared with the real number to output the number verification status.
2. The deep learning-based pre-dehydration inspection method for embedding cassettes according to claim 1, characterized in that, The specific process of step S1 is as follows: S11: Use a high-speed document scanner to capture images of dehydrated blue pathological specimens from above, showing that the embedding cassettes are densely and regularly arranged in the images; S12: After acquiring the image, use the Labelme tool to annotate the outline of each embedding box in the image one by one. The annotation format is polygon to accurately represent the edge of the embedding box. S13: Calculate the minimum bounding rectangle of each labeled area, and extract the coordinates of the center point, width, height and tilt angle of the rectangle as the position and orientation information of each embedding box; S14: Use the LabelImg tool to individually label the number above or on the surface of each embedding box. The label type is horizontal box. Save the labeling results in YOLO format for training the number recognition model. To ensure the generalization ability of the data, divide all samples into training set, validation set and test set in a ratio of 7:2:
1.
3. The deep learning-based pre-dehydration inspection method for embedding cassettes according to claim 1, characterized in that, The specific process of training the Eaglehost-YOLO model using the aforementioned embedding box dataset in step S2 is as follows: S21: Before inputting the embedding box dataset into the Eghost-YOLO model, the image is filled with a gray background square to maintain the original image ratio and adapt to the model input size. Then it is uniformly adjusted to 1024×1024 for embedding box detection. S22: The detection result is located on the original image, and the corresponding embedding box area is then cropped from the original high-resolution image to obtain a clearer embedding box image for number detection and recognition. The number detection model uses the same filling method, and the input size is 256×256. S23: Using the preset total loss function Loss Calculate the total loss of the Eaglet-YOLO model, where the total loss function is derived from the bounding box regression loss function. L box Category classification loss function L cls and distribution-focused loss function L DFL The composition, specifically the formula, is as follows: ; in, , , The values represent the weights of each part of the loss, and are 0.5, 7.5, and 1.5 during the training process. ; in, C It is the number of categories; N It is the number of positive samples; Is the i-th sample in the i-th order? c The true label of the class; No. i The sample at the th c The predicted value of the class; This represents the Sigmoid activation function. ; ; ; ; in, B It is a prediction box; It is a real frame; w It is the predicted width; h It is the predicted bounding box height; It is the actual frame width; It is the actual frame height; It is the squared Euclidean distance between the predicted bounding box and the ground truth bounding box; It is the square of the diagonal length of the smallest closure rectangle containing the two boxes; This constitutes a center point distance penalty term; This constitutes a length-to-width ratio penalty term; ; in, and It is the softmax output predicted by the model, corresponding to the... i The and the first i +1 discrete point probability; y For true continuous numerical labels; and This represents two adjacent discretization points after the continuous value has been discretized.
4. The deep learning-based pre-dehydration inspection method for embedding cassettes according to claim 1, characterized in that, The specific process of step S3 is as follows: S31: Input the dehydrated blue image to be analyzed into the trained Eghost-YOLO embedding box detection model to obtain the detection results of the rotated rectangle positioning box of each embedding box; S32: Map the detection results back to the original high-resolution image proportionally, crop the embedding box region on the original high-resolution image according to the rotating frame, and correct its angle to obtain a standardized embedding box image, providing a more stable and accurate input basis for subsequent numbering detection and recognition.
5. The deep learning-based pre-dehydration inspection method for embedding cassettes according to claim 4, characterized in that, In step S4, the cropped embedding box image is input into the Eagle-YOLO number detection model. The specific process of detecting and cropping the corresponding numbered regions is as follows: S41: Input the cropped embedding box image into the Eghost-YOLO number detection model, which accurately locates and crops the large number and multiple small number regions in the cropped embedding box image; S42: Resample the smaller numbered images cut out from the same embedding box proportionally to the height of the larger numbered images; S43: Connect the smaller number image and the larger number image using the "*" symbol to form a complete character image; S44: Input the stitched complete character image into the PP-OCRv4 character recognition model.
6. The deep learning-based pre-dehydration inspection method for embedding cassettes according to claim 1, characterized in that, Before inputting the stitched complete character image into the PP-OCRv4 text recognition model, its size is padded and uniformly adjusted to 320×48 to fit the structure of the PP-OCRv4 text recognition model. The specific process of step S5 is as follows: S51: The PP-OCRv4 text recognition model extracts the large and small number information corresponding to each embedding box; S52: The extracted large and small number information of each embedding box is compared with the corresponding number list to automatically complete the number verification; S53: Outputs a visual annotation result containing anomaly number information and the original image.
7. The deep learning-based pre-dehydration inspection method for embedding cassettes according to claim 6, characterized in that, The visualization annotation results in step S53 are as follows: The total number of detected embedding boxes is marked in the upper left corner of the image. Red dots are used to mark the locations of all identified embedding boxes. A large number of abnormal numbers is marked with a red box, a small number of redundant numbers is marked with a blue box, and a small number of duplicate numbers is marked with a green box.
8. A deep learning-based pre-dehydration inspection system for embedding cassettes, used to implement the deep learning-based pre-dehydration inspection method for embedding cassettes as described in any one of claims 1-7, characterized in that, include: The data acquisition and construction module is used to acquire images of dehydrated blue embedding boxes, and to construct an embedding box dataset after annotating the images. The model training module is used to improve the YOLOv11s model. It uses the EGhostC3 module and depthwise separable convolution as the main modules to build the Eghost-YOLO model. The Eghost-YOLO model is trained using the embedded box dataset to obtain two models, one for recognizing embedded boxes in images and the other for recognizing the corresponding embedded box numbers. The embedding box detection and cropping module is used to receive the dehydrated blue image to be analyzed, input it into the trained Eagle-YOLO embedding box detection model, and detect and crop out each embedding box. The numbered region detection and cropping module is used to receive the cropped embedding box image, input it into the Eagle-YOLO number detection model, and detect and crop out the corresponding numbered regions. The number recognition and verification module is used to identify the detected embedding box number region based on the pre-trained PP-OCRv4 text recognition model, compare the identified number with the real number, and output the number verification status.
Citation Information
Patent Citations
Embedded box label batch identification method and system based on deep learning and machine vision
CN115019306A
Lightweight container number identification method based on deep learning
CN117690137A
Long-term text line identification method based on word segmentation
CN118429989A
Multidirectional inclined steel grade detection and identification method based on deep learning
CN120340014A