Hepatoechinococcosis focus identification and classification method based on deep learning
By using deep learning methods to identify and classify echinococcosis lesions in the liver, the problems of low sensitivity of serum detection and easy confusion with ultrasound detection are solved, achieving efficient and reliable identification and classification of echinococcosis lesions in the liver, which is suitable for echinococcosis screening in resource-limited areas.
Patent Information
- Application Number
- CN202511937061.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-03-06
AI Technical Summary
In existing technologies, serum testing for hepatic echinococcosis has low sensitivity, making it difficult to meet the needs of early screening. Ultrasound testing is prone to confusing hepatic echinococcosis with other liver diseases. Furthermore, deep learning algorithms have weak generalization ability in hepatic echinococcosis detection and require high hardware computing power, making it difficult to meet actual testing needs.
A deep learning-based method for identifying and classifying echinococcosis lesions in the liver is adopted, including liver segmentation and identification, lesion target detection, coarse classification detection, fine sub-classification detection, and lesion contour segmentation. A lightweight model is used to identify and classify echinococcosis lesions in the liver in resource-constrained areas, thereby improving detection accuracy and reliability.
It can effectively identify and classify echinococcosis lesions in the liver, improve detection accuracy and reliability, and is suitable for real-time screening of echinococcosis in resource-limited areas.
Smart Images

Figure CN121616897A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a lesion identification and classification method, and more particularly to a deep learning-based method for identifying and classifying echinococcosis lesions in the liver. Background Technology
[0002] Hepatic echinococcosis is a prevalent zoonotic parasitic disease in pastoral areas. It has a long incubation period and can be life-threatening in its late stages, placing a heavy burden on social healthcare. While echinococcosis can be detected through serological testing, the sensitivity of serological tests is low and cannot meet the needs of early screening. Furthermore, primary healthcare resources are relatively scarce, and areas with a high incidence of hepatic echinococcosis lack specialized doctors and advanced equipment, making large-scale screening difficult. Currently, ultrasound detection can be used to screen for hepatic echinococcosis; however, existing ultrasound methods easily confuse hepatic echinococcosis with other liver diseases. In addition, ultrasound images cannot clearly show the spatial relationship between the lesions and blood vessels of hepatic echinococcosis, leading to misdiagnosis and missed diagnosis.
[0003] With the rapid development of artificial intelligence technology in the field of medical ultrasound, deep learning technology can address the aforementioned challenges. Specifically, deep learning models can automatically extract subtle features from ultrasound images and effectively detect early, minute lesions that are difficult to detect using traditional methods. However, current deep learning algorithms still have many shortcomings in the detection of echinococcosis of the liver, mainly due to weak model generalization ability, reliance on single image data, and high hardware computing power requirements, making it difficult to meet practical detection needs. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a deep learning-based method for identifying and classifying echinococcosis lesions in the liver, which can effectively identify and classify echinococcosis lesions in the liver and improve the accuracy and reliability of echinococcosis lesion detection.
[0005] According to the technical solution provided by this invention, a method for identifying and classifying hepatic echinococcosis lesions based on deep learning is provided, the method comprising: Acquire ultrasound images of the liver to be examined, and perform lesion target detection on the ultrasound images of the liver to be examined. When there are lesions in the ultrasound image of the liver to be examined, lesion detection status information of each lesion in the ultrasound image of the liver to be examined is generated. The lesion detection status information includes at least the lesion detection frame and the pathological type characterizing the pathological features of the lesion. For any lesion, when the pathological type of the lesion is benign, a coarse classification detection is performed on the lesion within the lesion detection frame to generate coarse classification information of the lesion, wherein the coarse classification information includes cystic echinococcosis, vesicular echinococcosis, or non-hepatic echinococcosis. For any lesion, if the coarse classification information of the lesion is cystic or alveolar echinococcosis, then the lesion is determined to be a hepatic echinococcosis lesion. The lesion is then subjected to echinococcosis sub-class detection and lesion contour segmentation to determine the corresponding echinococcosis sub-class when the lesion belongs to cystic or alveolar echinococcosis, as well as the target contour of the lesion.
[0006] When acquiring ultrasound images of the liver to be examined, the following steps are included: Abdominal ultrasound images generated by transabdominal ultrasound scans are acquired, and liver region segmentation and identification are performed on the abdominal ultrasound images. When the abdominal ultrasound scan image includes the liver region, liver identification status information corresponding to the liver region is generated, wherein the liver identification status information includes at least the liver region detection box in the abdominal ultrasound scan image; Based on liver identification status information, liver region extraction is performed on abdominal ultrasound images to generate ultrasound images of the liver to be examined. The liver region extraction process includes at least sequential image cropping and image denoising. After liver region extraction processing, the ultrasound image of the liver to be examined only includes the liver region.
[0007] When performing lesion target detection on ultrasound images of the liver to be examined, the following are included: The ultrasound image of the liver to be examined is loaded into the lesion target detection model, so that the lesion target detection model can be used to detect lesions in the ultrasound image of the liver to be examined. When a lesion is present in the ultrasound image of the liver to be examined and the lesion is liver cancer, the pathological type of the current lesion is identified as malignant; otherwise, the pathological type of the current lesion is identified as benign.
[0008] When constructing a lesion target detection model, the following are included: A YOLOv10 model is provided, and the backbone network within the YOLOv10 model is improved using EfficientFormerV2 to form a target detection backbone network within the lesion target detection model. The target detection backbone network includes a Stem Block module, a first Local Block module, a first Subsample module, a second Local Block module, a second Subsample module, a third Local Block module, a third Subsample module, a Local Global Block module, an SPPF Block module, and a PSA Block module connected in series. The second Local Block module, the third Local Block module, and the PSA Block module are adapted and connected to the target detection neck network of the lesion target detection model.
[0009] When constructing the lesion target detection model, the neck network within the YOLOv10 model was also improved using EfficientFormerV2 to form the target detection neck network for the lesion target detection model. When improving the neck network in the YOLOv10 model using EfficientFormerV2, the C2f module in the neck network is replaced with the Local Block module in EfficientFormerV2.
[0010] When the pathological type of a lesion is benign, a benign lesion image classification model is used to perform coarse classification detection on the current lesion to determine the coarse classification information of the lesion; When the coarse classification information of the lesion is cystic echinococcosis, the cystic echinococcosis image classification model is used to detect the sub-category of echinococcosis in the current lesion to determine the sub-category of echinococcosis in the current lesion, wherein the sub-category of echinococcosis is one of cystic echinococcosis CE1 to cystic echinococcosis CE5. When the coarse classification information of the lesion is vesicular echinococcosis, the vesicular echinococcosis image classification model is used to detect the sub-category of the echinococcosis of the current lesion in order to determine the sub-category of the echinococcosis of the current lesion. The sub-category of the echinococcosis is one of vesicular echinococcosis AE1 to vesicular echinococcosis AE4.
[0011] When the coarse classification information of the lesion is non-hepatic echinococcosis, the conventional lesion image classification model is used to perform conventional sub-category detection on the current lesion to determine the conventional sub-category of the current lesion, wherein the conventional sub-category includes at least hepatic cysts; When a lesion is routinely classified as a liver cyst, or classified as cystic echinococcosis CE1, a subclassification verification process is performed on the current lesion. When performing detailed category validation processing, the following are included: The current lesion is semantically segmented using a two-layer wall sign semantic segmentation model to generate wall sign information of the current lesion; If the wall sign information of the current lesion includes a double-wall sign, then the coarse classification information of the current lesion is identified as cystic echinococcosis, and the corresponding sub-category of echinococcosis is identified as cystic echinococcosis CE1. Otherwise, the coarse classification information of the current lesion is identified as non-hepatic echinococcosis, and the corresponding sub-category of echinococcosis is identified as hepatic cyst.
[0012] When segmenting the lesion outline, the following steps are included: Within the ultrasound image of the liver to be examined, the lesion detection frame of the current lesion is expanded and cropped to generate the lesion segmentation benchmark region; The lesion contour semantic segmentation model is used to perform contour semantic segmentation on the lesion segmentation benchmark region to generate the target contour of the lesion. The ultrasound image of the liver to be examined is completed and corrected based on the target contour of the lesion, so as to generate a complete liver contour after completion and correction. Based on the liver contour completion, the target contour of the lesion and the corresponding echinococcosis subcategories are mapped and output.
[0013] When completing and correcting the ultrasound image of the liver to be examined based on the target contour of the lesion, the following is included: The liver image to be examined and the target contour of the lesion are placed on the same mask. Then, the maximum contour search method is used to find the maximum connected region, and the found maximum connected region is configured as the liver completion contour.
[0014] When multiple consecutive ultrasound images of the liver to be examined are available, the process also includes lesion tracking, in which... When performing lesion tracking and treatment, the following are included: For any two adjacent time-series liver ultrasound images, the lesion detection box of each lesion in the previous time-series liver ultrasound image is used as the basic detection box, and the lesion detection box of each lesion in the next time-series liver ultrasound image is used as the tracking detection box. Motion prediction is performed on each basic detection box to generate a trajectory prediction box, and the similarity between the trajectory prediction box and each tracking detection box is calculated. Based on the calculated similarity, the Hungarian matching status of the trajectory prediction box and all tracking detection boxes is calculated, where, When a tracking detection box and a trajectory prediction box satisfy a Hungarian match, the current trajectory prediction box is configured as the tracking target box. The tracking target box and the corresponding tracking detection box are deduplicated, and the corresponding lesion target box is generated. The generated lesion target box is then configured as the lesion detection box for the corresponding lesion.
[0015] The advantages of this invention are as follows: A liver segmentation and recognition model is used to segment and identify the liver region in abdominal ultrasound images, generating corresponding ultrasound images of the liver to be examined. The lesion target detection model is used to detect lesions in the ultrasound images of the liver to be examined, obtaining the detection box for each lesion and its corresponding pathological type. A benign lesion image classification model can be used for coarse classification detection of lesions; cystic and vesicular echinococcosis image classification models can be used for sub-classification detection; a conventional lesion image classification model can be used for conventional sub-classification detection; and a lesion contour semantic segmentation model can be used to perform contour semantic segmentation of the lesion segmentation reference region. Since all the corresponding models use lightweight deep learning models, they can meet the real-time identification and classification needs of hepatic echinococcosis even in resource-constrained areas, effectively achieving the identification and classification of hepatic echinococcosis lesions and improving the accuracy and reliability of hepatic echinococcosis lesion detection. Attached Figure Description
[0016] Figure 1 This is a schematic flowchart of an embodiment of the present invention for identifying and classifying echinococcosis lesions in the liver.
[0017] Figure 2 This is a structural block diagram of an embodiment of the lesion target detection model of the present invention.
[0018] Figure 3 This is a schematic diagram of an embodiment of the present invention for coarse and fine classification detection of lesions.
[0019] Figure 4 This is a schematic diagram of one embodiment of an abdominal ultrasound scan.
[0020] Figure 5 This is a schematic diagram of an embodiment of the present invention that generates lesion identification and classification results after identifying and classifying echinococcosis lesions in the liver. Detailed Implementation
[0021] The present invention will be further described below with reference to specific accompanying drawings and embodiments.
[0022] To effectively identify and classify hepatic echinococcosis lesions, this invention provides a deep learning-based method for identifying and classifying hepatic echinococcosis lesions. Specifically, the method includes: Acquire ultrasound images of the liver to be examined, and perform lesion target detection on the ultrasound images of the liver to be examined. When there are lesions in the ultrasound image of the liver to be examined, lesion detection status information of each lesion in the ultrasound image of the liver to be examined is generated. The lesion detection status information includes at least the lesion detection frame and the pathological type characterizing the pathological features of the lesion. For any lesion, when the pathological type of the lesion is benign, a coarse classification detection is performed on the lesion within the lesion detection frame to generate coarse classification information of the lesion, wherein the coarse classification information includes cystic echinococcosis, vesicular echinococcosis, or non-hepatic echinococcosis. For any lesion, if the coarse classification information of the lesion is cystic or alveolar echinococcosis, then the lesion is determined to be a hepatic echinococcosis lesion. The lesion is then subjected to echinococcosis sub-class detection and lesion contour segmentation to determine the corresponding echinococcosis sub-class when the lesion belongs to cystic or alveolar echinococcosis, as well as the target contour of the lesion.
[0023] Since hydatid cysts can only exist within the liver, ultrasound images of the liver to be examined should be obtained when identifying and classifying hydatid cyst lesions. These liver ultrasound images can be obtained using existing technologies, such as performing an abdominal ultrasound scan and generating an abdominal ultrasound image, from which a corresponding liver ultrasound image can be generated. Of course, other methods can also be used to obtain liver ultrasound images, which will not be elaborated upon here.
[0024] Because some common focal liver diseases occupy a small area of the liver, such as intrahepatic hyperechoic lesions, when directly detecting lesions using abdominal ultrasound images, the abdominal ultrasound images undergo multi-layer downsampling, which greatly compresses the feature information of the lesion area in the feature map. Then, by upsampling and scaling back to the original size, some foreground targets are lost, resulting in missed lesions. Therefore, this invention generates an ultrasound image of the liver to be examined based on an abdominal ultrasound image, and uses the ultrasound image of the liver to be examined as a benchmark for subsequent identification and classification. This can effectively avoid missed lesions and improve the accuracy and reliability of identification and classification of hydatid cysts in the liver.
[0025] Depend on Figure 1 As can be seen, after acquiring the ultrasound image of the liver to be examined, lesion target detection should be performed on the ultrasound image to determine whether lesions exist within the liver ultrasound image. Specifically, when lesions are present in the ultrasound image of the liver to be examined, after lesion target detection, lesion detection status information for each lesion should be provided. In specific implementation, the lesion detection status information should at least include the lesion detection box of the current lesion and the pathological type of the lesion. The lesion detection box is generally the smallest bounding rectangle representing the location of the lesion. The pathological type of the lesion specifically refers to whether the current lesion is benign or malignant. Generally, when the lesion is likely to be liver cancer, it can be considered a malignant lesion, and others can be considered benign lesions. The method of lesion target detection on the ultrasound image of the liver to be examined will be explained in detail below, and you can refer to the corresponding content below for details.
[0026] It should be noted that if no lesions are found in the ultrasound image of the liver to be examined, no lesion detection status information will be generated after lesion target detection. This means that it can be assumed that there are no hydatid lesions in the current ultrasound image of the liver to be examined, and the identification and classification of hydatid lesions in this invention can be terminated. Furthermore, if the pathological type of a lesion is malignant, then the lesion is unlikely to be a hydatid lesion. Therefore, the identification and classification of the current lesion for hydatid lesions can also be terminated, thereby improving the efficiency and reliability of the identification and classification of hydatid lesions.
[0027] Because there are various types of hepatic echinococcosis lesions, and benign lesions may not necessarily be hepatic echinococcosis lesions, when the pathological type of a lesion is benign, a coarse classification test should be performed on each lesion to generate coarse classification information for the current lesion. This coarse classification information generally includes cystic echinococcosis, alveolar echinococcosis, or non-hepatic echinococcosis. It should be understood that "non-hepatic echinococcosis" specifically refers to lesions belonging to categories other than cystic and alveolar echinococcosis, such as hepatic cysts, hepatic hemangiomas, intrahepatic hyperechoic lesions, and other common benign liver diseases, which will not be listed here. In practice, when multiple lesions are present in the ultrasound image of the liver to be examined, lesion detection status information for each lesion will be generated. When the pathological type of all lesions is benign, coarse classification testing should be performed on all of them.
[0028] It should be noted that cystic and alveolar echinococcosis are broad categories of hepatic echinococcosis. Therefore, when the coarse classification information of the lesion is cystic or alveolar echinococcosis, the current lesion can be identified as a hepatic echinococcosis lesion. Further sub-classification of the lesion is then required. Figure 1 It is known that when the lesion is roughly classified as cystic or vesicular echinococcosis, the current lesion should be subjected to echinococcosis sub-classification detection and lesion contour segmentation. In practice, the current lesion can be subjected to echinococcosis sub-classification detection and lesion contour segmentation at the same time, or the echinococcosis sub-classification detection and lesion contour segmentation can be performed sequentially. The order of echinococcosis sub-classification detection and lesion contour segmentation can be selected as needed.
[0029] Specifically, when performing hydatid subclassification detection, the subclassification should be based on the image region corresponding to the lesion detection box of the current lesion. After performing hydatid subclassification detection, the main focus should be on determining whether the corresponding lesion belongs to cystic or vesicular hydatids. For example, if the coarse classification information of a lesion is cystic hydatids, the corresponding subclassification should be one of cystic hydatids CE1 to cystic hydatids CE5. If the coarse classification information of a lesion is vesicular hydatids, the corresponding subclassification should be one of vesicular hydatids AE1 to vesicular hydatids AE4. The method and process of hydatid subclassification detection can be referred to in the following description.
[0030] Since the lesion detection bounding box is used to mark the location area of the lesion, but cannot effectively mark the outline of the lesion, lesion outline segmentation is used to determine the target outline of the lesion. The method of lesion outline segmentation will be explained in detail below. After determining the target outline and echinococcosis subcategory of each lesion, the identification and classification of a liver echinococcosis lesion is completed. In practice, the target outline and echinococcosis subcategory of each lesion should be output, which can effectively realize the identification and classification of liver echinococcosis lesions.
[0031] In one embodiment of the present invention, acquiring an ultrasound image of the liver to be examined includes: Abdominal ultrasound images generated by transabdominal ultrasound scans are acquired, and liver region segmentation and identification are performed on the abdominal ultrasound images. When the abdominal ultrasound scan image includes the liver region, liver identification status information corresponding to the liver region is generated, wherein the liver identification status information includes at least the liver region detection box in the abdominal ultrasound scan image; Based on liver identification status information, liver region extraction is performed on abdominal ultrasound images to generate ultrasound images of the liver to be examined. The liver region extraction process includes at least sequential image cropping and image denoising. After liver region extraction processing, the ultrasound image of the liver to be examined only includes the liver region.
[0032] To improve the convenience of ultrasound scanning, a routine ultrasound scan of the abdomen is generally performed, generating corresponding abdominal ultrasound images. The methods for performing the abdominal ultrasound scan and generating abdominal ultrasound images are consistent with existing technologies and will not be elaborated here. It is understood that when performing an abdominal ultrasound scan, the abdominal ultrasound image may also include other abdominal organs, such as the gallbladder, kidneys, and pancreas. For the identification and classification of hepatic echinococcosis lesions mentioned above, an ultrasound image of the liver should be provided. Therefore, in order to generate the aforementioned ultrasound image of the liver to be examined, liver region segmentation and identification should be performed on the abdominal ultrasound image.
[0033] To effectively segment and identify the liver region, a liver segmentation and recognition model should be constructed. Subsequently, the abdominal ultrasound scan image is loaded into the liver segmentation and recognition model to perform the required liver region segmentation and recognition. It is understood that the abdominal ultrasound scan image may or may not contain a liver region. When a liver region is present in the abdominal ultrasound scan image, the liver segmentation and recognition model will generate corresponding liver recognition status information. If a liver region is absent, no liver recognition status information will be generated, and consequently, the ultrasound image of the liver to be examined cannot be generated.
[0034] In practice, the liver identification status information should generally include a liver region detection box. That is, after liver region segmentation and identification, a liver region detection box will be marked within the abdominal ultrasound image. Subsequently, based on the liver region detection box, the abdominal ultrasound image is processed to extract the liver region, generating a corresponding ultrasound image of the liver to be examined. It is understood that the generated ultrasound image of the liver to be examined only includes the liver and omits other abdominal organs, thus avoiding the influence of other abdominal organs on the identification and classification of hepatic echinococcosis lesions.
[0035] To obtain the desired ultrasound image of the liver, the liver region extraction process should include image cropping and image denoising. Specifically, during image cropping, the image can be cropped along the liver region detection bounding box. After cropping, the image will still include the abdominal imaging area other than the liver; therefore, image denoising is also necessary. During image denoising, the liver outline is used as a mask image to preserve the liver region, while the image of the region outside the liver is filled with areas with a pixel value of 0. This ensures that the generated ultrasound image of the liver only contains the liver region. Of course, other methods can also be used to achieve image denoising, depending on whether an ultrasound image containing only the liver region is generated; these will not be elaborated upon here.
[0036] To reduce the hardware computing power requirements, the liver segmentation and recognition model should adopt a lightweight semantic segmentation network model. In specific implementation, the liver segmentation and recognition model can adopt the LACTNet model. Specifically, the LACTNet model introduces a lightweight Transformer structure, which takes into account both local details and global context, achieving a balance between accuracy and real-time performance. It is suitable for use in real-time ultrasound scanning for liver diseases, enabling fine and fast segmentation of the liver contour, avoiding the exclusion of lesions close to the capsule or slightly protruding from the liver, which could lead to missed detections in subsequent tests.
[0037] To enable the liver segmentation and recognition model based on the LACTNet model to perform the aforementioned liver region segmentation and recognition, a liver segmentation training dataset should be constructed. This dataset should then be used to train the LACTNet model. Once the model training target state is reached, the required liver segmentation and recognition model can be generated. The following example illustrates how to construct the liver segmentation training dataset. Ultrasound images of the abdominal organs were acquired, including both images of patients with diseases and normal images without obvious abnormalities. The acquired ultrasound image data was then anonymized. One feasible anonymization method is to crop out information outside the ultrasound imaging area, such as the hospital name, examination time, and patient personal information.
[0038] In addition, when the acquired ultrasound image is in video format, keyframe annotations can be extracted. The number of frames extracted needs to be adjusted according to the video frame rate and motion changes. For video segments with small changes in consecutive frames, intermediate frame annotations can be automatically generated through annotation interpolation algorithms to reduce the workload of manual annotation.
[0039] Several abdominal organ training ultrasound images can be obtained through the above method. Then, each abdominal organ training ultrasound image is labeled. When labeling, if a frame of abdominal organ training ultrasound image includes different abdominal organs and there are lesions in the abdominal organs, the labeling content includes the type of organ included in each abdominal organ training ultrasound image, the outline of the organ, the disease type of the lesion, and the outline points of the lesion. It can be understood that when a frame of abdominal organ training ultrasound image does not contain lesions, the labeling content should include the type of organ and the outline of the organ. The type of organ can be the liver, gallbladder, kidney, pancreas, etc. mentioned above.
[0040] Since this invention targets the identification and detection of hepatic echinococcosis lesions, the lesions should include hepatic echinococcosis lesions. Therefore, the disease types of the lesions include, but are not limited to, the five cystic echinococcosis types (CE1-CE5), the four alveolar echinococcosis types (AE1-AE4), hepatic cysts, hepatic hemangiomas, intrahepatic hyperechoic lesions, and other common benign liver diseases. The malignant disease type is primarily hepatocellular carcinoma. Furthermore, for cystic echinococcosis CE1, the double-wall sign of the corresponding lesion should be marked, that is, the outline of the echinococcosis, the echinococcosis type, and the regional outline of the area containing the double-wall sign should be marked. In specific implementation, the method of marking the organ type, outline, disease type of the lesion, and the lesion outline points can be consistent with existing technologies and will not be elaborated here. It should be noted that when the lesion may be hepatocellular carcinoma, the label should be "Hepatocellular carcinoma possible."
[0041] After annotation, data cleaning is generally required to remove erroneous or invalid labeled samples, thereby improving data quality and the stability of model training. Furthermore, a bilateral filtering method can be used to denoise the abdominal organ training ultrasound images, effectively suppressing noise while preserving lesion edge information. Further, for all abdominal organ training ultrasound images, the aspect ratio can be maintained to preserve the original morphological characteristics of the lesions; channel and color space conversion can be performed, and contrast and brightness can be randomly enhanced to adapt to different devices and ultrasound gain scenarios; normalization and standardization can be used to ensure image consistency at the model level.
[0042] After the above-mentioned annotation and data cleaning steps, a liver segmentation training dataset can be generated based on the annotated abdominal organ training ultrasound images. The liver segmentation training dataset includes liver segmentation training samples, and each liver segmentation training sample includes a frame of abdominal organ training ultrasound image and the organ category and organ outline annotated in the abdominal organ training ultrasound image.
[0043] When training the model, necessary training conditions should be configured, such as the loss function and optimizer. These can be selected based on needs, ensuring the model training requirements are met. Specifically, when training the LACTNet model, the loss function can be chosen as cross-entropy loss. During training, a frame of abdominal organ training ultrasound image is loaded into the LACTNet model. The LACTNet model is then used to predict the organ categories and contours projected from the abdominal organ training ultrasound. Subsequently, the cross-entropy loss can be calculated using the standard organ categories and contours within the abdominal organ training ultrasound image, compared to the predicted categories and contours. The method for calculating the cross-entropy loss can be consistent with existing techniques and will not be elaborated here.
[0044] In practice, when the calculated loss function tends to stabilize, it can be considered that the training of the LACTNet model has reached the target state of model training, and the required liver segmentation and recognition model can be generated.
[0045] In one embodiment of the present invention, lesion target detection is performed on the ultrasound image of the liver to be examined, including: The ultrasound image of the liver to be examined is loaded into the lesion target detection model, so that the lesion target detection model can be used to detect lesions in the ultrasound image of the liver to be examined. When a lesion is present in the ultrasound image of the liver to be examined and the lesion is liver cancer, the pathological type of the current lesion is identified as malignant; otherwise, the pathological type of the current lesion is identified as benign.
[0046] Understandably, in order to perform lesion detection on the ultrasound image of the liver to be examined, a lesion detection model should be constructed. That is, when performing lesion detection, the lesion detection model should be used. As explained above, when using the lesion detection model for lesion detection, if a lesion is present in the ultrasound image of the liver to be examined, a lesion detection bounding box should be provided for each lesion, and the pathological type of each lesion should be given, which should be benign or malignant. In one embodiment of the present invention, if the lesion is liver cancer, the pathological type of the lesion is identified as malignant; the pathological types of other lesions are all identified as benign.
[0047] In practice, the color of each lesion detection box can indicate the pathological type. For example, if the pathological type is benign, the corresponding lesion detection box can be green, and if the pathological type is malignant, the corresponding lesion detection box can be yellow. This is to minimize the interference of the inference results on the real-time ultrasound image and provide possible benign or malignant results in real time for doctors' reference.
[0048] Currently, traditional convolutional neural networks (CNNs) acquire global information by gradually expanding the receptive field through stacked convolutions. In contrast, the Visual Transformer (ViT) uses a self-attention mechanism to directly capture global information of an image, achieving more efficient long-range dependency capture. However, ViT also suffers from drawbacks such as a large number of parameters, high inference latency, and difficulty in deployment on edge devices. To overcome the shortcomings of existing models, in one embodiment of this invention, the lesion target detection model includes: A YOLOv10 model is provided, and the backbone network within the YOLOv10 model is improved using EfficientFormerV2 to form a target detection backbone network within the lesion target detection model. The target detection backbone network includes a Stem Block module, a first Local Block module, a first Subsample module, a second Local Block module, a second Subsample module, a third Local Block module, a third Subsample module, a Local Global Block module, an SPPF Block module, and a PSA Block module connected in series. The second Local Block module, the third Local Block module, and the PSA Block module are adapted and connected to the target detection neck network of the lesion target detection model.
[0049] As described above, the lesion detection model of this invention adopts a hybrid architecture, which combines the local modeling capabilities of CNNs with the global attention mechanism of Transformers. This retains the real-time performance of CNNs while improving the ability to extract global information when the proportion of echinococcosis in the image is large. Since the lesions of some patients with hepatic echinococcosis are larger than those of conventional liver lesions, the lesion detection model with the hybrid architecture can acquire features of the lesions from the perspective of the global image, effectively reducing the false negative rate of larger lesions.
[0050] Figure 2The figure illustrates an embodiment of the lesion target detection model of the present invention. In the figure, Local Block 1 is the first Local Block module, Subsample 1 is the first Subsample module, Local Block 2 is the second Local Block module, Subsample 2 is the second Subsample module, and Local Block 3 is the third Local Block module. During inference, the ultrasound image of the liver to be examined is loaded into the Stem Block module. The feature map thereafter passes through the second Local Block module, the third Local Block module, and the final PSA Block module to perform downsampling at 8x, 16x, and 32x, obtaining feature maps at three scales, which are then fed into the target detection neck network of the lesion target detection model, thereby sharing features of various dimensions at different scales.
[0051] In practical implementation, the Stem Block module, Local Block module, and Subsample module can adopt existing commonly used forms. Specifically, A Local Block consists of N Local sub-modules. Each Local sub-module contains a pooling layer with a kernel size of 3*3, a convolutional layer with a kernel size of 1*1, a batch normalization layer, and a GeLU activation function layer. In one embodiment of the present invention, the number of N in the three Local Blocks is set to 3, 3, and 9, respectively.
[0052] The Subsample module is a dual-path attention downsampling architecture. Compared to traditional downsampling that only uses convolution or pooling, this structure uses a combination of static local path downsampling and global path downsampling through residual connections.
[0053] The Local Global Block module consists of N Local Global sub-modules. Each Local Global sub-module includes a multi-head self-attention (MHSA) module based on a transformer architecture, a layer normalization layer, a linear layer, and a GeLU activation function layer. Here, N is 6. The MHSA module injects local convolutions into the value, adds a depthwise separable convolution with a kernel size of 3*3, and introduces a cross-head communication layer to improve expressive power. For some difficult samples with hepatic echinococcosis lesions, it can improve the detection capability.
[0054] In one embodiment of the present invention, when constructing the lesion target detection model, the neck network in the YOLOv10 model is also improved using EfficientFormerV2 to form the target detection neck network of the lesion target detection model, wherein... When improving the neck network in the YOLOv10 model using EfficientFormerV2, the C2f module in the neck network is replaced with the Local Block module in EfficientFormerV2.
[0055] Figure 2 The figure also illustrates an embodiment of a target detection neck network, which includes an upsampling unit Upsample1, wherein... Upsampling unit Upsample1 is connected to PSA Block module. The output of upsampling unit Upsample1 and the third Local Block module are connected to stitcher Concat2. The output of stitcher Concat2 is connected to the fifth Local Block module. The fifth Local Block module is connected to upsampling unit Upsample2. The output of upsampling unit Upsample2 and the second Local Block module are both connected to stitcher Concat1. The output of stitcher Concat1 is connected to convolution block Conv and head module Head1 in target detection head. The output of the convolutional block Conv and the output of the fifth Local Block module are connected to the splicer Concat3. The output of the splicer Concat3 is connected to the sixth Local Block module. The sixth Local Block module is connected to the SCDown module. The output of the sixth Local Block module is also connected to the head module Head2 inside the target detection head. The SCDown module and PSA Block module are connected to the Concat4 splicer. The output of the Concat4 splicer is connected to the C2fCIB module, and the output of the C2fCIB module is connected to the Head3 detection head module inside the target detection head.
[0056] Figure 2 In this model, Local Block 5 is the fifth Local Block module, Local Block 6 is the sixth Local Block module, and the stitchers Concat1 through Concat4 all perform channel stitching. The head modules Head1 through Head3 within the target detection head can be consistent with the existing YOLOv10 model, and will not be described further here.
[0057] The lesion target detection model described above can be constructed and generated using the following method. In one feasible embodiment, Construct a basic model for lesion detection, and construct a lesion detection training dataset for training the basic model. Configure the model training conditions and train the basic lesion detection model on the lesion detection training dataset until the training reaches the target state. Then, generate the lesion target detection model based on the basic lesion detection model that has reached the target state.
[0058] In practical implementation, the basic model for lesion detection can be referenced from the description of the lesion target detection model above, that is, the basic model for lesion detection is the model state before it has been trained to the target state. The lesion detection training dataset should include several lesion detection training samples. Each lesion detection training sample may include the aforementioned frame of abdominal organ training ultrasound image, but the labels of the lesion detection training samples should only select the liver category, the category of lesions within the liver, and the outline of the lesion. Specifically, among the lesion categories, liver cancer is mapped to the malignant category, while other lesion categories are mapped to the benign category.
[0059] Configuring model training conditions generally refers to the necessary conditions for model training, such as the loss function, optimizer, and other essential parameters. When the loss value of the loss function tends to stabilize, it can generally be considered that the training has reached the target state. For the loss function used in training, we have:
[0060] in, For target detection training loss, For classifying losses, For the target bounding box regression loss, For confidence loss, The classification loss dynamic balance coefficient, The dynamic balance coefficient for the target box regression loss. This is the dynamic balance coefficient for confidence loss.
[0061] In practice, the dynamic balance coefficients of classification loss, target box regression loss, and confidence loss will be adaptively adjusted according to the training rounds (emphasis on regression loss in the early stages and classification loss in the later stages). The specific adaptive adjustment method can be consistent with existing technologies, and will not be elaborated here.
[0062] The classification loss uses cross-entropy loss. When calculating the classification loss, the category of each lesion in the liver region within each lesion detection training sample should be used, along with the category of the corresponding lesion predicted by the basic lesion detection model for the current lesion detection training sample. Then, the corresponding classification loss can be calculated according to the calculation expression of cross-entropy loss.
[0063] The target bounding box regression loss uses the CIOU loss. When calculating the target bounding box regression loss, the bounding boxes of each lesion within the liver region of each lesion detection training sample, along with the basic lesion detection model, are used to predict the corresponding lesion bounding boxes for the current lesion detection training sample. Then, the intersection-union ratio (IU) of the bounding boxes and the corresponding predicted boxes is calculated, and the corresponding CIOU loss can be calculated according to the CIOU loss calculation expression. It should be noted that the bounding box of each lesion can be the minimum bounding rectangle generated from the contour of each lesion.
[0064] The confidence loss uses the BCEWithLogitsLoss loss. When calculating the confidence loss, the foreground and background categories of the lesion within the liver region of each lesion detection training sample should be used, along with the foreground and background categories of the corresponding lesion predicted by the basic lesion detection model for the current lesion detection training sample. Then, the corresponding confidence loss can be calculated according to the BCEWithLogitsLoss loss formula. It should be noted that in the foreground and background categories, the area where the lesion is located is the foreground, and the other parts are the background.
[0065] The above provides an explanation of the calculation of classification loss, target box regression loss, and confidence loss. It is understood that other loss calculation methods are also possible, and the specific method can be selected according to the needs. These will not be illustrated here.
[0066] In one embodiment of the present invention, when the pathological type of a lesion is benign, a benign lesion image classification model is used to perform coarse classification detection on the current lesion to determine the coarse classification information of the lesion; When the coarse classification information of the lesion is cystic echinococcosis, the cystic echinococcosis image classification model is used to detect the sub-category of echinococcosis in the current lesion to determine the sub-category of echinococcosis in the current lesion, wherein the sub-category of echinococcosis is one of cystic echinococcosis CE1 to cystic echinococcosis CE5. When the coarse classification information of the lesion is vesicular echinococcosis, the vesicular echinococcosis image classification model is used to detect the sub-category of the echinococcosis of the current lesion in order to determine the sub-category of the echinococcosis of the current lesion. The sub-category of the echinococcosis is one of vesicular echinococcosis AE1 to vesicular echinococcosis AE4.
[0067] As explained above, after using the lesion target detection model to perform target detection on the ultrasound image of the liver under examination, the pathological type of each lesion can be determined. At this point, it can only be determined whether the lesion is benign or malignant. When the pathological type of the lesion is determined to be benign, the lesion may still be a non-hepatic echinococcosis case. In order to further determine the type of benign lesions, a benign lesion image classification model should be used to perform coarse classification detection of the lesions. After coarse classification detection, coarse classification information of the lesions can be obtained, such as... Figure 3As shown, the information regarding coarse classification can be found in the corresponding explanations above.
[0068] In practical implementation, the RepViT model can be used for the classification of benign lesion images. The RepViT model draws on the Transformer architecture pattern and uses a pure convolutional architecture. While retaining the high feature representation capability of ViT, it also utilizes the ability of convolutional neural networks to focus on local features. It is suitable for scenarios that need to focus on the morphological and texture feature differences of diseases, and maintains low latency to meet the requirements of real-time ultrasound scanning inference.
[0069] When constructing a classification model for benign lesion images based on the RepViT model, one feasible approach is as follows: Construct a benign lesion classification training dataset and configure the model training conditions until the RepViT model is trained to the target training state using the benign lesion classification training dataset. At this point, a benign lesion image classification model can be generated based on the RepViT model that has reached the target training state.
[0070] Specifically, the benign lesion classification training dataset includes several benign lesion component training samples. Each benign lesion classification training sample may include one frame of abdominal organ training ultrasound image mentioned above. The label of each benign lesion classification training sample may include cystic echinococcosis, vesicular echinococcosis, and non-hepatic echinococcosis. As can be seen from the above description, during the initial annotation, the specific category of the lesion is mainly annotated, such as the aforementioned cystic echinococcosis CE1~CE5, vesicular echinococcosis AE1~AE4, liver cysts, liver hemangiomas, and intrahepatic hyperechoic lesions. In order to generate labels for the benign lesion classification training samples, the annotated cystic echinococcosis CE1~CE5 should be mapped to the cystic echinococcosis category, the vesicular echinococcosis AE1~AE4 should be mapped to the vesicular echinococcosis category, and liver cysts, liver hemangiomas, and intrahepatic hyperechoic lesions should be mapped to the non-hepatic echinococcosis category.
[0071] Refer to the above description for configuring model training conditions. Specifically, the loss function in the model training conditions can be cross-entropy loss. When calculating cross-entropy loss, the labels of lesions located in the liver region within each benign lesion classification training sample are mainly used, along with the predicted labels of lesions located in the liver region within the current benign lesion classification training sample predicted by the RepViT model. Then, the corresponding training loss can be calculated according to the cross-entropy loss calculation expression. Generally, when the training loss tends to stabilize, the target training state can be considered reached.
[0072] After obtaining the coarse classification information for each lesion, in order to determine the sub-category of each lesion, a further sub-category test for echinococcosis should be performed. Figure 3It is known that when performing subcategorization detection of echinococcosis, at least two image classification models should be constructed: one for cystic echinococcosis and one for vesicular echinococcosis. Then, the image classification model for cystic echinococcosis is used to perform subcategorization detection of lesions with coarse classification information of cystic echinococcosis, and the image classification model for vesicular echinococcosis is used to perform subcategorization detection of lesions with coarse classification information of vesicular echinococcosis.
[0073] It should be noted that after using the lesion target detection model to detect lesions in the ultrasound image of the liver to be examined, a lesion detection box can be generated for each lesion. Therefore, when performing coarse classification detection and fine classification detection of echinococcosis, the focus is mainly on the lesion area corresponding to the lesion detection box, which is consistent with the existing technology.
[0074] In practice, both the cystic and vesicular echinococcosis image classification models can be built based on the RepViT model, but they use different training sets. For example, separate training datasets can be constructed for cystic and vesicular echinococcosis classification. The cystic echinococcosis classification training dataset includes several cystic echinococcosis classification training samples. Each cystic echinococcosis classification training sample includes one frame of abdominal organ training ultrasound image as mentioned above. The labels within each cystic echinococcosis classification training sample can be the corresponding annotations for the aforementioned cystic echinococcosis CE1 to CE5. Similarly, the vesicular echinococcosis classification training dataset includes several vesicular echinococcosis classification training samples. Each vesicular echinococcosis classification training sample includes one frame of abdominal organ training ultrasound image as mentioned above. The labels within each vesicular echinococcosis classification training sample can be the corresponding annotations for the aforementioned vesicular echinococcosis AE1 to AE4.
[0075] In practice, during model training, the cross-entropy loss function can be used. The cross-entropy loss is calculated primarily by utilizing the labels of lesions located in the liver region within each training sample and the predicted labels of lesions located in the liver region within the current training sample as predicted by the RepViT model. Subsequently, the corresponding training loss can be calculated according to the cross-entropy loss calculation expression. The training samples here specifically refer to the aforementioned training samples for either cystic or vesicular echinococcosis. The specific details of the training samples depend on the object being trained and will not be elaborated upon here. Generally, when the training loss stabilizes, the target training state can be considered reached.
[0076] In one embodiment of the present invention, when the coarse classification information of the lesion is non-hepatic echinococcosis, a conventional lesion image classification model is used to perform conventional sub-category detection on the current lesion to determine the conventional sub-category of the current lesion, wherein the conventional sub-category includes at least hepatic cysts; When a lesion is routinely classified as a liver cyst, or classified as cystic echinococcosis CE1, a subclassification verification process is performed on the current lesion. When performing detailed category validation processing, the following are included: The current lesion is semantically segmented using a two-layer wall sign semantic segmentation model to generate wall sign information of the current lesion; If the wall sign information of the current lesion includes a double-wall sign, then the coarse classification information of the current lesion is identified as cystic echinococcosis, and the corresponding sub-category of echinococcosis is identified as cystic echinococcosis CE1. Otherwise, the coarse classification information of the current lesion is identified as non-hepatic echinococcosis, and the corresponding sub-category of echinococcosis is identified as hepatic cyst.
[0077] Depend on Figure 3 It can be seen that when the coarse classification information is non-hepatic echinococcosis, when it is necessary to determine the sub-category of the lesion, the constructed conventional lesion image classification model should be used to perform conventional sub-category detection on the lesion to determine the conventional sub-category of the current lesion. The conventional sub-category should include the aforementioned liver cysts, hepatic hemangiomas, intrahepatic hyperechoic lesions, etc.
[0078] It should be noted that, because some category features are similar among cystic echinococcosis and alveolar echinococcosis, using image classification models for cystic echinococcosis, alveolar echinococcosis, and conventional lesion images for corresponding sub-category detection allows the model to focus more on the feature differences between these three major categories: cystic echinococcosis, alveolar echinococcosis, and conventional benign liver lesions. This improves the accuracy and reliability of the corresponding sub-category detection. Furthermore, further subdividing the sub-categories within the coarse classification information excludes other major categories with significant differences, thereby reducing the model's workload and allowing it to focus more on subtle inter-class feature differences. This also decouples the lesion subdivision problem, reducing the difficulty of model optimization.
[0079] In practice, the RepViT model can also be used for the conventional lesion image classification model. The construction method of the conventional lesion image classification model can refer to the construction process descriptions of the cystic echinococcosis image classification model and the vesicular echinococcosis image classification model mentioned above. The difference is that the predicted labels and the real labels are different during training. The conventional lesion image classification model is trained to pay more attention to the corresponding annotations of cysts, hepatic hemangiomas, and intrahepatic hyperechoic lesions mentioned above. The specific training process will not be detailed here.
[0080] It should be noted that the differences between liver cysts and cystic echinococcosis CE1 are small, with both having low internal echoes, making them difficult to distinguish for classification models and easily confused. Therefore, when performing echinococcosis sub-category detection to generate the corresponding sub-category as cystic echinococcosis CE1, or performing conventional sub-category detection to generate the corresponding conventional sub-category as liver cysts, sub-category verification processing should be performed to validate the sub-category obtained through sub-category verification processing.
[0081] Clinically, liver cysts and cystic echinococcosis (CE1) are usually distinguished by the thickness of the lesion wall. CE1 cystic echinococcosis exhibits a "double-wall sign," meaning its wall is slightly thicker than that of a liver cyst. Therefore, based on this characteristic, during subcategorization and verification, a double-wall sign semantic segmentation model can be used to semantically segment the current lesion to generate wall sign information. This information can determine whether a region with two walls exists within the lesion and near its edge. Understandably, if the current lesion's wall sign information includes a double-wall sign, the coarse classification is identified as cystic echinococcosis, and the corresponding subcategorization is identified as CE1 cystic echinococcosis. Otherwise, the coarse classification is identified as non-hepatic echinococcosis, and the corresponding subcategorization is identified as a liver cyst.
[0082] As explained in the above description of the subcategorization verification process, when the subcategorization of echinococcosis is determined to be cystic echinococcosis CE1 through subcategorization detection, after subcategorization verification, the subcategorization of the corresponding lesion can be changed to liver cyst, or the current cystic echinococcosis CE1 can be kept unchanged. Similarly, when the conventional subcategorization detection determines that the conventional subcategorization is liver cyst, after subcategorization verification, the subcategorization of the corresponding lesion can be changed to cystic echinococcosis CE1, or the current liver cyst can be kept unchanged.
[0083] As explained above, when performing detailed category verification, a double-layer wall sign semantic segmentation model should be constructed. In one embodiment of this invention, the double-layer wall sign semantic segmentation model can be a lightweight semantic segmentation model, PPLiteSeg-STDC2. Compared with the PPLiteSeg-STDC1 model, the PPLiteSeg-STDC2 model has a deeper network structure and stronger ability to extract high-level features from some difficult samples. During inference, the double-layer wall sign semantic segmentation model can be used to segment the double-layer wall sign region of the lesion and output whether the lesion contains a double-layer wall sign.
[0084] It should be understood that when constructing a two-layer wall feature semantic segmentation model based on the lightweight semantic segmentation model PPLiteSeg-STDC2, a two-layer wall feature segmentation training dataset should be constructed. This training dataset includes several two-layer wall feature segmentation training samples and their corresponding labels. The sample labels are two-layer wall features, and the details of the two-layer wall feature labels can be found in the aforementioned explanation. The following example illustrates the method and process of constructing the two-layer wall feature segmentation training dataset: Considering that the area containing the double-walled feature accounts for a relatively small proportion of the entire lesion, typically less than 20%, image processing of the lesion region is necessary to allow the semantic segmentation model of the double-walled feature to focus more on the differential features between the two disease types. One feasible approach is to perform morphological processing along the original lesion contour (ctr_ori). First, an erosion operation is performed with a kernel size of 7*7 and 4 iterations to obtain the eroded lesion contour (ctr_erode). Then, a dilation operation is performed on the original lesion contour (ctr_ori) with a kernel size of 7*7 and 3 iterations to obtain the dilated lesion contour (ctr_dilate). Based on the eroded ctr_erode and the dilated ctr_dilate, the annular region of the lesion edge can be obtained. Using a mask, the pixel values of the image areas outside the annular region are set to 0. This allows the model to focus on the contour edge information and ignore irrelevant information in other areas.
[0085] After obtaining the annular region of the contour edge, to further focus on edge features, cropping is required along the original contour points. The cropping method is as follows: traverse the coordinate points on ctr_ori, and every K coordinate points, take a rectangular region with a width and height of L as the center. K and L will adaptively change with the image size. In the cropped image, the proportion of double-walled features is increased compared to the proportion of double-walled feature regions in the original annular region, reducing the segmentation difficulty of the semantic segmentation model. It also reduces the problem of insufficient contour segmentation caused by upsampling, thus improving the accuracy of double-walled feature recognition.
[0086] After image processing, data augmentation is required for the aforementioned abdominal organ training ultrasound images. Due to the unique characteristics of the double-walled sign, it requires special protection. Therefore, the main data augmentation methods used are: random cropping to ensure that the original annotations containing the double-walled sign are not lost in the cropped image; random scaling with a scaling factor ranging from [0.7 to 1.3] to avoid the double-walled sign losing its gradient features with the background due to excessive scaling; random rotation with no limit on the rotation angle, as the original lesion outline is annular, and rotation at various angles can greatly expand the sample size; brightness and contrast adjustment with an adjustment range controlled within ±20% to prevent the double-walled sign from being assimilated by the background due to excessive darkness or losing details due to excessive brightness; and adding random noise and color jitter to adapt to ultrasound images from different brands, models, and gain parameters.
[0087] After the above data augmentation, a two-layer wall feature segmentation training sample can be generated. The loss function used during model training can be:
[0088] in, To train the loss function, For pixel prediction loss, For cross-entropy loss, The loss function for solving the contour matching optimization problem, , , Given the corresponding loss weights, we have: .
[0089] For pixel prediction loss, we have:
[0090] in, To predict the first semantic segmentation step of the lightweight semantic segmentation model PPLiteSeg-STDC2 The probability value of the category to which each pixel belongs. For the segmentation of training samples of double-walled features, the first The pixel value of each pixel.
[0091] In specific implementation, the first The probability value of the category to which the nth pixel belongs, specifically the probability value of the nth pixel's category. The probability that a pixel belongs to the double-walled feature region. Each pixel specifically refers to the pixels within the region of the double-wall feature outline marked in the training sample for double-wall feature segmentation.
[0092] When calculating the cross-entropy loss, the bilayer features labeled with each bilayer feature segmentation training sample should be used, along with the labels predicted by the lightweight semantic segmentation model PPLiteSeg-STDC2. Then, the corresponding cross-entropy loss can be calculated according to the calculation expression of the cross-entropy loss.
[0093] For loss function It can effectively penalize extreme deviations between model predictions and true contours, for the loss function Then we have:
[0094] in, This represents the maximum distance from each point in the X-contour to the nearest point in the Y-contour. This means that the 95th percentile is used for calculation, filtering out the most extreme 5% to mitigate the impact of extreme points. Similarly, This represents the maximum distance from each point in the Y-contour to the nearest point in the X-contour, also using... calculate.
[0095] In practice, the X contour represents the contour of the double-walled sign region predicted by the lightweight semantic segmentation model PPLiteSeg-STDC2, and x represents the location point on the X contour; the Y contour is the contour of the abdominal organ training ultrasound image labeled as the double-walled sign, and y represents the location point on the Y contour.
[0096] As can be seen from the above description, the lesion should also be segmented into its outline. In one embodiment of the present invention, segmenting the lesion into its outline includes: Within the ultrasound image of the liver to be examined, the lesion detection frame of the current lesion is expanded and cropped to generate the lesion segmentation benchmark region; The lesion contour semantic segmentation model is used to perform contour semantic segmentation on the lesion segmentation benchmark region to generate the target contour of the lesion. The ultrasound image of the liver to be examined is completed and corrected based on the target contour of the lesion, so as to generate a complete liver contour after completion and correction. Based on the liver contour completion, the target contour of the lesion and the corresponding echinococcosis subcategories are mapped and output.
[0097] In practice, the method of expanding the color image can be as follows: keep the center coordinates of the lesion detection box unchanged, and increase the width and height by 30 pixels respectively. The reason for the expansion is that there may be slight deviations between the position of the lesion detection box and the actual lesion position. If the lesion detection box is directly cropped, it may result in incomplete coverage of some larger disease outlines, such as liquefied vesicular hydatids. After the expansion and cropping, a lesion segmentation benchmark area can be generated.
[0098] After generating the lesion segmentation baseline region, the lesion contour semantic segmentation model can be used to perform contour semantic segmentation on the lesion segmentation baseline region to generate the target contour of the lesion. In specific implementation, the lesion contour semantic segmentation model can be built based on the PPLiteSeg-STDC1 model. As a lightweight semantic segmentation model, the PPLiteSeg-STDC1 model can balance speed and accuracy.
[0099] It should be understood that when constructing a semantic segmentation model for lesion contours, a training dataset for lesion contour segmentation should also be constructed. The training dataset for lesion contour segmentation may include several training samples for lesion contour segmentation. Each training sample for lesion contour segmentation may include a training ultrasound image for lesion contour segmentation and a corresponding sample label. The aforementioned training ultrasound image for lesion contour segmentation may be a frame of training ultrasound image of an abdominal organ mentioned above. The sample label may be the contour of the standard lesion corresponding to the current training ultrasound image of the abdominal organ. That is, at this time, the type of lesion is no longer distinguished, and only the lesion and its corresponding contour are considered.
[0100] During model training, the loss function used can be the same as that used in constructing the two-layer lesion semantic segmentation model described above. The calculation of the loss function can be referenced in the above explanation. The difference is that when calculating the loss function, the predicted contour is the contour of the lesion; all other parameters can be calculated using the same explanation. When the loss function trend stabilizes, the target training state can be considered reached, and the required lesion contour semantic segmentation model can then be constructed.
[0101] For lesions close to the liver edge or even partially protruding from the liver capsule, the liver contour segmentation may be inadequate, as the liver contour may not completely encompass the lesion's segmented contour. Therefore, it is necessary to complete the liver contour. In one embodiment of the present invention, when completing and correcting the ultrasound image of the liver to be examined based on the target contour of the lesion, the process includes: The liver image to be examined and the target contour of the lesion are placed on the same mask. Then, the maximum contour search method is used to find the maximum connected region, and the found maximum connected region is configured as the liver completion contour.
[0102] In practice, the mask is an image with a pixel value of 0, and the length and width of the mask are consistent with the length and width of the liver image to be examined. It should be noted that the method of finding the maximum connected region using the maximum contour search method is consistent with existing technology and will not be elaborated here. Figure 4 An example of an abdominal ultrasound scan is shown in the figure. Figure 5This image illustrates an embodiment of displaying the complete liver contour, the target contour of the lesion, and the sub-classification of the hydatid cyst on an abdominal ultrasound scan. In the image, "Liver" represents the liver, the outer contour is the complete liver contour, and the inner circular contour is the target contour of the lesion. The hydatid cyst sub-classification of the lesion is CE2, which is cystic hydatid cyst CE2. Furthermore, the image also provides a detailed explanation of CE2, namely, multi-daughter cystic cysts. Other mapping outputs can be found here.
[0103] It should be noted that during liver region segmentation and identification, the position coordinates of the liver to be examined in the abdominal ultrasound image can be determined simultaneously. At the same time, during lesion target detection, the position coordinates of each lesion in the liver to be examined in the ultrasound image can also be determined. During mapping output, the corresponding position coordinates can be accurately mapped into the abdominal ultrasound image. The specific method of determining the position coordinates can be consistent with the existing technology, and will not be elaborated here.
[0104] In one embodiment of the present invention, when multiple consecutive ultrasound images of the liver to be examined exist, the method further includes lesion tracking processing, wherein... When performing lesion tracking and treatment, the following are included: For any two adjacent time-series liver ultrasound images, the lesion detection box of each lesion in the previous time-series liver ultrasound image is used as the basic detection box, and the lesion detection box of each lesion in the next time-series liver ultrasound image is used as the tracking detection box. Motion prediction is performed on each basic detection box to generate a trajectory prediction box, and the similarity between the trajectory prediction box and each tracking detection box is calculated. Based on the calculated similarity, the Hungarian matching status of the trajectory prediction box and all tracking detection boxes is calculated, where, When a tracking detection box and a trajectory prediction box satisfy a Hungarian match, the current trajectory prediction box is configured as the tracking target box. The tracking target box and the corresponding tracking detection box are deduplicated, and the corresponding lesion target box is generated. The generated lesion target box is then configured as the lesion detection box for the corresponding lesion.
[0105] During abdominal ultrasound scans, multiple lesions often appear in the image. Due to the constantly changing position and angle of the probe, these lesions may shift, deform, or even be temporarily obscured. This phenomenon can lead to the loss of the original index of the lesions, and in severe cases, complete loss of lesion tracking. To minimize such problems, this invention performs lesion tracking processing, continuously tracking the dynamic changes of each lesion. It should be understood that lesion tracking processing should be performed on multiple consecutive frames of the liver ultrasound image to be examined. For example, in abdominal ultrasound scans, there is abdominal ultrasound data in video form. Subsequently, based on sampling frequency and other factors, continuous abdominal ultrasound images are formed. After liver region segmentation and recognition processing, multiple consecutive frames of the liver ultrasound image to be examined can be generated.
[0106] When performing lesion tracking, two temporally adjacent liver ultrasound images should be used. Temporally adjacent specifically refers to two liver ultrasound images generated sequentially in time. As explained above, each liver ultrasound image may contain one or more lesions. After lesion target detection, a lesion detection box and corresponding confidence score can be generated for each lesion. The confidence score represents the probability that the lesion belongs to the corresponding category. The meaning of the confidence score is consistent with existing technologies and will not be elaborated here.
[0107] It should be understood that during lesion tracking, the lesion detection box of each lesion in the previous frame of the liver ultrasound image to be examined is used as the basic detection box, and the lesion detection box of the corresponding lesion in the next frame of the liver ultrasound image to be examined is used as the tracking detection box. In the previous frame sequence and the next frame sequence, the lesion target detection is performed in the previous frame sequence before the next frame sequence.
[0108] Motion prediction is performed on each basic detection box, such as using a Kalman filter, to generate a corresponding trajectory prediction box. Then, the similarity between the trajectory prediction box and each tracking detection box is calculated. Specifically, the intersection-union ratio (IUU) between the trajectory prediction box and the tracking detection box can be used as the corresponding similarity. Based on the calculated similarity, the Hungarian matching state between the trajectory prediction box and all tracking detection boxes is calculated. It should be noted that the method for calculating the Hungarian matching state based on similarity can be consistent with existing techniques.
[0109] In practice, when a tracking detection box and a trajectory prediction box satisfy a Hungarian match, the current trajectory prediction box is configured as the tracking target box. This indicates that the same lesion exists in both the previous and subsequent frames, and the lesion is the lesion corresponding to the trajectory prediction box. It is understood that when no tracking detection box and a trajectory prediction box satisfy a Hungarian match, it indicates that there are no identical lesions in the two frames of liver ultrasound images.
[0110] Understandably, when Hungarian matching is satisfied, there will be a tracking target box and a corresponding tracking detection box. In order to prevent multiple detection boxes with close positions from appearing for the same lesion, the tracking target box and the corresponding tracking detection box should be deduplicated, and the corresponding lesion target box should be generated after deduplication. The generated lesion target box should be configured as the lesion detection box for the corresponding lesion.
[0111] During deduplication, a feasible approach is to calculate the intersection-union ratio (IUR) between the target bounding box and the detection bounding box. If the IUR is greater than the threshold thresh, the target bounding box is deleted, and the detection bounding box is used as the lesion detection box. Otherwise, the detection bounding box is deleted, and the target bounding box is used as the lesion detection box. In practice, the threshold thresh can be 0.7.
[0112] When a basic detection bounding box fails to find a tracking bounding box that satisfies a Hungarian match, it can continue tracking within subsequent frames of the liver ultrasound image to be examined. For example, if the previous frame is the first frame, and the same lesion is not tracked in the second frame, the same lesion can be tracked in the corresponding third and fourth frames of the liver ultrasound image to be examined. As the actual position of the probe changes significantly, it may lead to the tracking of some false positive lesion bounding boxes that do not actually exist. Therefore, a maximum waiting frame length `max_age` is set for long-term unmatched trajectories. For example, the maximum waiting frame length `max_age` can be 15. Of course, it can also be selected as needed, and examples will not be given here.
[0113] To improve the efficiency of lesion tracking, lesions can be stratified. The stratification method is based on the confidence score (Score) of the detected lesion bounding box. If the confidence score is greater than or equal to S_high, the lesion bounding box is assigned to the high-confidence set; if the confidence score is greater than or equal to S_low but less than S_high, the lesion bounding box is assigned to the low-confidence set. Typically, S_high is set in the range [0.5, 0.7], and S_low in the range [0.1, 0.3]. Since this is a medical scenario, low-confidence bounding boxes may be noise or normal human tissue structures. Therefore, a higher confidence score is required for the disease to minimize false positives. Thus, the high-score threshold S_high can be set to 0.7, and the low-score threshold S_low to 0.5.
[0114] When tracking lesions, the basic bounding box is first compared with the tracking bounding boxes in the high-confidence set using the similarity calculation and Hungarian matching process described above. If a tracking bounding box in the high-confidence set satisfies a Hungarian match, tracking of lesion bounding boxes in the low-confidence set can be stopped. If no tracking bounding box in the high-confidence set satisfies a Hungarian match, tracking of lesion bounding boxes in the low-confidence set can proceed.
Claims
1. A deep learning-based liver hydatid lesion identification and classification method, characterized in that, The liver hydatid lesion recognition classification method comprises: An ultrasound image of a liver to be detected is acquired, and lesion target detection is performed on the ultrasound image of the liver to be detected, wherein, When a lesion exists in the ultrasound image of the liver to be detected, lesion detection state information of each lesion in the ultrasound image of the liver to be detected is generated, and the lesion detection state information at least comprises a lesion detection frame of the lesion and a pathological type representing a pathological feature of the lesion; For any lesion, when the pathological type of the lesion is benign, coarse classification detection is performed on the lesion in the lesion detection frame to generate coarse classification information of the lesion, wherein the coarse classification information comprises a cystic hydatid type, a vesicular hydatid type or a non-liver hydatid type; For any lesion, when the coarse classification information of the lesion is the cystic hydatid type or the vesicular hydatid type, the lesion is determined to be a liver hydatid lesion, and hydatid fine classification detection and lesion contour segmentation are performed on the lesion to determine the corresponding hydatid fine classification when the lesion belongs to the cystic hydatid type or the vesicular hydatid type, and the target contour of the lesion. 2.The deep learning-based liver hydatid lesion identification and classification method according to claim 1, characterized in that, When acquiring the ultrasound image of the liver to be detected, the method comprises: An abdominal scan ultrasound image generated by abdominal ultrasound scanning is acquired, and liver region segmentation recognition is performed on the abdominal scan ultrasound image; When the liver region is included in the abdominal scan ultrasound image, liver recognition state information corresponding to the liver region is generated, wherein the liver recognition state information at least comprises a liver region detection frame of the liver region in the abdominal scan ultrasound image; Based on the liver recognition state information, liver region extraction processing is performed on the abdominal scan ultrasound image to generate the ultrasound image of the liver to be detected after the liver region extraction processing, wherein, The liver region extraction processing at least comprises image cropping processing and image denoising processing performed in sequence; After the liver region extraction processing, only the liver region is included in the ultrasound image of the liver to be detected. 3.The deep learning-based liver hydatid lesion identification and classification method according to claim 1, characterized in that, When performing lesion target detection on the ultrasound image of the liver to be detected, the method comprises: The ultrasound image of the liver to be detected is loaded into a lesion target detection model to perform lesion target detection on the ultrasound image of the liver to be detected by using the lesion target detection model, wherein, When a lesion exists in the ultrasound image of the liver to be detected and the lesion is a liver cancer type, the pathological type of the current lesion is identified as malignant, otherwise, the pathological type of the current lesion is identified as benign. 4.The deep learning-based liver hydatid lesion identification and classification method according to claim 3, characterized in that, When constructing the lesion target detection model, the method comprises: A YOLOv10 model is provided, and an EfficientFormerV2 is used to improve at least a backbone network of the YOLOv10 model to form a target detection backbone network in the lesion target detection model, wherein The target detection backbone network comprises a Stem Block module, a first Local Block module, a first Subsample module, a second Local Block module, a second Subsample module, a third Local Block module, a third Subsample module, a Local Golbal Block module, a SPPF Block module and a PSA Block module connected in sequence. The second Local Block module, the third Local Block module, and the PSA Block module are connected to the target detection neck network of the lesion target detection model. 5.The deep learning-based liver hydatid lesion identification and classification method according to claim 1, characterized in that, When constructing the lesion target detection model, the neck network in the YOLOv10 model is improved by using the EfficientFormerV2 to form the target detection neck network of the lesion target detection model, wherein, When the neck network in the YOLOv10 model is improved by using the EfficientFormerV2, the C2f module in the neck network is replaced by the Local Block module in the EfficientFormerV2. 6.The deep learning-based liver hydatid lesion identification and classification method according to claim 1, characterized in that, When the pathological type of a lesion is benign, a benign lesion image classification model is used to perform coarse classification detection on the current lesion to determine coarse classification information of the lesion; When the coarse classification information of the lesion is cystic echinococcosis, a cystic echinococcosis image classification model is used to perform echinococcosis fine classification detection on the current lesion to determine the echinococcosis fine classification of the current lesion, wherein the echinococcosis fine classification is one of cystic echinococcosis CE1-cystic echinococcosis CE5; When the coarse classification information of the lesion is bubble-type echinococcosis, a bubble-type echinococcosis image classification model is used to perform echinococcosis fine classification detection on the current lesion to determine the echinococcosis fine classification of the current lesion, wherein the echinococcosis fine classification is one of bubble-type echinococcosis AE1-bubble-type echinococcosis AE4. 7.The deep learning-based liver hydatid lesion identification and classification method according to claim 6, characterized in that, When the coarse classification information of the lesion is non-liver echinococcosis, a conventional lesion image classification model is used to perform conventional fine classification detection on the current lesion to determine the conventional fine classification of the current lesion, wherein the conventional fine classification at least includes liver cysts; When the conventional fine classification of a lesion is liver cysts, or the echinococcosis fine classification is cystic echinococcosis CE1, fine classification verification processing is performed on the current lesion, wherein, When performing fine classification verification processing, it includes: A double-layer wall sign semantic segmentation model is used to perform semantic segmentation on the current lesion to generate wall sign information of the current lesion; If the wall sign information of the current lesion contains a double-layer wall sign state, the coarse classification information of the current lesion is identified as cystic echinococcosis, and the corresponding echinococcosis fine classification is identified as cystic echinococcosis CE1, otherwise, the coarse classification information of the current lesion is identified as non-liver echinococcosis, and the corresponding echinococcosis fine classification is identified as liver cysts. 8.The deep learning-based liver hydatid lesion identification and classification method according to any one of claims 1 to 7, characterized in that, When performing lesion contour segmentation on the lesion, it includes: In the to-be-detected liver ultrasound image, based on the lesion detection frame of the current lesion, an external expansion crop is performed to generate a lesion segmentation reference area; A lesion contour semantic segmentation model is used to perform contour semantic segmentation on the lesion segmentation reference area to generate a target contour of the lesion; Based on the target contour of the lesion, the to-be-detected liver ultrasound image is completed and corrected to generate a liver completion contour; Based on the liver completion contour, the target contour of the lesion and the corresponding echinococcosis fine classification are mapped and output. 9.The deep learning-based liver hydatid lesion identification and classification method according to claim 8, characterized in that, When completing and correcting the to-be-detected liver ultrasound image based on the target contour of the lesion, it includes: The image of the liver to be detected and the target contour of the lesion are placed on the same mask, and then a maximum contour search method is used to search for a maximum connected region, and the found maximum connected region is configured as a liver completion contour.
10. The deep learning-based liver hydatid lesion identification and classification method according to any one of claims 1 to 7, characterized in that, when When there are multiple frames of continuous liver ultrasound images to be detected, the method further includes lesion tracking processing on the existing lesions, wherein, When the lesion tracking processing is performed, the method includes: For any two adjacent time sequence frames of the liver ultrasound images to be detected, the lesion detection frame of each lesion in the previous frame of the liver ultrasound images to be detected is taken as a basic detection frame, and the lesion detection frame of each lesion in the next frame of the liver ultrasound images to be detected is taken as a tracking detection frame; Motion prediction is performed on each basic detection frame to generate a trajectory prediction frame, and the similarity between the trajectory prediction frame and each tracking detection frame is calculated; Based on the calculated similarity, the Hungarian matching state of the trajectory prediction frame and all tracking detection frames is calculated, wherein, When there is one tracking detection frame that satisfies the Hungarian matching with the trajectory prediction frame, the current trajectory prediction frame is configured as a tracking target frame; The tracking target frame and the corresponding tracking detection frame are de-duplicated to generate a corresponding lesion target frame, and the generated lesion target frame is configured as the lesion detection frame of the corresponding lesion.