Lesion Size Evaluation Method, Device, Electronic Device, Storage Medium and Computer Program Product

By using lesion recognition models in capsule endoscopy, combining depth images and edge profile images to accurately evaluate lesion size, the problem that two-dimensional images in the prior art is difficult to accurately measure lesion size.

CN119295533BActive Publication Date: 2025-06-10GUANGZHOU SIDE MEDICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411817367.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-06-10
Estimated Expiration
2044-12-11

Smart Images

  • Figure CN119295533B_ABST
    Figure CN119295533B_ABST
Patent Text Reader

Abstract

The present application relates to the field of image processing technology, and provides a method, device, electronic device, storage medium, and computer program product for evaluating the size of a lesion. The method includes: inputting an image to be evaluated into a lesion recognition model to obtain a lesion depth image and a lesion edge contour image output by the lesion recognition model; evaluating the size of the lesion based on the lesion depth image and the lesion edge contour image; wherein, the lesion recognition model is obtained by a teacher model through knowledge distillation based on unlabeled sample data with pseudo-labels and labeled sample data; the teacher model is trained based on a synthetic image and a label corresponding to the synthetic image; the label corresponding to the synthetic image includes a lesion depth annotation image and a lesion edge contour annotation image; the pseudo-labels are generated by the teacher model based on the unlabeled sample data. The present application can accurately evaluate the size of the lesion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and particularly relates to a method, device, electronic device, storage medium, and computer program product for evaluating the size of a lesion. Background Art

[0002] As a non-invasive and anesthesia-free endoscopic examination technique, capsule endoscopy is commonly used to detect various lesions in the digestive tract. Currently, the images captured by capsule endoscopes are still two-dimensional images, and doctors need to identify the positions of the two-dimensional images in the stomach based on their knowledge and experience.

[0003] The development of artificial intelligence-assisted recognition technology for capsule endoscopes has been a major breakthrough in recent years. However, although significant progress has been made in lesion detection technology, there are still relatively few related studies on measuring the size of lesions under endoscopy. Accurate measurement of the size of lesions is of great significance for evaluating the condition, formulating treatment plans, and monitoring treatment effects. However, existing capsule endoscopes mainly use monocular cameras, and the captured images are all two-dimensional images, making it difficult to accurately evaluate the size of lesions. Summary of the Invention

[0004] The present application aims to solve at least one of the technical problems existing in the related art. For this purpose, the present application provides a method, device, electronic device, storage medium, and computer program product for evaluating the size of a lesion, so as to solve the problem that existing capsule endoscopes mainly use monocular cameras, and the captured images are all two-dimensional images, making it difficult to accurately evaluate the size of lesions, and to accurately evaluate the size of lesions.

[0005] According to an embodiment of the first aspect of the present application, the method for evaluating the size of a lesion includes:

[0006] Inputting the image to be evaluated into a lesion recognition model to obtain a lesion depth image and a lesion edge contour image output by the lesion recognition model;

[0007] Evaluating the size of the lesion based on the lesion depth image and the lesion edge contour image;

[0008] Wherein, the lesion recognition model is obtained by a teacher model through knowledge distillation based on unlabeled sample data with pseudo-labels and labeled sample data; the teacher model is trained based on synthetic images and corresponding labels of the synthetic images; the corresponding labels of the synthetic images include a lesion depth annotation image and a lesion edge contour annotation image; the pseudo-labels are generated by the teacher model based on the unlabeled sample data.

[0009] According to an embodiment of the present application, the evaluating the size of the lesion based on the lesion depth image and the lesion edge contour image includes:

[0010] Determine the size conversion coefficient between the microscopic size and the actual size;

[0011] Determine the parallax value of each pixel point in the depth image of the lesion;

[0012] Based on each of the parallax values and the size conversion coefficient, determine the actual physical coordinates of each pixel point of the lesion edge contour in the lesion edge contour image;

[0013] Evaluate the size of the lesion according to the actual physical coordinates of each pixel point in the lesion edge contour.

[0014] According to an embodiment of the present application, before knowledge distillation based on unlabeled sample data with pseudo-labels and labeled sample data, it further includes:

[0015] Perform data augmentation processing on the unlabeled sample data with pseudo-labels.

[0016] According to an embodiment of the present application, the synthetic image is generated based on a preset three-dimensional physical model; the preset three-dimensional physical model is constructed based on a virtual image engine.

[0017] According to an embodiment of the present application, the loss function of the teacher model includes depth estimation loss, image segmentation loss, and affine invariant loss.

[0018] According to an embodiment of the present application, the loss function of the lesion recognition model includes depth estimation loss and image segmentation loss.

[0019] According to the lesion size evaluation device of the second aspect embodiment of the present application, it includes:

[0020] An identification module, configured to input an image to be evaluated into a lesion recognition model, and obtain a lesion depth image and a lesion edge contour image output by the lesion recognition model;

[0021] A determination module, configured to evaluate the size of the lesion based on the lesion depth image and the lesion edge contour image;

[0022] Wherein, the lesion recognition model is obtained by a teacher model performing knowledge distillation based on unlabeled sample data with pseudo-labels and labeled sample data; the teacher model is trained based on a synthetic image and the corresponding label of the synthetic image; the corresponding label of the synthetic image includes a lesion depth annotation image and a lesion edge contour annotation image; the pseudo-label is generated by the teacher model based on the unlabeled sample data.

[0023] An electronic device according to an embodiment of the third aspect of the present application includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for evaluating the lesion size as described in any one of the above is implemented.

[0024] A storage medium according to an embodiment of the fourth aspect of the present application is a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for evaluating the lesion size as described in any one of the above is implemented.

[0025] A computer program product according to an embodiment of the fifth aspect of the present application includes a computer program. When the computer program is executed by a processor, the method for evaluating the lesion size as described in any one of the above is implemented.

[0026] One or more of the above technical solutions in the embodiments of the present application have at least the following technical effects:

[0027] A teacher model is trained in advance based on a synthetic image, a lesion depth annotation image corresponding to the synthetic image, and a lesion edge contour annotation image. After generating pseudo-labels based on unlabeled sample data through the teacher model, a lesion recognition model is obtained through knowledge distillation of the unlabeled sample data with pseudo-labels and labeled sample data by the teacher model. Therefore, when lesion recognition is required, the image to be evaluated is input into the lesion recognition model, and a lesion depth image and a lesion edge contour image output by the lesion recognition model can be obtained. Thus, the lesion size can be accurately evaluated through the lesion depth image and the lesion edge contour image.

[0028] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Description of the Drawings

[0029] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0030] Figure 1 It is a schematic flowchart of the method for evaluating the lesion size provided by the embodiment of the present application.

[0031] Figure 2 It is a schematic diagram of multi-task learning framework and knowledge distillation in the method for evaluating the lesion size provided by the embodiment of the present application.

[0032] Figure 3It is a schematic diagram of the pseudo-label generation and self-training process in the lesion size evaluation method provided by the embodiments of the present application.

[0033] Figure 4 It is a schematic diagram of the input and output structures of lesion recognition in the lesion size evaluation method provided by the embodiments of the present application.

[0034] Figure 5 It is a schematic diagram of the lesion size evaluation device provided by the embodiments of the present application.

[0035] Figure 6 It is a schematic diagram of the structure of the electronic device provided by the present application. Detailed implementation manners

[0036] The following further describes the implementation manners of the present application in detail with reference to the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but cannot be used to limit the scope of the present application.

[0037] In the description of the embodiments of the present application, it should be noted that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the embodiments of the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the embodiments of the present application. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0038] In the description of the embodiments of the present application, it should be noted that unless otherwise clearly specified and limited, the terms "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific situations.

[0039] In the embodiments of the present application, unless otherwise clearly specified and limited, the first feature being "on" or "under" the second feature can be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on" the second feature can be that the first feature is directly above or obliquely above the second feature, or simply means that the first feature has a higher horizontal height than the second feature. The first feature being "under", "below" and "beneath" the second feature can be that the first feature is directly below or obliquely below the second feature, or simply means that the first feature has a lower horizontal height than the second feature.

[0040] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0041] It should be noted that traditional lesion detection and size measurement techniques mainly rely on two-dimensional images, and identify the lesion area and measure the size through image processing and analysis algorithms. These techniques are widely used in medical imaging, but have the following limitations:

[0042] 1. Lack of depth information: Due to the limitations of two-dimensional images, it is difficult to accurately obtain the three-dimensional information of the lesion, resulting in errors in size measurement. Many existing techniques rely on planar images for measurement and lack effective calibration of the depth of the lesion. The capsule endoscope moves in the digestive tract and is limited by the narrow and curved space, making it difficult to obtain sufficient perspectives for depth reconstruction.

[0043] 2. Influence of illumination and texture: The illumination changes and tissue texture complexity in the image will affect the accuracy of lesion detection and measurement, resulting in false detections and missed detections. The quality of endoscopic images is limited and is often affected by the interference of light source intensity and shooting angle.

[0044] 3. Data dependence: Lesion detection and measurement algorithms usually require a large amount of high-quality labeled data for training, and it is difficult to obtain such data in the clinical environment. The image quality is limited by factors such as the cleanliness of the stomach, the capsule hardware configuration, and the image compression ratio, resulting in a lot of noise in the data set, which affects the training results.

[0045] 4. Computational complexity and real-time performance: The process of multi-view imaging and three-dimensional reconstruction is complex and computationally intensive, making it difficult to perform real-time three-dimensional reconstruction processing. During the real-time detection process, the limitations of processing speed and computing power make it difficult for the system to provide accurate lesion detection and size measurement results in a short time.

[0046] Based on this, this application proposes a method, device, electronic device, storage medium, and computer program product for evaluating lesion size.

[0047] Figure 1 It is a schematic flowchart of the method for evaluating lesion size provided by the embodiments of this application, asFigure 1 As shown in Figure 1 , the method for evaluating the size of a lesion includes:

[0048] Step 110: Input the image to be evaluated into the lesion recognition model to obtain the lesion depth image and the lesion edge contour image output by the lesion recognition model.

[0049] Step 120: Evaluate the size of the lesion based on the lesion depth image and the lesion edge contour image.

[0050] Among them, the lesion recognition model is obtained by a teacher model through knowledge distillation based on unlabeled sample data with pseudo-labels and labeled sample data; the teacher model is trained based on synthetic images and the corresponding labels of the synthetic images; the corresponding labels of the synthetic images include the lesion depth annotation image and the lesion edge contour annotation image; the pseudo-labels are generated by the teacher model based on unlabeled sample data.

[0051] Furthermore, the synthetic images are generated based on a preset three-dimensional physical model; the preset three-dimensional physical model is constructed based on a virtual image engine.

[0052] The loss function of the teacher model includes depth estimation loss, image segmentation loss, and affine invariant loss.

[0053] The loss function of the lesion recognition model includes depth estimation loss and image segmentation loss.

[0054] It should be noted that the execution subject of the method for evaluating the size of a lesion provided in the embodiments of the present application can be a server, a computer device, etc., such as a mobile phone, a tablet computer, a notebook computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an Ultra-mobile Personal Computer (UMPC), a netbook, or a Personal Digital Assistant (PDA), etc.

[0055] A lesion size evaluation device can be set or connected in the server or computer device of the present application, so as to control the lesion size evaluation device to execute the method for evaluating the size of a lesion of the present application.

[0056] The present application can use a virtual image engine such as Unreal Engine to construct three-dimensional models of objects such as the esophagus and gastric cavity, and further generate high-quality synthetic images with accurate depth information according to the three-dimensional models. These synthetic images can make up for the deficiencies of real medical image data, provide a large amount of high-quality data for the training of depth estimation, and in addition, can reduce the image noise in real images caused by differences in image quality, capsule hardware, or compression ratio.

[0057] Specifically, this application can use Unreal Engine to build a 3D model and simulate the anatomical structures of the esophagus and gastric cavity. The material parameters are set using a physically based rendering model, including reflectivity, roughness, and subsurface scattering, etc.

[0058] Use the Depth Pass rendering function of Unreal Engine to export the depth map, ensuring that the value range of the depth map is proportional to the actual distance.

[0059] Normalize the value of the depth map to range, and the calculation formula is:

[0060] ;

[0061] where, is the original depth value, and are the minimum and maximum values of the depth value respectively.

[0062] Therefore, when constructing the dataset, a virtual engine (such as Unreal Engine) can be used to build a 3D model, simulate the esophagus and gastric cavity structures, generate high-quality color images and depth maps, and use these images to train models for depth estimation and lesion segmentation.

[0063] Furthermore, the synthetic images can be information-annotated manually or by other means to generate labels for the synthetic images. Among them, the information annotation can include lesion depth annotation and lesion edge contour annotation. Therefore, the labels of the synthetic images can include lesion depth annotation images and lesion edge contour annotation images.

[0064] At the same time, a complex and powerful model architecture (such as Vision Transformer or ResNet-50) can be selected, the encoder is initialized with DINOv2 pre-trained weights, and the model is trained using the dataset composed of synthetic images and the corresponding labels of the synthetic images to generate a teacher model that can accurately identify the lesion depth and accurately segment the lesion. Vision Transformer is a deep learning model specifically designed for computer vision tasks. ResNet-50 is a convolutional neural network and is a model in the residual network series. DINOv2 is a method for training computer vision models using self-supervised learning.

[0065] Specifically, during the pre-training process of the teacher model, a pre-trained semantic segmentation model is applied to detect the farthest area in the gastric image under a wide-angle view. The specific approach can be to set the disparity value of the segmented farthest area to as the reference benchmark for depth estimation.

[0066] The loss function of the teacher model is as follows:

[0067] ;

[0068] where, is the depth estimation loss, is the image segmentation loss, is the affine-invariant loss, are the weight coefficients of the corresponding losses respectively;

[0069] The affine-invariant loss is specifically:

[0070] ;

[0071] where,

[0072] ;

[0073] ;

[0074] ;

[0075] ;

[0076] HW represents the total number of pixels in the image, which is the product of the image width (W) and height (H). This is used to normalize the loss value so that the calculation of the loss is not affected by the image resolution. The purpose of doing this is to obtain a consistent loss evaluation for images of different sizes. represents the true depth value of the i-th pixel in the image, that is, the target depth value (Ground Truth Depth) of this pixel. represents the depth value predicted by the model for the i-th pixel (Predicted Depth). In the calculation of the loss, by comparing and the difference, the deviation between the model prediction and the true value is calculated, which is used to optimize the depth prediction ability of the model.

[0077] Furthermore, this application can use the trained teacher model to generate pseudo-labels for unlabeled real medical images. These pseudo-labels are used as additional supervision information for the further training of the student model.

[0078] Furthermore, this application can select a deep convolutional neural network model as the student model, such as selecting EfficientNet-B2. EfficientNet-B2 is a deep convolutional neural network model mainly constructed based on a compound scaling strategy.

[0079] Furthermore, the unlabeled data with pseudo-labels (i.e., the above-mentioned unlabeled real medical images with pseudo-labels) is mixed with the labeled data (synthetic images and their labels) to construct an enhanced training set, and self-training of the student model is carried out through knowledge distillation technology. Through multiple iterations of self-training, the generalization ability and robustness of the student model are continuously optimized, and finally a lesion recognition model that can output a lesion depth image (subsequently abbreviated as depth map) and a lesion edge contour image (subsequently abbreviated as edge contour map) according to the input image is obtained.

[0080] Specifically, when performing knowledge distillation, the distillation loss function can be:

[0081] ;

[0082] where and are the outputs of the teacher model and the student model respectively, is the softmax function, is the distillation temperature, is the loss weight, is the loss of the hard label.

[0083] More specifically, and are the output logits of the teacher model and the student model respectively;

[0084] denotes the softmax function, which is used to convert logits into a probability distribution;

[0085] is the distillation temperature, which is used to smooth the logits so that the model can learn more subtle probability distributions from the soft labels of the teacher model. Usually .

[0086] is the loss weight, which determines the contribution ratio of the soft label (distillation part) and the hard label (true label) to the total loss. The value of can be adjusted according to the specific task;

[0087] KL is the divergence, is a function that measures the difference between two probability distributions and :

[0088] ;

[0089] During the distillation process, .

[0090] It should be noted that during the training process, the data augmentation technique CutMix is used.

[0091] The specific implementation is as follows: The disparity space of the real image with pseudo-labels is regionally mixed according to a randomly generated binary mask, and then the affine-invariant loss is calculated:

[0092] ;

[0093] Among them,

[0094] are the output logits of the student model;

[0095] and are the output logits of the teacher model respectively;

[0096] is the binary mask, which is used to mix different image regions;

[0097] is , usually Smooth L1 Loss, which is used to measure the difference between the outputs of the student model and the teacher model; Smooth L1 Loss is a loss function;

[0098] is the spatial resolution of the image (i.e., the product of height and width), which is used to normalize the loss.

[0099] Figure 2 is the schematic diagram of multi-task learning framework and knowledge distillation in the lesion size assessment method provided by the embodiments of the present application. As Figure 2 shown, the present application provides a multi-task learning framework: It supports both lesion segmentation and depth estimation tasks simultaneously. A dual-decoder structure is used to process depth map generation and lesion area segmentation respectively to ensure efficiency and accuracy in the multi-task scenario.

[0100] Use the lightweight EfficientNet-B2 trained by knowledge distillation technology as the shared encoder. After knowledge distillation, EfficientNet-B2 has the ability to inherit depth estimation and lesion segmentation from the teacher model, and is more efficient in terms of inference speed and model size. EfficientNet-B2 is initialized with the weights distilled from the teacher model. These weights combine the efficient feature extraction ability of the teacher model and are also optimized in terms of computational efficiency.

[0101] Dual decoder architecture design: It includes a depth estimation decoder and a lesion segmentation decoder. The depth estimation decoder includes: Multi-scale upsampling: From the multi-scale feature maps extracted and output by the EfficientNet-B2 encoder, gradually perform upsampling through the task execution path to restore the spatial resolution of the image, generate and output a depth map; Skip connections: During the upsampling process, use skip connections to directly transfer the high-resolution features in the EfficientNet-B2 encoder to the decoder to improve the detail performance of depth estimation; Loss function: Use Smooth L1 Loss to optimize depth prediction:

[0102] ;

[0103] where, is the depth prediction generated by the student model, is the depth in the pseudo-label or real annotation data.

[0104] Thus, the depth estimation decoder can complete the depth estimation task and output the lesion depth image D1 for 3D reconstruction and lesion size estimation.

[0105] The lesion segmentation decoder includes: Convolution and transposed convolution layers: From the multi-scale feature maps extracted and output by the EfficientNet-B2 encoder, gradually restore the spatial resolution of the image through convolution and transposed convolution operations in the task execution path, generate and output a lesion segmentation map; Context attention mechanism: Introduce a context attention mechanism to enable the model to better focus on the lesion area during segmentation, reducing missed detections and false detections; Loss function: Use Dice Loss to optimize lesion segmentation. Dice Loss is a loss function commonly used in image segmentation tasks:

[0106] ;

[0107] where, is the segmentation map generated by the student model, is the segmentation label in the real annotation data.

[0108] Thus, the lesion segmentation decoder can complete the lesion segmentation task and output the lesion edge contour image E1 for lesion area detection and diagnosis.

[0109] Multi-task joint loss function: To effectively train the model when performing depth estimation and lesion segmentation simultaneously, design a multi-task joint loss function:

[0110] ;

[0111] where, and They are the loss weights for depth estimation and lesion segmentation respectively, and the initial recommended values are , which can be adjusted according to the experimental results later.

[0112] Furthermore, it should be noted that Dropout and weight decay (L2 regularization) can be added during model training to prevent model overfitting. Dropout is a regularization technique used to prevent neural network overfitting.

[0113] Use the gradient clipping technique to prevent the gradient explosion problem, and the maximum gradient norm can be set to 1.0.

[0114] Learning rate and optimizer: Use the learning rate scheduler CosineAnnealingLR to dynamically adjust the learning rate during training. The initial learning rate can be set to (encoder part), while a higher learning rate (such as 10 times the encoder learning rate) can be used for the decoder part.

[0115] Optimizer selection: Use AdamW as the optimizer, combined with L2 regularization to enhance the generalization ability of the model. AdamW is an optimizer.

[0116] Model validation and deployment: (1) Model validation: Conduct cross-validation on multiple datasets to evaluate the performance of the student model in depth estimation and lesion segmentation, and use metrics such as MAE, RMSE, IoU, and Dice Coefficient for evaluation. MAE, RMSE, IoU, and Dice Coefficient are commonly used evaluation metrics and are applicable to different types of models and tasks.

[0117] Real-time inference and deployment: Deploy the trained EfficientNet-B2 in a computing environment that supports multi-task parallel processing to ensure efficient operation in multi-task scenarios (such as processing multiple images simultaneously).

[0118] Use tools such as TensorFlow Serving or PyTorch Serve to deploy the model inference service to ensure the response speed and accuracy in clinical applications. TensorFlow Serving and PyTorch Serve are tools for deploying machine learning models.

[0119] Model testing and validation: Validate using multiple performance metrics on real clinical images: precision, recall, F1-score, etc. F1-score is a metric used to evaluate the performance of binary classification models, which comprehensively considers the precision and recall of the model.

[0120] Real-time test: Conduct inference speed tests to ensure that the response time of the model in actual applications can meet real-time requirements while ensuring the accuracy of inference results.

[0121] The teacher model can perform knowledge distillation based on a shared encoder. Specifically, according to the knowledge distillation path, knowledge transfer can be carried out between the pseudo-labels and the student model, thereby completing the self-training of the student model.

[0122] Figure 3 It is a schematic diagram of the pseudo-label generation and self-training process in the lesion size assessment method provided by the embodiments of the present application. As Figure 3 shown, after the input of unannotated medical images, the present application can generate pseudo-labels as initial pseudo-labels through the prediction of the teacher model, and at the same time obtain manually annotated data. Furthermore, the pseudo-label data and the annotated data can be mixed, and the merged dataset can be used for the training of the student model. Specifically, self-training iterations can be performed to achieve self-training output. If the optimized student model does not meet the requirements, it continues to be optimized; when the optimized student model meets the requirements, the optimized student model can be deployed, and a self-training loop can be carried out during the self-training process to optimize the model in each iteration.

[0123] On this basis, the present application can use the images obtained by the capsule endoscope device as the images to be evaluated, and input them into the lesion recognition model. The lesion recognition model performs in-depth evaluation and lesion segmentation on the images to be evaluated, and obtains the lesion depth image and the lesion edge contour image output after the lesion recognition model completes in-depth evaluation and lesion segmentation.

[0124] Figure 4 It is a schematic diagram of the input and output structure of lesion recognition in the lesion size assessment method provided by the embodiments of the present application. As Figure 4 shown, Figure 4 It shows the workflow of the lesion recognition model in the embodiments of the present application for depth estimation and edge contour detection: input the digestive tract image into the lesion recognition model, and perform depth estimation and edge contour detection through this model to obtain the output depth map and edge contour map. Among them, the depth map represents depth information with different gray values, such as 0.2, 0.4, 0.6, 0.8, etc.; the unit of depth can be set according to actual needs, such as centimeters (cm) or millimeters (mm).

[0125] Furthermore, the present application can determine the size conversion coefficient between the microscopic size and the actual size.

[0126] At the same time, determine the disparity value of each pixel point in the lesion depth image.

[0127] Furthermore, based on each disparity value and the size conversion coefficient, determine the actual physical coordinates of each pixel point of the lesion edge contour in the lesion edge contour image.

[0128] Further, according to the actual physical coordinates of each pixel point in the lesion edge contour, the lesion size is evaluated.

[0129] After obtaining the lesion size, a lesion treatment plan can be further formulated based on the lesion size, which can be specifically implemented by manual evaluation or deep learning, and is not specifically limited in this application. In a feasible embodiment, the lesion size can be input into a plan prediction model obtained by training based on sample lesion sizes and corresponding sample treatment plans, and the lesion treatment plan output by the plan prediction model can be obtained, so as to implement it after further feasibility evaluation based on the lesion treatment plan.

[0130] According to the lesion size evaluation method of the embodiments of the present application, a teacher model is trained in advance based on a synthetic image, a lesion depth annotation image corresponding to the synthetic image, and a lesion edge contour annotation image. After the teacher model generates pseudo-labels based on unlabeled sample data, a lesion recognition model is obtained through knowledge distillation by the teacher model based on the unlabeled sample data with pseudo-labels and labeled sample data. Therefore, when lesion recognition is required, the image to be evaluated is input into the lesion recognition model, and the lesion depth image and lesion edge contour image output by the lesion recognition model can be obtained. Thus, through the lesion depth image and the lesion edge contour image, the lesion size can be accurately evaluated.

[0131] Based on the above embodiments, based on the lesion depth image and the lesion edge contour image, evaluating the lesion size includes:

[0132] Determine the size conversion coefficient between the microscopic size and the actual size;

[0133] Determine the disparity value of each pixel point in the lesion depth image;

[0134] Based on each disparity value and the size conversion coefficient, determine the actual physical coordinates of each pixel point of the lesion edge contour in the lesion edge contour image;

[0135] According to the actual physical coordinates of each pixel point in the lesion edge contour, evaluate the lesion size.

[0136] Specifically, the present application can pre-construct a standardized physical model (such as an esophagus and stomach model) to establish an accurate size reference system.

[0137] Further, a capsule endoscope is used to simulate the real environment (underwater) to photograph a target object with a known size, measure the size of the target object in the image, and convert the microscopic size to the actual size through the size conversion coefficient. This calibration process ensures that when photographing in vivo and underwater, the influence of the air-underwater refractive index difference and the irregular cavity can be avoided, and the actual size of the lesion can still be accurately estimated.

[0138] More specifically, the following steps may be included:

[0139] Data collection and model making: Use transparent materials to make standardized physical models of the esophagus and stomach. The model size should be based on standard anatomical data and ensure accuracy at the millimeter level.

[0140] Simulate the shooting environment: Add an appropriate amount of liquid (such as water) to the physical model and use a fixed-resolution capsule endoscope (e.g., 480 by 480 pixels) to shoot the target object (of known size) from different angles and distances.

[0141] Record parameters such as the shooting distance d, angle θ, and liquid environment conditions to ensure the repeatability of the data.

[0142] Furthermore, let the projected size of the target object in the image be , the actual size be , and the distance between the lens and the target object be d. Then the calculation formula for the size conversion coefficient C is:

[0143] ;

[0144] Furthermore, store the size conversion coefficient C in the database for depth estimation and size calculation.

[0145] Therefore, when evaluating the size of the lesion, based on the parallax information of the lesion depth image and the lesion edge contour image, combined with the size conversion coefficient, the size of the lesion in the two-dimensional image is converted into the actual physical size. The specific steps are as follows:

[0146] Parallax information and lesion edge contour extraction: Provide parallax values for each pixel point through the lesion depth image, that is, the relative distance from each pixel point to the camera of the capsule endoscope. Perform edge detection on the lesion edge contour image through an edge detection algorithm (such as Canny edge detection) to extract the edge contour of the lesion. Each pixel point of the edge contour represents the boundary of the lesion. Canny edge detection is a commonly used edge detection algorithm.

[0147] Parallax conversion based on the depth map: In the data calibration stage, determine the size conversion coefficient C through a simulated object of known size, so that each unit of parallax value corresponds to a physical distance (unit: millimeter). Therefore, for each pixel point in the area formed by the lesion edge contour, convert the parallax value in the lesion depth image into the actual physical depth value. The specific formula is as follows:

[0148] d actual = C·d disparity ;

[0149] where, d actual is the actual distance (millimeter or centimeter), ddisparity is the parallax of the corresponding pixel point, and C is the size conversion coefficient.

[0150] Calculation of the lesion size: Obtain the edge pixel points of the lesion from the edge contour of the lesion, and convert the parallax of these pixel points into actual physical coordinates. Further calculate the actual width and height of the lesion through these coordinates.

[0151] For example: Assume that the set of edge contour points of the lesion is P = {(x1, y1, z1), (x2, y2, z2),..., (xn, yn, zn)}, then:

[0152] Width w: Calculate the maximum and minimum coordinate differences of the lesion edge contour in the x direction;

[0153] Height h: Calculate the maximum and minimum coordinate differences of the lesion edge contour in the y direction;

[0154] The lesion area can be calculated based on the distribution and spacing of each pixel point. The specific formula is as follows:

[0155] A lesion = N pixels ·a pixel ;

[0156] Where, N pixels is the number of pixels within the contour, and a pixel is the physical area of each pixel (calculated through the size conversion coefficient).

[0157] And, three-dimensional size and volume estimation: In the case of lacking specific thickness information, the thickness of the lesion can be assumed based on experience. Assume the thickness of the lesion is t.

[0158] If the lesion can be regarded as a regular shape (such as an ellipsoid), its volume V can be estimated by the following formula:

[0159] ;

[0160] Where, w and h are the actual width and height of the lesion, and t is the assumed thickness of the lesion.

[0161] Thus, the lesion size can be obtained.

[0162] In some cases, the edge of the lesion may be unclear due to the occlusion of surrounding tissues. This application uses the lesion depth image and the edge contour image for comprehensive evaluation. Since the depth image provides more comprehensive three-dimensional information, it can capture the actual size and shape of the lesion more accurately, reducing the possible errors in the contour image. Therefore, the lesion size can be accurately evaluated.

[0163] The lesion size evaluation device provided by the present application will be described below. The lesion size evaluation device described below can be correspondingly referred to the lesion size evaluation method described above.

[0164] Furthermore, the present application also provides a lesion size evaluation device.

[0165] The lesion size evaluation device includes:

[0166] An identification module, configured to input an image to be evaluated into a lesion identification model, and obtain a lesion depth image and a lesion edge contour image output by the lesion identification model;

[0167] A determination module, configured to evaluate the lesion size based on the lesion depth image and the lesion edge contour image;

[0168] Wherein, the lesion identification model is obtained by a teacher model through knowledge distillation based on unlabeled sample data with pseudo-labels and labeled sample data; the teacher model is trained based on synthetic images and corresponding labels of the synthetic images; the corresponding labels of the synthetic images include a lesion depth annotation image and a lesion edge contour annotation image; the pseudo-labels are generated by the teacher model based on the unlabeled sample data.

[0169] For the lesion size evaluation device of the present application, a teacher model is pre-trained based on synthetic images and corresponding lesion depth annotation images and lesion edge contour annotation images of the synthetic images, and after generating pseudo-labels by the teacher model based on unlabeled sample data, the lesion identification model is obtained by the teacher model through knowledge distillation based on unlabeled sample data with pseudo-labels and labeled sample data. Therefore, when lesion identification is required, the image to be evaluated is input into the lesion identification model, and a lesion depth image and a lesion edge contour image output by the lesion identification model can be obtained. Thus, through the lesion depth image and the lesion edge contour image, the lesion size can be accurately evaluated.

[0170] In one embodiment, the determination module is specifically configured to:

[0171] Determine a size conversion coefficient between the microscopic size and the actual size;

[0172] Determine the parallax value of each pixel point in the lesion depth image;

[0173] Based on the parallax values and the size conversion coefficient, determine the actual physical coordinates of each pixel point of the lesion edge contour in the lesion edge contour image;

[0174] Evaluate the lesion size according to the actual physical coordinates of each pixel point in the lesion edge contour.

[0175] Figure 5 It is a schematic diagram of a lesion size evaluation device provided by an embodiment of the present application. As Figure 5 shown, in some embodiments, the recognition module in the lesion size evaluation device may include an input layer and a parallel processing layer. The determination module may include a result integration layer and an output layer. The result integration layer may include a size estimation module, and the output layer may include a report generation module. Among them, the input layer may include an image processing module, which can be used to input the two-dimensional image into the image processing module after acquiring image data through a capsule endoscope device, perform basic processing such as noise removal and edge enhancement on the input image through the image processing module, and then input the preprocessed image into the depth estimation module and the lesion segmentation module in the lesion recognition model deployed in the parallel processing layer respectively. Thus, the depth information of the lesion can be deduced from the image through the depth estimation module, and the lesion area can be segmented from the preprocessed image through the lesion segmentation module.

[0176] Furthermore, the depth information and the lesion edge contour can be input into the size estimation module in the result integration layer, and the size estimation module can calculate the size of the lesion according to the segmentation contour and depth information of the lesion. Further, the physical size of the lesion output by the size estimation module is input into the report generation module in the output layer, and the report generation module generates a diagnostic report according to the physical size of the lesion, which summarizes the segmentation information and size of the lesion.

[0177] Figure 6 illustrates a schematic diagram of the physical structure of an electronic device. As Figure 6 shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 complete mutual communication through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the following method: input the image to be evaluated into the lesion recognition model to obtain the lesion depth image and the lesion edge contour image output by the lesion recognition model;

[0178] evaluate the lesion size based on the lesion depth image and the lesion edge contour image;

[0179] wherein, the lesion recognition model is obtained by knowledge distillation of a teacher model based on unlabeled sample data with pseudo-labels and labeled sample data; the teacher model is trained based on synthetic images and the corresponding labels of the synthetic images; the corresponding labels of the synthetic images include lesion depth annotation images and lesion edge contour annotation images; the pseudo-labels are generated by the teacher model based on the unlabeled sample data.

[0180] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs, Read-Only Memories), random access memories (RAMs, Random Access Memories), magnetic disks, or optical discs that can store program codes.

[0181] In another aspect, an embodiment of the present application further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above-mentioned various embodiments. For example, it includes: inputting an image to be evaluated into a lesion recognition model to obtain a lesion depth image and a lesion edge contour image output by the lesion recognition model;

[0182] Evaluating the lesion size based on the lesion depth image and the lesion edge contour image;

[0183] Wherein, the lesion recognition model is obtained by a teacher model through knowledge distillation based on unlabeled sample data with pseudo-labels and labeled sample data; the teacher model is trained based on synthetic images and corresponding labels of the synthetic images; the corresponding labels of the synthetic images include a lesion depth annotation image and a lesion edge contour annotation image; the pseudo-labels are generated by the teacher model based on the unlabeled sample data.

[0184] In another aspect, an embodiment of the present application further provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above-mentioned various embodiments. For example, it includes: inputting an image to be evaluated into a lesion recognition model to obtain a lesion depth image and a lesion edge contour image output by the lesion recognition model;

[0185] Evaluating the lesion size based on the lesion depth image and the lesion edge contour image;

[0186] Among them, the lesion recognition model is obtained by knowledge distillation of a teacher model based on unlabeled sample data with pseudo-labels and labeled sample data; the teacher model is trained based on synthetic images and corresponding labels of the synthetic images; the corresponding labels of the synthetic images include lesion depth annotation images and lesion edge contour annotation images; the pseudo-labels are generated by the teacher model based on the unlabeled sample data.

[0187] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0188] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0189] Finally, it should be noted that the above embodiments are only used to illustrate the present application, rather than to limit the present application. Although the present application has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that various combinations, modifications, or equivalent replacements of the technical solutions of the present application do not depart from the spirit and scope of the technical solutions of the present application.

Claims

1. A method for assessing lesion size, characterized in that: include: Inputting the image to be evaluated into a lesion recognition model to obtain a lesion depth image and a lesion edge contour image output by the lesion recognition model; Based on the lesion depth image and the lesion edge contour image, evaluating the lesion size; Among them, the lesion recognition model is obtained by performing knowledge distillation based on unlabeled sample data and labeled sample data with pseudo labels by the teacher model; the teacher model is trained based on synthetic images and labels corresponding to the synthetic images; the labels corresponding to the synthetic images include lesion depth labeled images and lesion edge contour labeled images; the pseudo labels are generated by the teacher model based on the unlabeled sample data; the synthetic images are generated based on a preset three-dimensional physical model; the preset three-dimensional physical model is constructed based on a virtual image engine; the loss function of the teacher model includes depth estimation loss, image segmentation loss and affine invariant loss; The step of evaluating the lesion size based on the lesion depth image and the lesion edge contour image comprises: Determine the size conversion factor between the under-the-mirror size and the actual size; Determine the disparity value of each pixel in the lesion depth image; Determining the actual physical coordinates of each pixel point of the lesion edge contour in the lesion edge contour image based on each of the disparity values ​​and the size conversion coefficient; The lesion size is evaluated according to the actual physical coordinates of each pixel point in the lesion edge contour.

2. The method for evaluating lesion size according to claim 1, characterized in that: Before performing knowledge distillation based on unlabeled sample data and labeled sample data with pseudo labels, it also includes: Perform data augmentation on unlabeled sample data with pseudo labels.

3. The method for evaluating lesion size according to claim 1, characterized in that: The loss function of the lesion recognition model includes depth estimation loss and image segmentation loss.

4. A lesion size assessment device, characterized in that: include: A recognition module, used for inputting the image to be evaluated into a lesion recognition model to obtain a lesion depth image and a lesion edge contour image output by the lesion recognition model; A determination module, configured to evaluate the size of a lesion based on the lesion depth image and the lesion edge contour image; The determination module is specifically used to determine the size conversion coefficient between the microscopic size and the actual size; determine the disparity value of each pixel in the lesion depth image; determine the actual physical coordinates of each pixel of the lesion edge contour in the lesion edge contour image based on each disparity value and the size conversion coefficient; and evaluate the lesion size according to the actual physical coordinates of each pixel in the lesion edge contour; Among them, the lesion recognition model is obtained by performing knowledge distillation on the teacher model based on unlabeled sample data and labeled sample data with pseudo labels; the teacher model is trained based on synthetic images and corresponding labels of the synthetic images; the corresponding labels of the synthetic images include lesion depth labeled images and lesion edge contour labeled images; the pseudo labels are generated by the teacher model based on the unlabeled sample data; the synthetic images are generated based on a preset three-dimensional physical model; the preset three-dimensional physical model is constructed based on a virtual image engine; the loss function of the teacher model includes depth estimation loss, image segmentation loss and affine invariant loss.

5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for assessing lesion size as described in any one of claims 1 to 3 is implemented.

6. A storage medium, the storage medium being a non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that: When the computer program is executed by a processor, the method for evaluating the size of a lesion as described in any one of claims 1 to 3 is implemented.

7. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the lesion size assessment method according to any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Establishment method of intelligent esophageal focus detection model based on electronic endoscope

    CN114511728A

  • Target area determination method, device and equipment based on depth image features

    CN116958147A