Training Method of Endoscope Image Recognition Model, Endoscope Image Recognition Method and Device

By combining pseudo-notation and gradient vector selection favorable samples in the endoscopic image recognition model for training, the problems of high labor and time cost and insufficient model robustness in ileocecal site recognition are solved, and more efficient endoscopic image recognition is achieved.

CN114240867BActive Publication Date: 2025-07-08XIAOHE MEDICAL EQUIP (HAINAN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111501503.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-09
Publication Date
2025-07-08
Estimated Expiration
2041-12-09

AI Technical Summary

Technical Problem

The prior art has problems such as high labor and time cost and insufficient model robustness in ileocecal site recognition, especially the endoscopic image recognition model trained on low-proportion ileocecal image frame datasets cannot guarantee accuracy.

Method used

By using a combination of pseudo-notation and gradient vectors in the endoscopic image recognition model, we can independently select favorable samples for training, reducing the impact of noise data and improving the robustness of the model.

Benefits of technology

It reduces the cost of manual labeling, improves the accuracy and robustness of the endoscopic image recognition model, and reduces the calculation cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114240867B_ABST
    Figure CN114240867B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a training method for an endoscopic image recognition model, an endoscopic image recognition method and device, which automatically select target sample images from unlabeled endoscopic images for model training to improve the training efficiency and recognition efficiency of the endoscopic image recognition model. The training method includes: inputting a first sample image into the endoscopic image recognition model to obtain a first predicted ileocecal recognition result, and generating a pseudo-labeled ileocecal recognition result corresponding to the first sample image based on the first predicted ileocecal recognition result, where the first sample image is an endoscopic image without a labeled ileocecal recognition result; determining a first gradient vector of the endoscopic image recognition model based on the predicted ileocecal recognition result and the pseudo-labeled ileocecal recognition result; obtaining a second gradient vector of the endoscopic image recognition model; and selecting a target sample image from the first sample images to train the endoscopic image recognition model based on the length of the first gradient vector and the similarity between the first gradient vector and the second gradient vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of medical image technology, and in particular, to a method for training an endoscopic image recognition model, an endoscopic image recognition method, and an apparatus. Background Art

[0002] The ileocecal part refers to the part where the terminal ileum and the cecum of the human body intersect. A colonoscope can reach the ileocecal part using an electronic colonoscope to observe the colon from the mucosal side. During endoscopic examination, the recognition of the ileocecal part is crucial.

[0003] In practical applications, the image frames of the ileocecal region account for a low proportion in the entire endoscopic image. Therefore, related technologies mainly perform fully supervised training on a fixed small-scale dataset through a convolutional neural network to obtain an endoscopic image recognition model for ileocecal recognition. Among them, fully supervised training means that the training samples are first manually labeled, and then the model is trained based on the labeled training samples. In this way, on the one hand, it takes a lot of manpower and time to label the samples, and the labeling cost is relatively high. On the other hand, due to the scale limitation of the dataset, the robustness of the trained endoscopic image recognition model cannot be guaranteed, thus affecting the accuracy of ileocecal recognition. Summary of the Invention

[0004] This Summary of the Invention section is provided to introduce concepts in a brief form that will be described in detail in the subsequent Detailed Description section. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] In a first aspect, the present disclosure provides a method for training an endoscopic image recognition model, where the endoscopic image recognition model is used to recognize the ileocecal part, and the method includes:

[0006] Input a first sample image into the endoscopic image recognition model to obtain a first predicted ileocecal recognition result, and generate a pseudo-labeled ileocecal recognition result corresponding to the first sample image based on the first predicted ileocecal recognition result, where the first sample image is an endoscopic image without a labeled ileocecal recognition result;

[0007] Based on the predicted ileocecal recognition result and the pseudo-labeled ileocecal recognition result, determine a first gradient vector of the endoscopic image recognition model, where the first gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting the first sample image;

[0008] Obtain a second gradient vector of the endoscopic image recognition model, where the second gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting a second sample image, and the second sample image is an endoscopic image labeled with a sample ileocecal recognition result;

[0009] Determine the sampling probability of the first sample image based on the length of the first gradient vector and the similarity between the first gradient vector and the second gradient vector, where the sampling probability is positively correlated with the length of the first gradient vector, and the sampling probability is positively correlated with the similarity between the first gradient vector and the second gradient vector;

[0010] Determine a target sample image in the first sample image whose sampling probability is greater than a probability threshold;

[0011] Train the endoscopic image recognition model based on the target sample image.

[0012] In a second aspect, the present disclosure provides an endoscopic image recognition method, the method comprising:

[0013] Obtain an endoscopic image to be recognized;

[0014] Input the endoscopic image into the endoscopic image recognition model to obtain a cecum recognition result corresponding to the endoscopic image, where the endoscopic image recognition model is trained by the training method of the endoscopic image recognition model described in the first aspect.

[0015] In a third aspect, the present disclosure provides a training device for an endoscopic image recognition model, the endoscopic image recognition model being used to recognize the cecum part, the device comprising:

[0016] A prediction module, configured to input a first sample image into the endoscopic image recognition model, obtain a predicted cecum recognition result corresponding to the first sample image, and generate a pseudo-labeled cecum recognition result corresponding to the first sample image based on the predicted cecum recognition result, where the first sample image is an endoscopic image without a labeled sample cecum recognition result;

[0017] A determination module, configured to determine a first gradient vector of the endoscopic image recognition model based on the predicted cecum recognition result and the pseudo-labeled cecum recognition result, the first gradient vector being used to characterize the parameter change of the endoscopic image recognition model after inputting the first sample image;

[0018] A first acquisition module, configured to acquire a second gradient vector of the endoscopic image recognition model, the second gradient vector being used to characterize the parameter change of the endoscopic image recognition model after inputting a second sample image, the second sample image being an endoscopic image labeled with a sample cecum recognition result;

[0019] A first training module, configured to determine a sampling probability of the first sample image based on the length of the first gradient vector and the similarity between the first gradient vector and the second gradient vector, where the sampling probability is positively correlated with the length of the first gradient vector, and the sampling probability is positively correlated with the similarity between the first gradient vector and the second gradient vector, determine a target sample image in the first sample image whose sampling probability is greater than a probability threshold, and train the endoscopic image recognition model based on the target sample image.

[0020] In a fourth aspect, the present disclosure provides an endoscopic image recognition device, where the device includes:

[0021] A second acquisition module, configured to acquire an endoscopic image to be recognized;

[0022] A recognition module, configured to input the endoscopic image into an endoscopic image recognition model to obtain a cecum recognition result corresponding to the endoscopic image, where the endoscopic image recognition model is trained by the training method of the endoscopic image recognition model described in the first aspect.

[0023] In a fifth aspect, the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of the method described in the first aspect are implemented.

[0024] In a sixth aspect, the present disclosure provides an electronic device, including:

[0025] A storage device, on which a computer program is stored;

[0026] A processing device, configured to execute the computer program in the storage device to implement the steps of the method described in the first aspect.

[0027] Through the above technical solutions, samples beneficial to the learning of the endoscopic image recognition model can be discovered from a large amount of unlabeled data, and the robustness of the endoscopic image recognition model can be improved. Moreover, by calculating the model change through the pseudo-labeling of samples, it is not necessary to traverse all possible labels of the samples to calculate the model change expectation, thereby reducing the calculation cost and manual labeling cost in the training process of the endoscopic image recognition model. In addition, during the sample selection process, the similarity between the unlabeled sample and the labeled sample (i.e., the second sample image) is calculated by combining the gradient (i.e., the second gradient vector) of the labeled sample, and the target sample image is determined in the first sample image based on this similarity for training the endoscopic image recognition model, which can reduce the influence of noise data on the training of the endoscopic image recognition model and improve the accuracy of the output result of the trained endoscopic image recognition model.

[0028] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the original elements and elements are not necessarily drawn to scale. In the drawings:

[0030] Figure 1 is a flowchart of a method for training an endoscopic image recognition model shown according to an exemplary embodiment of the present disclosure;

[0031] Figure 2 is a schematic diagram of the process of a method for training an endoscopic image recognition model shown according to an exemplary embodiment of the present disclosure;

[0032] Figure 3 is a flowchart of a method for endoscopic image recognition shown according to an exemplary embodiment of the present disclosure;

[0033] Figure 4 is a block diagram of a device for training an endoscopic image recognition model shown according to an exemplary embodiment of the present disclosure;

[0034] Figure 5 is a block diagram of an endoscopic image recognition device shown according to an exemplary embodiment of the present disclosure;

[0035] Figure 6 is a block diagram of an electronic device shown according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0037] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0038] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0039] It should be noted that the concepts such as "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0040] It should be noted that the modification of "one" and "multiple" mentioned in this disclosure is illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".

[0041] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0042] First, the technical terms that may be involved in the embodiments of this disclosure will be described.

[0043] Active Learning refers to automatically selecting some data from a dataset and requesting manual annotation. Active Learning, by designing a reasonable sample selection strategy (i.e., query function), continuously selects data from unlabeled data, manually annotates it, and then puts it into the training set.

[0044] Model Change (MC) refers to the change and impact brought by a sample to the current model, and usually the gradient is used to represent the model change.

[0045] Pseudo Label means that the model is first trained with labeled data and then used to predict the labels of unlabeled data, thereby generating pseudo labels.

[0046] Self-training refers to selecting appropriate (such as the samples with the highest confidence) pseudo-labeled samples and adding them to the training set, and iteratively and continuously expanding the training set with unlabeled data in this way.

[0047] As described in the background art, in practical applications, the image frames of the ileocecal region account for a low proportion in the entire endoscopic image. Therefore, related technologies mainly perform fully supervised training on a fixed small-scale dataset through a convolutional neural network to obtain an endoscopic image recognition model for ileocecal recognition. Among them, fully supervised training means that the training samples are first manually labeled, and then the model is trained based on the labeled training samples. In this way, on the one hand, it takes a lot of manpower and time to label the samples, and the labeling cost is relatively high. On the other hand, due to the scale limitation of the dataset, the robustness of the trained endoscopic image recognition model cannot be guaranteed, thus affecting the accuracy of ileocecal recognition.

[0048] The inventors' research found that a direct way to improve the model's robustness is to expand the scale of the training data. However, in the scenario of ileocecal recognition tasks, the image frames of the ileocecal region account for a very low proportion in the entire endoscopic image. Therefore, in order to collect sufficient data, the dataset inevitably comes from multiple medical centers. The medical devices in different medical centers come from different manufacturers, and the patient populations in different medical centers are also different. The data in the dataset cannot meet the assumption of identical distribution, that is, the data in the dataset cannot satisfy the same data distribution, resulting in the generalization ability and robustness of the endoscopic image recognition model not being guaranteed. In addition, after obtaining the endoscopic image data from multiple medical centers, it is also necessary to manually discover the image data that is beneficial to the learning of the endoscopic image recognition model from the massive unwashed data, and require multiple experts to provide labels for these image data, which brings a high labeling cost.

[0049] To reduce the data labeling cost, related technologies have proposed a semi-supervised training method of learning from unlabeled data with a small amount of labels. Specifically, first, a model is trained using a small amount of labeled data, then this trained model is used to predict the unlabeled data, and the prediction results are used as the labels of the unlabeled data to retrain the model. However, this semi-supervised method relies on the assumption of identical distribution between the unlabeled data and the labeled data. The ileocecal dataset from multiple medical centers cannot meet this assumption of identical distribution. When introducing low-quality noise data (i.e., the prediction results of the unlabeled data), it will reduce the performance of the endoscopic image recognition model.

[0050] In addition, related technologies also propose an active learning method that continuously selects useful images from unlabeled data, requests manual annotation, and then puts the samples into the training set. However, the active learning method still requires a large number of requests for expert annotation and cannot automatically learn from unlabeled data. Moreover, related technologies mainly preferentially select data with high uncertainty or the largest gradient norm from unlabeled data. For the former data selection method, in the ileocecal recognition scenario, the uncertainty is related to the imaging effect of endoscopic images. The worse the imaging effect, the higher the uncertainty, so it is easy to select unlabeled samples with poor imaging effects, which affects the training effect of the endoscopic image recognition model. For the latter data selection method, it is necessary to traverse all possible label values of each sample to calculate the expected change (i.e., gradient) of the model, and the calculation cost is relatively high, which affects the training efficiency and recognition efficiency of the endoscopic image recognition model.

[0051] In view of this, the present disclosure provides a training method for an endoscopic image recognition model, which can discover samples beneficial to the learning of the endoscopic image recognition model from a large amount of unlabeled data and improve the robustness of the endoscopic image recognition model. Moreover, by calculating the model change through the pseudo-labeling of samples, it is not necessary to traverse all possible labels of the samples to calculate the expected model change, thereby reducing the calculation cost and manual annotation cost in the training process of the endoscopic image recognition model. In addition, during the sample selection process, the gradient of the labeled samples is combined to calculate the similarity between the unlabeled samples and them, and the target sample image is determined in the first sample image based on this similarity for the training of the endoscopic image recognition model, reducing the influence of noise data on the training of the endoscopic image recognition model and improving the accuracy of the output result of the trained endoscopic image recognition model.

[0052] Figure 1 is a flowchart of a training method for an endoscopic image recognition model shown according to an exemplary embodiment of the present disclosure. Refer to Figure 1 , the endoscopic image recognition model is used to recognize the ileocecal region, and the training method includes:

[0053] Step 101, input the first sample image into the endoscopic image recognition model to obtain the first predicted ileocecal recognition result, and generate a pseudo-labeled ileocecal recognition result corresponding to the first sample image based on the first predicted ileocecal recognition result. Among them, the first sample image is an endoscopic image without a labeled ileocecal recognition result.

[0054] Step 102, based on the predicted ileocecal recognition result and the pseudo-labeled ileocecal recognition result, determine the first gradient vector of the endoscopic image recognition model. Among them, the first gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting the first sample image.

[0055] Step 103: Obtain the second gradient vector of the endoscopic image recognition model. The second gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting the second sample image, and the second sample image is an endoscopic image labeled with the sample ileocecal recognition result.

[0056] Step 104: Determine the sampling probability of the first sample image based on the length of the first gradient vector and the similarity between the first gradient vector and the second gradient vector. The sampling probability is positively correlated with the length of the first gradient vector and is also positively correlated with the similarity between the first gradient vector and the second gradient vector.

[0057] Step 105: Determine the target sample images in the first sample image whose sampling probability is greater than the probability threshold.

[0058] Step 106: Train the endoscopic image recognition model based on the target sample images.

[0059] In a possible way, before step 101, the endoscopic image recognition model can also be initially trained based on a third sample image, which is an endoscopic image labeled with the sample ileocecal recognition result, and the third sample image can be the same as or different from the second sample image. Correspondingly, step 101 can be to input the first sample image into the initially trained endoscopic image recognition model.

[0060] Exemplarily, the endoscopic image recognition model can be any classification model. For example, it can be a Transformer network model. After connecting and pooling multi-level features, they are sent to a fully connected layer for classification. The type and structure of the endoscopic image recognition model in the embodiments of the present disclosure are not limited.

[0061] Through the above method, in the initialization training stage, the endoscopic image recognition model can be trained with a small-scale ileocecal recognition data set with manual annotations first. After that, pseudo-annotations can be generated for a large-scale unlabeled endoscopic image through the initially trained endoscopic image recognition model.

[0062] For example, the small-scale ileocecal recognition data set with annotations is D s ={(x n ,y n )|1≤n≤N}, where x n represents the nth endoscopic image in the data set D s , y n represents the sample ileocecal recognition result labeled for the nth endoscopic image, and N represents the number of endoscopic images included in the data set D s . The large-scale unlabeled ileocecal recognition data set is D u ={x m|1 ≤ m ≤ M}, where x m represents the m-th endoscopic image in the dataset D u and M represents the number of endoscopic images included in the dataset D u If the ileocecal recognition network is f, then the initial training of the endoscopic image recognition model based on the third sample image can be: first input the third sample image x q into the endoscopic image recognition model to obtain the corresponding predicted ileocecal recognition result: Then, based on the ileocecal recognition result labeled for the third sample image and the predicted ileocecal recognition result y q , calculate the loss function Finally, adjust the parameters of the endoscopic image recognition model based on the calculation result of the loss function. Thus, the endoscopic image recognition model can be initially trained based on a small-scale ileocecal recognition dataset with manual labels. After that, based on the initially trained endoscopic image recognition model, pseudo-labels can be generated for a large number of unlabeled endoscopic images.

[0063] It should be understood that the endoscopic image recognition model is equivalent to a classification model. For example, the ileocecal recognition result can be used to classify whether the endoscopic image includes ileocecal valve information. During the labeling process, it can be indicated by labeling 0 that the endoscope does not include ileocecal valve information, and by labeling 1 that the endoscopic image includes ileocecal valve information. Then, the predicted ileocecal recognition result output by the endoscopic image recognition model can include the predicted probabilities corresponding to each type of ileocecal recognition result (for example, including ileocecal valve information is one type of result, and not including ileocecal valve information is another type of result). Thus, according to the following formula, the pseudo-labeled ileocecal recognition result corresponding to the first sample image can be generated based on the first predicted ileocecal recognition result:

[0064]

[0065] where represents the pseudo-labeled ileocecal recognition result corresponding to the j-th first sample image, represents the first predicted ileocecal recognition result corresponding to the j-th first sample image, and argmax(·) represents taking the class label corresponding to the maximum predicted probability in the first predicted ileocecal recognition result.

[0066] Taking the first predicted ileocecal recognition result including the predicted probabilities corresponding to two types of ileocecal recognition results as an example, the first type of ileocecal recognition result indicates no ileocecal valve information, and the class label is 0. The second type of ileocecal recognition result indicates the presence of ileocecal valve information, and the class label is 1. If the predicted probability corresponding to the first type of ileocecal recognition result for the first sample image is 0.3, and the predicted probability corresponding to the second type of ileocecal recognition result is 0.7, then according to the above formula, the class label corresponding to the predicted probability of 0.7 can be taken as the pseudo-labeled ileocecal recognition result, that is, the pseudo-labeled ileocecal recognition result is 1.

[0067] After obtaining the pseudo-labeled ileocecal recognition result, the pseudo-labeled ileocecal recognition result can be regarded as the recognition label corresponding to the first sample image. Thus, based on the predicted ileocecal recognition result and the pseudo-labeled ileocecal recognition result, the first gradient vector of the endoscopic image recognition model can be determined. This first gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting the first sample image. That is, the change and influence brought by the first sample image to the endoscopic image recognition model are measured by the first gradient vector.

[0068] Exemplarily, the loss function can be calculated first based on the predicted ileocecal recognition result and the pseudo-labeled ileocecal recognition result, and then the gradient vector of this loss function is calculated to obtain the first gradient vector. That is, the first gradient vector of the endoscopic image recognition model can be determined according to the following formula:

[0069]

[0070] where, represents the first gradient vector of the endoscopic image recognition model generated by the j-th first sample image, represents the calculation of the gradient vector, represents based on the predicted ileocecal recognition result and the pseudo-labeled ileocecal recognition result the result of the loss function calculated.

[0071] It should be understood that the length of the first gradient vector can be determined by calculating the 2-norm of the first gradient vector, that is, the length of the first gradient vector is determined according to the following formula: Thus, the change scale brought by the first sample image to the endoscopic image recognition model is determined, and then the first sample image with a larger model change scale is selected for model training to improve the robustness of the endoscopic image recognition model.

[0072] However, since the pseudo-label is predicted by the endoscopic image recognition model and may be incorrect, only selecting unlabeled data through the first gradient vector may affect the accuracy of the output result of the trained model. Related technologies usually filter out potential potentially incorrect pseudo-labels according to the confidence (or uncertainty) of the predicted ileocecal recognition result . However, there are still a large number of high-confidence incorrect pseudo-labels used for subsequent training, affecting the robust learning of the endoscopic image recognition model. And the embodiments of the present disclosure can combine the gradient of the labeled samples, that is, the second gradient vector can be obtained in step 103, thereby reducing the influence of pseudo-label noise data on model training and improving the accuracy of the output result of the trained endoscopic image recognition model.

[0073] Exemplarily, the second gradient vector may be the average gradient vector of the endoscopic image recognition model for multiple image pairs annotated with ileocecal recognition results. In a possible manner, the second gradient vector can be obtained as follows: First, input multiple second sample images into the endoscopic image recognition model to obtain multiple second predicted ileocecal recognition results. Then, for each second sample image, based on the second predicted ileocecal recognition result corresponding to the second sample image and the sample ileocecal recognition result, determine the gradient vector of the endoscopic image recognition model after inputting the second sample image. Finally, based on multiple gradient vectors, determine the average gradient vector and use the average gradient vector as the second gradient vector.

[0074] For example, the second sample image is represented as where x k represents the k-th second sample image, y k represents the sample ileocecal recognition result annotated for the k-th second sample image, and K represents the number of second sample images. In this case, the second gradient vector can be determined according to the following formula:

[0075]

[0076] where, g s represents the second gradient vector, g k represents the gradient vector of the endoscopic image recognition model after inputting the k-th second sample image, represents the second predicted ileocecal recognition result of the endoscopic image recognition model for the k-th second sample image.

[0077] After obtaining the first gradient vector and the second gradient vector, the target sample image can be selected from the first sample images based on the length of the first gradient vector and the similarity between the first gradient vector and the second gradient vector.

[0078] Exemplarily, first, based on the length of the first gradient vector and the similarity between the first gradient vector and the second gradient vector, determine the sampling probability of the first sample image, where the sampling probability is positively correlated with the length of the first gradient vector and the sampling probability is positively correlated with the similarity between the first gradient vector and the second gradient vector. Then, determine the target sample image in the first sample images whose sampling probability is greater than the probability threshold. Among them, the probability threshold can be set according to the actual situation, and the embodiments of the present disclosure do not limit this.

[0079] Exemplarily, the similarity can be obtained by calculating the cosine similarity between the first gradient vector and the second gradient vector to obtain the similarity between the first gradient vector and the second gradient vector, or the similarity between the first gradient vector and the second gradient vector can also be determined by other means, and the embodiments of the present disclosure do not limit this. Taking the cosine similarity as an example, the sampling probability of the first sample image can be determined according to the following formula:

[0080]

[0081] Among them, s(j) represents the sampling probability of the j-th first sample image, and τ represents the confidence threshold. represents the predicted ileocecal recognition probability corresponding to the j-th first sample image, I(·) represents the indicator function. If the maximum predicted recognition probability corresponding to the j-th first sample image is greater than or equal to the confidence threshold, the value of this indicator function is 1; otherwise, the value of this indicator function is 0. ||·||2 represents the calculation of the 2-norm. represents the length of the first gradient vector. represents the first gradient vector, g s represents the second gradient vector.

[0082] It should be understood that the maximum predicted recognition probability corresponding to the j-th first sample image (i.e., ) can be understood as the confidence of the j-th first sample image. Then, the indicator function I(·) can take the value of 1 when the confidence of the j-th first sample image is greater than or equal to the confidence threshold τ, and take the value of 0 otherwise. Among them, the confidence threshold τ can be set according to the actual situation. For example, it can be set to 0.95, and the embodiments of the present disclosure do not limit this.

[0083] Thus, according to the above formula for determining the sampling probability, the target sample image selected from the first sample images needs to satisfy that the confidence is greater than or equal to the confidence threshold τ. Moreover, the higher the similarity between the first gradient vector and the second gradient vector, and the larger the scale of the first gradient vector, the higher the corresponding sampling probability. Therefore, the pseudo-labeled endoscopic images that are more conducive to the learning of the endoscopic image recognition model can be selected for model training, reducing the influence of noise pseudo-labels on model training and improving the robustness of the endoscopic image recognition model.

[0084] It should be understood that can be expanded as: represents g s and 's dot product. Then, the above formula for determining the sampling probability can be transformed into:

[0085]

[0086] Among them, represents g s Τ and 's vector inner product.

[0087] After obtaining the sampling probability of the first sample image, if the sampling probability is greater than or equal to the probability threshold, the first sample image can be selected as the target sample image for subsequent model training. Conversely, if the sampling probability is less than the probability threshold, the first sample image is not selected as the target sample image for subsequent model training. Thus, pseudo-labeled endoscopic images that are more conducive to the learning of the endoscopic image recognition model can be selected for model training, reducing the impact of noisy pseudo-labels on model training and improving the robustness of the endoscopic image recognition model.

[0088] In a possible way, training the endoscopic image recognition model based on the target sample image can be as follows: First, input the target sample image into the endoscopic image recognition model to obtain the ileocecal recognition result corresponding to the target sample image, and input the second sample image into the endoscopic image recognition model to obtain the predicted ileocecal recognition result of the second sample image. Then, based on the sampling probability of the target sample image, the predicted ileocecal recognition result corresponding to the target sample image, and the pseudo-labeled ileocecal recognition result, calculate the first loss function, and based on the predicted ileocecal recognition result corresponding to the second sample image and the sample ileocecal recognition result, calculate the second loss function. Finally, based on the calculation results of the first loss function and the second loss function, adjust the parameters of the endoscopic image recognition model.

[0089] That is to say, the training process of the endoscopic image recognition model includes two types of loss functions. One is the first loss function of the pseudo-labeled samples, and the other is the second loss function of the labeled samples. Thus, the parameters of the endoscopic image recognition model can be adjusted by combining the first loss function and the second loss function.

[0090] Exemplarily, the first loss function and the second loss function can be the cross entropy loss (CE), or can be other types of loss functions, which are not limited in the embodiments of the present disclosure.

[0091] Taking the cross entropy loss as an example, each time B labeled second sample images and B pseudo-labeled target sample images are selected from the mixed sample images for training, then the first loss function can be calculated according to the following formula:

[0092]

[0093] where, L u represents the calculation result of the first loss function, and CE(·) represents the cross entropy loss calculation.

[0094] At the same time, the second loss function can be calculated according to the following formula:

[0095]

[0096] where, L sRepresents the calculation result of the first loss function, represents the predicted ileocecal recognition result corresponding to the i-th second sample image, y i represents the sample ileocecal recognition result of the i-th second sample image.

[0097] After calculating the first loss function and the second loss function, based on the calculation results of the first loss function and the second loss function, the parameters of the endoscopic image recognition model can be adjusted. For example, the overall loss function of the endoscopic image recognition model can be determined first based on the calculation results of the first loss function and the second function, and then the parameters of the endoscopic image recognition model can be adjusted based on the overall loss function. Among them, the overall loss function can be the sum of the calculation results of the first loss function and the second loss function, or the overall loss function of the endoscopic image recognition model can be calculated according to the following formula based on the calculation results of the first loss function and the second loss function:

[0098] L = L s + λL u

[0099] where λ is a fixed scalar hyperparameter used to adjust the relative weight of the first loss function, which can be set according to the actual situation, and the embodiments of the present disclosure do not limit this.

[0100] The training method of the endoscopic image recognition model provided by the present disclosure will be described below through another exemplary embodiment.

[0101] Referring to Figure 2 , the training method of the endoscopic image recognition model mainly includes three processes. The first process is to train the endoscopic image recognition model through a small-scale ileocecal recognition dataset with manual annotations. The second process is to generate pseudo-annotations for a large-scale unannotated endoscopic image, obtain the first gradient vector based on the pseudo-annotations, and at the same time, combine the pseudo-annotations and the second gradient vector corresponding to the annotated data to determine the sampling probability of the unannotated endoscopic image. Then, target sample images are selected from the unannotated endoscopic images based on the sampling probability. The third process is to train the endoscopic image recognition model together with the annotated ileocecal recognition dataset and the selected target sample images. Among them, the second process and the third process can be alternately iterated until the iteration stop condition is met. The iteration stop condition can be, for example, reaching a preset number of iterations, and the embodiments of the present disclosure do not limit this.

[0102] In the above manner, samples beneficial to model learning can be discovered from a large amount of unlabeled data to improve the robustness of the endoscopic image recognition model. Moreover, by calculating the model change through the pseudo-labeling of the samples, it is not necessary to traverse all possible labels of the samples to calculate the expected value of the model change, thereby reducing the computational cost and manual labeling cost in the model training process. In addition, by combining the gradients of the labeled samples during the sample selection process, the influence of noisy data on model training can be reduced, and the accuracy of the output result of the trained endoscopic image recognition model can be improved.

[0103] Based on the same concept, an embodiment of the present disclosure further provides an endoscopic image recognition method. Referring to Figure 3 , the endoscopic image recognition method includes the following steps:

[0104] Step 301, obtain an endoscopic image to be recognized;

[0105] Step 302, input the endoscopic image into the endoscopic image recognition model to obtain a cecum recognition result corresponding to the endoscopic image, where the endoscopic image recognition model is trained by any endoscopic image recognition model training method provided by the present disclosure.

[0106] Exemplarily, obtaining the endoscopic image can be from an endoscopic device. In a specific implementation, the endoscopic image recognition method provided by the present disclosure can be applied to the control unit of the endoscopic device. After the control unit obtains the endoscopic image collected by the image acquisition unit of the endoscopic device, it can execute the endoscopic image recognition method provided by the present disclosure, so as to determine the cecum recognition result corresponding to the endoscopic image through the trained endoscopic image recognition model. Or, the polyp classification method provided by the present disclosure can be applied to a medical system including an endoscopic device. The control device in the medical system can communicate with the endoscopic device in a wired or wireless manner, so as to obtain the endoscopic image from the endoscopic device and execute the endoscopic image recognition method provided by the present disclosure, so as to determine the cecum recognition result corresponding to the endoscopic image through the trained endoscopic image recognition model.

[0107] Therefore, since the model change is calculated through the pseudo-labeling of the samples during the training process of the endoscopic image recognition model, it is not necessary to traverse all possible labels of the samples to calculate the expected value of the model change, reducing the computational cost and manual labeling cost in the model training process, thereby improving the efficiency of cecum recognition. In addition, by combining the gradients of the labeled samples during the sample selection process, the influence of noisy data on model training can be reduced, and the accuracy of the output result of the trained endoscopic image recognition model can be improved, that is, the accuracy of cecum recognition can be improved.

[0108] Based on the same concept, an embodiment of the present disclosure further provides a training device for an endoscopic image recognition model. Referring to Figure 4, the endoscopic image recognition model is used to recognize the ileocecal region, and the training device 400 includes:

[0109] A prediction module 401, configured to input a first sample image into the endoscopic image recognition model, obtain a predicted ileocecal recognition result corresponding to the first sample image, and generate a pseudo-labeled ileocecal recognition result corresponding to the first sample image based on the predicted ileocecal recognition result, where the first sample image is an endoscopic image without a labeled sample ileocecal recognition result;

[0110] A determination module 402, configured to determine a first gradient vector of the endoscopic image recognition model based on the predicted ileocecal recognition result and the pseudo-labeled ileocecal recognition result, where the first gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting the first sample image;

[0111] A first acquisition module 403, configured to acquire a second gradient vector of the endoscopic image recognition model, where the second gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting a second sample image, and the second sample image is an endoscopic image labeled with a sample ileocecal recognition result;

[0112] A first training module 404, configured to determine a sampling probability of the first sample image based on the length of the first gradient vector and the similarity between the first gradient vector and the second gradient vector, where the sampling probability is positively correlated with the length of the first gradient vector, and the sampling probability is positively correlated with the similarity between the first gradient vector and the second gradient vector, determine a target sample image in the first sample image with a sampling probability greater than a probability threshold, and train the endoscopic image recognition model based on the target sample image.

[0113] Optionally, the second gradient vector is obtained through the following modules:

[0114] A first processing module, configured to input a plurality of second sample images into the endoscopic image recognition model to obtain a plurality of second predicted ileocecal recognition results;

[0115] A second processing module, configured to, for each of the second sample images, determine a gradient vector of the endoscopic image recognition model after inputting the second sample image based on the second predicted ileocecal recognition result and the sample ileocecal recognition result corresponding to the second sample image;

[0116] A third processing module, configured to determine an average gradient vector based on the plurality of gradient vectors and use the average gradient vector as the second gradient vector.

[0117] Optionally, the first training module 404 is configured to:

[0118] Determine the sampling probability of the first sample image according to the following formula:

[0119]

[0120] where s(j) represents the sampling probability of the j-th first sample image, τ represents the confidence threshold, represents the predicted ileocecal recognition probability corresponding to the j-th first sample image, I(·) represents the indicator function. If the maximum predicted recognition probability corresponding to the j-th first sample image is greater than or equal to the confidence threshold, the value of the indicator function is 1; otherwise, the value of the indicator function is 0. ||·||2 represents the 2-norm calculation, represents the length of the first gradient vector, represents the first gradient vector, g s represents the second gradient vector.

[0121] Optionally, the first training module 404 is configured to:

[0122] Input the target sample image into the endoscopic image recognition model to obtain the ileocecal recognition result corresponding to the target sample image, and input the second sample image into the endoscopic image recognition model to obtain the predicted ileocecal recognition result of the second sample image;

[0123] Calculate a first loss function based on the sampling probability of the target sample image, the predicted ileocecal recognition result and the pseudo-labeled ileocecal recognition result corresponding to the target sample image, and calculate a second loss function based on the predicted ileocecal recognition result and the sample ileocecal recognition result corresponding to the second sample image;

[0124] Adjust the parameters of the endoscopic image recognition model based on the calculation results of the first loss function and the second loss function.

[0125] Optionally, the apparatus 400 further includes:

[0126] A second training module, configured to perform initial training on the endoscopic image recognition model based on a third sample image before inputting the first sample image into the endoscopic image recognition model. The third sample image is an endoscopic image labeled with a sample ileocecal recognition result, and the third sample image is the same as or different from the second sample image;

[0127] The prediction module 401 is configured to:

[0128] Input the first sample image into the endoscopic image recognition model after initial training.

[0129] Based on the same concept, an embodiment of the present disclosure further provides an endoscopic image recognition apparatus. Refer to Figure 5, the endoscopic image recognition device 500 includes:

[0130] A second acquisition module 501, configured to acquire an endoscopic image to be recognized;

[0131] A recognition module 502, configured to input the endoscopic image into an endoscopic image recognition model to obtain a cecum recognition result corresponding to the endoscopic image, where the endoscopic image recognition model is trained by any endoscopic image recognition model training method provided by the present disclosure.

[0132] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0133] Based on the same concept, an embodiment of the present disclosure further provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of any of the above endoscopic image recognition model training methods or any of the endoscopic image recognition methods are implemented.

[0134] Based on the same concept, an embodiment of the present disclosure further provides an electronic device, including:

[0135] A storage device, on which a computer program is stored;

[0136] A processing device, configured to execute the computer program in the storage device to implement the steps of any of the above endoscopic image recognition model training methods or any of the endoscopic image recognition methods.

[0137] Next, refer to Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.

[0138] As Figure 6As shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0139] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wirelesly to exchange data. Although Figure 6 an electronic device 600 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included.

[0140] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0141] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0142] In some embodiments, communication can be carried out using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (for example, a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (for example, the Internet), and end-to-end networks (for example, ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0143] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or it can exist separately without being assembled into the electronic device.

[0144] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: input a first sample image into an endoscopic image recognition model to obtain a first predicted ileocecal recognition result, and generate a pseudo-labeled ileocecal recognition result corresponding to the first sample image based on the first predicted ileocecal recognition result, where the first sample image is an endoscopic image without a labeled ileocecal recognition result; determine a first gradient vector of the endoscopic image recognition model based on the predicted ileocecal recognition result and the pseudo-labeled ileocecal recognition result, where the first gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting the first sample image; obtain a second gradient vector of the endoscopic image recognition model, where the second gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting a second sample image, and the second sample image is an endoscopic image labeled with a sample ileocecal recognition result; determine a sampling probability of the first sample image based on the length of the first gradient vector and the similarity between the first gradient vector and the second gradient vector, where the sampling probability is positively correlated with the length of the first gradient vector and is positively correlated with the similarity between the first gradient vector and the second gradient vector; determine a target sample image in the first sample image with a sampling probability greater than a probability threshold; and train the endoscopic image recognition model based on the target sample image.

[0145] Alternatively, the above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: obtain an endoscopic image to be recognized; input the endoscopic image into an endoscopic image recognition model to obtain an ileocecal recognition result corresponding to the endoscopic image, where the endoscopic image recognition model is trained by any one of the endoscopic image recognition model training methods provided by the present disclosure.

[0146] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or may be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0147] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0148] The modules described in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of a module does not constitute a limitation on the module itself.

[0149] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0150] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0151] According to one or more embodiments of the present disclosure, Example 1 provides a method for training an endoscopic image recognition model for recognizing the ileocecal region, including:

[0152] Input a first sample image into the endoscopic image recognition model to obtain a first predicted ileocecal recognition result, and generate a pseudo-labeled ileocecal recognition result corresponding to the first sample image based on the first predicted ileocecal recognition result, where the first sample image is an endoscopic image without a labeled ileocecal recognition result;

[0153] Based on the predicted ileocecal recognition result and the pseudo-labeled ileocecal recognition result, determine a first gradient vector of the endoscopic image recognition model, where the first gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting the first sample image;

[0154] Obtain a second gradient vector of the endoscopic image recognition model, where the second gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting a second sample image, and the second sample image is an endoscopic image labeled with a sample ileocecal recognition result;

[0155] Based on the length of the first gradient vector and the similarity between the first gradient vector and the second gradient vector, determine the sampling probability of the first sample image, where the sampling probability is positively correlated with the length of the first gradient vector, and the sampling probability is positively correlated with the similarity between the first gradient vector and the second gradient vector;

[0156] Determine a target sample image in the first sample image whose sampling probability is greater than a probability threshold;

[0157] Train the endoscopic image recognition model based on the target sample image.

[0158] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, and the second gradient vector is obtained by the following method:

[0159] Input a plurality of second sample images into the endoscopic image recognition model to obtain a plurality of second predicted ileocecal recognition results;

[0160] For each of the second sample images, based on the second predicted ileocecal recognition result corresponding to the second sample image and the sample ileocecal recognition result, determine the gradient vector of the endoscopic image recognition model after inputting the second sample image;

[0161] Based on the plurality of gradient vectors, determine an average gradient vector, and use the average gradient vector as the second gradient vector.

[0162] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 1. Determining the sampling probability of the first sample image based on the length of the first gradient vector and the similarity between the first gradient vector and the second gradient vector includes:

[0163] Determine the sampling probability of the first sample image according to the following formula:

[0164]

[0165] where s(j) represents the sampling probability of the j-th first sample image, τ represents the confidence threshold, represents the predicted ileocecal recognition probability corresponding to the j-th first sample image, I(·) represents the indicator function. If the maximum predicted recognition probability corresponding to the j-th first sample image is greater than or equal to the confidence threshold, the value of the indicator function is 1; otherwise, the value of the indicator function is 0. ||·||2 represents the 2-norm calculation, represents the length of the first gradient vector, represents the first gradient vector, g s represents the second gradient vector.

[0166] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 1. Training the endoscopic image recognition model based on the target sample image includes:

[0167] Input the target sample image into the endoscopic image recognition model to obtain the ileocecal recognition result corresponding to the target sample image, and input the second sample image into the endoscopic image recognition model to obtain the predicted ileocecal recognition result of the second sample image;

[0168] Calculate the first loss function based on the sampling probability of the target sample image, the predicted ileocecal recognition result and the pseudo-labeled ileocecal recognition result corresponding to the target sample image, and calculate the second loss function based on the predicted ileocecal recognition result and the sample ileocecal recognition result corresponding to the second sample image;

[0169] Adjust the parameters of the endoscopic image recognition model based on the calculation results of the first loss function and the second loss function.

[0170] According to one or more embodiments of the present disclosure, Example 5 provides the method according to any one of Examples 1-4. Before inputting the first sample image into the endoscopic image recognition model, the method further includes:

[0171] Initial training is performed on the endoscopic image recognition model based on a third sample image, where the third sample image is an endoscopic image labeled with a sample ileocecal recognition result, and the third sample image is the same as or different from the second sample image;

[0172] The inputting the first sample image into the endoscopic image recognition model includes:

[0173] Input the first sample image into the endoscopic image recognition model after initial training.

[0174] According to one or more embodiments of the present disclosure, Example 6 provides an endoscopic image recognition method, the method includes:

[0175] Obtain an endoscopic image to be recognized;

[0176] Input the endoscopic image into the endoscopic image recognition model to obtain an ileocecal recognition result corresponding to the endoscopic image, where the endoscopic image recognition model is trained by the training method of the endoscopic image recognition model according to any one of Examples 1-5.

[0177] According to one or more embodiments of the present disclosure, Example 7 provides a training device for an endoscopic image recognition model, where the endoscopic image recognition model is used to recognize the ileocecal part, and the device includes:

[0178] A prediction module, configured to input a first sample image into the endoscopic image recognition model, obtain a predicted ileocecal recognition result corresponding to the first sample image, and generate a pseudo-labeled ileocecal recognition result corresponding to the first sample image based on the predicted ileocecal recognition result, where the first sample image is an endoscopic image without a labeled sample ileocecal recognition result;

[0179] A determination module, configured to determine a first gradient vector of the endoscopic image recognition model based on the predicted ileocecal recognition result and the pseudo-labeled ileocecal recognition result, where the first gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting the first sample image;

[0180] A first acquisition module, configured to acquire a second gradient vector of the endoscopic image recognition model, where the second gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting a second sample image, and the second sample image is an endoscopic image labeled with a sample ileocecal recognition result;

[0181] The first training module is configured to determine the sampling probability of the first sample image based on the length of the first gradient vector and the similarity between the first gradient vector and the second gradient vector, where the sampling probability is positively correlated with the length of the first gradient vector and the sampling probability is positively correlated with the similarity between the first gradient vector and the second gradient vector; determine a target sample image in the first sample image whose sampling probability is greater than a probability threshold, and train the endoscopic image recognition model based on the target sample image.

[0182] According to one or more embodiments of the present disclosure, Example 8 provides an endoscopic image recognition device, the device includes:

[0183] A second acquisition module, configured to acquire an endoscopic image to be recognized;

[0184] A recognition module, configured to input the endoscopic image into an endoscopic image recognition model to obtain a cecum recognition result corresponding to the endoscopic image, where the endoscopic image recognition model is trained by the training method of the endoscopic image recognition model according to any one of Examples 1-5.

[0185] According to one or more embodiments of the present disclosure, Example 9 provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of the method according to any one of Examples 1-6 are implemented.

[0186] According to one or more embodiments of the present disclosure, Example 10 provides an electronic device, including:

[0187] A storage device, on which a computer program is stored;

[0188] A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of Examples 1-6.

[0189] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0190] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0191] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. With regard to the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

Claims

1. A training method for an endoscopic image recognition model, characterized in that, The endoscopic image recognition model is used to recognize the ileocecal region, and the method includes: Inputting a first sample image into the endoscopic image recognition model to obtain a first predicted ileocecal recognition result, and generating a pseudo-labeled ileocecal recognition result corresponding to the first sample image based on the first predicted ileocecal recognition result, where the first sample image is an endoscopic image without a labeled ileocecal recognition result; Determining a first gradient vector of the endoscopic image recognition model based on the first predicted ileocecal recognition result and the pseudo-labeled ileocecal recognition result, where the first gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting the first sample image; Obtaining a second gradient vector of the endoscopic image recognition model, where the second gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting a second sample image, and the second sample image is an endoscopic image labeled with a sample ileocecal recognition result; Determining the sampling probability of the first sample image based on the length of the first gradient vector and the similarity between the first gradient vector and the second gradient vector, where the sampling probability is positively correlated with the length of the first gradient vector, and the sampling probability is positively correlated with the similarity between the first gradient vector and the second gradient vector; Determining a target sample image in the first sample image with a sampling probability greater than a probability threshold; Training the endoscopic image recognition model based on the target sample image; Wherein, the second gradient vector is obtained by the following method: Inputting a plurality of second sample images into the endoscopic image recognition model to obtain a plurality of second predicted ileocecal recognition results; For each of the second sample images, determining the gradient vector of the endoscopic image recognition model after inputting the second sample image based on the second predicted ileocecal recognition result and the sample ileocecal recognition result corresponding to the second sample image; Determining an average gradient vector based on the plurality of gradient vectors and using the average gradient vector as the second gradient vector.

2. The method according to claim 1, wherein The determining the sampling probability of the first sample image based on the length of the first gradient vector and the similarity between the first gradient vector and the second gradient vector includes: Determining the sampling probability of the first sample image according to the following formula: Among them, represents the sampling probability of the th first sample image, represents the confidence threshold, represents the predicted ileocecal recognition probability corresponding to the th first sample image, represents the indicator function. If the maximum predicted recognition probability corresponding to the th first sample image is greater than or equal to the confidence threshold, the value of the indicator function is 1; otherwise, the value of the indicator function is 0. represents the calculation of the 2-norm, represents the length of the first gradient vector, represents the first gradient vector, represents the second gradient vector.

3. The method according to claim 1, characterized in that, The training the endoscopic image recognition model based on the target sample image includes: Inputting the target sample image into the endoscopic image recognition model to obtain the ileocecal recognition result corresponding to the target sample image, and inputting the second sample image into the endoscopic image recognition model to obtain the predicted ileocecal recognition result of the second sample image; Calculating a first loss function based on the sampling probability of the target sample image, the predicted ileocecal recognition result and the pseudo-labeled ileocecal recognition result corresponding to the target sample image, and calculating a second loss function based on the predicted ileocecal recognition result and the sample ileocecal recognition result corresponding to the second sample image; Adjusting the parameters of the endoscopic image recognition model based on the calculation results of the first loss function and the second loss function.

4. The method according to any one of claims 1 to 3, characterized in that, Before inputting the first sample image into the endoscopic image recognition model, the method further includes: Initially training the endoscopic image recognition model based on a third sample image, where the third sample image is an endoscopic image labeled with a sample ileocecal recognition result, and the third sample image is the same as or different from the second sample image; The step of inputting the first sample image into the endoscopic image recognition model includes: Inputting the first sample image into the initially trained endoscopic image recognition model.

5. An endoscopic image recognition method, characterized in that, The method includes: Obtaining an endoscopic image to be recognized; Inputting the endoscopic image into the endoscopic image recognition model to obtain an ileocecal recognition result corresponding to the endoscopic image, where the endoscopic image recognition model is obtained by the training method of the endoscopic image recognition model according to any one of claims 1-4.

6. A training device for an endoscopic image recognition model, characterized in that, The endoscopic image recognition model is used to recognize the ileocecal region, and the device includes: A prediction module, configured to input a first sample image into the endoscopic image recognition model to obtain a first predicted ileocecal recognition result, and generate a pseudo-labeled ileocecal recognition result corresponding to the first sample image based on the first predicted ileocecal recognition result, where the first sample image is an endoscopic image without a labeled sample ileocecal recognition result; A determination module, configured to determine a first gradient vector of the endoscopic image recognition model based on the first predicted ileocecal recognition result and the pseudo-labeled ileocecal recognition result, where the first gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting the first sample image; A first acquisition module, configured to acquire a second gradient vector of the endoscopic image recognition model, where the second gradient vector is used to characterize the parameter change of the endoscopic image recognition model after inputting a second sample image, and the second sample image is an endoscopic image labeled with a sample ileocecal recognition result; A first training module, configured to determine a sampling probability of the first sample image based on the length of the first gradient vector and the similarity between the first gradient vector and the second gradient vector, where the sampling probability is positively correlated with the length of the first gradient vector, and the sampling probability is positively correlated with the similarity between the first gradient vector and the second gradient vector, determine a target sample image in the first sample image with a sampling probability greater than a probability threshold, and train the endoscopic image recognition model based on the target sample image; Wherein, the second gradient vector is obtained through the following modules: A first processing module, configured to input a plurality of second sample images into the endoscopic image recognition model to obtain a plurality of second predicted ileocecal recognition results; A second processing module, configured to, for each of the second sample images, determine a gradient vector of the endoscopic image recognition model after inputting the second sample image based on the second predicted ileocecal recognition result corresponding to the second sample image and the sample ileocecal recognition result; A third processing module, configured to determine an average gradient vector based on the plurality of gradient vectors, and use the average gradient vector as the second gradient vector.

7. An endoscopic image recognition device, characterized in that, The device includes: A second acquisition module, configured to acquire an endoscopic image to be recognized; An identification module, configured to input the endoscopic image into an endoscopic image recognition model to obtain a cecum recognition result corresponding to the endoscopic image, wherein the endoscopic image recognition model is trained by the training method of the endoscopic image recognition model according to any one of claims 1-4.

8. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by a processing device, the program implements the steps of the method according to any one of claims 1-5.

9. An electronic device, characterized in that, Comprising: A storage device storing a computer program thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Method for pylorus and ileocecal valve positioning through wireless capsule endoscope images

    CN110367913A

  • Training method and device of endoscope image feature learning model and classification model

    CN113706526A