Training methods, detection methods, devices, media and equipment for polyp detection models
By extracting foreground nodes and their connections in the source and target domains of the polyp detection model and updating features to improve detection accuracy, the problem of insufficient detection accuracy caused by the inconsistency between the distribution of training and test data is solved, and high efficiency and accuracy of cross-domain polyp detection are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2022-08-23
- Publication Date
- 2026-05-26
AI Technical Summary
Existing polyp detection models based on deep learning technology have insufficient detection accuracy when the training data and test data are not distributed in a consistent manner, and they are particularly difficult to meet the detection requirements when using data from different modalities, devices or hospitals.
By extracting foreground and background features from source and target domain images based on a polyp detection model, the foreground nodes and their connections in the source and target domains are determined. Features are updated through domain connectivity, and the model is trained using labeled source domain data, reducing the need for labeled target domain data and improving detection accuracy.
It effectively reduces the workload of manual annotation, improves the accuracy of polyp detection models in cross-domain detection, and is suitable for application scenarios where polyps have strong foreground camouflage.
Smart Images

Figure CN115375657B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing, and more specifically, to a training method, detection method, apparatus, medium, and device for a polyp detection model. Background Technology
[0002] Thanks to the rapid development of deep learning technology and the massive amount of clinical endoscopic data in recent years, computer-assisted endoscopic polyp detection has received widespread attention from academia and industry.
[0003] In related technologies, polyp detection based on deep learning is usually based on the assumption that training data and test data are under the same distribution. However, in practical applications, this assumption is often difficult to meet, as training data and test data may come from different modalities, devices, or hospitals, resulting in insufficient accuracy of deep learning models in detecting polyps. Summary of the Invention
[0004] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] Firstly, this disclosure provides a method for training a polyp detection model, the method comprising:
[0006] Based on the polyp detection model, foreground and background features are extracted from source domain images in the source domain dataset and target domain images in the target domain dataset to obtain first source domain features and first target domain features. The images in the source domain dataset are labeled with polyp tags, while the images in the target domain dataset are not labeled. The data distributions of the source domain dataset and the target domain dataset are different.
[0007] Based on the first source domain features and the first target domain features, source domain foreground nodes in the source domain image and target domain foreground nodes in the target domain image are determined.
[0008] Based on the domain foreground nodes under each target type, determine the domain connection relationship under the target type, wherein the target type includes a source domain type and a target domain type, and the domain connection relationship includes the source domain connection relationship between source domain foreground nodes under the source domain type and the target domain connection relationship between target domain foreground nodes under the target domain type;
[0009] The first domain feature under each target type is updated according to the domain connection relationship under each target type to obtain the second domain feature. The first domain feature includes the first source domain feature under the source domain type and the first target domain feature under the target domain type. The second domain feature includes the second source domain feature obtained by updating the first source domain feature according to the source domain connection relationship and the second target domain feature obtained by updating the first target domain feature according to the target domain connection relationship.
[0010] The prediction result corresponding to the source domain image is obtained based on the second source domain feature;
[0011] Based on the prediction results corresponding to the source domain image, the polyp label, the second source domain features, and the second target domain features, the target loss of the polyp detection model is determined, and the polyp detection model is trained based on the target loss.
[0012] Secondly, this disclosure provides a method for detecting polyps, the method comprising:
[0013] Received the target image for detection;
[0014] The target image is input into the trained polyp detection model to obtain the polyp detection result corresponding to the target image, wherein the polyp detection model is trained based on the training method of the polyp detection model described in the first aspect.
[0015] Thirdly, this disclosure provides a training apparatus for a polyp detection model, the apparatus comprising:
[0016] The feature extraction module is used to extract foreground and background features from source domain images in the source domain dataset and target domain images in the target domain dataset based on the polyp detection model, to obtain first source domain features and first target domain features. The images in the source domain dataset are labeled with polyp tags, while the images in the target domain dataset are not labeled. The data distributions of the source domain dataset and the target domain dataset are different.
[0017] The first determining module is used to determine the source domain foreground node in the source domain image and the target domain foreground node in the target domain image based on the first source domain features and the first target domain features.
[0018] The second determining module is used to determine the domain connection relationship under the target type based on the domain foreground nodes under each target type, wherein the target type includes a source domain type and a target domain type, and the domain connection relationship includes the source domain connection relationship between source domain foreground nodes under the source domain type and the target domain connection relationship between target domain foreground nodes under the target domain type.
[0019] An update module is used to update the first domain feature under each target type according to the domain connection relationship under each target type to obtain a second domain feature, wherein the first domain feature includes the first source domain feature under the source domain type and the first target domain feature under the target domain type, and the second domain feature includes the second source domain feature obtained by updating the first source domain feature according to the source domain connection relationship and the second target domain feature obtained by updating the first target domain feature according to the target domain connection relationship;
[0020] The acquisition module is used to obtain the prediction result corresponding to the source domain image based on the second source domain features;
[0021] The training module is used to determine the target loss of the polyp detection model based on the prediction results corresponding to the source domain image, the polyp label, the second source domain features, and the second target domain features, and to train the polyp detection model based on the target loss.
[0022] Fourthly, this disclosure provides a polyp detection device, the device comprising:
[0023] The receiving module is used to receive the detected target image;
[0024] The processing module is used to input the target image into the trained polyp detection model to obtain the polyp detection result corresponding to the target image, wherein the polyp detection model is trained based on the training method of the polyp detection model described in the first aspect.
[0025] Fifthly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.
[0026] Sixthly, this disclosure provides an electronic device, comprising:
[0027] A storage device on which computer programs are stored;
[0028] A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect.
[0029] The above technical solution allows for training of a polyp detection model based on labeled source domain data and unlabeled target domain data. On the one hand, it effectively reduces the workload and technical requirements of manual labeling. On the other hand, during training, by determining the correlation between foreground features in the source domain image and the correlation between foreground features in the target domain image, the foreground features of the source and target domain images are enhanced respectively, thereby improving the discriminative power of polyp features and ensuring the detection accuracy of polyp detection across domains. This approach is suitable for application scenarios where polyps have strong foreground camouflage.
[0030] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0031] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:
[0032] Figure 1 This is a flowchart of a training method for a polyp detection model provided according to one embodiment of the present disclosure;
[0033] Figure 2 This is a block diagram of a training apparatus for a polyp detection model provided according to one embodiment of the present disclosure;
[0034] Figure 3 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation
[0035] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0036] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0037] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0038] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0039] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0040] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0041] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0042] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0043] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0044] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0045] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0046] Figure 1 The diagram shows a flowchart of a training method for a polyp detection model according to one embodiment of this disclosure. Figure 1 As shown, the method may include:
[0047] In step 11, foreground and background features are extracted from source domain images in the source domain dataset and target domain images in the target domain dataset based on the polyp detection model to obtain first source domain features and first target domain features. The images in the source domain dataset are labeled with polyp tags, while the images in the target domain dataset are not labeled. The data distributions of the source domain dataset and the target domain dataset are different.
[0048] The polyp detection model can be implemented based on a detection network, such as the Faster-RCNN network. In the Faster-RCNN network: the feature extraction subnetwork can be used to extract features, such as through the backbone network, or by extracting feature maps of the input image based on a set of conv+relu+pooling layers; the region proposal network (RPN) can obtain a large number of foreground regions by inputting the features extracted by the feature extraction subnetwork into the region proposal network. The foreground regions and the input features are then input into the Region of Interest (ROI) Align subnetwork, thereby obtaining the region of interest (ROI) features corresponding to a large number of foreground features.
[0049] For example, a source domain image can be input into the polyp detection model to obtain the ROI features corresponding to the source domain image, i.e., the first source domain features, through the above method. Similarly, a target domain image can be input into the polyp detection model to obtain the ROI features corresponding to the target domain image, i.e., the first target domain features, through the above steps. For instance, 128 foreground and background features corresponding to the source domain image and 128 foreground and background features corresponding to the target domain image can be obtained respectively. The foreground and background features can include foreground and background features in the image. The number of features can be set based on the actual application scenario, and this disclosure does not limit this.
[0050] In this process, the source image can be annotated by a physician with some experience to obtain corresponding polyp labels. In the field of polyp detection, the polyp label can include a classification label and a location label to indicate what type of polyp is contained in the source image and its specific location in the source image.
[0051] As described in the background section of this disclosure, data from different modalities, devices, or hospitals may have different distributions. Therefore, a polyp detection model trained on images from the source domain dataset may struggle to directly detect polyps in images from the target domain. Furthermore, in the field of polyp detection, due to the relatively small amount of data and high annotation requirements, the workload of annotating training images is substantial. Therefore, this disclosure allows for simultaneous training of the model on images from both the source and target domain datasets, enabling the trained polyp detection model to be adapted for polyp detection in images from the target domain without requiring annotation of the data in the target domain dataset.
[0052] For example, the source domain dataset could be image data from Hospital Hs1 with corresponding polyp labels, while the target domain dataset could be image data from Hospital Hs2.
[0053] In step 12, based on the first source domain features and the first target domain features, source domain foreground nodes in the source domain image and target domain foreground nodes in the target domain image are determined.
[0054] In this process, the polyp is usually located in the foreground of the image in both the source and target domain images. Accordingly, in this step, the source domain foreground node is used to represent the predicted feature points corresponding to the foreground in the source domain image, and the target domain foreground node is used to represent the predicted feature points corresponding to the foreground in the target domain image, so as to accurately detect the polyp.
[0055] In step 13, based on the domain foreground nodes under each target type, the domain connection relationships under the target type are determined. The target type includes a source domain type and a target domain type, and the domain connection relationships include the source domain connection relationships between source domain foreground nodes under the source domain type and the target domain connection relationships between target domain foreground nodes under the target domain type.
[0056] For example, in this step, for the source domain type, the source domain connectivity relationships between source domain foreground nodes can be determined based on the various source domain foreground nodes under the source domain type; similarly, for the target domain type, the target domain connectivity relationships between target domain foreground nodes can be determined based on the various target domain foreground nodes under the target domain type. The multiple source domain foreground nodes corresponding to the determined source domain image can represent foreground features in the source domain image. Further determining the connectivity relationships between these source domain foreground nodes in this step can further determine whether the various foreground features are correlated, thereby improving the feature contrast of the foreground features to a certain extent. Similarly, the same processing can be performed on the target domain foreground nodes to ensure consistency in the processing of source and target domain data.
[0057] In step 14, the first domain feature under each target type is updated according to the domain connection relationship to obtain the second domain feature. The first domain feature includes the first source domain feature under the source domain type and the first target domain feature under the target domain type. The second domain feature includes the second source domain feature obtained by updating the first source domain feature according to the source domain connection relationship and the second target domain feature obtained by updating the first target domain feature according to the target domain connection relationship.
[0058] For example, in this step, for the source domain type, the first source domain feature can be updated according to the source domain connection relationship under the source domain type to obtain the second source domain feature; and for the target domain type, the first target domain feature can be updated according to the target domain connection relationship under the target domain type to obtain the second target domain feature.
[0059] In the case of polyp detection, the polyp's color and texture are very similar to the surrounding normal tissue, resulting in low contrast and giving the polyp a strong camouflage characteristic, which leads to insufficient detection accuracy.
[0060] In the steps, the source domain connectivity includes the correlation between foreground features in the source domain image. Based on this correlation, the first source domain features are updated, which can improve the foreground feature representation in the feature image corresponding to the source domain image to a certain extent. This is suitable for application scenarios where polyps have strong foreground camouflage.
[0061] In step 15, the prediction result corresponding to the source domain image is obtained based on the second source domain features. This step can be based on the classifier and regressor in the Faster-RCNN network to obtain the prediction result. For example, foreground features can be extracted from the second source domain features and input into the classifier of the fully connected layer to obtain the corresponding classification prediction result, and input into the regressor of the fully connected layer to obtain the corresponding location prediction result.
[0062] In step 16, the target loss of the polyp detection model is determined based on the prediction results and polyp labels corresponding to the source domain image, as well as the second source domain features and the second target domain features, and the polyp detection model is trained based on the target loss.
[0063] The detection accuracy of the polyp detection model can be determined based on the prediction results and polyp labels corresponding to the source domain image. Based on the second source domain features and the second target domain features, cross-domain detection of polyps can be achieved by constraining the distributions of the two to be close. Therefore, the target loss can include the accuracy of the polyp detection model and the distance between the distributions of the source domain features and the target domain features.
[0064] For example, training can be stopped when the target loss is less than the loss threshold, or when the number of training iterations reaches the threshold. Otherwise, the parameters in the polyp detection model can be updated using gradient descent based on the target loss, and the training can be repeated through the steps described above until the model training is complete.
[0065] Therefore, the above technical solution can be used to train a polyp detection model based on labeled source domain data and unlabeled target domain data. On the one hand, it can effectively reduce the workload and technical requirements of manual annotation. On the other hand, during the training process, by determining the correlation between foreground features in the source domain image and the correlation between foreground features in the target domain image respectively, the foreground features of the source domain image and the target domain image can be enhanced respectively, thereby improving the discriminativeness of polyp features and ensuring the detection accuracy of polyp detection across domains. This is suitable for application scenarios where polyps have strong foreground camouflage.
[0066] In one possible embodiment, an exemplary implementation of determining the source domain foreground node in the source domain image and the target domain foreground node in the target domain image based on the first source domain features and the first target domain features in step 12 is as follows: This step may include:
[0067] The set of candidate foreground nodes corresponding to the first source domain feature is determined based on the polyp label.
[0068] The polyp label includes a location label, which allows us to obtain features at the corresponding location. For example, the first source domain feature is f. s Then, the features corresponding to the position labels in the first source domain features can be used as the candidate foreground node set f. s obj That is, f s obj ∈f s .
[0069] For each candidate foreground node in the candidate foreground node set, determine the target domain matching foreground node in the first target domain feature corresponding to the candidate foreground node, and determine the source domain matching foreground node in the first source domain feature of the target domain matching foreground node.
[0070] For the i-th candidate foreground node f s obj,i The candidate foreground node f can be calculated separately. s obj,i Features f of the first target domain t Each feature f in t j The cosine similarity between the nodes is used to determine the feature with the highest cosine similarity as the candidate foreground node f. s obj,i The corresponding target domain matches the foreground node, and the index representation of the target domain matching foreground node is as follows:
[0071]
[0072] That is, with candidate foreground node f s obj,i The corresponding target domain matching foreground node is the j'th feature point in the first target domain features. This is used to represent the number of features in the first target domain feature set, continuing from the example above. The value is 128.
[0073] Furthermore, for the target domain matching foreground node, i.e., the j'-th feature point in the first target domain features, the cosine similarity between each feature in the first source domain features can be calculated, and the feature with the largest corresponding cosine similarity is determined as the source domain matching foreground node of the target domain matching foreground node in the first source domain features. For example, the j'-th feature point f in the first target domain features is determined as follows: t j' The source domain matching foreground node in the first source domain feature is the i'th feature point in the first source domain feature.
[0074] Subsequently, if the source domain matching foreground node belongs to the candidate foreground node set, the candidate foreground node is determined as the source domain foreground node, and the target domain matching foreground node is determined as the target domain foreground node.
[0075] If the i'-th feature point in the first source domain features belongs to f s obj At this point, the candidate foreground node f in the first source domain feature can be considered as s obj,iAs a source domain foreground node, it can also determine feature f in the first target domain features. t j' Foreground nodes of the target domain.
[0076] Therefore, through the above technical solution, the source domain foreground nodes in the source domain features and the target domain foreground nodes in the target domain features can be filtered based on the feature similarity between the source domain features and the target domain features. This can accurately determine the target domain foreground nodes without labeling the target domain image, thus improving the application scope of the method.
[0077] In one possible embodiment, an exemplary implementation of determining the domain connectivity relationship under each target type based on the domain foreground node is as follows, and this step may include:
[0078] For each domain foreground node under each target type, the overlap between the first predicted bounding box corresponding to the domain foreground node under the target type and the second predicted bounding box corresponding to each candidate domain foreground node is determined. The candidate domain foreground nodes are different from the domain foreground nodes and correspond to the same target type.
[0079] If the overlap is greater than or equal to a preset threshold, it is determined that there is a connection between the domain foreground node and the candidate domain foreground node. The domain connection relationship under the target type includes the connection relationship corresponding to each domain foreground node under the target type.
[0080] Wherein, determining the overlap between the first predicted bounding box corresponding to the domain foreground node under the target type and the second predicted bounding box corresponding to each candidate domain foreground node includes:
[0081] Based on the first predicted bounding box and the domain image to which the domain foreground node belongs, a first domain predicted bounding box corresponding to the first predicted bounding box in the domain image is determined, wherein the domain image to which the domain foreground node belongs is the image of the extracted domain foreground node, the domain image to which the source domain foreground node belongs is the source domain image, and the domain image to which the target domain foreground node belongs is the target domain image.
[0082] Based on the second predicted bounding box and the domain image, determine the second domain predicted bounding box corresponding to the second predicted bounding box in the domain image;
[0083] The overlap ratio corresponding to the first domain predicted bounding box and the second domain predicted bounding box is determined as the overlap degree.
[0084] The following section details the implementation of the above steps using the method for determining the source domain connection relationship under the source domain type.
[0085] For each source domain foreground node under the source domain type, the overlap degree between the first predicted bounding box corresponding to the source domain foreground node and the second predicted bounding box corresponding to each candidate source domain foreground node is determined, wherein the candidate source domain foreground node is different from the source domain foreground node.
[0086] For example, if m source domain foreground nodes are identified, for the first source domain foreground node, the overlap between its corresponding first predicted bounding box and the second predicted bounding boxes of the remaining m-1 candidate source domain foreground nodes can be calculated sequentially.
[0087] The predicted bounding box corresponding to the source domain foreground node can be the predicted bounding box determined based on the Regional Candidate Subnetwork (RPN) described above.
[0088] As an example, an exemplary implementation of determining the overlap between the first predicted bounding box corresponding to the source domain foreground node and the second predicted bounding box corresponding to each candidate source domain foreground node is as follows, which may include:
[0089] Based on the first predicted bounding box and the source domain image, a first source domain predicted bounding box is determined in the source domain image corresponding to the first predicted bounding box.
[0090] For example, if a pooling operation is performed during the process of determining the predicted bounding box corresponding to the source domain foreground node, then in this embodiment, the first predicted bounding box can be mapped onto the source domain image based on the size of the source domain image and the size corresponding to the first source domain feature, thereby obtaining the first source domain predicted bounding box corresponding to the source domain foreground node, that is, obtaining the coordinate position of the first source domain predicted bounding box in the source domain image.
[0091] Based on the second predicted bounding box and the source domain image, a second source domain predicted bounding box corresponding to the second predicted bounding box in the source domain image is determined. Similarly, the second predicted bounding box can be mapped onto the source domain image in the manner described above to obtain the coordinate position of the second source domain predicted bounding box in the source domain image.
[0092] The overlap ratio corresponding to the first source domain prediction bounding box and the second source domain prediction bounding box is determined as the overlap degree.
[0093] The predicted bounding box p of the first source domain can be calculated using the following formula. i Second source domain predicted bounding box p j Corresponding IOU ij :
[0094] Subsequently, if the overlap is greater than or equal to a preset threshold, it is determined that there is a connection between the source domain foreground node and the candidate source domain foreground node, and the source domain connection relationship includes the connection relationship corresponding to each source domain foreground node.
[0095] The preset threshold can be set based on the actual application scenario, and this disclosure does not limit it. In this step, if the overlap is greater than or equal to the preset threshold, the first source domain prediction border and the second source domain prediction border have a high overlap, indicating that the source domain foreground node and the candidate source domain foreground node are close in position. At this time, a connection relationship between the source domain foreground node and the candidate source domain foreground node can be constructed to facilitate joint analysis of adjacent nodes.
[0096] For example, source domain connectivity can be represented by an m*m connectivity matrix. An element with a value of 1 in the matrix indicates that there is a connectivity between the two source domain foreground nodes at that position, while an element with a value of 0 indicates that there is no connectivity between the two source domain foreground nodes at that position.
[0097] Accordingly, the method for determining the target domain connection relationship under the target domain type is the same as the method for determining the source domain connection relationship under the source domain type described above, and will not be repeated here.
[0098] Therefore, by using the above technical solution, after determining the source domain foreground nodes in the source domain image, the connection relationship between the source domain foreground nodes can be further constructed based on the geometric relationship corresponding to the source domain foreground nodes, which further improves the correlation of polyp foreground features during feature extraction and provides effective and reliable data support for subsequent feature enhancement of foreground features.
[0099] In one possible embodiment, another exemplary implementation of determining the domain connectivity relationship under each target type based on the domain foreground node is as follows, which may include:
[0100] For each domain foreground node under each target type, determine the overlap between the first domain prediction bounding box corresponding to the domain foreground node under the target type and the second prediction bounding box corresponding to each candidate domain foreground node, as well as the distance between the center points of the first prediction bounding box and the second prediction bounding box. The candidate domain foreground nodes are different from the domain foreground nodes and correspond to the same target type.
[0101] The method for determining the overlap between predicted bounding boxes has been detailed above and will not be repeated here. For example, the distance between the center points of the first and second predicted bounding boxes can be calculated based on the squared Euclidean distance.
[0102] Then, based on each domain foreground node under the target type, and the overlap and distance between the domain foreground node and the candidate domain foreground node, the geometric adjacency matrix and semantic adjacency matrix corresponding to the domain foreground node under the target type are determined.
[0103] Perform a dot product operation on the geometric adjacency matrix and the semantic adjacency matrix, and determine the resulting matrix as the domain connectivity under the target type.
[0104] Accordingly, an exemplary implementation of determining the geometric adjacency matrix and semantic adjacency matrix corresponding to the domain foreground nodes under the target type based on each domain foreground node under the target type, and the overlap degree and distance between the domain foreground nodes and the candidate domain foreground nodes, may include:
[0105] For each of the domain foreground nodes, if the overlap between the domain foreground node and the candidate domain foreground node is greater than a first threshold, or if the overlap between the domain foreground node and the candidate domain foreground node is zero and the distance is less than a second threshold, then it is determined that there is a connection between the domain foreground node and the candidate domain foreground node. The similarity between the domain foreground node and the candidate domain foreground node is determined as the geometric adjacency value between the domain foreground node and the candidate domain foreground node, so as to obtain the geometric adjacency matrix.
[0106] For example, taking source domain foreground nodes as an example, the connections between various source domain foreground nodes can be determined by the following formula:
[0107]
[0108] Wherein, θ1 represents the first threshold and θ2 represents the second threshold;
[0109] φ ij Used to represent the first source domain prediction bounding box p i Second source domain predicted bounding box p j The distance between the center points;
[0110] e ij This is used to represent the connection between the i-th source domain foreground node and the j-th source domain foreground node. A value of 1 indicates that there is a connection between the two nodes, and a value of 0 indicates that there is no connection between the two nodes.
[0111] Furthermore, the geometric adjacency matrix corresponding to the foreground nodes in the domain can be determined based on this connection, as shown in the following formula:
[0112]
[0113] in, The value used to represent the geometric adjacency relationship between the i-th source foreground node and the j-th source foreground node is cos(x). i ,x j) is used to represent the cosine similarity between the i-th source domain foreground node and the j-th source domain foreground node.
[0114] For each of the domain foreground nodes, if the similarity between the domain foreground node and the candidate domain foreground node is greater than a third threshold, then the similarity is determined as the semantic adjacency value between the domain foreground node and the candidate domain foreground node, so as to obtain the semantic adjacency matrix.
[0115] Similarly, the semantic adjacency matrix corresponding to the foreground nodes in the domain can be determined using the following formula:
[0116]
[0117] τ is used to represent the third threshold. This value represents the semantic adjacency relationship between the i-th source domain foreground node and the j-th source domain foreground node.
[0118] Then, the source domain connection relationship A corresponding to the source domain foreground node can be represented as follows:
[0119]
[0120] in, Used to indicate a dot product operation.
[0121] Similarly, for the target domain connectivity, the geometric adjacency matrix and semantic adjacency matrix corresponding to the foreground nodes of the target domain can be determined in the same way as above, thereby determining the target domain connectivity between the foreground nodes of the target domain, and then performing subsequent calculations.
[0122] Therefore, by using the above technical solution, when determining the domain connectivity relationship under the target type, not only the geometric relationship between the domain foreground nodes under the target type is included, but also the semantic relationship between the domain foreground nodes under the target type is included. The semantic relationship can be used to represent the category dependency relationship in the high-level feature, thereby further improving the accuracy of the domain connectivity relationship and providing accurate data support for subsequent updates to the domain features.
[0123] In one possible embodiment, an exemplary implementation of updating the first domain feature under each target type according to the domain connection relationship to obtain the second domain feature is as follows, and this step may include:
[0124] Under each target type, the first domain feature under the target type is convolved based on the domain connectivity relationship under the target type and at least one attention feature layer to obtain the second domain feature under the target type, wherein the attention weights corresponding to the attention feature layer under the target type are determined based on the domain connectivity relationship under the target type.
[0125] The following section will take updating the first source domain feature of a source domain type based on the source domain connection relationship as an example to explain in detail.
[0126] For example, an exemplary implementation of updating the first source domain features based on source domain connectivity to obtain the second source domain features includes:
[0127] Based on the source domain connectivity and at least one attention feature layer, the first source domain features are convolved to obtain the second source domain features. In this step, the attention weights in the attention feature layer are determined based on the source domain connectivity, so that when determining the attention weights, more attention can be paid to the features of feature points that are connected to the current feature point, thereby enhancing the features of each foreground feature and further improving the discriminativeness of the foreground features, providing accurate data support for subsequent polyp classification and location regression.
[0128] For example, convolutional processing can be performed layer by layer using one or more attention feature layers. The attention weights corresponding to each attention feature layer under the target type are determined based on the domain connectivity under the target type using the following formula:
[0129]
[0130] in, This is used to represent the attention weight between the i-th domain foreground node and the j-th domain foreground node under the target type in the l-th attention feature layer. The i-th domain foreground node and the j-th domain foreground node under the target type are connected. If the target type is a source domain type, then the i-th domain foreground node under the target type is the i-th source domain foreground node. If the target type is a target domain type, then the i-th domain foreground node under the target type is the i-th target domain foreground node. The representation of other nodes is similar and will not be repeated here.
[0131] A ij This value represents the connection relationship between the i-th domain foreground node and the j-th domain foreground node under the target type.
[0132] Used to represent the features of the i-th domain foreground node under the target type in the l-th attention feature layer;
[0133] Ni is used to represent the set of domain foreground nodes that have a connection relationship with the i-th domain foreground node under the target type;
[0134] W is used to represent learnable feature transformation operations. It can be a projection matrix used to transform the features of different nodes into a unified space to facilitate feature computation and processing.
[0135] Used to represent learnable weight vectors;
[0136] || is used to represent feature splicing operations.
[0137] Accordingly, after determining the attention weights corresponding to the current attention feature layer, the attention features output by the current attention feature layer can be obtained using the following formula:
[0138]
[0139] in, This is used to represent the feature of the i-th domain foreground node under the target type at layer l. After performing attention feature extraction, the attention features at the (l+1)th layer, where the features at the lth layer are the first domain features under the target type;
[0140] δ() is used to represent nonlinear mapping.
[0141] Each attention feature layer can determine the attention weight corresponding to the current layer in the above manner, thereby enhancing and updating the input features based on the attention weight. The attention features corresponding to the last attention feature layer can be used as the second domain features under the target type.
[0142] Therefore, through the above technical solution, the attention weights for extracting attention features can be determined based on the domain connectivity under the target type. Then, the node feature values can be updated by calculating the attention weights between nodes, thus completing the foreground feature enhancement based on the multi-layer graph attention mechanism. This further ensures the identifiability and discriminability of the second domain features under the target type, and guarantees the recognition accuracy of the polyp recognition model trained in this way.
[0143] It should be noted that determining the target domain connectivity between foreground nodes in the target domain is similar to determining the source domain connectivity between foreground nodes in the source domain. Similarly, updating the first target domain features based on the target domain connectivity to obtain the second target domain features is similar to updating the first source domain features based on the source domain connectivity to obtain the second source domain features. Both can be implemented using the methods described above, and will not be repeated here. Therefore, it is possible to further associate foreground nodes in the target domain image under the target domain, and update the target domain features.
[0144] In one possible embodiment, the polyp label includes a location label and a classification label;
[0145] Accordingly, an exemplary implementation of determining the target loss of the polyp detection model based on the prediction result corresponding to the source domain image, the polyp label, the second source domain features, and the second target domain features is as follows, which may include:
[0146] The regression loss is determined based on the predicted location information and the location label in the prediction results;
[0147] The classification loss is determined based on the predicted classification information and the classification label in the prediction results;
[0148] The methods for determining regression loss and classification loss can be determined based on the corresponding loss calculation methods in the Faster-RCNN network in this field. For example, classification loss can be calculated based on cross-entropy loss, and regression loss can be calculated based on smooth L1 loss, which will not be elaborated here.
[0149] The feature loss is determined based on the second source domain features and the second target domain features.
[0150] For example, the feature loss L can be determined based on the second source domain features and the second target domain features using the following formula. cst :
[0151]
[0152] Where, N s Used to represent the number of foreground nodes in the source domain;
[0153] N t Used to represent the number of foreground nodes in the target domain;
[0154] x i s Used to represent the feature corresponding to the i-th source domain foreground node in the second source domain features;
[0155] x j t This is used to represent the feature corresponding to the j-th target domain foreground node in the second target domain features.
[0156] Therefore, this feature loss can be used to narrow down the feature distribution between source domain features and target domain features, thereby adjusting the accuracy of the cross-domain detection results of the polyp detection model based on this loss.
[0157] Then, the target loss can be determined based on the regression loss, the classification loss, and the feature loss.
[0158] For example, the target loss can be obtained by weighted summation based on the weights corresponding to the regression loss, classification loss, and feature loss, respectively. Thus, through the above technical solution, the target loss can include both a loss constraining the accuracy of polyp detection results and a loss constraining the feature distribution between source and target domain features. This improves the model's detection accuracy during training while also making the model applicable to polyp detection images within the target domain, enhancing the accuracy of polyp detection in cross-domain situations, and broadening the scope of application of the polyp detection model.
[0159] As another example, the method may also include:
[0160] Based on the feature maps obtained by feature extraction from the source domain image and the target domain image using the backbone network, the adversarial loss is calculated. This adversarial loss can be implemented based on the adversarial loss commonly used in generative adversarial networks in this field, and this disclosure does not limit it.
[0161] Furthermore, determining the target loss for the polyp detection model can include:
[0162] The target loss is determined based on the regression loss, the classification loss, the feature loss, and the adversarial loss, which can further enhance the loss constraint in the target loss, ensure the accuracy and effectiveness of parameter adjustment of the polyp detection model based on the target loss, and improve the training efficiency of the polyp detection model to a certain extent.
[0163] This disclosure also provides a method for detecting polyps, the method comprising:
[0164] The target image to be detected is received, wherein the target image may be an image acquired from an endoscope for real-time detection.
[0165] The target image is input into the trained polyp detection model to obtain the polyp detection result corresponding to the target image. The polyp detection model is trained based on any of the polyp detection model training methods described above.
[0166] The above technical solution enables polyp detection in target images based on a trained polyp detection model, ensuring both accuracy and efficiency. Furthermore, this polyp detection model is trained using both source and target domain images as training samples. During this process, the model enhances foreground feature representation by leveraging the connectivity relationships between polyp foreground features in the image. This reduces detection errors caused by the strong camouflage of polyp foreground features, further broadening the application scope and accuracy of polyp detection methods.
[0167] This disclosure also provides a training device for a polyp detection model, such as... Figure 2 As shown, the device 10 includes:
[0168] The feature extraction module 101 is used to extract foreground and background features from source domain images in the source domain dataset and target domain images in the target domain dataset based on the polyp detection model, to obtain first source domain features and first target domain features. The images in the source domain dataset are labeled with polyp tags, while the images in the target domain dataset are not labeled. The data distribution of the source domain dataset and the target domain dataset is different.
[0169] The first determining module 102 is used to determine the source domain foreground node in the source domain image and the target domain foreground node in the target domain image based on the first source domain features and the first target domain features.
[0170] The second determining module 103 is used to determine the domain connection relationship under the target type based on the domain foreground nodes under each target type, wherein the target type includes a source domain type and a target domain type, and the domain connection relationship includes the source domain connection relationship between source domain foreground nodes under the source domain type and the target domain connection relationship between target domain foreground nodes under the target domain type.
[0171] Update module 104 is used to update the first domain feature under the target type according to the domain connection relationship under each target type to obtain the second domain feature, wherein the first domain feature includes the first source domain feature under the source domain type and the first target domain feature under the target domain type, and the second domain feature includes the second source domain feature obtained by updating the first source domain feature according to the source domain connection relationship and the second target domain feature obtained by updating the first target domain feature according to the target domain connection relationship;
[0172] The acquisition module 105 is used to obtain the prediction result corresponding to the source domain image based on the second source domain features;
[0173] The training module 106 is used to determine the target loss of the polyp detection model based on the prediction result corresponding to the source domain image, the polyp label, the second source domain feature, and the second target domain feature, and to train the polyp detection model based on the target loss.
[0174] Optionally, the first determining module includes:
[0175] The first determining submodule is used to determine the set of candidate foreground nodes corresponding to the first source domain features based on the polyp label;
[0176] The second determining submodule is used to determine, for each candidate foreground node in the candidate foreground node set, the target domain matching foreground node corresponding to the candidate foreground node in the first target domain features, and to determine the source domain matching foreground node in the first source domain features of the target domain matching foreground node.
[0177] The third determining submodule is used to determine the candidate foreground node as the source domain foreground node and the target domain matching foreground node as the target domain foreground node if the source domain matching foreground node belongs to the candidate foreground node set.
[0178] Optionally, the second determining module includes:
[0179] The fourth determination submodule is used to determine the overlap between the first predicted bounding box corresponding to the domain foreground node under each target type and the second predicted bounding box corresponding to each candidate domain foreground node for each domain foreground node under each target type, wherein the candidate domain foreground node is different from the domain foreground node and corresponds to the same target type;
[0180] The fifth determining submodule is used to determine that there is a connection relationship between the domain foreground node and the candidate domain foreground node if the overlap is greater than or equal to a preset threshold. The domain connection relationship under the target type includes the connection relationship corresponding to each domain foreground node under the target type.
[0181] Optionally, the second determining module includes:
[0182] The fourth determination submodule is used to determine, for each domain foreground node under each target type, the overlap between the first predicted bounding box corresponding to the domain foreground node under the target type and the second predicted bounding box corresponding to each candidate domain foreground node, as well as the distance between the center points of the first predicted bounding box and the second predicted bounding box, wherein the candidate domain foreground node is different from the domain foreground node and corresponds to the same target type;
[0183] The sixth determining submodule is used to determine the geometric adjacency matrix and semantic adjacency matrix corresponding to the domain foreground nodes under the target type based on each domain foreground node under the target type, as well as the overlap degree and distance between the domain foreground nodes and the candidate domain foreground nodes.
[0184] The calculation submodule is used to perform a dot product operation on the geometric adjacency matrix and the semantic adjacency matrix, and determine the resulting matrix as the domain connection relationship under the target type.
[0185] Optionally, the sixth determining submodule includes:
[0186] The seventh determining submodule is used to determine, for each of the domain foreground nodes, if the overlap between the domain foreground node and the candidate domain foreground node is greater than a first threshold, or the overlap between the domain foreground node and the candidate domain foreground node is zero and the distance is less than a second threshold, that there is a connection relationship between the domain foreground node and the candidate domain foreground node, and to determine the similarity between the domain foreground node and the candidate domain foreground node as the geometric adjacency relationship value corresponding to the domain foreground node and the candidate domain foreground node, so as to obtain the geometric adjacency relationship matrix;
[0187] The eighth determining submodule is used to determine the similarity between the domain foreground node and the candidate domain foreground node as the semantic adjacency relationship value between the domain foreground node and the candidate domain foreground node for each domain foreground node, so as to obtain the semantic adjacency relationship matrix.
[0188] Optionally, the fourth determining submodule includes:
[0189] The ninth determining submodule is used to determine the first domain prediction border corresponding to the first prediction border in the domain image based on the first prediction border and the domain image to which the domain foreground node belongs.
[0190] The tenth determining submodule is used to determine the second domain predictive bounding box corresponding to the second predictive bounding box in the domain image based on the second predictive bounding box and the domain image;
[0191] The eleventh determination submodule is used to determine the intersection-union ratio (IUU) of the first domain predicted border and the second domain predicted border as the overlap degree.
[0192] Optionally, the update module includes:
[0193] The update submodule is used to perform convolution processing on the first domain features of the target type under each target type based on the domain connectivity relationship under the target type and at least one attention feature layer to obtain the second domain features under the target type, wherein the attention weights corresponding to the attention feature layer under the target type are determined based on the domain connectivity relationship under the target type.
[0194] Optionally, the attention weights in the attention feature layer under the target type are determined based on the domain connectivity under the target type using the following formula:
[0195]
[0196] in, This is used to represent the attention weight between the i-th domain foreground node and the j-th domain foreground node under the target type in the l-th attention feature layer, wherein the i-th domain foreground node and the j-th domain foreground node under the target type have a connection relationship;
[0197] A ij This value represents the connection relationship between the i-th domain foreground node and the j-th domain foreground node under the target type.
[0198] Used to represent the features of the i-th domain foreground node under the target type in the l-th attention feature layer;
[0199] Ni is used to represent the set of domain foreground nodes that have a connection relationship with the i-th domain foreground node under the target type;
[0200] W is used to represent learnable feature transformation operations;
[0201] Used to represent learnable weight vectors;
[0202] || is used to represent feature splicing operations.
[0203] Optionally, the polyp label includes a location label and a classification label;
[0204] The training module includes:
[0205] The regression loss determination submodule is used to determine the regression loss based on the predicted location information and the location label in the prediction results;
[0206] The classification loss determination submodule is used to determine the classification loss based on the predicted classification information in the prediction results and the classification label;
[0207] The feature loss determination submodule is used to determine the feature loss based on the second source domain features and the second target domain features;
[0208] The twelfth determination submodule is used to determine the target loss based on the regression loss, the classification loss, and the feature loss.
[0209] Optionally, the feature loss determination submodule determines the feature loss based on the second source domain features and the second target domain features using the following formula:
[0210]
[0211] Among them, L cst Used to represent the feature loss;
[0212] N sUsed to represent the number of foreground nodes in the source domain;
[0213] N t Used to represent the number of foreground nodes in the target domain;
[0214] x i s Used to represent the feature corresponding to the i-th source domain foreground node in the second source domain features;
[0215] x j t This is used to represent the feature corresponding to the j-th target domain foreground node in the second target domain features.
[0216] This disclosure also provides a polyp detection device, the device comprising:
[0217] The receiving module is used to receive the detected target image;
[0218] The processing module is used to input the target image into the trained polyp detection model to obtain the polyp detection result corresponding to the target image, wherein the polyp detection model is trained based on any of the polyp detection models described above.
[0219] The following is for reference. Figure 3 This diagram illustrates a structural schematic of an electronic device 600 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0220] like Figure 3 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0221] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0222] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0223] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0224] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0225] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0226] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: extract foreground and background features from source domain images in a source domain dataset and target domain images in a target domain dataset using a polyp detection model, obtaining first source domain features and first target domain features, wherein the images in the source domain dataset are labeled with polyp tags, the images in the target domain dataset are not labeled, and the data distributions of the source domain dataset and the target domain dataset are different; determine source domain foreground nodes in the source domain images and target domain foreground nodes in the target domain images based on the first source domain features and the first target domain features; and determine domain connectivity relationships under each target type based on the domain foreground nodes, wherein the target type includes source domain type and target domain type, and the domain connectivity relationships include relationships between source domain foreground nodes under the source domain type. The source domain connectivity relationship and the target domain connectivity relationship between the foreground nodes of the target domain under the target domain type are defined. The first domain feature under each target type is updated according to the domain connectivity relationship to obtain a second domain feature. The first domain feature includes a first source domain feature under the source domain type and a first target domain feature under the target domain type. The second domain feature includes a second source domain feature obtained by updating the first source domain feature according to the source domain connectivity relationship and a second target domain feature obtained by updating the first target domain feature according to the target domain connectivity relationship. A prediction result corresponding to the source domain image is obtained based on the second source domain feature. The target loss of the polyp detection model is determined based on the prediction result corresponding to the source domain image, the polyp label, the second source domain feature, and the second target domain feature. The polyp detection model is then trained based on the target loss.
[0227] Alternatively, the aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: receive a detected target image; input the target image into a trained polyp detection model to obtain a polyp detection result corresponding to the target image, wherein the polyp detection model is trained based on the polyp detection model training method described above.
[0228] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0229] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0230] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the modules do not necessarily limit the module itself; for example, a feature extraction module can also be described as "a module that extracts foreground and background features from source domain images in a source domain dataset and target domain images in a target domain dataset based on a polyp detection model, to obtain first source domain features and first target domain features."
[0231] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0232] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0233] According to one or more embodiments of this disclosure, Example 1 provides a method for training a polyp detection model, wherein the method includes:
[0234] Based on the polyp detection model, foreground and background features are extracted from source domain images in the source domain dataset and target domain images in the target domain dataset to obtain first source domain features and first target domain features. The images in the source domain dataset are labeled with polyp tags, while the images in the target domain dataset are not labeled. The data distributions of the source domain dataset and the target domain dataset are different.
[0235] Based on the first source domain features and the first target domain features, source domain foreground nodes in the source domain image and target domain foreground nodes in the target domain image are determined.
[0236] Based on the domain foreground nodes under each target type, determine the domain connection relationship under the target type, wherein the target type includes a source domain type and a target domain type, and the domain connection relationship includes the source domain connection relationship between source domain foreground nodes under the source domain type and the target domain connection relationship between target domain foreground nodes under the target domain type;
[0237] The first domain feature under each target type is updated according to the domain connection relationship under each target type to obtain the second domain feature. The first domain feature includes the first source domain feature under the source domain type and the first target domain feature under the target domain type. The second domain feature includes the second source domain feature obtained by updating the first source domain feature according to the source domain connection relationship and the second target domain feature obtained by updating the first target domain feature according to the target domain connection relationship.
[0238] The prediction result corresponding to the source domain image is obtained based on the second source domain feature;
[0239] Based on the prediction results corresponding to the source domain image, the polyp label, the second source domain features, and the second target domain features, the target loss of the polyp detection model is determined, and the polyp detection model is trained based on the target loss.
[0240] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein determining the source domain foreground node in the source domain image and the target domain foreground node in the target domain image based on the first source domain features and the first target domain features includes:
[0241] The set of candidate foreground nodes corresponding to the first source domain feature is determined based on the polyp label;
[0242] For each candidate foreground node in the candidate foreground node set, determine the target domain matching foreground node in the first target domain feature corresponding to the candidate foreground node, and determine the source domain matching foreground node in the first source domain feature of the target domain matching foreground node;
[0243] If the source domain matching foreground node belongs to the candidate foreground node set, the candidate foreground node is determined as the source domain foreground node, and the target domain matching foreground node is determined as the target domain foreground node.
[0244] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 1, wherein,
[0245] The step of determining the domain connectivity relationship under each target type based on the domain foreground nodes includes:
[0246] For each domain foreground node under each target type, the overlap between the first predicted bounding box corresponding to the domain foreground node under the target type and the second predicted bounding box corresponding to each candidate domain foreground node is determined. The candidate domain foreground nodes are different from the domain foreground nodes and correspond to the same target type.
[0247] If the overlap is greater than or equal to a preset threshold, it is determined that there is a connection between the domain foreground node and the candidate domain foreground node. The domain connection relationship under the target type includes the connection relationship corresponding to each domain foreground node under the target type.
[0248] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 1, wherein determining the domain connectivity relationship under the target type based on the domain foreground nodes under each target type includes:
[0249] For each domain foreground node under each target type, determine the overlap between the first predicted bounding box corresponding to the domain foreground node under the target type and the second predicted bounding box corresponding to each candidate domain foreground node, as well as the distance between the center points of the first predicted bounding box and the second predicted bounding box. The candidate domain foreground node is different from the domain foreground node and corresponds to the same target type.
[0250] Based on each domain foreground node under the target type, and the overlap and distance between the domain foreground node and the candidate domain foreground node, determine the geometric adjacency matrix and semantic adjacency matrix corresponding to the domain foreground node under the target type;
[0251] Perform a dot product operation on the geometric adjacency matrix and the semantic adjacency matrix, and determine the resulting matrix as the domain connectivity under the target type.
[0252] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 4, wherein determining the geometric adjacency matrix and semantic adjacency matrix corresponding to the domain foreground nodes under the target type based on each domain foreground node under the target type, and the overlap degree and distance between the domain foreground nodes and the candidate domain foreground nodes, includes:
[0253] For each of the domain foreground nodes, if the overlap between the domain foreground node and the candidate domain foreground node is greater than a first threshold, or if the overlap between the domain foreground node and the candidate domain foreground node is zero and the distance is less than a second threshold, then it is determined that there is a connection between the domain foreground node and the candidate domain foreground node, and the similarity between the domain foreground node and the candidate domain foreground node is determined as the geometric adjacency value between the domain foreground node and the candidate domain foreground node, so as to obtain the geometric adjacency matrix;
[0254] For each of the domain foreground nodes, if the similarity between the domain foreground node and the candidate domain foreground node is greater than a third threshold, then the similarity is determined as the semantic adjacency value between the domain foreground node and the candidate domain foreground node, so as to obtain the semantic adjacency matrix.
[0255] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 3 or 4, wherein,
[0256] Determining the overlap between the first predicted bounding box corresponding to the foreground node of the target type and the second predicted bounding box corresponding to each candidate foreground node includes:
[0257] Based on the first predicted bounding box and the domain image to which the domain foreground node belongs, determine the first domain predicted bounding box corresponding to the first predicted bounding box in the domain image.
[0258] Based on the second predicted bounding box and the domain image, determine the second domain predicted bounding box corresponding to the second predicted bounding box in the domain image;
[0259] The overlap ratio corresponding to the first domain predicted bounding box and the second domain predicted bounding box is determined as the overlap degree.
[0260] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 1, wherein,
[0261] The step of updating the first domain feature under each target type according to the domain connection relationship to obtain the second domain feature includes:
[0262] Under each target type, the first domain feature under the target type is convolved based on the domain connectivity relationship under the target type and at least one attention feature layer to obtain the second domain feature under the target type, wherein the attention weights corresponding to the attention feature layer under the target type are determined based on the domain connectivity relationship under the target type.
[0263] According to one or more embodiments of this disclosure, Example 8 provides the method of Example 7, wherein the attention weights corresponding to the attention feature layer under the target type are determined based on the domain connectivity under the target type by the following formula:
[0264]
[0265] in, This is used to represent the attention weight between the i-th domain foreground node and the j-th domain foreground node under the target type in the l-th attention feature layer, wherein the i-th domain foreground node and the j-th domain foreground node under the target type have a connection relationship;
[0266] A ij This value represents the connection relationship between the i-th domain foreground node and the j-th domain foreground node under the target type.
[0267] Used to represent the features of the i-th domain foreground node under the target type in the l-th attention feature layer;
[0268] Ni is used to represent the set of domain foreground nodes that have a connection relationship with the i-th domain foreground node under the target type;
[0269] W is used to represent learnable feature transformation operations;
[0270] Used to represent learnable weight vectors;
[0271] || is used to represent feature splicing operations.
[0272] According to one or more embodiments of this disclosure, Example 9 provides the method of Example 1, wherein the polyp label includes a location label and a classification label;
[0273] The step of determining the target loss of the polyp detection model based on the prediction result corresponding to the source domain image, the polyp label, the second source domain features, and the second target domain features includes:
[0274] The regression loss is determined based on the predicted location information and the location label in the prediction results;
[0275] The classification loss is determined based on the predicted classification information and the classification label in the prediction results;
[0276] The feature loss is determined based on the second source domain features and the second target domain features;
[0277] The target loss is determined based on the regression loss, the classification loss, and the feature loss.
[0278] According to one or more embodiments of this disclosure, Example 10 provides the method of Example 9, wherein the feature loss is determined based on the second source domain features and the second target domain features using the following formula:
[0279]
[0280] Among them, L cst Used to represent the feature loss;
[0281] N s Used to represent the number of foreground nodes in the source domain;
[0282] N t Used to represent the number of foreground nodes in the target domain;
[0283] x i s Used to represent the feature corresponding to the i-th source domain foreground node in the second source domain features;
[0284] x j t This is used to represent the feature corresponding to the j-th target domain foreground node in the second target domain features.
[0285] According to one or more embodiments of this disclosure, Example 11 provides a method for detecting polyps, the method comprising:
[0286] Received the target image for detection;
[0287] The target image is input into the trained polyp detection model to obtain the polyp detection result corresponding to the target image. The polyp detection model is trained based on the training method of the polyp detection model in any one of Examples 1-10.
[0288] According to one or more embodiments of this disclosure, Example 12 provides a training apparatus for a polyp detection model, the apparatus comprising:
[0289] The feature extraction module is used to extract foreground and background features from source domain images in the source domain dataset and target domain images in the target domain dataset based on the polyp detection model, to obtain first source domain features and first target domain features. The images in the source domain dataset are labeled with polyp tags, while the images in the target domain dataset are not labeled. The data distributions of the source domain dataset and the target domain dataset are different.
[0290] The first determining module is used to determine the source domain foreground node in the source domain image and the target domain foreground node in the target domain image based on the first source domain features and the first target domain features.
[0291] The second determining module is used to determine the domain connection relationship under the target type based on the domain foreground nodes under each target type, wherein the target type includes a source domain type and a target domain type, and the domain connection relationship includes the source domain connection relationship between source domain foreground nodes under the source domain type and the target domain connection relationship between target domain foreground nodes under the target domain type.
[0292] An update module is used to update the first domain feature under each target type according to the domain connection relationship under each target type to obtain a second domain feature, wherein the first domain feature includes the first source domain feature under the source domain type and the first target domain feature under the target domain type, and the second domain feature includes the second source domain feature obtained by updating the first source domain feature according to the source domain connection relationship and the second target domain feature obtained by updating the first target domain feature according to the target domain connection relationship;
[0293] The acquisition module is used to obtain the prediction result corresponding to the source domain image based on the second source domain features;
[0294] The training module is used to determine the target loss of the polyp detection model based on the prediction results corresponding to the source domain image, the polyp label, the second source domain features, and the second target domain features, and to train the polyp detection model based on the target loss.
[0295] According to one or more embodiments of this disclosure, Example 13 provides a polyp detection device, the device comprising:
[0296] The receiving module is used to receive the detected target image;
[0297] The processing module is used to input the target image into the trained polyp detection model to obtain the polyp detection result corresponding to the target image, wherein the polyp detection model is trained based on the training method of the polyp detection model in any one of Examples 1-10.
[0298] According to one or more embodiments of the present disclosure, Example 14 provides a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processing device, implements the steps of the method described in any one of Examples 1-11.
[0299] According to one or more embodiments of this disclosure, Example 15 provides an electronic device, including:
[0300] A storage device on which computer programs are stored;
[0301] A processing device for executing the computer program in the storage device to implement the steps of any one of the methods in Examples 1-11.
[0302] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0303] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0304] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. A training method for a polyp detection model, characterized in that, The method includes: Based on the polyp detection model, foreground and background features are extracted from source domain images in the source domain dataset and target domain images in the target domain dataset to obtain first source domain features and first target domain features. The images in the source domain dataset are labeled with polyp tags, while the images in the target domain dataset are not labeled. The data distributions of the source domain dataset and the target domain dataset are different. Based on the first source domain features and the first target domain features, source domain foreground nodes in the source domain image and target domain foreground nodes in the target domain image are determined. Based on the domain foreground nodes under each target type, determine the domain connection relationship under the target type, wherein the target type includes a source domain type and a target domain type, and the domain connection relationship includes the source domain connection relationship between source domain foreground nodes under the source domain type and the target domain connection relationship between target domain foreground nodes under the target domain type; Under each target type, the first domain feature under the target type is convolved based on the domain connectivity relationship under the target type and at least one attention feature layer to obtain the second domain feature under the target type. The first domain feature includes the first source domain feature under the source domain type and the first target domain feature under the target domain type. The second domain feature includes the second source domain feature obtained by updating the first source domain feature according to the source domain connectivity relationship and the second target domain feature obtained by updating the first target domain feature according to the target domain connectivity relationship. The prediction result corresponding to the source domain image is obtained based on the second source domain feature; Based on the prediction results corresponding to the source domain image, the polyp label, the second source domain features, and the second target domain features, the target loss of the polyp detection model is determined, and the polyp detection model is trained based on the target loss. The attention weights in the attention feature layer under the target type are determined based on the domain connectivity under the target type using the following formula: in, Used to indicate the first l The first attention feature layer describes the target type under the [number]th [item]. i Foreground nodes in each domain and the first j Attention weights among foreground nodes in the domain, wherein the first node under the target type... i Foreground nodes in each domain and the first j There are connections between the foreground nodes in each domain; Used to represent the first under the target type i Foreground nodes in each domain and the first j The connection relationship values between foreground nodes in each domain; Used to represent the first under the target type i The foreground node in the domain is at the l Features in an attention feature layer; Used to represent the first under the target type i A set of foreground nodes that have connections to each other; Used to represent learnable feature transformation operations; Used to represent learnable weight vectors; || is used to represent feature splicing operations.
2. The method according to claim 1, characterized in that, The step of determining the source domain foreground nodes in the source domain image and the target domain foreground nodes in the target domain image based on the first source domain features and the first target domain features includes: The set of candidate foreground nodes corresponding to the first source domain feature is determined based on the polyp label; For each candidate foreground node in the candidate foreground node set, determine the target domain matching foreground node in the first target domain feature corresponding to the candidate foreground node, and determine the source domain matching foreground node in the first source domain feature of the target domain matching foreground node; If the source domain matching foreground node belongs to the candidate foreground node set, the candidate foreground node is determined as the source domain foreground node, and the target domain matching foreground node is determined as the target domain foreground node.
3. The method according to claim 1, characterized in that, The step of determining the domain connectivity relationship under each target type based on the domain foreground nodes includes: For each domain foreground node under each target type, the overlap between the first predicted bounding box corresponding to the domain foreground node under the target type and the second predicted bounding box corresponding to each candidate domain foreground node is determined. The candidate domain foreground nodes are different from the domain foreground nodes and correspond to the same target type. If the overlap is greater than or equal to a preset threshold, it is determined that there is a connection between the domain foreground node and the candidate domain foreground node. The domain connection relationship under the target type includes the connection relationship corresponding to each domain foreground node under the target type.
4. The method according to claim 1, characterized in that, The step of determining the domain connectivity relationship under each target type based on the domain foreground nodes includes: For each domain foreground node under each target type, determine the overlap between the first predicted bounding box corresponding to the domain foreground node under the target type and the second predicted bounding box corresponding to each candidate domain foreground node, as well as the distance between the center points of the first predicted bounding box and the second predicted bounding box. The candidate domain foreground node is different from the domain foreground node and corresponds to the same target type. Based on each domain foreground node under the target type, and the overlap and distance between the domain foreground node and the candidate domain foreground node, determine the geometric adjacency matrix and semantic adjacency matrix corresponding to the domain foreground node under the target type; Perform a dot product operation on the geometric adjacency matrix and the semantic adjacency matrix, and determine the resulting matrix as the domain connectivity under the target type.
5. The method according to claim 4, characterized in that, The step of determining the geometric adjacency matrix and semantic adjacency matrix corresponding to the domain foreground nodes under the target type based on each domain foreground node under the target type, and the overlap degree and distance between the domain foreground nodes and the candidate domain foreground nodes, includes: For each of the domain foreground nodes, if the overlap between the domain foreground node and the candidate domain foreground node is greater than a first threshold, or if the overlap between the domain foreground node and the candidate domain foreground node is zero and the distance is less than a second threshold, then it is determined that there is a connection between the domain foreground node and the candidate domain foreground node, and the similarity between the domain foreground node and the candidate domain foreground node is determined as the geometric adjacency value between the domain foreground node and the candidate domain foreground node, so as to obtain the geometric adjacency matrix; For each of the domain foreground nodes, if the similarity between the domain foreground node and the candidate domain foreground node is greater than a third threshold, then the similarity is determined as the semantic adjacency value between the domain foreground node and the candidate domain foreground node, so as to obtain the semantic adjacency matrix.
6. The method according to claim 3 or 4, characterized in that, Determining the overlap between the first predicted bounding box corresponding to the foreground node of the target type and the second predicted bounding box corresponding to each candidate foreground node includes: Based on the first predicted bounding box and the domain image to which the domain foreground node belongs, determine the first domain predicted bounding box corresponding to the first predicted bounding box in the domain image. Based on the second predicted bounding box and the domain image, determine the second domain predicted bounding box corresponding to the second predicted bounding box in the domain image; The overlap ratio corresponding to the first domain predicted bounding box and the second domain predicted bounding box is determined as the overlap degree.
7. The method according to claim 1, characterized in that, The polyp label includes a location label and a classification label; The step of determining the target loss of the polyp detection model based on the prediction result corresponding to the source domain image, the polyp label, the second source domain features, and the second target domain features includes: The regression loss is determined based on the predicted location information and the location label in the prediction results; The classification loss is determined based on the predicted classification information and the classification label in the prediction results; The feature loss is determined based on the second source domain features and the second target domain features; The target loss is determined based on the regression loss, the classification loss, and the feature loss.
8. The method according to claim 7, characterized in that, The feature loss is determined using the following formula based on the features of the second source domain and the features of the second target domain: in, Used to represent the feature loss; Used to represent the number of foreground nodes in the source domain; Used to represent the number of foreground nodes in the target domain; Used to represent the second source domain feature of the first i Features corresponding to each source domain foreground node; Used to represent the first feature in the second target domain j Features corresponding to each foreground node in the target domain.
9. A method for detecting polyps, characterized in that, The method includes: Received the target image for detection; The target image is input into the trained polyp detection model to obtain the polyp detection result corresponding to the target image, wherein the polyp detection model is trained based on the training method of the polyp detection model according to any one of claims 1-8.
10. A training device for a polyp detection model, characterized in that, The device includes: The feature extraction module is used to extract foreground and background features from source domain images in the source domain dataset and target domain images in the target domain dataset based on the polyp detection model, to obtain first source domain features and first target domain features. The images in the source domain dataset are labeled with polyp tags, while the images in the target domain dataset are not labeled. The data distributions of the source domain dataset and the target domain dataset are different. The first determining module is used to determine the source domain foreground node in the source domain image and the target domain foreground node in the target domain image based on the first source domain features and the first target domain features. The second determining module is used to determine the domain connection relationship under the target type based on the domain foreground nodes under each target type, wherein the target type includes a source domain type and a target domain type, and the domain connection relationship includes the source domain connection relationship between source domain foreground nodes under the source domain type and the target domain connection relationship between target domain foreground nodes under the target domain type. An update module is configured to, under each target type, perform convolution processing on the first domain features of the target type based on the domain connectivity relationship under the target type and at least one attention feature layer to obtain the second domain features of the target type. The first domain features include the first source domain features under the source domain type and the first target domain features under the target domain type. The second domain features include the second source domain features obtained by updating the first source domain features according to the source domain connectivity relationship and the second target domain features obtained by updating the first target domain features according to the target domain connectivity relationship. The acquisition module is used to obtain the prediction result corresponding to the source domain image based on the second source domain features; The training module is used to determine the target loss of the polyp detection model based on the prediction results corresponding to the source domain image, the polyp label, the second source domain features, and the second target domain features, and to train the polyp detection model based on the target loss. The attention weights in the attention feature layer under the target type are determined based on the domain connectivity under the target type using the following formula: in, Used to indicate the first l The first attention feature layer describes the target type under the [number]th [item]. i Foreground nodes in each domain and the first j Attention weights among foreground nodes in the domain, wherein the first node under the target type... i Foreground nodes in each domain and the first j There are connections between the foreground nodes in each domain; Used to represent the first under the target type i Foreground nodes in each domain and the first j The connection relationship values between foreground nodes in each domain; Used to represent the first under the target type i The foreground node in the domain is at the l Features in an attention feature layer; Used to represent the first under the target type i A set of foreground nodes that have connections to each other; Used to represent learnable feature transformation operations; Used to represent learnable weight vectors; || is used to represent feature splicing operations.
11. A polyp detection device, characterized in that, The device includes: The receiving module is used to receive the detected target image; The processing module is used to input the target image into the trained polyp detection model to obtain the polyp detection result corresponding to the target image, wherein the polyp detection model is trained based on the training method of the polyp detection model according to any one of claims 1-8.
12. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the method according to any one of claims 1-9.
13. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-9.