Image processing method, image processing device, endoscope and biliary navigation system

Images are captured through an endoscopy and feature position relationship library and image processing model are used to automatically identify and update the position information of the biliary system, solving the positioning problem in biliary system shooting and improving shooting accuracy and efficiency.

CN118691580BActive Publication Date: 2025-08-15BEIJING UNIV OF CHINESE MEDICINE THIRD AFFILIATED HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410821378.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2025-08-15
Estimated Expiration
2044-06-24

AI Technical Summary

Technical Problem

During the shooting of complex visceral structures such as biliary duct systems, it is difficult to accurately locate each branch, and it is easy to miss or repeat the shooting.

Method used

By acquiring the images taken by the endoscopy, using the feature position relationship library and position naming rules to generate target position information, update and display the current and entrance positions in real time, combine the image processing model to identify key points and feature vectors, and automatically locate various entrances of the biliary system.

Benefits of technology

It improves the accuracy and efficiency of the shooting of the biliary system, avoids missed or repeated shooting, and achieves real-time positioning and accurate positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118691580B_ABST
    Figure CN118691580B_ABST
Patent Text Reader

Abstract

The present application relates to the field of image processing, and more specifically, to an image processing method, an image processing device, an endoscope, and a biliary navigation system. The method comprises acquiring an image to be processed; generating target position information corresponding to the image to be processed according to a preset position naming rule, updating the target position information into the feature position relationship library, and displaying the target position information on the image to be processed. The feature position relationship library can be automatically updated, so that when a person photographs the bile duct, he or she can not only know the current shooting position in real time, but also automatically update the feature position relationship library to facilitate the subsequent identification of the various entrances of the bile duct.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and more specifically, to an image processing method, an image processing device, an endoscope, and a biliary navigation system. Background Art

[0002] With the advancement of technology, image capture technology is becoming increasingly sophisticated. For example, in the medical field, miniature cameras can capture images of internal organs, allowing medical personnel to more intuitively observe their condition. However, the current imaging process often relies on medical personnel to determine the shooting route independently. For complex internal organs, such as the biliary system, which has numerous branches, it is very easy to miss a branch during the imaging process or to capture the same branch repeatedly, making it difficult to locate the location. Summary of the Invention

[0003] The present application provides an image processing method, an image processing device, an endoscope and a biliary navigation system to at least solve the technical problem of difficult positioning.

[0004] According to a first aspect of an embodiment of the present application, there is provided an image processing method, comprising:

[0005] Acquire an image to be processed captured by an endoscope within a cavity of a target detection object, wherein the image to be processed is an image in a video captured within the cavity of the target detection object;

[0006] When the preset feature position relationship library is empty, target position information corresponding to the image to be processed is generated according to the preset position naming rules and the target position information is updated to the feature position relationship library, and the target position information is displayed on the image to be processed, wherein the target position information includes current position information and / or entrance position information, and the feature position relationship library is used to store the feature vectors of each entrance in the target detection object cavity and the position information of each entrance in the target detection object cavity, wherein the current position information represents the shooting position of the endoscope in the target detection object cavity, and the entrance position information represents the position of the entrance appearing in the image to be processed in the target detection object cavity, and the target detection object cavity includes several branches, and the entrance refers to the connecting port of each branch.

[0007] Optionally, the method further includes:

[0008] When the feature position relationship library is not empty, the target position information corresponding to the image to be processed is determined according to the shooting direction of the endoscope in the cavity of the target detection object, the current position information of the image to be processed and the feature position relationship library, so as to display the target position information on the image to be processed.

[0009] Optionally, determining the target position information corresponding to the image to be processed according to the shooting direction of the endoscope in the cavity of the target detection object, the current position information of the image to be processed, and the feature position relationship library includes:

[0010] When the current position information does not belong to the current first-layer position and last-layer position in the feature position relationship library, determining whether the image to be processed contains an entrance, wherein the position information of each entrance in the feature position relationship library is stored according to the hierarchy of the branch to which each entrance belongs, with the top-layer position information being the first-layer position and the bottom-layer position information being the last-layer position;

[0011] If no entrance is included, determining the current position information as the target position information and displaying it on the image to be processed;

[0012] If an entrance is included, the entrance position information of each entrance is determined according to the feature position relationship library.

[0013] Optionally, determining the target position information corresponding to the image to be processed according to the shooting direction of the endoscope in the cavity of the target detection object, the current position information of the image to be processed, and the feature position relationship library includes:

[0014] When the current position information belongs to the first layer position or the last layer position in the feature position relationship library, determining whether the image to be processed contains an entrance;

[0015] If no entrance is included, determining the current position information as the target position information and displaying it on the image to be processed;

[0016] If an entrance is included, when the shooting direction is away from the first-floor position or the last-floor position, the entrance position information of each entrance is determined according to the feature position relationship library; when the shooting direction is close to the first-floor position or the last-floor position, the entrance position information of each entrance is generated according to the position naming rule.

[0017] Optionally, determining the entrance position information of each entrance according to the feature position relationship library includes:

[0018] Determining a local image capable of representing the entrance position of the target detection object from the image to be processed;

[0019] Determining key points in the local image and generating feature vectors for the key points based on position information of the key points;

[0020] According to the feature vector, the feature position relationship library is retrieved to obtain position information matching the feature vector as the entry position information.

[0021] Optionally, determining a local image capable of representing the entrance position of the target detection object from the image to be processed includes:

[0022] The image to be processed is processed using a trained image processing model to determine the local image, wherein the image processing model determines the completion of training according to a first loss value during training, wherein the first loss value includes a second loss value for selecting a rectangular box for the local image and a third loss value for the confidence of the rectangular box.

[0023] Optionally, before retrieving position information matching the feature vector from the feature position relationship library according to the feature vector as the entry position information, the method further includes:

[0024] If the feature vector identical to the feature vector is not stored in the feature position relationship library, the entry position information is generated according to the current position information and the position naming rule.

[0025] Optionally, before processing the image to be processed using the trained image processing model to determine the local image, the method further includes:

[0026] Processing the training image using the image processing model to be trained to obtain a prediction rectangular box and prediction information of the prediction rectangular box, wherein the image within the prediction rectangular box is a local image determined by the image processing model to be trained that can represent the entrance position of the target detection object, and the prediction information includes the predicted position and the prediction confidence;

[0027] calculating a distance between the predicted position and a reference position of a reference rectangular frame of the training image using a first loss function to obtain a second loss value, wherein the reference rectangular frame is a rectangular frame pre-determined from the training image, and the reference rectangular frame has reference information, the reference information including a reference position and a reference confidence;

[0028] Calculating the prediction confidence of the prediction rectangular box and the reference confidence of the reference rectangular box using the second loss function to obtain the third loss value;

[0029] Calculate the first loss value according to the second loss value and the third loss value;

[0030] The training of the image processing model is completed according to the matching between the first loss value and the preset training completion condition to obtain a trained image processing model.

[0031] Optionally, the prediction information further includes a prediction classification probability, and the reference information further includes a reference classification probability;

[0032] Before processing the image to be processed using the trained image processing model to determine the local image, the method further includes:

[0033] Processing the training image using the image processing model to be trained to obtain a prediction rectangular box and prediction information of the prediction rectangular box, wherein the image within the prediction rectangular box is a local image determined by the image processing model to be trained that can represent the entrance position of the target detection object, and the prediction information includes the predicted position and the prediction confidence;

[0034] calculating a distance between the predicted position and a reference position of a reference rectangular frame of the training image using a first loss function to obtain a second loss value, wherein the reference rectangular frame is a rectangular frame pre-determined from the training image, and the reference rectangular frame has reference information, the reference information including a reference position and a reference confidence;

[0035] Calculating the prediction confidence of the prediction rectangular box and the reference confidence of the reference rectangular box using the second loss function to obtain the third loss value;

[0036] Calculating the predicted classification probability and the reference classification probability using the second loss function to obtain a fourth loss value;

[0037] Performing a weighted summation on the second loss value, the third loss value, and the fourth loss value to obtain the first loss value;

[0038] The training of the image processing model is completed according to the matching between the first loss value and the preset training completion condition to obtain a trained image processing model.

[0039] Optionally, determining key points in the local image and generating feature vectors for the key points according to position information of the key points includes:

[0040] Detecting points of interest in the local image using a scale-invariant feature variation algorithm;

[0041] Determining the key points from the points of interest using the scale-invariant feature change algorithm and obtaining the point information, wherein the point information includes the position of the key points within a rectangular box, the scale of the rectangular box to which the key points belong, and the direction of the key points, wherein the rectangular box is used to frame an area in the image to be processed to obtain the local image;

[0042] The feature vector is generated according to the point position, scale and point direction using the scale-invariant feature change algorithm.

[0043] Optionally, the retrieving, according to the feature vector, from the feature position relationship library the position information that matches the feature vector as the entry position information includes:

[0044] Calculating the distances between the feature vector of the key point and the feature vectors of each entry of the target detection object in the feature position relationship library, and determining the target detection object positions closest and second closest to the key point in the feature position relationship library;

[0045] If the ratio of the feature vector of the closest target detection object position to the feature vector of the second closest target detection object position is less than a preset ratio threshold, it is determined that the feature vector of the closest target detection object position matches the feature vector of the key point;

[0046] The position information corresponding to the feature vector of the target detection object position closest to the target detection object is retrieved as the entry position information.

[0047] Optionally, the calculating of the distances between the feature vector of the key point and the feature vector of each position of the target detection object in the feature position relationship library includes:

[0048] Calculate the Euclidean distance between the feature vector of the key point and the feature vector of each position of the target detection object.

[0049] Optionally, the target detection object includes the bile duct.

[0050] According to a second aspect of an embodiment of the present application, there is provided an image processing device, comprising an acquisition module for acquiring an image to be processed captured by an endoscope within a cavity of a target detection object, wherein the image to be processed is an image in a video captured within the cavity of the target detection object;

[0051] a first judgment module, configured to generate target position information corresponding to the image to be processed according to a preset position naming rule and update the target position information to the feature position relationship library when the preset feature position relationship library is empty, and display the target position information on the image to be processed, wherein the target position information includes the current position information and / or entrance position information, and the feature position relationship library includes the feature vector of each entrance in the target detection object cavity and the position information of each entrance in the target detection object cavity, wherein the current position information represents the shooting position of the endoscope in the target detection object cavity, and the entrance position information represents the position of the entrance appearing in the image to be processed in the target detection object cavity;

[0052] The second judgment module is used to determine the target position information corresponding to the image to be processed based on the shooting direction of the video in the cavity of the target detection object and the current position information of the image to be processed when the feature position relationship library is not empty, so as to display the target position information on the image to be processed.

[0053] Optionally, the second judgment module includes a second judgment unit configured to judge whether the image to be processed contains an entrance when the current position information does not belong to the current first-layer position and last-layer position in the feature position relationship library;

[0054] If there is no information capable of characterizing the entrances in the cavity of the target detection object, determining the current position information as the target position information and displaying it on the image to be processed;

[0055] If there is information that can characterize each entrance in the cavity of the target detection object, determining the entrance position information of each entrance according to the characteristic position relationship library;

[0056] When the current position information belongs to the current first layer position and the last layer position in the feature position relationship library, determining whether the image to be processed contains an entrance;

[0057] If there is no information capable of characterizing the entrances in the cavity of the target detection object, determining the current position information as the target position information and displaying it on the image to be processed;

[0058] If there is information that can characterize the entrances in the cavity of the target detection object, then when the shooting direction is away from the first-layer position or the last-layer position, the entrance position information of each entrance is determined according to the feature position relationship library; when the shooting direction is close to the first-layer position or the last-layer position, the entrance position information of each entrance is generated according to the position naming rule.

[0059] Optionally, the second judgment unit includes a local unit, configured to determine a local image capable of representing the entrance position of the target detection object from the image to be processed;

[0060] A vector unit, configured to determine key points in the local image and generate a feature vector for the key points based on position information of the key points;

[0061] The location unit is used to retrieve location information matching the feature vector from the feature location relationship library according to the feature vector as the entry location information.

[0062] Optionally, the local unit includes a model unit, which is used to process the image to be processed using a trained image processing model to determine the local image, wherein the image processing model determines the training completion status based on a first loss value during training, wherein the first loss value includes a second loss value for selecting a rectangular box for the local image and a third loss value for the confidence of the rectangular box.

[0063] Optionally, the model unit includes a first subunit, configured to process a training image using the image processing model to be trained to obtain a prediction rectangular box and prediction information of the prediction rectangular box, wherein the image within the prediction rectangular box is a local image determined by the image processing model to be trained that can characterize the entrance position of the target detection object, and the prediction information includes a predicted position and a prediction confidence;

[0064] a second subunit, configured to calculate a distance between the predicted position and a reference position of a reference rectangular frame of the training image using a first loss function to obtain a second loss value, wherein the reference rectangular frame is a rectangular frame pre-determined from the training image, and the reference rectangular frame has reference information, the reference information including a reference position and a reference confidence;

[0065] a third subunit, configured to calculate the prediction confidence of the prediction rectangular box and the reference confidence of the reference rectangular box using a second loss function to obtain the third loss value;

[0066] A fourth subunit, configured to calculate the first loss value according to the second loss value and the third loss value;

[0067] The fifth subunit is used to complete the training of the image processing model according to the matching between the first loss value and the preset training completion condition to obtain a trained image processing model.

[0068] Optionally, the prediction information further includes a prediction classification probability, and the reference information further includes a reference classification probability;

[0069] The apparatus further includes a sixth subunit, configured to calculate the predicted classification probability and the reference classification probability using the second loss function to obtain a fourth loss value;

[0070] The seventh subunit is used to perform weighted summation on the second loss value, the third loss value and the fourth loss value to obtain the first loss value.

[0071] The eighth subunit is used to complete the training of the image processing model according to the matching between the first loss value and the preset training completion condition to obtain a trained image processing model.

[0072] Optionally, the vector unit includes an interest point unit, configured to detect interest points in the local image using a scale-invariant feature change algorithm;

[0073] a determining unit, configured to determine the key point from the points of interest using the scale-invariant feature change algorithm and obtain the point position information, wherein the point position information includes the point position of the key point within a rectangular frame, the scale of the rectangular frame to which the key point belongs, and the point direction of the key point, wherein the rectangular frame is used to frame an area in the image to be processed to obtain the local image;

[0074] A vector unit is used to generate the feature vector according to the point position, scale and point direction using the scale-invariant feature change algorithm.

[0075] Optionally, the position unit includes a distance unit, which is used to calculate the distance between the feature vector of the key point and the feature vector of each entry position of the target detection object in the feature position relationship library, and determine the target detection object position closest and second closest to the key point in the feature position relationship library;

[0076] a ratio unit, configured to determine that the feature vector of the closest target detection object position matches the feature vector of the key point if a ratio of the feature vector of the closest target detection object position to the feature vector of the second closest target detection object position is less than a preset ratio threshold;

[0077] The retrieving unit is used to retrieve the position information corresponding to the feature vector of the position of the target detection object that is closest to the target detection object as the entry position information.

[0078] Optionally, the distance unit includes a distance subunit, which is used to calculate the Euclidean distance between the feature vector of the key point and the feature vector of each position of the target detection object.

[0079] According to a third aspect of the embodiments of the present application, an image processing device is provided, which adopts the above-mentioned image processing method.

[0080] According to a fourth aspect of an embodiment of the present application, an endoscope is provided, the endoscope being used to capture a video within the cavity of a bile duct and using the above-described image processing method to display target position information on each frame of the video;

[0081] The endoscope includes a display or is connected to a display, and the display is used to display a video with the target position information.

[0082] According to a fifth aspect of the embodiments of the present application, there is provided a biliary navigation system, comprising the above-mentioned image processing device and endoscope;

[0083] The endoscope is used to capture video within the bile duct cavity and transmit the video to the image processing device;

[0084] The image processing device is used to display target position information on each frame of the video.

[0085] According to a sixth aspect of an embodiment of the present application, a medical robot is provided, which is used to control an endoscope to move in the cavity of a bile duct and shoot a video in the cavity of the bile duct, determine the shooting position of each frame image in the video in the cavity of the bile duct according to a preset feature position relationship library, and display the shooting position on the corresponding frame image; wherein, the bile duct includes multiple layers of branches, and there are connecting ports between the branches of adjacent layers. When the frame image is the first frame image of the video, if the first frame image does not include the connecting port, the shooting position of the first frame image is generated according to a preset position naming rule, and the shooting position is displayed on the first frame image; if the first frame image includes the connecting port, the feature vector of the connecting port is determined, and the hierarchical position of the connecting port is generated according to the moving direction of the endoscope and the position naming rule, the feature vector and the hierarchical position of the connecting port in the first frame image are stored in the feature position relationship library, and the shooting position and the hierarchical position of the first frame image are displayed on the first frame image.

[0086] When the frame image is not the first frame image, if the frame image does not include a connecting port, the shooting position of the first frame image is used as the shooting position of the frame image and is displayed on the frame image; if the frame image includes a connecting port, the feature vector of the connecting port is determined, and it is judged whether there is an identical feature vector in the feature position relationship library; if there is an identical feature vector, the hierarchical position corresponding to the identical feature vector in the feature position relationship library is used as the hierarchical position of the connecting port in the frame image, and the shooting position of the frame image and the hierarchical position of the connecting port are displayed on the frame image; if there is no identical feature vector, the hierarchical position of the connecting port is generated according to the moving direction of the endoscope and the position naming rule, the feature vector of the connecting port and the hierarchical position of the connecting port are associated and stored in the feature position relationship library, and the shooting position of the frame image and the hierarchical position of the connecting port in the frame image are displayed on the frame image.

[0087] Among them, the position naming rule includes: when the moving direction is forward, the hierarchical position of the connecting port is determined to be the next level of the shooting position, and the hierarchical position is generated as the next level of the shooting position; when the moving direction is backward, the hierarchical position of the connecting port is determined to be the previous level of the shooting position, and the hierarchical position is generated as the previous level of the shooting position.

[0088] In an embodiment of the present application, after obtaining an image of a target object to be processed, if no position information is generated in the feature position relationship library, indicating that the image to be processed is shot at the initial position, target position information is generated according to a preset position naming rule and the current position information is displayed on the image to be processed, allowing personnel photographing the cavity of the target object to know where the current shooting position belongs within the cavity of the target object. If position information is generated in the feature position relationship library, on the one hand, it proves that the current position information is displayed on the image to be processed, and on the other hand, it can determine the position information of the entrance, such as the entrance, in the image to be processed, thereby generating target position information, allowing personnel to locate the position of each entrance in the cavity of the target object. This makes it difficult to miss or repeat the image, achieves real-time positioning, and improves the shooting effect. Taking the biliary system as an example, when medical personnel are photographing the biliary system, if the image contains the entrance of a biliary branch, the position of the entrance will be automatically located, making it difficult for medical personnel to miss images of various locations in the biliary system and to repeat the image, thereby improving the accuracy and efficiency of the shooting. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] Figure 1 The figure is a flowchart of an image processing method in an embodiment.

[0090] Figure 2 is an architectural diagram of an image processing model in an embodiment.

[0091] Figure 3 The present invention is a flowchart of an image processing method for processing an image of a biliary system in an embodiment.

[0092] Figure 4 It is a structural block diagram of an image processing device in an embodiment. DETAILED DESCRIPTION

[0093] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0094] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0095] According to an embodiment of the present application, an embodiment of an image processing method, an image processing device, an endoscope and a biliary navigation system is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0096] like Figure 1 As shown, the image processing method provided in the embodiment of the present application includes:

[0097] S101: Acquire an image to be processed captured by an endoscope within a cavity of a target detection object.

[0098] The image to be processed is an image in a video captured from inside the cavity of the target detection object. In other words, the image to be processed is an image in the video, and by performing frame processing on the video, a number of images to be processed can be obtained in units of frames.

[0099] S102: When the preset feature position relationship library is empty, generate target position information corresponding to the image to be processed according to a preset position naming rule, update the target position information to the feature position relationship library, and display the target position information on the image to be processed.

[0100] In which, the target position information includes the current position information and / or entrance position information, and the feature position relationship library is used to store the feature vectors of each entrance in the target detection object cavity and the position information of each entrance in the target detection object cavity, wherein the current position information represents the shooting position of the endoscope in the target detection object cavity, and the entrance position information represents the position of the entrance appearing in the image to be processed in the target detection object cavity.

[0101] In another embodiment of the present application, the method further comprises:

[0102] When the feature position relationship library is not empty, the target position information corresponding to the image to be processed is determined according to the shooting direction of the endoscope in the cavity of the target detection object, the current position information of the image to be processed and the feature position relationship library, so as to display the target position information on the image to be processed.

[0103] Among them, the endoscope refers to a device that can extend into the cavity of the target detection object and take pictures to form a video. The feature position relationship library is used to obtain the feature vectors and position information of each entrance in the target detection object. If the feature position relationship library is empty, it means that there is no position information and feature vectors in the feature position relationship library. If it is not empty, it means that there is position information and / or feature vectors. The position naming rule is used to name each channel in the cavity of the target detection object. It is necessary to be able to distinguish between the channels at different levels. The specific name is not limited.

[0104] Taking the bile duct as an example, if the position information is not generated in the feature position relationship library, it proves that the image to be processed is the initial image taken of the bile duct. At this time, there is no position information and feature vector for the bile duct in the feature position relationship library. Therefore, the shooting position of the image to be processed is named according to the preset position naming rules, such as 1-1. At the same time, the feature position relationship library is updated so that there is position information in the feature position relationship library, that is, 1-1. In this way, the shooting personnel can know the position of the bile duct currently being shot, and the feature position relationship library is updated. It should be noted that for the first image in the video of the bile duct shooting, the corresponding feature vector may not be matched in the feature position relationship library.

[0105] If the feature position relationship library contains position information, it proves that the image to be processed does not belong to the initial image. At this time, since the previous images in the video have already been marked with position information, the target position information of the image to be processed can be determined based on the shooting direction and current position information. For example, if the image to be processed is the second frame of the video, and the bile duct entrance appears in the image to be processed, the tracker can determine that the image to be processed has not entered the bile duct entrance. Therefore, the current position information of the image to be processed is the same as the current position information of the image in the first frame of the video, both of which are 1-1. If the shooting direction is forward, that is, moving towards the depth of the bile duct, the position information of the entrance that appears can be named 1-1-1, 1-1-2, 1-1-3, and so on. If the shooting direction is backward, that is, moving towards the shallow part of the bile duct, the position information of the entrance that appears can be named 1, 1`, etc., thereby generating the target position information.

[0106] Through the above content, after obtaining the image of the target detection object to be processed, if the feature position relationship library does not generate position information, it proves that the shooting position of the image to be processed is the initial position. At this time, target position information is generated according to the preset position naming rules and the current position information is displayed on the image to be processed, so that the person photographing the cavity of the target detection object can know where the current shooting position belongs in the cavity of the target detection object. If position information is generated in the feature position relationship library, on the one hand, it proves that the current position information is displayed on the image to be processed, and on the other hand, it can determine the position information of the entrance in the image to be processed, thereby generating target position information, so that the photographer can know the position of each entrance in the cavity of the target detection object. In this way, it is less likely to miss or repeat the shooting, thereby improving the shooting effect.

[0107] In one embodiment, determining the target position information corresponding to the image to be processed based on the shooting direction of the endoscope in the cavity of the target detection object, the current position information of the image to be processed, and the feature position relationship library includes:

[0108] When the current position information does not belong to the current first-layer position and last-layer position in the feature position relationship library, determining whether the image to be processed contains an entrance, wherein the position information of each entrance in the feature position relationship library is stored according to the hierarchy of the branch to which each entrance belongs, with the top-layer position information being the first-layer position and the bottom-layer position information being the last-layer position;

[0109] If no entrance is included, determining the current position information as the target position information and displaying it on the image to be processed;

[0110] If an entrance is included, the entrance position information of each entrance is determined according to the feature position relationship library.

[0111] The determining of the target position information corresponding to the image to be processed according to the shooting direction of the endoscope in the cavity of the target detection object, the current position information of the image to be processed, and the feature position relationship library includes:

[0112] When the current position information belongs to the current first layer position and the last layer position in the feature position relationship library, determining whether the image to be processed contains an entrance;

[0113] If the entry information is not included, the current position information is determined as the target position information and displayed on the image to be processed;

[0114] If an entrance is included, when the shooting direction is away from the first-floor position or the last-floor position, the entrance position information of each entrance is determined according to the feature position relationship library; when the shooting direction is close to the first-floor position or the last-floor position, the entrance position information of each entrance is generated according to the position naming rule.

[0115] For ease of understanding, let's take the bile duct as an example. Since the bile duct has multiple layers, when determining the target location, we first determine whether the current location is the first or last layer in the feature position relationship library. If not, we determine whether the image to be processed contains information representing the bile duct entrance. If not, we directly determine the current location as the target location. If so, we determine the entrance location information for each entrance based on the feature position relationship library. Because the location is not the first or last layer, the entrance location information can be found in the feature position relationship library.

[0116] For the case where the current position information of the image to be processed is the first layer position or the last layer position in the feature position relationship library, for example, if it is the last layer position, if the video continues to shoot deep into the bile duct, the entrance in the image to be processed has not been shot. At this time, the entrance position information is generated according to the position naming rules, and the feature position relationship library is updated so that the feature position relationship library stores the feature vector and position information of the newly shot bile duct entrance.

[0117] Through the above content, the feature position relationship library can be updated automatically, so that when personnel are photographing the bile duct, they can not only know the current shooting position in real time, but also automatically update the feature position relationship library to facilitate subsequent identification of the various entrances of the bile duct.

[0118] It should be noted that the target detection object refers to the object that needs to be detected, which can be the internal organs of the human body, such as the biliary system, lungs, intestines, stomach, etc. This embodiment does not specify the equipment for taking the image. As long as it can enter the cavity of the target detection object, it can obtain an image inside the cavity of the target detection object, such as using endoscopes such as dermatoscopes, ophthalmoscopes, gastroscopes, colonoscopes, and esophagoscopes.

[0119] Among them, when shooting the target detection object, it can be shot one by one, or the image of the target detection object can be obtained in the form of a video. Specifically, for example, the video is disassembled into frames to obtain each frame image. The image to be processed refers to the image that needs to be positioned. For ease of understanding, taking the image of the bile duct as an example, since the bile duct has multiple levels of branches, for images containing branch entrances, it is necessary to determine the position of each branch entrance. At this time, the image containing the branch entrance can be used as the image to be processed. In addition, images that do not contain branch entrances can also be used as images to be processed, so that the position of the image in the bile duct system can be determined based on the characteristics of the bile duct reflected in the image.

[0120] Furthermore, for example, the biliary system includes a first-layer biliary entrance that initially enters the biliary system. After entering the first-layer biliary entrance, multiple second-layer biliary entrances appear. After entering one of the second-layer biliary entrances, multiple third-layer biliary entrances appear. This is represented in a tree diagram as follows:

[0121]

[0122] In one embodiment, determining the entrance location information of each entrance according to the feature position relationship library includes:

[0123] A local image capable of representing the entrance position of the target detection object is determined from the image to be processed.

[0124] Since the image to be processed needs to determine the shooting position of the target detection object, that is, to which position within the target detection object the image to be processed belongs, a local image can be determined from the image to be processed. The local image contains features that can reflect the shooting position, such as the branch entrance of the bile duct. It should be noted that the bile duct branch entrances at different levels have different characteristics, such as the size of the entrance, the location of the entrance, the number of entrances, etc., which can be determined based on the actual situation of the target detection object.

[0125] Key points are determined in the local image and feature vectors for the key points are generated according to position information of the key points.

[0126] A key point is a point in a local image that indicates the location of an entrance. A key point may have a certain area, but the area of the key point is smaller than the area of the local image. For example, in the biliary system, a key point may be a point in a local image that indicates the entrance of a branch of the biliary tract.

[0127] The point information is related to the key point, such as the position of the key point in the local image, the size of the key point, etc., which is not specifically limited in this embodiment. It should be noted that since it is necessary to generate a feature vector of the key point based on the point information, and the feature vector needs to be able to reflect the characteristics of the key point, when determining the point information, it is necessary to be able to generate a unique feature vector based on the point information to distinguish it from other key points. In other words, since the feature vector is unique, once the feature vector is known, it is possible to know which key point the feature vector belongs to, and thus the position of the key point in the target detection object can be known, thereby achieving the positioning of the shooting position.

[0128] According to the feature vector, the feature position relationship library is retrieved to obtain position information matching the feature vector as the entry position information.

[0129] Through the above content, after obtaining the image to be processed of the target detection object, the local image can be automatically identified. Since the local image contains content representing the shooting position of the target detection object, by generating feature vectors of key points in the local image and then retrieving the target position information from the feature position relationship library, the shooting position can be determined based on the target position information. In this way, it is less likely to miss or repeat shots, improving the shooting effect. Taking the biliary system as an example, when medical personnel are shooting a biliary image, if the image contains the entrance of a biliary branch, a local image containing the entrance of the biliary branch will be automatically obtained. Then, the target position information is obtained based on the local image, allowing medical personnel to know whether the first-layer biliary branch or the second-layer biliary branch is being photographed. This facilitates medical personnel to locate the shooting position, making it difficult for medical personnel to miss images of various locations in the biliary system and to repeat shots, thereby improving shooting accuracy and efficiency.

[0130] In one embodiment, determining a local image capable of representing the entrance position of the target detection object from the image to be processed includes:

[0131] The image to be processed is processed using the trained image processing model to determine the local image.

[0132] In which, the image processing model determines the completion status of training based on a first loss value during training, wherein the first loss value includes a second loss value for selecting a rectangular box for the local image and a third loss value for the confidence of the rectangular box.

[0133] The image processing model may be a convolutional neural network (CNN) model, specifically a Yolov5 model, such as Figure 2 As shown in the figure, it is the network architecture diagram of the Yolov5 model, where lossbj is the third loss value, lossrect is the second loss value, lossclc is the fourth loss value, and loss is the first loss value. Backbone refers to the backbone network, which is a convolutional neural network used to aggregate and form image features at different image granularities, and is generally used for front-end extraction of image information; Neck refers to a series of network layers that mix and combine image features and pass image features to the prediction layer; Conv refers to convolutional neural network; Prediction refers to prediction network; objectness refers to confidence network; boxregression refers to box position network; classification refers to classification probability network.

[0134] Through the above content, the image to be processed is processed using the image processing model to obtain a local image, which is beneficial to improving processing efficiency and processing accuracy.

[0135] In one embodiment, before retrieving, according to the feature vector, from the feature position relationship library the position information matching the feature vector as the entry position information, the method further includes:

[0136] If the feature vector identical to the feature vector is not stored in the feature position relationship library, the entry position information is generated according to the current position information and the position naming rule.

[0137] In other words, when repeatedly photographing a bile duct, the entrance may have been missed due to a difference in the previous photographing angle. Consequently, the feature vector and location information for that entrance may not be stored in the feature position relationship library. In this case, the entrance location information for that entrance is generated based on the current location information and the location naming rules and stored in the feature position relationship library.

[0138] For ease of understanding, for example, the feature position relationship library includes the following feature vectors and position information:

[0139] 1-1, x1;

[0140] 1-1-1, x2;

[0141] 1-1-1-1, x3;

[0142] 1-1-1-2, x4;

[0143] 1-1-2, x5;

[0144] 1-1-2-1, x6;

[0145] 1-1-2-2, x7.

[0146] The current location information is 1-1-2. There are three entrances in the image to be processed. The feature vector for entrance 1 is x6, so the location information for entrance 1 is 1-1-2-1. The feature vector for entrance 2 is x7, so the location information for entrance 2 is 1-1-2-2. The feature vector for entrance 3 is x8, which does not appear in the feature location relationship library. Since the current location information is 1-1-2, according to the location naming rules, the location information for entrance 3 is named 1-1-2-3.

[0147] In one embodiment, before processing the image to be processed using the trained image processing model to determine the local image, the method further includes:

[0148] S501: Process a training image using the image processing model to be trained to obtain a predicted rectangular frame and prediction information of the predicted rectangular frame.

[0149] The image within the predicted rectangular frame is a local image determined by the image processing model to be trained that can characterize the entrance position of the target detection object, and the prediction information includes the predicted position and the prediction confidence.

[0150] That is, when the image processing model processes the training image, it automatically identifies the local image in the training image that represents the shooting location, thereby forming a predicted rectangular frame so that the local image is located within the predicted rectangular frame. It should be noted that a circular frame or an irregular frame may also be used, and this embodiment does not limit the shape of the frame.

[0151] After selecting a portion of the image, information related to the predicted rectangular frame, i.e., prediction information, can be obtained. The prediction information includes, for example, the position of the predicted rectangular frame in the image to be processed, the size of the predicted rectangular frame, and the aspect ratio. In this embodiment, the prediction information includes the predicted position and prediction confidence. The predicted position refers to the position of the predicted rectangular frame in the image to be processed, and the prediction confidence refers to the degree of confidence that the predicted rectangular frame contains content representing the shooting location.

[0152] Among them, the prediction information can reflect the specific information of the predicted rectangular box, and aims to determine the accuracy of the predicted rectangular box based on the prediction information, so as to compensate and optimize the image processing model, so that the image processing model can accurately select the local image after training is completed, and then determine the shooting position based on the local image.

[0153] S502: Calculate the distance between the predicted position and the reference position of the reference rectangular frame of the training image using a first loss function to obtain a second loss value.

[0154] The reference rectangular frame is a rectangular frame determined in advance from the training image, and the reference rectangular frame has reference information, and the reference information includes a reference position and a reference confidence.

[0155] Specifically, the first loss function is used to measure the accuracy of the predicted rectangular box and the reference rectangular box. The first loss function includes the CIOULoss function. Since the training image is used to train the image processing model, the accurate position of the local image in the training image is known, and the reference rectangular box and reference information can be obtained based on the accurate local image position. Therefore, by calculating the second loss value, the difference between the predicted position and the reference position can be reflected, and the smaller the difference, the better. Therefore, the second loss value can be used to compensate for the position of the predicted rectangular box determined by the image processing model, thereby improving the position accuracy of the predicted rectangular box determined by the image processing model.

[0156] In one embodiment, the second loss value is calculated by:

[0157] L CIOU =1-CIOU

[0158]

[0159] Among them, L CIOU Represents the second loss value, IOU represents the intersection-over-union ratio between the predicted rectangular box and the reference rectangular box, which measures the degree of overlap between the two boxes; ρ is used to measure the distance between the center points of the predicted rectangular box and the reference rectangular box, and c represents the diagonal length of the minimum outer bounding box between the two; v is used to measure the difference in the length-to-width ratio between the predicted rectangular box and the reference rectangular box to ensure that the predicted result is closer to the true result, and α is the balance factor between the length-to-width ratio difference and IOU; among them, the two refer to the predicted rectangular box and the reference rectangular box.

[0160] S503 : Calculate the prediction confidence of the prediction rectangular box and the reference confidence of the reference rectangular box using the second loss function to obtain the third loss value.

[0161] Among them, the prediction confidence is presented in the form of a matrix, and the reference confidence is presented in the form of a matrix. The second loss function includes the BCEloss function, and the calculation method of the third loss value includes:

[0162] loss BCE (z,x,y)=-L(z,x,y)*logP(z,x,y)-(1-L(z,x,y))*log(1-P(z,x,y));

[0163] Among them, x and y represent the confidence of the neural network prediction rectangle around (x, y), z represents the number of neural network prediction confidences; L is the reference confidence matrix; P is the prediction confidence matrix.

[0164] S504: Calculate the first loss value according to the second loss value and the third loss value.

[0165] The second loss value and the third loss value may be weighted and summed to obtain the first loss value.

[0166] S505. Complete the training of the image processing model according to the matching between the first loss value and the preset training completion condition to obtain a trained image processing model.

[0167] Among them, the training completion condition can be to limit the range of the first loss value, so that when the first loss value falls into this range, it is determined that the first loss value matches the training completion condition, thereby determining that the image processing model has completed training. This embodiment does not limit the specific content of the training completion condition, and the training completion condition can be set according to actual needs. For example, when a high-precision image processing model is not required, the restriction of the training completion condition on the first loss value can be relaxed. On the contrary, if the accuracy requirement is relatively high, the restriction of the training completion condition on the first loss value can be narrowed, and the training completion condition can even be set to a specific value. Only when the first loss value is the same as the value of the training completion condition, the training is determined to be completed. Otherwise, the parameters of the image processing model continue to be optimized or the image processing model continues to be compensated, so that the first loss value continues to approach the value of the training completion condition.

[0168] Through the above content, when training the image processing model, restricting the position and confidence of the rectangular box helps to improve the processing accuracy of the image processing model on the processed image, thereby improving the accuracy of local features.

[0169] In one embodiment, the prediction information further includes a prediction classification probability, and the reference information further includes a reference classification probability;

[0170] The reference classification probability refers to the probability that the image processing model recognizes the type of the target detection object through the image within the rectangular frame.

[0171] Before processing the image to be processed using the trained image processing model to determine the local image, the method further includes:

[0172] Processing the training image using the image processing model to be trained to obtain a prediction rectangular box and prediction information of the prediction rectangular box, wherein the image within the prediction rectangular box is a local image determined by the image processing model to be trained that can represent the entrance position of the target detection object, and the prediction information includes the predicted position and the prediction confidence;

[0173] calculating a distance between the predicted position and a reference position of a reference rectangular frame of the training image using a first loss function to obtain a second loss value, wherein the reference rectangular frame is a rectangular frame pre-determined from the training image, and the reference rectangular frame has reference information, the reference information including a reference position and a reference confidence;

[0174] Calculating the prediction confidence of the prediction rectangular box and the reference confidence of the reference rectangular box using the second loss function to obtain the third loss value;

[0175] The predicted classification probability and the reference classification probability are calculated using the second loss function to obtain a fourth loss value.

[0176] The fourth loss value is used to measure the distance between the predicted classification probability and the reference classification probability. In this embodiment, the predicted classification probability is presented in the form of a matrix, and the reference classification probability is presented in the form of a matrix.

[0177] The calculation method of the fourth loss value includes:

[0178] loss BCE (z,x,yt)=-L smooth (z,x,yt)*logQ(z,x,yt)-(1-L smooth (z,x,yt))*log(1-Q(z,x,yt));

[0179] Among them, t represents the category, L smooth is the matrix of predicted classification probabilities; Q is the matrix of reference classification probabilities.

[0180] The second loss value, the third loss value, and the fourth loss value are weighted and summed to obtain the first loss value.

[0181] Specifically, the calculation method of the first loss value includes:

[0182] First loss value = a*second loss value + b*third loss value + c*fourth loss value.

[0183] Among them, a is the weight of the second loss value, b is the weight of the third loss value, and c is the weight of the fourth loss value.

[0184] The training of the image processing model is completed according to the matching between the first loss value and the preset training completion condition to obtain a trained image processing model.

[0185] Based on the above, during training, the image processing model divides the training image into an N*N grid, generates prediction boxes, and predicts three metrics for each box: the box's position, confidence level, and classification probability, thereby calculating the first loss value. The image processing model also performs the same steps when processing the target image, which helps improve the model's accuracy in processing the target image.

[0186] In one embodiment, determining key points in the partial image and generating feature vectors for the key points based on position information of the key points includes:

[0187] S701: Detect interest points in the local image using a scale-invariant feature change algorithm.

[0188] Among them, the scale-invariant feature transform (SIFT) algorithm includes an interest point detection module, a local feature selection module, a local feature description module, a local feature compression module, and a local feature fusion module.

[0189] The interest point detection module detects points of interest in a local image, the local feature selection module identifies key points from these points of interest, the local feature description module generates feature vectors for key points, the local feature compression module removes unstable key points from different angles, and the local feature fusion module fuses feature vectors for key points from different angles. The specific functions and usage of each module can be found in existing documentation on the scale-invariant feature transform (SIFT) algorithm and will not be detailed here.

[0190] In an application scenario, for a local image, the SIFT algorithm constructs a Gaussian scale pyramid, uses continuous Gaussian blurring to find points by image differences, then performs Laplacian filtering to generate a Laplacian pyramid. A low-degree polynomial (ALP) detection method is then applied to the LoG graph to find points of interest, thereby obtaining points of interest.

[0191] S702: Determine the key points from the points of interest using the scale-invariant feature change algorithm and obtain the point information.

[0192] The point information includes the point position of the key point in the rectangular box, the scale of the rectangular box to which the key point belongs, and the point direction of the key point, wherein the rectangular box is used to frame the area in the image to be processed to obtain the local image.

[0193] Specifically, in the SIFT algorithm, the interest point detection process generates a certain number of key points. Since some key points may be relatively difficult to identify or easily interfered with by noise, it is necessary to eliminate key points that are easily interfered with by using the information of pixels near the key points, the size of the key points, the curvature of the key points, etc. Specific methods include: (1) accurate positioning of key points: in order to improve the stability of key points, the scale space DoG function curve is interpolated; (2) unclear key points are discarded; (3) edge response is eliminated: (edge instability means that the gradient value in one direction is too large, while the gradient value in another direction is too small, and this type of edge response point needs to be proposed) obtain the Hessian matrix at the feature point, calculate the ratio of the eigenvalues, and remove the edge response points corresponding to the larger ratio.

[0194] To ensure that the keypoint's feature vector is rotationally invariant, a reference direction is assigned to each keypoint using the local image, and an image gradient method is used to determine the stable direction. This involves calculating the gradient values of pixels near the keypoint and plotting a histogram. The peak direction of the histogram represents the keypoint's primary direction.

[0195] S703: Generate the feature vector according to the point position, scale, and point direction using the scale-invariant feature change algorithm.

[0196] Through the above content, the scale-invariant feature change algorithm is used to calculate the feature vector of the key point, which is fast and accurate and helps to improve the accuracy of the feature vector.

[0197] In one embodiment, the retrieving, from the feature position relationship library according to the feature vector, position information matching the feature vector as the entry position information includes:

[0198] S801. Calculate the distances between the feature vector of the key point and the feature vectors of each entry position of the target detection object in the feature position relationship library, and determine the target detection object positions closest and second closest to the key point in the feature position relationship library.

[0199] S802. If the ratio of the feature vector of the target detection object position that is closest to the target detection object position to the feature vector of the target detection object position that is second closest to the target detection object position is less than a preset ratio threshold, it is determined that the feature vector of the target detection object position that is closest to the target detection object position matches the feature vector of the key point.

[0200] S803: Retrieve the position information corresponding to the feature vector of the position of the target detection object that is closest to the target detection object as the entry position information.

[0201] Specifically, for example, the feature vector of the key point is:

[0202] R i =(r i1 ,r i2 ,r i3 ...r i128 );

[0203] Calculate the feature vector R of the key point i Each feature vector R of the feature position relationship library j The distance between the two methods include:

[0204]

[0205] In this way, R can be calculated i With each R j The distance between them can be determined by R i The nearest eigenvector R n and the next closest eigenvector R m .

[0206] Among them, the closest refers to the shortest distance, and the second closest refers to the second shortest distance.

[0207] R n / R m The ratio can be obtained if R n / R m < the ratio threshold, then it proves a match.

[0208] In one embodiment, the calculating of the distances between the feature vector of the key point and the feature vector of each position of the target detection object in the feature position relationship library includes:

[0209] Calculate the Euclidean distance between the feature vector of the key point and the feature vector of each position of the target detection object.

[0210] For ease of understanding, Figure 3 As shown, the process of this method is explained by taking the bile duct image as an example of the image to be processed. First, the bile duct image is input into the CNN model, that is, the image processing model. The CNN model will automatically extract the target and obtain the image to be processed marked with a rectangular frame, wherein the content in the rectangular frame is the local image. Then, the image to be processed with the rectangular frame is calculated using the SIFT algorithm to extract the feature descriptor, that is, to obtain the feature vector of the key point, and then the feature vector of the key point is stored in a pre-set local descriptor database and compared with the feature vector in the global relationship library. The target position information is output and the target position information is displayed in the form of text on the image to be processed, so that medical personnel can clearly know which position of the bile duct the current image is, and thus decide which bile duct branch entrance to enter, thereby improving the efficiency and accuracy of bile duct shooting.

[0211] In one embodiment, the target detection object includes the bile duct and / or the biliary system.

[0212] The present application also provides an image processing device, such as Figure 4 As shown, it includes obtaining an image to be processed taken by an endoscope in the cavity of a target detection object, wherein the image to be processed is an image in a video taken in the cavity of the target detection object;

[0213] A first judgment module 2 is configured to generate target position information corresponding to the image to be processed according to a preset position naming rule and update the target position information to the feature position relationship library when the preset feature position relationship library is empty, and display the target position information on the image to be processed, wherein the target position information includes the current position information and / or entrance position information, and the feature position relationship library includes the feature vector of each entrance in the target detection object cavity and the position information of each entrance in the target detection object cavity, wherein the current position information represents the shooting position of the endoscope in the target detection object cavity, and the entrance position information represents the position of the entrance appearing in the image to be processed in the target detection object cavity;

[0214] The second judgment module 3 is used to determine the target position information corresponding to the image to be processed according to the shooting direction of the video in the cavity of the target detection object and the current position information of the image to be processed when the feature position relationship library is not empty, so as to display the target position information on the image to be processed.

[0215] Optionally, the second judgment module includes a second judgment unit configured to judge whether the image to be processed contains an entrance when the current position information does not belong to the current first-layer position and last-layer position in the feature position relationship library;

[0216] If there is no information capable of characterizing the entrances in the cavity of the target detection object, determining the current position information as the target position information and displaying it on the image to be processed;

[0217] If there is information that can characterize each entrance in the cavity of the target detection object, determining the entrance position information of each entrance according to the characteristic position relationship library;

[0218] When the current position information belongs to the current first layer position and the last layer position in the feature position relationship library, determining whether the image to be processed contains an entrance;

[0219] If there is no information capable of characterizing the entrances in the cavity of the target detection object, determining the current position information as the target position information and displaying it on the image to be processed;

[0220] If there is information that can characterize the entrances in the cavity of the target detection object, then when the shooting direction is away from the first-layer position or the last-layer position, the entrance position information of each entrance is determined according to the feature position relationship library; when the shooting direction is close to the first-layer position or the last-layer position, the entrance position information of each entrance is generated according to the position naming rule.

[0221] Optionally, the second judgment unit includes a local unit, configured to determine a local image capable of representing the entrance position of the target detection object from the image to be processed;

[0222] A vector unit, configured to determine key points in the local image and generate a feature vector for the key points based on position information of the key points;

[0223] The location unit is used to retrieve location information matching the feature vector from the feature location relationship library according to the feature vector as the entry location information.

[0224] Optionally, the local unit includes a model unit, which is used to process the image to be processed using a trained image processing model to determine the local image, wherein the image processing model determines the training completion status based on a first loss value during training, wherein the first loss value includes a second loss value for selecting a rectangular box for the local image and a third loss value for the confidence of the rectangular box.

[0225] Optionally, the model unit includes a first subunit, configured to process a training image using the image processing model to be trained to obtain a prediction rectangular box and prediction information of the prediction rectangular box, wherein the image within the prediction rectangular box is a local image determined by the image processing model to be trained that can characterize the entrance position of the target detection object, and the prediction information includes a predicted position and a prediction confidence;

[0226] a second subunit, configured to calculate a distance between the predicted position and a reference position of a reference rectangular frame of the training image using a first loss function to obtain a second loss value, wherein the reference rectangular frame is a rectangular frame pre-determined from the training image, and the reference rectangular frame has reference information, the reference information including a reference position and a reference confidence;

[0227] a third subunit, configured to calculate the prediction confidence of the prediction rectangular box and the reference confidence of the reference rectangular box using a second loss function to obtain the third loss value;

[0228] A fourth subunit, configured to calculate the first loss value according to the second loss value and the third loss value;

[0229] The fifth subunit is used to complete the training of the image processing model according to the matching between the first loss value and the preset training completion condition to obtain a trained image processing model.

[0230] Optionally, the prediction information further includes a prediction classification probability, and the reference information further includes a reference classification probability;

[0231] The apparatus further includes a sixth subunit, configured to calculate the predicted classification probability and the reference classification probability using the second loss function to obtain a fourth loss value;

[0232] The seventh subunit is used to perform weighted summation on the second loss value, the third loss value and the fourth loss value to obtain the first loss value.

[0233] The eighth subunit is used to complete the training of the image processing model according to the matching between the first loss value and the preset training completion condition to obtain a trained image processing model.

[0234] Optionally, the vector unit includes an interest point unit, configured to detect interest points in the local image using a scale-invariant feature change algorithm;

[0235] a determining unit, configured to determine the key point from the points of interest using the scale-invariant feature change algorithm and obtain the point position information, wherein the point position information includes the point position of the key point within a rectangular frame, the scale of the rectangular frame to which the key point belongs, and the point direction of the key point, wherein the rectangular frame is used to frame an area in the image to be processed to obtain the local image;

[0236] A vector unit is used to generate the feature vector according to the point position, scale and point direction using the scale-invariant feature change algorithm.

[0237] Optionally, the position unit includes a distance unit, which is used to calculate the distance between the feature vector of the key point and the feature vector of each entry position of the target detection object in the feature position relationship library, and determine the target detection object position closest and second closest to the key point in the feature position relationship library;

[0238] a ratio unit, configured to determine that the feature vector of the closest target detection object position matches the feature vector of the key point if a ratio of the feature vector of the closest target detection object position to the feature vector of the second closest target detection object position is less than a preset ratio threshold;

[0239] The retrieving unit is used to retrieve the position information corresponding to the feature vector of the position of the target detection object that is closest to the target detection object as the entry position information.

[0240] Optionally, the distance unit includes a distance subunit, which is used to calculate the Euclidean distance between the feature vector of the key point and the feature vector of each position of the target detection object.

[0241] The embodiment of the present application also provides an image processing method using the above-mentioned image processing method.

[0242] The embodiment of the present application further provides an endoscope, which is used to shoot a video in the cavity of the bile duct and uses the above-mentioned image processing method to display target position information on each frame of the video;

[0243] The endoscope includes a display or is connected to a display, and the display is used to display a video with the target position information.

[0244] The embodiment of the present application further provides a biliary navigation system, comprising the above-mentioned image processing device and endoscope;

[0245] The endoscope is used to capture video within the bile duct cavity and transmit the video to the image processing device;

[0246] The image processing device is used to display target position information on each frame of the video.

[0247] An embodiment of the present application also provides a medical robot, which is used to control an endoscope to move in the cavity of a bile duct and shoot a video in the cavity of the bile duct, determine the shooting position of each frame image in the video in the cavity of the bile duct according to a preset feature position relationship library, and display the shooting position on the corresponding frame image; wherein, the bile duct includes multiple layers of branches, and there are connecting ports between the branches of adjacent layers. When the frame image is the first frame image of the video, if the first frame image does not include the connecting port, the shooting position of the first frame image is generated according to a preset position naming rule, and the shooting position is displayed on the first frame image; if the first frame image includes the connecting port, the feature vector of the connecting port is determined, and the hierarchical position of the connecting port is generated according to the moving direction of the endoscope and the position naming rule, the feature vector and hierarchical position of the connecting port in the first frame image are stored in the feature position relationship library, and the shooting position and hierarchical position of the first frame image are displayed on the first frame image.

[0248] When the frame image is not the first frame image, if the frame image does not include a connecting port, the shooting position of the first frame image is used as the shooting position of the frame image and is displayed on the frame image; if the frame image includes a connecting port, the feature vector of the connecting port is determined, and it is judged whether there is an identical feature vector in the feature position relationship library; if there is an identical feature vector, the hierarchical position corresponding to the identical feature vector in the feature position relationship library is used as the hierarchical position of the connecting port in the frame image, and the shooting position of the frame image and the hierarchical position of the connecting port are displayed on the frame image; if there is no identical feature vector, the hierarchical position of the connecting port is generated according to the moving direction of the endoscope and the position naming rule, the feature vector of the connecting port and the hierarchical position of the connecting port are associated and stored in the feature position relationship library, and the shooting position of the frame image and the hierarchical position of the connecting port in the frame image are displayed on the frame image.

[0249] Among them, the position naming rule includes: when the moving direction is forward, the hierarchical position of the connecting port is determined to be the next level of the shooting position, and the hierarchical position is generated as the next level of the shooting position; when the moving direction is backward, the hierarchical position of the connecting port is determined to be the previous level of the shooting position, and the hierarchical position is generated as the previous level of the shooting position.

[0250] An embodiment of the present application also provides a storage medium storing a computer program that can be loaded by a processor and execute the above-described method.

[0251] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0252] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0253] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0254] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0255] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0256] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0257] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. An image processing method, characterized in that: include: Acquire an image to be processed captured by an endoscope within a cavity of a target detection object, wherein the image to be processed is an image in a video captured within the cavity of the target detection object; When the preset feature position relationship library is empty, generating target position information corresponding to the image to be processed according to a preset position naming rule, updating the target position information to the feature position relationship library, and displaying the target position information on the image to be processed; The target position information includes current position information and / or entrance position information, and the feature position relationship library is used to store the feature vectors of each entrance in the target detection object cavity and the position information of each entrance in the target detection object cavity; The current position information represents the shooting position of the endoscope in the cavity of the target detection object, the entrance position information represents the position of the entrance appearing in the image to be processed in the cavity of the target detection object, the cavity of the target detection object includes several branches, and the entrance refers to the connecting port of each branch; The method further comprises: When the feature position relationship library is not empty, determining the target position information corresponding to the image to be processed according to the shooting direction of the endoscope in the cavity of the target detection object, the current position information of the image to be processed and the feature position relationship library, so as to display the target position information on the image to be processed; The determining of the target position information corresponding to the image to be processed according to the shooting direction of the endoscope in the cavity of the target detection object, the current position information of the image to be processed, and the feature position relationship library includes: When the current position information does not belong to the current first-layer position and last-layer position in the feature position relationship library, determining whether the image to be processed contains an entrance, wherein the position information of each entrance in the feature position relationship library is stored according to the hierarchy of the branch to which each entrance belongs, with the top-layer position information being the first-layer position and the bottom-layer position information being the last-layer position; If no entrance is included, determining the current position information as the target position information and displaying it on the image to be processed; If an entrance is included, the entrance position information of each entrance is determined according to the feature position relationship library.

2. The image processing method according to claim 1, wherein: The determining of the target position information corresponding to the image to be processed according to the shooting direction of the endoscope in the cavity of the target detection object, the current position information of the image to be processed, and the feature position relationship library includes: When the current position information belongs to the first layer position or the last layer position in the feature position relationship library, determining whether the image to be processed contains an entrance; If no entrance is included, determining the current position information as the target position information and displaying it on the image to be processed; If an entrance is included, when the shooting direction is away from the first-floor position or the last-floor position, the entrance position information of each entrance is determined according to the feature position relationship library; when the shooting direction is close to the first-floor position or the last-floor position, the entrance position information of each entrance is generated according to the position naming rule.

3. The image processing method according to claim 2, wherein: The determining of the entrance position information of each entrance according to the feature position relationship library includes: Determining a local image capable of representing the entrance position of the target detection object from the image to be processed; Determining key points in the local image and generating feature vectors for the key points based on position information of the key points; According to the feature vector, the feature position relationship library is retrieved to obtain position information matching the feature vector as the entry position information.

4. The image processing method according to claim 3, wherein: Before retrieving position information matching the feature vector from the feature position relationship library according to the feature vector as the entry position information, the method further includes: If the feature vector identical to the feature vector is not stored in the feature position relationship library, the entry position information is generated according to the current position information and the position naming rule.

5. The image processing method according to claim 3, wherein: The determining of a local image capable of representing the entrance position of the target detection object from the image to be processed includes: The image to be processed is processed using a trained image processing model to determine the local image, wherein the image processing model determines the completion of training according to a first loss value during training, wherein the first loss value includes a second loss value for selecting a rectangular box for the local image and a third loss value for the confidence of the rectangular box.

6. The image processing method according to claim 5, characterized in that Before processing the image to be processed using the trained image processing model to determine the local image, the method further includes: Processing the training image using the image processing model to be trained to obtain a prediction rectangular box and prediction information of the prediction rectangular box, wherein the image within the prediction rectangular box is a local image determined by the image processing model to be trained that can represent the entrance position of the target detection object, and the prediction information includes a predicted position, a prediction confidence, and a predicted classification probability; calculating a distance between the predicted position and a reference position of a reference rectangular frame of the training image using a first loss function to obtain a second loss value, wherein the reference rectangular frame is a rectangular frame pre-determined from the training image, and the reference rectangular frame has reference information, the reference information including a reference position, a reference confidence, and a reference classification probability; Calculating the prediction confidence of the prediction rectangular box and the reference confidence of the reference rectangular box using the second loss function to obtain the third loss value; Calculating the predicted classification probability and the reference classification probability using the second loss function to obtain a fourth loss value; Performing a weighted summation on the second loss value, the third loss value, and the fourth loss value to obtain the first loss value; The training of the image processing model is completed according to the matching between the first loss value and the preset training completion condition to obtain a trained image processing model.

7. The image processing method according to claim 3, wherein: The determining of key points in the local image and generating a feature vector for the key points according to position information of the key points includes: Detecting points of interest in the local image using a scale-invariant feature variation algorithm; Determining the key points from the points of interest using the scale-invariant feature change algorithm and obtaining the point information, wherein the point information includes the position of the key points within a rectangular box, the scale of the rectangular box to which the key points belong, and the direction of the key points, wherein the rectangular box is used to frame an area in the image to be processed to obtain the local image; The feature vector is generated according to the point position, scale and point direction using the scale-invariant feature change algorithm.

8. The image processing method according to claim 3, wherein: The retrieving, according to the feature vector, from the feature position relationship library the position information matching the feature vector as the entry position information includes: Calculating the distances between the feature vector of the key point and the feature vectors of each entry position of the target detection object in the feature position relationship library, and determining the target detection object positions closest and second closest to the key point in the feature position relationship library; If the ratio of the feature vector of the closest target detection object position to the feature vector of the second closest target detection object position is less than a preset ratio threshold, it is determined that the feature vector of the closest target detection object position matches the feature vector of the key point; The position information corresponding to the feature vector of the target detection object position closest to the target detection object is retrieved as the entry position information.

9. The image processing method according to claim 8, characterized in that: The calculating of the distances between the feature vector of the key point and the feature vector of each entrance position of the target detection object in the feature position relationship library includes: Calculate the Euclidean distance between the feature vector of the key point and the feature vector of each position of the target detection object.

10. The image processing method according to any one of claims 1 to 9, characterized in that: The target detection object includes the bile duct.

11. An image processing device, characterized in that: The image processing method according to any one of claims 1 to 10 is adopted.

12. An endoscope, characterized in that: The endoscope is used to capture a video within the cavity of the bile duct and uses the image processing method according to any one of claims 1 to 10 to display target position information on each frame of the video; The endoscope includes a display or is connected to a display, and the display is used to display a video with the target position information.

13. A biliary navigation system, characterized in that: comprising an endoscope and the image processing device according to claim 11; The endoscope is used to capture video within the bile duct cavity and transmit the video to the image processing device; The image processing device is used to display target position information on each frame of the video.

14. A medical robot, characterized in that: The medical robot is used to control the movement of the endoscope in the cavity of the bile duct and to shoot a video of the cavity of the bile duct, and to determine the shooting position of each frame image in the video in the cavity of the bile duct according to a preset feature position relationship library and to display the shooting position on the corresponding frame image; wherein, the bile duct includes multiple layers of branches, and there are connecting ports between the branches of adjacent layers. When the frame image is the first frame image of the video, if the first frame image does not include the connecting port, the shooting position of the first frame image is generated according to a preset position naming rule, and the shooting position is displayed on the first frame image; if the first frame image includes the connecting port, the feature vector of the connecting port is determined, and the hierarchical position of the connecting port is generated according to the moving direction of the endoscope and the position naming rule, and the feature vector and hierarchical position of the connecting port in the first frame image are stored in the feature position relationship library, and the shooting position and hierarchical position of the first frame image are displayed on the first frame image; When the frame image is not the first frame image, if the frame image does not include a connecting port, the shooting position of the first frame image is used as the shooting position of the frame image and is displayed on the frame image. If the frame image includes a connecting port, the feature vector of the connecting port is determined, and it is judged whether there is an identical feature vector in the feature position relationship library. If there is an identical feature vector, the hierarchical position corresponding to the identical feature vector in the feature position relationship library is used as the hierarchical position of the connecting port in the frame image and the shooting position of the frame image and the hierarchical position of the connecting port are displayed on the frame image. If there is no identical feature vector, the hierarchical position corresponding to the identical feature vector in the feature position relationship library is used as the hierarchical position of the connecting port in the frame image and the shooting position of the frame image and the hierarchical position of the connecting port are displayed on the frame image. If there is no identical feature vector, the hierarchical position of the connecting port is determined according to the endoscope. The traveling direction and the position naming rule generate the hierarchical position of the connecting port, the feature vector of the connecting port and the hierarchical position of the connecting port are associated and stored in the feature position relationship library, and the shooting position of the frame image and the hierarchical position of the connecting port in the frame image are displayed on the frame image, wherein the position naming rule includes when the traveling direction is forward, determining the hierarchical position of the connecting port as the next level of the shooting position, and generating the hierarchical position of the next level of the shooting position; when the traveling direction is backward, determining the hierarchical position of the connecting port as the previous level of the shooting position, and generating the hierarchical position of the previous level of the shooting position.

Citation Information

Patent Citations

  • Endoscope image processing device

    US20230172428A1

  • Tissue cavity locating method and apparatus for endoscope, medium and device

    WO2023029741A1