Endoscope rotation angle determination method and device, endoscope system and storage medium
By acquiring the high and low texture feature point sets and feature weights in the endoscopic image, combining the optical flow method and the endoscopic motion model, the accuracy of rotation angle judgment in the endoscopic operation is solved, and the operation convenience and robustness are improved.
Patent Information
- Application Number
- CN202510838295.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In endoscopic operation, it is difficult for users to accurately judge the rotational transformation amount of the image, resulting in misjudgment of orientation and clinical operation errors. Especially under the complex operation path, it is difficult for physicians with insufficient experience to accurately judge the rotational transformation amount of the real-time image relative to the initial position.
By obtaining the set of high-texture feature points and low-texture feature points in the current frame field image, the feature weights of each feature point are determined, and the pyramid-layered Lucas-Canard optical flow method is used to determine the optical flow field. Combined with the endoscopic motion model, the rotation angle difference of the current frame field image relative to the previous frame field image is calculated, and the target rotation angle is finally determined.
It improves the convenience of endoscopic operation, provides an accurate basis for judging the rotation angle, enhances the robustness of the optical flow field, and reduces the risk of operation errors.
Smart Images

Figure CN120370537A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular to a method and device for determining the rotation angle of an endoscope, an endoscope system, and a storage medium. Background Art
[0002] The head of an endoscope is usually equipped with a light source, a camera, and a working channel. The user needs to rotate the operation handle to control the rotation of the front end to adjust the viewing angle or the operation direction. When the handle rotates, the field-of-view image captured by the endoscope lens also rotates accordingly. Therefore, the operator needs to judge the relative position of the target area according to the rotation of the handle. During this process, the real-time image captured by the endoscope imaging system will generate a coaxial rotation transformation with the rotation of the active bending section of the endoscope, and the image presented on the display lacks an absolute spatial orientation reference system. When there are mirror-symmetric anatomical structures in the exploration area, the operator needs to manually infer the image orientation based on the mechanical rotation angle of the handle.
[0003] However, the complex operation path leads to the degradation of spatial orientation perception, making it difficult for inexperienced physicians to accurately judge the rotation transformation amount of the real-time image relative to the initial pose, which is extremely easy to cause misjudgment of orientation and thus lead to the risk of clinical operation errors. Therefore, there is an urgent need to develop a method for determining the rotation angle of an endoscope. Summary of the Invention
[0004] This application provides a method and device for determining the rotation angle of an endoscope, an endoscope system, and a storage medium, which can determine the target rotation angle of the current frame field-of-view image of the endoscope compared with the initial frame field-of-view image in real time, can provide a relatively accurate judgment basis for users, and can improve the convenience of endoscope operation.
[0005] A method for determining the rotation angle of an endoscope provided by this application includes: Obtain a current frame field-of-view image, where the current frame field-of-view image is an image obtained by the endoscope lens at the current moment; Based on the current frame field-of-view image, determine a high-texture feature point set and a low-texture feature point set in the current frame field-of-view image; the high-texture feature point set includes a plurality of ORB feature points, and the low-texture feature point set includes a plurality of low-texture region feature points; the visual saliency of the ORB feature points in the current frame field-of-view image is greater than that of the low-texture region feature points. Based on the high-texture feature point set and the low-texture feature point set, determine the feature weights corresponding to each feature point; the feature points include ROB feature points and low-texture region feature points; Based on the high-texture feature point set, the low-texture feature point set, and the feature weights, use the pyramid hierarchical Lucas-Kanade optical flow method to determine the optical flow field; Determine the rotation angle difference corresponding to the current frame field of view image based on the previous frame field of view image, the optical flow field, and the endoscope motion model. The rotation angle difference corresponding to the current frame field of view image is the rotation angle difference of the current frame field of view image compared to the previous frame field of view image; Determine the target rotation angle of the current frame field of view image compared to the initial frame field of view image based on the rotation angle differences corresponding to each frame field of view image.
[0006] To achieve the above and other related purposes, the present application provides an endoscope rotation angle determination device, including: A data acquisition module for acquiring the current frame field of view image, where the current frame field of view image is the image acquired by the endoscope lens at the current moment; A point set determination module for determining a high-texture feature point set and a low-texture feature point set in the current frame field of view image based on the current frame field of view image; the high-texture feature point set includes a plurality of ORB feature points, and the low-texture feature point set includes a plurality of low-texture region feature points; the visual saliency of the ORB feature points in the current frame field of view image is greater than the visual saliency of the low-texture region feature points; A weight determination module for determining the feature weights corresponding to each feature point based on the high-texture feature point set and the low-texture feature point set; the feature points include ROB feature points and low-texture region feature points; A data processing module for determining the optical flow field using the pyramid hierarchical Lucas-Kanade optical flow method based on the high-texture feature point set, the low-texture feature point set, and the feature weights; A first angle determination module for determining the rotation angle difference corresponding to the current frame field of view image based on the previous frame field of view image, the optical flow field, and the endoscope motion model. The rotation angle difference corresponding to the current frame field of view image is the rotation angle difference of the current frame field of view image compared to the previous frame field of view image; A second angle determination module for determining the target rotation angle of the current frame field of view image compared to the initial frame field of view image based on the rotation angle differences corresponding to each frame field of view image.
[0007] To achieve the above and other related purposes, the present application also provides an endoscope system, including an endoscope host, and the endoscope host includes: One or more processors; A memory for storing the executable program code of the processor; Wherein, the processor is configured to execute the program code to implement the above endoscope rotation angle determination method.
[0008] To achieve the above and other related objectives, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor of a computer, the computer is caused to execute the foregoing endoscopic rotation angle determination method of one or more of the foregoing.
[0009] As described above, an endoscopic rotation angle determination method, device, endoscopic system, and storage medium provided by the present application have the following beneficial effects: An endoscopic rotation angle determination method in the present application. This method determines the feature weights corresponding to each feature point in different feature point sets by obtaining a high-texture feature point set and a low-texture feature point set in the current frame field of view image, combines the Lucas-Kanade optical flow method with pyramid layering to determine the optical flow field, and determines the rotation angle difference between the current frame field of view image and the previous frame field of view image through the optical flow field and the endoscopic motion model, and then determines the target rotation angle. By combining high-texture features, low-texture features, and the weights of each feature point, the contribution of each feature point to the calculation of the optical flow field is balanced, which can enhance the robustness of the optical flow field, and further improve the accuracy of the rotation angle difference; by determining the rotation angle difference to determine the target rotation angle of the current frame field of view image, it can assist the user in operating the endoscope, and can provide a relatively accurate judgment basis for the user, achieving the effect of improving the convenience of endoscope operation.
[0010] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings: Figure 1 is a flowchart of an endoscopic rotation angle determination method shown in an exemplary embodiment of the present application; Figure 2 is a structural block diagram of an endoscopic rotation angle determination device shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0012] The embodiments of the present application will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for illustrating the present application, rather than for limiting the protection scope of the present application.
[0013] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present application in a schematic manner. Therefore, only the components related to the present application are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in its actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0014] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.
[0015] Please refer to Figure 1 , Figure 1 which is a flowchart of the method for determining the rotation angle of an endoscope shown in an exemplary embodiment of the present application. Referring to Figure 1 it can be seen that the method for determining the rotation angle of the endoscope may include: Step S110, obtaining the current frame field of view image.
[0016] Wherein, the current frame field of view image is the image obtained by the endoscope lens at the current moment.
[0017] In an embodiment of the present application, the endoscope host can obtain the current frame field of view image. A camera is provided at the front end of the active bending section of the endoscope. The active bending section of the endoscope can penetrate into the body, and the camera at the front end of the active bending section of the endoscope can capture the current frame field of view image at a preset frame rate and transmit the captured current frame field of view image to the endoscope host.
[0018] Step S120, based on the current frame field of view image, determining a high-texture feature point set and a low-texture feature point set in the current frame field of view image.
[0019] Wherein, the high-texture feature point set includes a plurality of ORB feature points, and the low-texture feature point set includes a plurality of low-texture region feature points; the visual saliency of the ORB feature points in the current frame field of view image is greater than that of the low-texture region feature points.
[0020] In one embodiment of the present application, the endoscope host can determine a high-texture feature point set and a low-texture feature point set in the current frame field of view image based on the current frame field of view image. ORB feature points are easier to identify compared to low-texture region feature points. In the current frame field of view image, the high-texture region has more details and higher regional complexity, while the low-texture region is relatively smooth and has fewer details. During the processing, the high-texture region includes more edges, corner points, and larger gradient changes, while the low-texture region has fewer edges, gentle color and brightness changes, and a simple structure. That is, the feature points in the high-texture region are easier to track.
[0021] Exemplarily, when the endoscope is a medical endoscope, the high-texture region may include bronchial bifurcations or intestinal folds. The ORB feature points may be the feature points in the bronchial bifurcation region, and the ORB feature points may also be the feature points in the intestinal fold region. The low-texture region may include smooth mucous membranes or mucus-covered regions. The low-texture region feature points may be the feature points in the smooth mucous membrane region, and the low-texture region feature points may also be the feature points in the mucus-covered region.
[0022] In one embodiment, determining a high-texture feature point set and a low-texture feature point set in the current frame field of view image based on the current frame field of view image includes: downsampling the current frame field of view image to obtain a multi-scale pyramid image; the multi-scale pyramid image includes multiple pyramid layer images with different resolutions; using the ORB algorithm to extract features from the multi-scale pyramid image to determine the high-texture feature point set in the current frame field of view image; based on the multi-scale pyramid image and a pre-trained convolutional neural network model, determining a class probability map corresponding to the current frame field of view image, and determining the low-texture feature point set based on the pixels in the class probability map whose confidence is greater than a preset confidence threshold; the pre-trained convolutional neural network model is used to obtain an output class probability map based on the input current frame field of view image, and the class probability map includes the confidence of each pixel in the current frame field of view image belonging to the low-texture class. The data extracted from the second-to-last layer before CNN global average pooling can be determined as the feature vector corresponding to the low-texture region feature points.
[0023] The multi-scale pyramid can be a preprocessed multi-scale pyramid, and the preprocessing can be Gaussian filtering. In the embodiments of the present application, the number of layers of the multi-scale pyramid can be 3 layers, namely layer L0, layer L1, and layer L2, and the resolutions of the images of each layer are different. The low-resolution image layer can capture global motion, and the high-resolution image layer can optimize local details. Exemplarily, the resolution of layer L0 can be 640×480, the resolution of layer L1 can be 320×240, and the resolution of layer L2 can be 160×120. The Gaussian filtering can be performed on the current frame field of view image to obtain the L0 layer image; the bilinear interpolation downsampling can be performed on the L0 layer image to obtain the L1 layer image; the downsampling can be performed on the L1 layer image to obtain the L2 layer image.
[0024] The ORB (Oriented FAST and Rotated BRIEF) algorithm is a feature extraction and matching algorithm used in the field of computer vision. By using the ORB algorithm to perform feature extraction on the current frame field of view image to obtain ORB feature points, the coordinates corresponding to each ORB feature point can be obtained. The ORB algorithm uses the FAST algorithm to detect key points in the multi-scale pyramid image, calculates the direction of the key points through the first moment, and uses the direction-corrected BRIEF to generate feature descriptors.
[0025] In one embodiment, a training sample set can be obtained. The training sample set can include multiple training sample pairs. Each training sample pair includes an endoscopic lens field of view image and a corresponding sample label, and the sample label can be the low-texture area in the endoscopic lens field of view image. With the goal of minimizing the cross-entropy loss function, a convolutional neural network model is trained based on the training sample set to obtain a pre-trained convolutional neural network model.
[0026] Step S130, based on the high-texture feature point set and the low-texture feature point set, determine the feature weights corresponding to each feature point.
[0027] Among them, the feature points include ROB feature points and low-texture area feature points.
[0028] In an embodiment of the present application, the endoscopic host can determine the feature weights corresponding to each feature point based on the high-texture feature point set and the low-texture feature point set. During the process of the endoscope advancing in the body, the proportions of high-texture features and low-texture features at different positions are different. For example, when the active bending section of the medical endoscope is located in the human body, there are more ORB feature points when the endoscopic lens is located in the bronchus, and there are more low-texture area feature points when the endoscopic lens is located in the stomach or bladder. Based on the current frame field of view images obtained when the endoscopic lens is located at different positions, the number of ORB feature points and low-texture area feature points in the current frame field of view image can be recognized, and then the weights of the ORB feature points and low-texture area feature points can be dynamically adjusted.
[0029] Based on the high-texture feature point set and the low-texture feature point set, determine the feature weights corresponding to each feature point, including: converting the current frame field of view image into a single-channel grayscale image; for each ORB feature point: in the single-channel grayscale image, take a first image to be processed within a first preset range centered on the ORB feature point; based on a preset sliding window, calculate the variance of the pixel grayscale values within the sliding window in the first image to be processed; based on the variances corresponding to each ORB feature point, determine the first normalization value corresponding to each ORB feature point, and obtain an average texture score based on all the first normalization values; for each low-texture region feature point: centered on the low-texture region feature point, take a second image to be processed within a second preset range around it in the class probability map; determine the average confidence of the low-texture class in the second image to be processed; based on the average confidence corresponding to all the low-texture region feature points, determine an average semantic score; based on the average texture score and the average semantic score, determine the global modality weight; based on the probability class map and the first normalization values of each ORB feature point, respectively determine the individual weight of the ORB feature point and the individual weight of the low-texture region feature point; based on the individual weight of the ORB feature point, the individual weight of the low-texture region feature point, and the global modality weight, determine the first weight corresponding to each ORB feature point and the second weight corresponding to each low-texture region feature point; determine the feature weight of the ORB feature point as the first weight corresponding to the ORB feature point, and determine the feature weight of the low-texture region feature point as the difference between a preset value and the second weight corresponding to the low-texture region feature point.
[0030] Exemplarily, the current frame field of view image can be converted into a single-channel grayscale image. For each ORB feature point: in the single-channel grayscale image, centered on the ORB feature point, take a first image to be processed within a first preset range (32×32) around it, and based on the preset sliding window, calculate the variance of the pixel grayscale values within the sliding window in the first image to be processed; based on the variances corresponding to each ORB feature point, determine the first normalization value corresponding to each ORB feature point; and obtain an average texture score based on all the first normalization values.
[0031] Exemplarily, the process of determining the variance corresponding to the ORB feature point may include: ; where is the variance corresponding to the i-th ORB feature point, is the total number of pixels within the sliding window, is the -th pixel grayscale value in the sliding window, is the grayscale mean value within the sliding window.
[0032] Exemplarily, the process of determining the first normalization value may include: ; Among them, is the minimum value among the variances corresponding to all ORB feature points, is the maximum value among the variances corresponding to all ORB feature points, is the first normalized value corresponding to the i-th ORB feature point.
[0033] Exemplarily, the determination process of the average texture score may include: ; Among them, is the average texture score, is the number of valid ORB feature points in the current frame field of view image. can be the number of valid ORB feature points obtained after RANSAC screening of the high-texture feature point set.
[0034] It should be noted that if is close to 1, it indicates that most areas in the current frame field of view image are rich in high texture and are suitable for relying on ORB features; being close to 0 indicates that the scene is smooth and mostly low in texture, and CNN features need to be relied on.
[0035] Exemplarily, for each low-texture region feature point: taking the low-texture region feature point as the center, a second image to be processed with a second preset range (32×32) around it is taken in the class probability map; the average confidence of the low-texture class in the second image to be processed is determined; based on the average confidence corresponding to all low-texture region feature points, the average semantic score is determined. The class probability map has been normalized by the Sigmoid function.
[0036] Exemplarily, the determination process of the average confidence corresponding to the low-texture region feature point may include: ; Among them, is the average confidence corresponding to the j-th low-texture region feature point, is the number of pixels in the second image to be processed, is the confidence that the
[0037] Exemplarily, the determination process of the average semantic score may include: ; Among them, is the average semantic score, is the number of valid low-texture region feature points in the current frame field of view image. can be the number of valid low-texture region feature points obtained after RANSAC screening of the high-texture feature point set.
[0038] Based on the average texture score and the average semantic score, determine the global modality weight: ; wherein, is the global modality weight, is the empirical balance factor, and its value range is 0.6 - 0.8.
[0039] Exemplarily, based on the probability category map and the first normalization value of each ORB feature point, determine the individual weight of the ORB feature point and the individual weight of the feature point in the low-texture area respectively; based on the individual weight of the ORB feature point, the individual weight of the feature point in the low-texture area, and the global modality weight, determine the first weight corresponding to each ORB feature point and the second weight corresponding to each feature point in the low-texture area.
[0040] Exemplarily, the individual weight of the ORB feature point can be expressed as: ; wherein, is the individual weight of the i-th ORB feature.
[0041] Exemplarily, the individual weight of the feature point in the low-texture area can be expressed as: ; wherein, is the individual weight of the j-th feature point in the low-texture area.
[0042] The first weight can be expressed as: ; wherein, is the first weight corresponding to the i-th ORB feature point.
[0043] The second weight can be expressed as: ; wherein, is the second weight corresponding to the j-th feature point in the low-texture area.
[0044] Determine the first weight corresponding to the ORB feature point as the feature weight of the ORB feature point, and determine the difference between the preset value and the second weight corresponding to the feature point in the low-texture area as the feature weight of the feature point in the low-texture area.
[0045] Exemplarily, the preset value can be 1. When the feature point is an ORB feature point, , represents the feature weight of the feature point, represents the first weight of the ORB feature point; when the feature point is a feature point in the low-texture area, , Represents the second weight of the feature points in the low-texture area.
[0046] Step S140, based on the high-texture feature point set, the low-texture feature point set, and the feature weights, use the pyramid-layered Lucas-Kanade optical flow method to determine the optical flow field.
[0047] In an embodiment of the present application, the optical flow field can be determined based on the high-texture feature point set, the low-texture feature point set, and the feature weights by using the pyramid-layered Lucas-Kanade optical flow method. The pyramid-layered Lucas-Kanade optical flow method estimates the rough motion in the low-resolution layer of the multi-scale pyramid image and gradually transfers and refines it to the high-resolution layer: The L2 layer can be defined as the top layer, and the displacement of the top layer is small. The LK optical flow method can be used to calculate the optical flow vector of this layer ; Upsample the optical flow result of the upper layer to the lower layer and use it as the initial estimate of the lower layer for local optimization; finally, optimize the optical flow at the original resolution layer to obtain the accurate result. The original resolution layer is also the L0 layer; in each layer of the multi-scale pyramid image, the optical flow vector is solved by minimizing the error function .
[0048] Before determining the optical flow field, it is necessary to first perform feature point matching on the feature points corresponding to the previous frame of the field of view image and the feature points corresponding to the current frame of the field of view image. For ORB feature points, the Hamming distance can be calculated for the ORB feature points corresponding to the previous frame of the field of view image and the current frame of the field of view image through the BRIEF descriptor. If the Hamming distance is less than the first threshold, it can be determined that the two ORB feature points are the same physical point, and the ORB feature points corresponding to the current frame of the field of view image can be determined as target feature points; for the feature points in the low-texture area, the cosine similarity can be calculated for the feature vectors corresponding to the low-texture area feature points of the previous frame of the field of view image and the current frame of the field of view image. If the cosine similarity is greater than the second threshold, it can be determined that the two low-texture area feature points are the same physical point, and the low-texture area feature points corresponding to the current frame of the field of view image can be determined as target feature points. The first threshold and the second threshold can be set by the user according to the actual situation.
[0049] Exemplarily, the initial displacement estimate of each target feature point is obtained using the LK optical flow method at the top layer , and the specific formula can be expressed as: ; where is the feature weight of the target feature point, is the gray-scale gradient of the L2 layer image in the x direction, is the gray-scale gradient of the L2 layer image in the y direction, is the time gradient corresponding to the L2 layer image. It can be calculated by the Sobel operator and , is the time gradient, which is the gray difference between the current frame and the previous frame. is the displacement in the x direction in the initial displacement estimation, is the displacement in the y direction in the initial displacement estimation. Based on the difference between the gray value corresponding to the L2 layer image of the current frame and the gray value corresponding to the L2 layer image of the previous frame.
[0050] Select a local window of a preset size around each pixel point, and obtain the initial displacement estimation of each feature point by solving the above formula . It should be noted that within the local window, if the pixel point is the target feature point, the corresponding feature weight based on the target feature point participates in the optical flow calculation; if the pixel point is not the target feature point, it participates in the optical flow calculation based on the preset weight, and the preset weight can be 0.1. By assigning a higher weight to the target feature point and a lower preset weight to other pixel points, the contribution rate of the known target feature point to the optical flow calculation is increased, and the effect of improving the accuracy of the optical flow calculation can be achieved.
[0051] Map the initial displacement estimation of the L2 layer to the L1 layer to obtain the intermediate displacement of the L1 layer , where represents the displacement in the x direction in the intermediate displacement, represents the displacement in the y direction in the intermediate displacement, and a is the multiple of the resolution between the L0 layer and the L1 layer. When the resolution of the L0 layer is 640×480, the resolution of the L1 layer is 320×240, and the resolution of the L2 layer is 160×120, , .
[0052] Align the L1 layer of the previous frame of the field of view image with the initial displacement estimation to obtain the first aligned image. Based on the first aligned image and the L1 layer image, determine the corrected displacement corresponding to the L1 layer. The process of determining the corrected displacement may include: ; where is the gray gradient of the L1 layer image in the x direction, is the gray gradient of the L1 layer image in the y direction, is the time gradient residual corresponding to the L1 layer image, is the residual in the x direction, is the residual in the y direction; ; ; where is the time gradient corresponding to the L1 layer image of the current frame, is the gray value of the first aligned image, is the gray value of the L2 layer image corresponding to the previous frame of the field of view image after intermediate displacement alignment.
[0053] The correction displacement can be expressed as , , is the displacement in the x direction in the correction displacement, is the displacement in the y direction in the correction displacement.
[0054] Finally, map the correction displacement of the L1 layer to the L0 layer to obtain the basic displacement of the L0 layer , where represents the displacement in the x direction in the basic displacement, represents the displacement in the y direction in the basic displacement, and a is the multiple of the resolutions of the L1 layer and the L0 layer. When the resolution of the L0 layer is 640×480, the resolution of the L1 layer is 320×240, and the resolution of the L2 layer is 160×120, , .
[0055] Align the L0 layer of the previous frame of the field of view image with the basic displacement to align the L0 layer of the previous field of view image to obtain the second aligned image, and determine the optical flow field based on the second aligned image and the L1 layer image. The process of determining the optical flow field can include: ; where is the gray gradient of the L0 layer image in the x direction, is the gray gradient of the L0 layer image in the y direction, is the time gradient residual corresponding to the L0 layer image, is the residual in the x direction, is the residual in the y direction; ; ; where is the time gradient corresponding to the current frame L0 layer image, is the gray value of the second aligned image, is the gray value of the L1 layer image corresponding to the previous frame of the field of view image after intermediate displacement alignment.
[0056] The displacement of each pixel point in the optical flow field can be expressed as , , is the displacement in the x direction in the optical flow field, is the displacement in the y direction in the optical flow field. The optical flow field is a displacement field covering all pixels of the entire current frame field of view image.
[0057] It should be noted that in this application, first, feature points are matched to determine target feature points. Feature point matching provides sparse but reliable displacement estimation, which serves as the initial value for optical flow calculation, reducing the number of iterations. In low-texture areas, the optical flow field may fail due to insufficient gradients. At this time, the semantic information of CNN feature points can provide supplementary constraints. Feature point matching can detect and eliminate outliers in the optical flow field, such as incorrect displacements caused by reflection or motion blur.
[0058] Step S150: Determine the rotation angle difference corresponding to the current frame of the field of view image based on the previous frame of the field of view image, the optical flow field, and the endoscope motion model.
[0059] Wherein, the rotation angle difference corresponding to the current frame of the field of view image is the rotation angle difference of the current frame of the field of view image compared to the previous frame of the field of view image.
[0060] In an embodiment of this application, the rotation angle difference of the current frame of the field of view image compared to the previous frame of the field of view image can be determined based on the optical flow field and the endoscope motion model. The endoscope motion model can characterize the motion process of the endoscope. The endoscope motion model can combine the rotation angle of the endoscope, the scaling ratio of the front lens of the endoscope when approaching the same point in the field of view, and predict the position of a pixel point in the current frame of the field of view image based on the pixel point in the previous frame of the field of view image.
[0061] In an embodiment, the process of determining the rotation angle difference of the current frame of the field of view image compared to the previous frame of the field of view image based on the optical flow field and the endoscope motion model may include: predicting the predicted coordinates of the pixel points in the current frame of the field of view image based on the previous frame of the field of view image and the endoscope motion model; determining the actual coordinates of the pixel points in the current frame of the field of view image based on the previous frame of the field of view image and the optical flow field; and determining the rotation angle difference using the nonlinear least squares method based on the predicted coordinates and the actual coordinates of the pixel points.
[0062] Exemplarily, the endoscope motion model may include: ; Wherein, is the scaling factor, is the rotation angle, is the translational compensation in the x-axis direction, is the translational compensation in the y-axis direction, is the fixed offset between the camera optical axis and the axis of the bending section, is the x-axis coordinate of the q-th pixel in the previous frame of the field of view image, is the y-axis coordinate of the q-th pixel in the previous frame of the field of view image, is the predicted x-axis coordinate, is the predicted y-axis coordinate.
[0063] Solve for each parameter by non - linear least squares method: Objective function: ; wherein, is the actual coordinate on the x - axis, is the actual coordinate on the y - axis; Initialization: Use the parameters of the previous frame or assumptions , , , ; Construct the Jacobian matrix and use the Levenberg - Marquardt algorithm to iteratively update the parameters to approximate the optimal solution.
[0064] Based on the coordinates of each pixel point corresponding to the previous - frame field - of - view image and the optical flow field of each pixel point, the coordinates of each pixel point corresponding to the current - frame field - of - view image can be determined. For any pixel point in the previous - frame field - of - view image, the coordinates of the previous - frame field - of - view image can be added to the displacement of the pixel point in the optical flow field to obtain the actual coordinates of the pixel point in the current - frame field - of - view image.
[0065] Step S160, determine the target rotation angle of the current - frame field - of - view image relative to the initial - frame field - of - view image based on the rotation - angle differences corresponding to each frame of the field - of - view image.
[0066] In an embodiment of the present application, after determining the rotation - angle difference between the current - frame field - of - view image and the previous - frame field - of - view image in step S150, the rotation - angle differences of all frames of the field - of - view image starting from the initial position of the endoscope lens can be accumulated to obtain the target rotation angle of the current - frame field - of - view image relative to the initial - frame field - of - view image. The angle corresponding to the initial - frame field - of - view image can be defined as 0.
[0067] It should be noted that the initial - frame field - of - view image can be an image selected by the user after the endoscope is inserted into the body. The initial - frame field - of - view image can also be a field - of - view image obtained after the endoscope is inserted a preset distance.
[0068] In one embodiment, the process of determining the target rotation angle of the current frame field of view image relative to the initial frame field of view image based on the rotation angle differences corresponding to each frame field of view image may include: obtaining the angle initial value corresponding to the current key frame field of view image; the angle initial value is the difference between the absolute rotation angles of the current key frame field of view image and the previous key frame field of view image; based on the rotation angles corresponding to each frame image and the angle initial value, determining the absolute rotation angles of each frame image relative to the current key frame field of view image; based on each absolute rotation angle, obtaining the initial smoothing angle corresponding to the current frame field of view image; based on the absolute rotation angles corresponding to the previous frame field of view image and the frame field of view image before the previous one, as well as the initial smoothing angle, using the Kalman filtering algorithm to obtain the corrected angle corresponding to the current frame field of view image, and determining the corrected angle as the target rotation angle of the current frame field of view image relative to the initial frame field of view image.
[0069] Obtain the angle initial value corresponding to the current key frame field of view image; based on the rotation angles corresponding to each frame image, determine the absolute rotation angles of each frame image relative to the key frame field of view image. Taking the angle initial value corresponding to the key frame field of view image as a reference, based on the rotation angles corresponding to each frame field of view image, obtain the absolute rotation angles of each frame field of view image including the current frame field of view image and a preset number of frames relative to the current key frame field of view image. The angle initial value of the key frame field of view image may be the difference between the absolute rotation angles of the current key frame field of view image and the previous key frame field of view image. When the endoscopic lens starts to penetrate into the human body, the current key frame field of view image may be the initial frame field of view image, the initial frame field of view image is the image obtained when the endoscope is in the initial position, and the absolute rotation angle corresponding to the initial frame field of view image is 0 degree.
[0070] Exemplarily, the absolute rotation angles of 5 frames of images including the current frame field of view image may be obtained to form an absolute rotation angle sequence ; represents the absolute rotation angle corresponding to the current frame field of view image, and the current frame field of view image is the t-th frame field of view image. The determination formula of the absolute rotation angle may include: ; where, is the absolute rotation angle corresponding to the t-th frame field of view image, is the absolute rotation angle corresponding to the current key frame field of view image, is the rotation angle difference corresponding to the a-th frame field of view image.
[0071] Obtain the initial smoothing angle corresponding to the current frame field of view image based on the absolute rotation angle sequence: ; where, is the initial smoothing angle corresponding to the current frame field of view image.
[0072] Based on the previous-frame field of view image, the absolute rotation angle corresponding to the frame-before-previous-frame field of view image, and the initial smoothing angle, the Kalman filter algorithm is used to obtain the corrected angle corresponding to the current-frame field of view image, and the corrected angle is determined as the target rotation angle of the current-frame field of view image relative to the initial-frame field of view image.
[0073] First, the angular velocity can be determined based on the absolute rotation angle of the previous-frame field of view image and the absolute rotation angle of the frame-before-previous-frame field of view image: ; where is the angular velocity, is the preset frame rate, is the absolute rotation angle of the previous-frame field of view image, is the absolute rotation angle of the frame-before-previous-frame field of view image.
[0074] The predicted angle can be determined based on the absolute rotation angle of the previous-frame field of view image and the angular velocity: ; where is the predicted angle.
[0075] The predicted covariance can be determined: ; where is the process noise, is the previous-frame covariance, is the predicted covariance. The value of can be 0.1. The initial covariance
[0076] The Kalman gain can be determined: ; where is the Kalman gain, is the observation noise. The value of
[0077] The corrected angle can be determined based on the initial smoothing angle and the predicted angle: ; where is the target rotation angle of the current-frame field of view image relative to the initial-frame field of view image.
[0078] Finally, the covariance can be updated: ; where is the covariance of the current frame.
[0079] It should be noted that determining the initial smoothing angle can suppress short-term noise, such as lens jitter; Kalman filtering performs global correction, filters out sudden angles, and enhances the continuity of angle changes; the denoised observations are provided through the rotation angle sequence and further optimized by Kalman filtering through model prediction, and the two complement each other to enhance robustness. For abnormal jumps, the Kalman gain automatically reduces the trust in the observations and relies on model prediction for smooth transition.
[0080] The reference frame can be reset regularly or conditionally by setting the key-frame field of view image, which can solve the problems of cumulative error and environmental mutation in long-term navigation. Since the user needs to know the total rotation angle of the endoscope's current orientation relative to the initial position, it is necessary to determine the target rotation angle of the current frame field of view image compared to the initial frame field of view image. In addition, if only relying on the rotation angle difference between adjacent frames, long-term integration will lead to error accumulation. By regularly resetting the reference frame through the key-frame field of view image, the error can be limited within the interval of the key-frame field of view image, avoiding the angle error accumulation caused by infinite accumulation.
[0081] In one embodiment, the process of obtaining the angle initial value corresponding to the current key-frame field of view image may include: when the number of image frames after the current key-frame field of view image is equal to the preset number of frames, or the cumulative rotation angle of each frame field of view image after the current key-frame field of view image is greater than the preset angle, or the feature point matching rate between the current frame field of view image and the current key-frame field of view image is less than the first preset ratio, the current key-frame field of view image is determined as the previous key-frame field of view image; this step preliminarily determines whether the key-frame field of view image needs to be updated. Then, feature matching can be performed on the ORB feature points and low-texture region feature points of the current frame field of view image and the current key-frame field of view image to determine the feature point pairs in the two frame fields of view images, and RANSAC verification is performed based on the feature point pairs to determine the current inlier ratio. When the current inlier ratio is greater than or equal to the preset inlier ratio, the current frame field of view image is determined as the current key-frame field of view image; when the current inlier ratio is less than the preset inlier ratio, the previous key-frame field of view image is determined as the current key-frame field of view image. The preset inlier ratio can be 30%. Feature point matching is performed on the current frame field of view image and the key-frame field of view image to verify the geometric consistency between the current frame field of view image and the key-frame field of view image, ensuring the reliability of motion parameter calculation.
[0082] It should be noted that the cumulative rotation angle can be the difference between the absolute rotation angle of the current frame field of view image and the absolute rotation angle of the current key-frame field of view image, the preset angle can be 15 degrees, the first preset ratio can be 30%, and the preset number of frames can be 30 frames. Feature point matching rate = number of matching points / number of key-frame feature points.
[0083] By continuously updating the key-frame field of view image, the accumulation of errors can be avoided, and the accuracy of determining the rotation angle can be improved.
[0084] Figure 2 It is a block diagram of an endoscope rotation angle determination device shown in an exemplary embodiment of the present application. As Figure 2 shown, the exemplary endoscope rotation angle determination device 200 includes: A data acquisition module 210, configured to acquire a current-frame field of view image, where the current-frame field of view image is an image acquired by the endoscope lens at the current moment.
[0085] A point set determination module 220, configured to determine a high-texture feature point set and a low-texture feature point set in the current-frame field of view image based on the current-frame field of view image; the high-texture feature point set includes a plurality of ORB feature points, and the low-texture feature point set includes a plurality of low-texture region feature points; the visual saliency of the ORB feature points in the current-frame field of view image is greater than that of the low-texture region feature points.
[0086] A weight determination module 230, configured to determine a feature weight corresponding to each feature point based on the high-texture feature point set and the low-texture feature point set; the feature points include ROB feature points and low-texture region feature points.
[0087] A data processing module 240, configured to determine an optical flow field by using the Lucas-Kanade optical flow method with pyramid layering based on the high-texture feature point set, the low-texture feature point set, and the feature weights.
[0088] A first angle determination module 250, configured to determine a rotation angle difference corresponding to the current-frame field of view image based on the previous-frame field of view image, the optical flow field, and the endoscope motion model, where the rotation angle difference corresponding to the current-frame field of view image is the rotation angle difference of the current-frame field of view image compared with the previous-frame field of view image.
[0089] A second angle determination module 260, configured to determine a target rotation angle of the current-frame field of view image compared with the initial-frame field of view image based on the rotation angle differences corresponding to each frame of the field of view image.
[0090] It should be noted that the endoscope rotation angle determination device provided in the above embodiment and the endoscope rotation angle determination method provided in the above embodiment belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiment, and will not be repeated here. In practical applications, the endoscope rotation angle determination device provided in the above embodiment can, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the system into different functional modules to complete all or part of the functions described above, and this is not limited here either.
[0091] Embodiments of the present application also provide an endoscope system, including an endoscope host, where the endoscope host includes: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, cause the endoscope system to implement the endoscope rotation angle determination method provided in each of the above embodiments.
[0092] On the other hand, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor of a computer, the computer is caused to execute the endoscope rotation angle determination method provided in each of the above embodiments. The computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist alone without being assembled into the electronic device.
[0093] On the other hand, the present application also provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the endoscope rotation angle determination method provided in each of the above embodiments.
[0094] In the embodiments of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. The terms "comprising" and "including" mentioned throughout the specification and claims are open-ended terms and should be interpreted as "including but not limited to".
[0095] The above embodiments are only used to exemplarily illustrate the principles and effects of the present application, rather than to limit the present application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed in the present application should still be covered by the claims of the present application.
Claims
1. An endoscope rotation angle determination method, characterized in that Including: Obtain the current frame field of view image, where the current frame field of view image is the image obtained by the endoscope lens at the current moment; Based on the current frame field of view image, determine the high-texture feature point set and the low-texture feature point set in the current frame field of view image; The high-texture feature point set includes multiple ORB feature points, and the low-texture feature point set includes multiple low-texture region feature points; in the current frame field of view image, the visual saliency of the ORB feature points is greater than that of the low-texture region feature points; Based on the high-texture feature point set and the low-texture feature point set, determine the feature weights corresponding to each feature point; the feature points include ROB feature points and low-texture region feature points; Based on the high-texture feature point set, the low-texture feature point set, and the feature weights, use the pyramid hierarchical Lucas-Kanade optical flow method to determine the optical flow field; Based on the previous frame field of view image, the optical flow field, and the endoscope motion model, determine the rotation angle difference corresponding to the current frame field of view image. The rotation angle difference corresponding to the current frame field of view image is the rotation angle difference of the current frame field of view image compared to the previous frame field of view image; Based on the rotation angle differences corresponding to each frame field of view image, determine the target rotation angle of the current frame field of view image compared to the initial frame field of view image.
2. The endoscopic rotation angle determination method according to claim 1, characterized in that, Based on the current frame field of view image, determining the high-texture feature point set and the low-texture feature point set in the current frame field of view image includes: Downsample the current frame field of view image to obtain a multi-scale pyramid image; the multi-scale pyramid image includes multiple pyramid layer images with different resolutions; Use the ORB algorithm to extract features from the multi-scale pyramid image to determine the high-texture feature point set in the current frame field of view image; Based on the multi-scale pyramid image and a pre-trained convolutional neural network model, determine the class probability map corresponding to the current frame field of view image, and based on the pixels in the class probability map whose confidence is greater than the preset confidence threshold, determine the low-texture feature point set; the pre-trained convolutional neural network model is used to obtain the output class probability map based on the input current frame field of view image, and the class probability map includes the confidence of each pixel in the current frame field of view image belonging to the low-texture class.
3. The endoscopic rotation angle determination method according to claim 2, wherein Based on the high-texture feature point set and the low-texture feature point set, determining the feature weights corresponding to each feature point includes: Convert the current frame field of view image into a single-channel grayscale image; For each ORB feature point: in the single-channel grayscale image, take the first image to be processed within the first preset range centered on the ORB feature point; based on the preset sliding window, calculate the variance of the pixel grayscale values within the sliding window in the first image to be processed; Based on the variances corresponding to each ORB feature point, determine the first normalization value corresponding to each ORB feature point, and based on all the first normalization values, obtain the average texture score; For each low-texture region feature point: take the second image to be processed within the second preset range around it in the class probability map; determine the average confidence of the low-texture class in the second image to be processed; Based on the average confidences corresponding to all low-texture region feature points, determine the average semantic score; Based on the average texture score and the average semantic score, determine the global modality weight; Based on the probability category map and the first normalization value of each ORB feature point, determine the individual weights of the ORB feature points and the individual weights of the feature points in the low-texture area respectively; based on the individual weights of the ORB feature points, the individual weights of the feature points in the low-texture area, and the global modality weight, determine the first weight corresponding to each ORB feature point and the second weight corresponding to each feature point in the low-texture area. Determine the feature weight of the ORB feature point as the first weight corresponding to the ORB feature point, and determine the feature weight of the feature point in the low-texture area as the difference between the preset value and the second weight corresponding to the feature point in the low-texture area.
4. The endoscopic rotation angle determination method according to claim 1, wherein Based on the high-texture feature point set, the low-texture feature point set, and the feature weights, use the pyramid-layered Lucas-Kanade optical flow method to determine the optical flow field, including: Based on the high-texture feature point set and the low-texture feature point set corresponding to the previous frame of the field of view image and the current frame of the field of view image, perform feature point matching to determine the target feature points that belong to the same physical point in the previous frame of the field of view image and the current frame of the field of view image. Based on the target feature points and the feature weights corresponding to the target feature points, use the pyramid-layered Lucas-Kanade optical flow method to determine the optical flow field.
5. The endoscopic rotation angle determination method according to claim 1, characterized in that Based on the optical flow field and the endoscope motion model, determine the rotation angle difference of the current frame of the field of view image compared to the previous frame of the field of view image, including: Based on the previous frame of the field of view image and the endoscope motion model, predict the predicted coordinates of the pixel points in the current frame of the field of view image. Based on the previous frame of the field of view image and the optical flow field, determine the actual coordinates of the pixel points in the current frame of the field of view image. Based on the predicted coordinates and the actual coordinates of the pixel points, use the nonlinear least squares method to determine the rotation angle difference.
6. The method for determining the rotation angle of the endoscope according to any one of claims 1-5, characterized in that, Based on the rotation angle differences corresponding to each frame of the field of view image, determine the target rotation angle of the current frame of the field of view image compared to the initial frame of the field of view image, including: Obtain the angle initial value corresponding to the current key frame of the field of view image; the angle initial value is the difference between the absolute rotation angles of the current key frame of the field of view image and the previous key frame of the field of view image. Based on the rotation angles corresponding to each frame of the image and the angle initial value, determine the absolute rotation angles of each frame of the image compared to the current key frame of the field of view image. Based on each absolute rotation angle, obtain the initial smoothed angle corresponding to the current frame of the field of view image. Based on the previous frame of the field of view image, the absolute rotation angles corresponding to the frame before the previous frame of the field of view image, and the initial smoothed angle, use the Kalman filter algorithm to obtain the corrected angle corresponding to the current frame of the field of view image, and determine the corrected angle as the target rotation angle of the current frame of the field of view image compared to the initial frame of the field of view image.
7. The method for determining the rotation angle of the endoscope according to claim 6, wherein Obtain the angle initial value corresponding to the current key frame of the field of view image, including: When the number of image frames after the current key frame of the field of view image is equal to the preset number of frames, or, the cumulative rotation angle of each frame of the field of view image after the current key frame of the field of view image is greater than the preset angle, or, the feature point matching rate between the current frame of the field of view image and the current key frame of the field of view image is less than the first preset ratio, determine the current key frame of the field of view image as the previous key frame of the field of view image. Perform feature matching on the feature points of the previous key frame of the field of view image and the current frame of the field of view image to determine the feature point pairs in the two frames of the field of view image. Perform RANSAC verification based on feature point pairs to determine the current inlier ratio; When the current inlier ratio is greater than or equal to the preset inlier ratio, determine the current frame field of view image as the current key frame field of view image; When the current inlier ratio is less than the preset inlier ratio, determine the previous key frame field of view image as the current key frame field of view image.
8. An endoscope rotation angle determination device, characterized in that It includes: A data acquisition module for acquiring the current frame field of view image, where the current frame field of view image is the image acquired by the endoscope lens at the current moment; A point set determination module for determining a high-texture feature point set and a low-texture feature point set in the current frame field of view image based on the current frame field of view image; The high-texture feature point set includes multiple ORB feature points, and the low-texture feature point set includes multiple low-texture region feature points; in the current frame field of view image, the visual saliency of the ORB feature points is greater than that of the low-texture region feature points; A weight determination module for determining the feature weights corresponding to each feature point based on the high-texture feature point set and the low-texture feature point set; the feature points include ROB feature points and low-texture region feature points; A data processing module for determining the optical flow field by using the pyramid hierarchical Lucas-Kanade optical flow method based on the high-texture feature point set, the low-texture feature point set, and the feature weights; A first angle determination module for determining the rotation angle difference corresponding to the current frame field of view image based on the previous frame field of view image, the optical flow field, and the endoscope motion model, where the rotation angle difference corresponding to the current frame field of view image is the rotation angle difference of the current frame field of view image compared to the previous frame field of view image; A second angle determination module for determining the target rotation angle of the current frame field of view image compared to the initial frame field of view image based on the rotation angle differences corresponding to each frame field of view image.
9. An endoscope system, characterized in that, It includes an endoscope host, and the endoscope host includes: One or more processors; A memory for storing the executable program code of the processor; Wherein, the processor is configured to execute the program code to implement the endoscope rotation angle determination method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by the processor of the computer, the computer is made to execute the endoscope rotation angle determination method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Item position change detection method and device, storage medium, and electronic device
CN109472824A
Target tracking method, device and equipment and storage medium
CN111709973A
Single target tracking method based on sparse optical flow motion enhancement
CN116543017A
Optical flow estimation method and device, computer equipment and storage medium
CN118279352A
Method, device and equipment for optimizing visual inertial odometer of unmanned aerial vehicle and medium
CN119124215A
Cited By
Optical flow point screening method and device, electronic equipment and storage medium
CN121120707A
Intelligent following shot method based on optical flow method and target behavior prediction
CN121665116A
Real-time video stream image stabilization and multi-magnification clear imaging method and device
CN122137986A