Endoscope rotation angle determination method, device, endoscope system and storage medium
By acquiring the set of high and low texture feature points in the endoscopic image and the calculation of optical flow field, combined with the endoscopic motion model, the accuracy of rotation angle judgment in the endoscopic operation is solved, and the operation convenience and accuracy are improved.
Patent Information
- Application Number
- CN202510838295.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In endoscopic operation, it is difficult for users to accurately judge the rotational transformation amount of real-time images relative to the initial position, resulting in a risk of misjudgment of orientation and clinical operation errors.
By obtaining the high-texture feature point set and low-texture feature point set in the current frame field image, combining the Lucas-Canard optical flow method of pyramid stratification, the optical flow field is determined, and the endoscopic motion model is used to calculate the rotation angle difference value to provide an accurate basis for judging the rotation angle.
The convenience of endoscope operation is improved, the accuracy of the rotation angle difference is enhanced, and a reliable judgment basis for user operation is provided.
Smart Images

Figure CN120370537B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method and device for determining an endoscope rotation angle, an endoscope system, and a storage medium. Background Art
[0002] The head of an endoscope is usually equipped with a light source, a camera, and a working channel. The user needs to adjust the observation angle or operation direction by rotating the operating handle to control the rotation of the front end. When the handle rotates, the field of view image captured by the endoscope lens will also rotate. Therefore, the operator needs to judge the relative position of the target area based on the rotation of the handle. In this process, the real-time image captured by the endoscope imaging system will undergo a coaxial rotation transformation with the rotation of the active bending section of the endoscope, while the image presented on the display lacks an absolute spatial orientation reference system. When there is a mirror-symmetrical anatomical structure in the exploration area, the operator needs to manually infer the image orientation based on the mechanical rotation angle of the handle.
[0003] However, complex operation paths lead to a degradation of spatial orientation perception, making it difficult for inexperienced physicians to accurately judge the rotation transformation of the real-time image relative to the initial position, which can easily lead to misjudgment of orientation and the risk of clinical operation errors. Therefore, it is urgent to develop a method to determine the rotation angle of the endoscope. Summary of the Invention
[0004] The present application provides a method, device, endoscope system and storage medium for determining the rotation angle of an endoscope, which can determine in real time the target rotation angle of the current frame field of view image of the endoscope compared to the initial frame field of view image, provide users with a more accurate judgment basis, and improve the convenience of endoscope operation.
[0005] The present application provides a method for determining an endoscope rotation angle, comprising:
[0006] Acquire a current frame field of view image, where the current frame field of view image is the image acquired by the endoscope lens at the current moment;
[0007] Determining, based on the current frame field of view image, a high texture feature point set and a low texture feature point set in the current frame field of view image; the high texture feature point set includes a plurality of ORB feature points, and the low texture feature point set includes a plurality of low texture area feature points; and visual saliency of the ORB feature points in the current frame field of view image is greater than visual saliency of the low texture area feature points;
[0008] Based on the high texture feature point set and the low texture feature point set, the feature weight corresponding to each feature point is determined; the feature points include ROB feature points and low texture area feature points;
[0009] Based on the high-texture feature point set, low-texture feature point set and feature weights, the pyramid-layered Lucas-Kanad optical flow method is used to determine the optical flow field;
[0010] Determine the rotation angle difference corresponding to the current frame field of view image based on the previous frame field of view image, the optical flow field and the endoscope motion model. The rotation angle difference corresponding to the current frame field of view image is the rotation angle difference between the current frame field of view image and the previous frame field of view image.
[0011] The target rotation angle of the current frame of the field of view image compared with the initial frame of the field of view image is determined based on the rotation angle difference corresponding to each frame of the field of view image.
[0012] To achieve the above-mentioned and other related purposes, the present application provides a device for determining the rotation angle of an endoscope, comprising:
[0013] A data acquisition module is used to acquire a current frame field of view image, where the current frame field of view image is the image acquired by the endoscope lens at the current moment;
[0014] a point set determination module, configured to determine, based on the current frame field of view image, a high-texture feature point set and a low-texture feature point set in the current frame field of view image; the high-texture feature point set includes a plurality of ORB feature points, and the low-texture feature point set includes a plurality of low-texture region feature points; and the visual saliency of the ORB feature points in the current frame field of view image is greater than the visual saliency of the low-texture region feature points;
[0015] A weight determination module is used to determine the feature weight corresponding to each feature point based on a high-texture feature point set and a low-texture feature point set; the feature points include ROB feature points and low-texture area feature points;
[0016] A data processing module is used to determine the optical flow field using a pyramid-layered Lucas-Kanad optical flow method based on a high-texture feature point set, a low-texture feature point set, and feature weights;
[0017] A first angle determination module is used to determine a rotation angle difference corresponding to a current frame of view image based on a previous frame of view image, an optical flow field, and an endoscope motion model, where the rotation angle difference corresponding to the current frame of view image is the rotation angle difference between the current frame of view image and the previous frame of view image;
[0018] The second angle determination module is configured to determine a target rotation angle of the current frame of the field of view image compared to the initial frame of the field of view image based on the rotation angle difference corresponding to each frame of the field of view image.
[0019] To achieve the above-mentioned and other related purposes, the present application further provides an endoscope system, including an endoscope host, wherein the endoscope host includes:
[0020] one or more processors;
[0021] a memory for storing program code executable by the processor;
[0022] The processor is configured to execute the program code to implement the above-mentioned method for determining the endoscope rotation angle.
[0023] To achieve the above-mentioned purpose and other related purposes, the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by the computer processor, the computer executes one or more of the aforementioned methods for determining the endoscope rotation angle.
[0024] As described above, the present application provides a method, device, endoscope system, and storage medium for determining an endoscope rotation angle, which have the following beneficial effects:
[0025] The present application discloses a method for determining the rotation angle of an endoscope. The method obtains a set of high-texture feature points and a set of low-texture feature points in the current frame field of view image, determines the feature weights corresponding to each feature point in different feature point sets, determines the optical flow field in combination with the pyramid-layered Lucas-Kanad optical flow method, and determines the rotation angle difference between the current frame field of view image and the previous frame field of view image through the optical flow field and the endoscope motion model, thereby determining the target rotation angle. By combining high-texture features, low-texture features, and the weights of each feature point, the contribution of each feature point to the optical flow field calculation is balanced, thereby enhancing the robustness of the optical flow field and improving the accuracy of the rotation angle difference. By determining the rotation angle difference, the target rotation angle of the current frame field of view image is determined, thereby assisting the user in operating the endoscope, providing the user with a more accurate basis for judgment, and achieving the effect of improving the convenience of endoscope operation.
[0026] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, serving to explain the principles of the present application. It is obvious that the drawings described below are merely some embodiments of the present application, and a person of ordinary skill in the art can derive other drawings based on these drawings without inventive effort. In the drawings:
[0028] Figure 1 is a flow chart of a method for determining an endoscope rotation angle shown in an exemplary embodiment of the present application;
[0029] Figure 2 It is a structural block diagram of an endoscope rotation angle determination device shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0030] The following will describe the embodiments of the present application with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand the other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for the purpose of illustrating the present application and are not intended to limit the scope of protection of the present application.
[0031] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application. Therefore, the illustrations only show components related to the present application and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.
[0032] In the following description, a large number of details are discussed to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.
[0033] See also Figure 1 , Figure 1 FIG. 1 is a flow chart of a method for determining an endoscope rotation angle according to an exemplary embodiment of the present application. Figure 1 It can be seen that the method for determining the endoscope rotation angle may include:
[0034] Step S110: Acquire the current frame field of view image.
[0035] The current frame field of view image is the image captured by the endoscope lens at the current moment.
[0036] In one embodiment of the present application, an endoscope host can obtain a current frame field of view image. A camera is provided at the front end of the active bending section of the endoscope, which can be inserted into the body. The camera at the front end of the active bending section of the endoscope can capture the current frame field of view image at a preset frame rate and transmit the captured current frame field of view image to the endoscope host.
[0037] Step S120 : determining a high texture feature point set and a low texture feature point set in the current frame field of view image based on the current frame field of view image.
[0038] Among them, the high texture feature point set includes multiple ORB feature points, and the low texture feature point set includes multiple low texture area feature points; the visual significance of the ORB feature points in the current frame field of view image is greater than the visual significance of the low texture area feature points.
[0039] In one embodiment of the present application, the endoscope host can determine a set of high-texture feature points and a set of low-texture feature points in the current frame field of view image based on the current frame field of view image. ORB feature points are easier to identify than feature points in low-texture areas. In the current frame field of view image, the high-texture area has more details and higher regional complexity, while the low-texture area is a relatively smooth area with fewer details. During the processing, the high-texture area includes more edges, corners, and large gradient changes, while the low-texture area has fewer edges, smoother color and brightness changes, and a simpler structure. That is, feature points in the high-texture area are easier to track.
[0040] For example, when the endoscope is a medical endoscope, the high-texture region may include a bronchial bifurcation or an intestinal fold, and the ORB feature point may be a feature point in the bronchial bifurcation region, or the ORB feature point may also be a feature point in the intestinal fold region. The low-texture region may include a smooth mucosa or a mucus-covered region, and the low-texture region feature point may be a feature point in the smooth mucosa region, or the low-texture region feature point may also be a feature point in the mucus-covered region.
[0041] In one embodiment, based on a current frame field of view image, a set of high-texture feature points and a set of low-texture feature points in the current frame field of view image are determined, including: downsampling the current frame field of view image to obtain a multi-scale pyramid image; the multi-scale pyramid image includes multiple pyramid layer images of different resolutions; extracting features from the multi-scale pyramid image using an ORB algorithm to determine the set of high-texture feature points in the current frame field of view image; determining a category probability map corresponding to the current frame field of view image based on the multi-scale pyramid image and a pre-trained convolutional neural network model, and determining a set of low-texture feature points based on pixels in the category probability map having confidence levels greater than a preset confidence threshold; the pre-trained convolutional neural network model is used to output a category probability map based on the input current frame field of view image, the category probability map including the confidence level that each pixel in the current frame field of view image belongs to a low-texture category. Data extracted by the penultimate layer of the CNN before global average pooling can be determined as feature vectors corresponding to feature points in the low-texture region.
[0042] The multi-scale pyramid may be a pre-processed multi-scale pyramid, and the pre-processing may be a Gaussian filter. In an embodiment of the present application, the number of layers of the multi-scale pyramid may be 3 layers, namely, the L0 layer, the L1 layer, and the L2 layer, and the resolution of each layer of the image is different. The low-resolution layer can capture global motion, and the high-resolution layer can optimize local details. Exemplarily, the resolution of the L0 layer may be 640×480, the resolution of the L1 layer may be 320×240, and the resolution of the L2 layer may be 160×120. The current frame field of view image may be Gaussian filtered to obtain the L0 layer image; the L0 layer image may be downsampled by bilinear difference to obtain the L1 layer image; and the L1 layer image may be downsampled to obtain the L2 layer image.
[0043] The ORB (Oriented FAST and Rotated BRIEF) algorithm is a feature extraction and matching algorithm used in computer vision. It extracts features from the current frame's field of view image to generate ORB feature points, which can be used to determine the coordinates of each ORB feature point. The ORB algorithm uses the FAST algorithm to detect key points in a multi-scale pyramid image, calculates the orientation of key points using first-order moments, and generates feature descriptors using orientation-corrected BRIEF.
[0044] In one embodiment, a training sample set may be obtained. The training sample set may include multiple training sample pairs, each of which includes an endoscope lens field of view image and a corresponding sample label. The sample label may be a low-texture region in the endoscope lens field of view image. A convolutional neural network model is trained based on the training sample set with the goal of minimizing a cross-entropy loss function, thereby obtaining a pre-trained convolutional neural network model.
[0045] Step S130 : determining a feature weight corresponding to each feature point based on the high-texture feature point set and the low-texture feature point set.
[0046] Among them, the feature points include ROB feature points and low texture area feature points.
[0047] In one embodiment of the present application, the endoscope host can determine the feature weight corresponding to each feature point based on the high texture feature point set and the low texture feature point set. During the process of the endoscope advancing in the body, the proportion of high texture features and low texture features at different positions is different. For example, when the active bending section of the endoscope of a medical endoscope is located in the human body, there are more ORB feature points when the endoscope lens is located in the bronchus, and there are more low texture area feature points when the endoscope lens is located in the stomach or bladder. Based on the current frame field of view image obtained when the endoscope lens is located at different positions, the number of ORB feature points and low texture area feature points in the current frame field of view image can be identified, and the weights of the ORB feature points and low texture area feature points can be dynamically adjusted.
[0048] Based on the high texture feature point set and the low texture feature point set, the feature weight corresponding to each feature point is determined, including: converting the current frame field of view image into a single-channel grayscale image; for each ORB feature point: taking the first image to be processed within a first preset range with the ORB feature point as the center in the single-channel grayscale image; calculating the variance of the pixel grayscale values within the sliding window in the first image to be processed based on the preset sliding window; determining the first normalized value corresponding to each ORB feature point based on the variance corresponding to each ORB feature point, and obtaining an average texture score based on all the first normalized values; for each low texture area feature point: taking the low texture area feature point as the center, taking the second image to be processed within a second preset range in the category probability map; determining the average low texture category in the second image to be processed confidence; determine the average semantic score based on the average confidence corresponding to all low texture area feature points; determine the global modal weight based on the average texture score and the average semantic score; determine the individual weights of the ORB feature points and the individual weights of the low texture area feature points based on the probability category map and the first normalized value of each ORB feature point; determine the first weight corresponding to each ORB feature point and the second weight corresponding to each low texture area feature point based on the individual weights of the ORB feature points, the individual weights of the low texture area feature points and the global modal weight; determine the first weight corresponding to the ORB feature point as the feature weight of the ORB feature point, and determine the difference between the preset value and the second weight corresponding to the low texture area feature point as the feature weight of the low texture area feature point.
[0049] For example, the current frame field of view image can be converted into a single-channel grayscale image. For each ORB feature point in the single-channel grayscale image, the first image to be processed is taken from a first preset range (32×32) around the ORB feature point. Based on a preset sliding window, the variance of the pixel grayscale values within the sliding window in the first image to be processed is calculated. Based on the variance corresponding to each ORB feature point, the first normalized value corresponding to each ORB feature point is determined. The average texture score is obtained based on all the first normalized values.
[0050] Exemplarily, the process of determining the variance corresponding to the ORB feature point may include:
[0051] ;
[0052] in, is the variance corresponding to the i-th ORB feature point, is the total number of pixels in the sliding window, The first Pixel grayscale value, is the grayscale mean within the sliding window.
[0053] Exemplarily, the process of determining the first normalized value may include:
[0054] ;
[0055] in, is the minimum value of the variance corresponding to all ORB feature points, is the maximum value of the variance corresponding to all ORB feature points, is the first normalized value corresponding to the i-th ORB feature point.
[0056] Exemplarily, the process of determining the average texture score may include:
[0057] ;
[0058] in, is the average texture score, The number of valid ORB feature points in the current frame field of view image. It can be the number of valid ORB feature points obtained after RANSAC screening of the high-texture feature point set.
[0059] It should be noted that if A value close to 1 indicates that most areas in the current frame field of view are high-textured and suitable for relying on ORB features; a value close to 0 indicates that the scene is smooth and mostly low-textured, and needs to rely on CNN features.
[0060] For each low-texture feature point, the following steps are used: A second image to be processed is taken from the class probability map, centered at the low-texture feature point and encompassing a second preset range (32×32 pixels); the average confidence level of the low-texture category in the second image to be processed is determined; and the average semantic score is determined based on the average confidence level corresponding to all low-texture feature points. The class probability map is normalized using a sigmoid function.
[0061] Exemplarily, the process of determining the average confidence level corresponding to the feature points in the low-texture area may include:
[0062] ;
[0063] in, is the average confidence corresponding to the j-th low texture area feature point, is the number of pixels of the second image to be processed, For the The confidence that a pixel belongs to the low texture class.
[0064] Exemplarily, the process of determining the average semantic score may include:
[0065] ;
[0066] in, is the average semantic score, is the number of effective low-texture feature points in the current frame field of view image. It can be the number of effective low-texture area feature points obtained after the high-texture feature point set is screened by RANSAC.
[0067] Based on the average texture score and the average semantic score, the global modality weight is determined:
[0068] ;
[0069] in, is the global modal weight, It is the empirical balance factor, and its value range is 0.6-0.8.
[0070] Exemplarily, based on the probability category map and the first normalized value of each ORB feature point, the individual weights of the ORB feature points and the individual weights of the low texture area feature points are determined respectively; based on the individual weights of the ORB feature points, the individual weights of the low texture area feature points and the global modal weight, the first weight corresponding to each ORB feature point and the second weight corresponding to each low texture area feature point are determined.
[0071] For example, the individual weights of ORB feature points can be expressed as:
[0072] ;
[0073] in, is the individual weight of the i-th ORB feature.
[0074] For example, the individual weights of feature points in low-texture areas can be expressed as:
[0075] ;
[0076] in, is the individual weight of the j-th feature point in the low texture area.
[0077] The first weight can be expressed as:
[0078] ;
[0079] in, is the first weight corresponding to the i-th ORB feature point.
[0080] The second weight can be expressed as:
[0081] ;
[0082] in, is the second weight corresponding to the j-th low texture area feature point.
[0083] The first weight corresponding to the ORB feature point is determined as the feature weight of the ORB feature point, and the difference between the preset value and the second weight corresponding to the low texture area feature point is determined as the feature weight of the low texture area feature point.
[0084] For example, the preset value may be 1. When the feature point is an ORB feature point, , represents the feature weight of the feature point, Indicates the first weight of the ORB feature point; when the feature point is a low-texture area feature point, , Indicates the second weight of the feature point in the low texture area.
[0085] Step S140 : determining an optical flow field using a pyramid-layered Lucas-Kanad optical flow method based on the high-texture feature point set, the low-texture feature point set, and the feature weights.
[0086] In one embodiment of the present application, the optical flow field can be determined by using a pyramid-layered Lucas-Kanade optical flow method based on a high-texture feature point set, a low-texture feature point set, and feature weights. The pyramid-layered Lucas-Kanade optical flow method estimates the rough motion at the low-resolution layer in the multi-scale pyramid image, and gradually transfers and refines it to the high-resolution layer: the L2 layer can be defined as the top layer, where the displacement is small, and the LK optical flow method can be used to calculate the optical flow vector of this layer. ; Upsample the optical flow results of the upper layer to the next layer and use them as the initial estimate for local optimization of the next layer; finally optimize the optical flow at the original resolution layer to obtain accurate results. The original resolution layer is also the L0 layer; in each layer of the multi-scale pyramid image, the optical flow vector is solved by minimizing the error function .
[0087] Before determining the optical flow field, it is necessary to first perform feature point matching on the feature points corresponding to the previous frame field of view image and the feature points corresponding to the current frame field of view image. For ORB feature points, the Hamming distance of the ORB feature points corresponding to the previous frame field of view image and the current frame field of view image can be calculated using the BRIEF descriptor. If the Hamming distance is less than the first threshold, the two ORB feature points can be determined to be the same physical point, and the ORB feature point corresponding to the current frame field of view image can be determined as the target feature point; for low-texture area feature points, the cosine similarity of the feature vectors corresponding to the low-texture area feature points corresponding to the previous frame field of view image and the current frame field of view image can be calculated. If the cosine similarity is greater than the second threshold, the two low-texture area feature points can be determined to be the same physical point, and the low-texture area feature point corresponding to the current frame field of view image can be determined as the target feature point. The first threshold and the second threshold can be set by the user according to actual conditions.
[0088] For example, the LK optical flow method is used on the top layer to obtain the initial displacement estimate of each target feature point , the specific formula can be expressed as:
[0089] ;
[0090] in, is the feature weight of the target feature point, is the grayscale gradient of the L2 layer image in the x direction, is the grayscale gradient of the L2 layer image in the y direction, is the time gradient corresponding to the L2 layer image. It can be calculated by the Sobel operator and , is the temporal gradient, which is the grayscale difference between the current frame and the previous frame. is the displacement in the x direction in the initial displacement estimation, is the displacement in the y direction in the initial displacement estimate. Based on the difference between the grayscale value corresponding to the L2 layer image of the current frame and the grayscale value corresponding to the L2 layer image of the previous frame.
[0091] A local window of preset size is selected around each pixel point, and the initial displacement estimate of each feature point is obtained by solving the above formula It should be noted that within a local window, if a pixel is a target feature point, it participates in the optical flow calculation based on the corresponding feature weight of the target feature point. If a pixel is not a target feature point, it participates in the optical flow calculation based on a preset weight, which can be 0.1. By assigning a higher weight to the target feature point and a lower preset weight to other pixels, the contribution rate of the known target feature point to the optical flow calculation is increased, which can achieve the effect of improving the accuracy of the optical flow calculation.
[0092] The initial displacement estimate of the L2 layer Mapped to the L1 layer, the intermediate displacement of the L1 layer is obtained ,in, represents the displacement in the x direction in the intermediate displacement, Indicates the displacement in the y direction in the intermediate displacement, and a is a multiple of the resolution of the L0 and L1 layers. When the resolution of the L0 layer is 640×480, the resolution of the L1 layer is 320×240, and the resolution of the L2 layer is 160×120, , .
[0093] Aligning the L1 layer of the previous frame of view image based on the L1 layer and the initial displacement estimate to obtain a first aligned image, and determining a corrected displacement corresponding to the L1 layer based on the first aligned image and the L1 layer image. The process of determining the corrected displacement may include:
[0094] ;
[0095] in, is the grayscale gradient of the L1 layer image in the x direction, is the grayscale gradient of the L1 layer image in the y direction, is the temporal gradient residual corresponding to the L1 layer image, is the residual in the x direction, is the residual in the y direction;
[0096] ;
[0097] ;
[0098] in, The temporal gradient corresponding to the L1 layer image of the current frame, is the grayscale value of the first aligned image, It is the grayscale value of the L2 layer image corresponding to the previous frame of field of view image after intermediate displacement alignment.
[0099] The corrected displacement can be expressed as , , is the displacement in the x direction in the correction displacement, is the displacement in the y direction in the correction displacement.
[0100] Finally, the corrected displacement of the L1 layer is mapped to the L0 layer to obtain the base displacement of the L0 layer ,in, represents the displacement in the x direction of the foundation displacement, Indicates the displacement in the y direction of the basic displacement, and a is a multiple of the resolution of the L1 layer and the L0 layer. When the resolution of the L0 layer is 640×480, the resolution of the L1 layer is 320×240, and the resolution of the L2 layer is 160×120, , .
[0101] The L0 layer of the previous frame of visual field image is aligned with the basic displacement to obtain a second aligned image, and the optical flow field is determined based on the second aligned image and the L1 layer image. The determination process of the optical flow field may include:
[0102] ;
[0103] in, is the grayscale gradient of the L0 layer image in the x direction, is the grayscale gradient of the L0 layer image in the y direction, is the time gradient residual corresponding to the L0 layer image, is the residual in the x direction, is the residual in the y direction;
[0104] ;
[0105] ;
[0106] in, The time gradient corresponding to the L0 layer image of the current frame, is the grayscale value of the second aligned image, It is the grayscale value of the L1 layer image corresponding to the previous frame of field of view image after intermediate displacement alignment.
[0107] The displacement of each pixel in the optical flow field can be expressed as , , is the displacement in the x direction in the optical flow field, is the displacement in the y direction in the optical flow field. The optical flow field is a displacement field covering all pixels in the entire field of view of the current frame.
[0108] It should be noted that in this application, feature points are first matched to determine the target feature points. Feature point matching provides sparse but reliable displacement estimates, which serve as the initial value for optical flow calculation and reduce the number of iterations. In low-texture areas, the optical flow field may fail due to insufficient gradients. At this time, the semantic information of CNN feature points can provide supplementary constraints. Feature point matching can detect and eliminate outliers in the optical flow field, such as erroneous displacements caused by reflections or motion blur.
[0109] Step S150 : determining a rotation angle difference corresponding to a current frame of the field of view image based on the previous frame of the field of view image, the optical flow field, and the endoscope motion model.
[0110] The corresponding degree rotation angle difference of the current frame field of view image is the rotation angle difference between the current frame field of view image and the previous frame field of view image.
[0111] In one embodiment of the present application, the rotation angle difference between the current frame field of view image and the previous frame field of view image can be determined based on the optical flow field and the endoscope motion model. The endoscope motion model can characterize the movement process of the endoscope. The endoscope motion model can combine the rotation angle of the endoscope and the zoom ratio of the endoscope front lens to the same point in the field of view when the endoscope is moving forward, and predict the position of the pixel point in the current frame field of view image based on the pixel point in the previous frame field of view image.
[0112] In one embodiment, the process of determining the rotation angle difference between the current frame field of view image and the previous frame field of view image based on the optical flow field and the endoscope motion model may include: predicting the predicted coordinates of the pixel points in the current frame field of view image based on the previous frame field of view image and the endoscope motion model; determining the actual coordinates of the pixel points in the current frame field of view image based on the previous frame field of view image and the optical flow field; and determining the rotation angle difference using a nonlinear least squares method based on the predicted coordinates and actual coordinates of the pixel points.
[0113] Exemplarily, the endoscope motion model may include:
[0114] ;
[0115] in, is the scaling factor, is the rotation angle, is the translation compensation in the x-axis direction, is the translation compensation in the y-axis direction, is the fixed offset between the camera optical axis and the axis of the bending section, is the x-axis coordinate of the qth pixel in the previous frame of field of view, is the y-axis coordinate of the qth pixel in the previous frame of field of view, is the predicted coordinate of the x-axis, Predict the coordinates for the y-axis.
[0116] By using the nonlinear least squares method, the parameters are solved:
[0117] Objective function: ;
[0118] in, is the actual x-axis coordinate, is the actual coordinate of the y-axis;
[0119] Initialization: Using the previous frame parameters or assumptions , , , ;
[0120] The Jacobian matrix is constructed, and the Levenberg-Marquardt algorithm is used to iteratively update the parameters to approach the optimal solution.
[0121] The coordinates of each pixel in the current frame can be determined based on the coordinates of each pixel in the previous frame and the optical flow field of each pixel. For any pixel in the previous frame, the coordinates of the previous frame can be added to the displacement of the pixel in the optical flow field to obtain the actual coordinates of the pixel in the current frame.
[0122] Step S160 : determining a target rotation angle of the current frame of the viewing image compared with the initial frame of the viewing image based on the rotation angle difference corresponding to each frame of the viewing image.
[0123] In one embodiment of the present application, after determining the rotation angle difference between the current frame field of view image and the previous frame field of view image in step S150, the rotation angle differences of all frame field of view images starting from the initial position of the endoscope lens can be accumulated to obtain the target rotation angle of the current frame field of view image compared to the initial frame field of view image. The angle corresponding to the initial frame field of view image can be defined as 0.
[0124] It should be noted that the initial frame field of view image may be an image selected by the user after the endoscope is inserted into the body. The initial frame field of view image may also be a field of view image acquired after the endoscope is inserted into a preset distance.
[0125] In one embodiment, step S160 is a process of determining the target rotation angle of the current frame field of view image compared to the initial frame field of view image based on the rotation angle difference corresponding to each frame field of view image, which may include: obtaining the initial angle value corresponding to the current key frame field of view image; the initial angle value is the difference in absolute rotation angle between the current key frame field of view image and the previous key frame field of view image; based on the rotation angle corresponding to each frame image and the initial angle value, determining the absolute rotation angle of each frame image compared to the current key frame field of view image; based on each absolute rotation angle, obtaining the initial smoothing angle corresponding to the current frame field of view image; based on the absolute rotation angles corresponding to the previous frame field of view image and the previous previous frame field of view image, and the initial smoothing angle, using the Kalman filtering algorithm to obtain the correction angle corresponding to the current frame field of view image, and determining the correction angle as the target rotation angle of the current frame field of view image compared to the initial frame field of view image.
[0126] Obtain the initial value of the angle corresponding to the current key frame field of view image; based on the rotation angle corresponding to each frame image, determine the absolute rotation angle of each frame image compared to the key frame field of view image. Based on the initial value of the angle corresponding to the key frame field of view image, and based on the rotation angle corresponding to each frame field of view image, obtain the absolute rotation angle of each frame field of view image including a preset number of frames of the current frame field of view image compared to the current key frame field of view image. The initial value of the angle of the key frame field of view image can be the difference in absolute rotation angle between the current key frame field of view image and the previous key frame field of view image. When the endoscope lens begins to penetrate the human body, the current key frame field of view image can be the initial frame field of view image, which is the image obtained when the endoscope is in the initial position, and the absolute rotation angle corresponding to the initial frame field of view image is 0 degrees.
[0127] For example, the absolute rotation angles of the five frames including the current frame field of view image can be obtained to form an absolute rotation angle sequence ; represents the absolute rotation angle corresponding to the current frame field of view image, and the current frame field of view image is the t-th frame field of view image. The formula for determining the absolute rotation angle may include:
[0128] ;
[0129] in, is the absolute rotation angle corresponding to the t-th frame field of view image, is the absolute rotation angle corresponding to the current key frame field of view image, is the rotation angle difference corresponding to the a-th frame field of view image.
[0130] Based on the absolute rotation angle sequence, the initial smoothing angle corresponding to the current frame field of view image is obtained:
[0131] ;
[0132] in, It is the initial smoothing angle corresponding to the current frame field of view image.
[0133] Based on the previous frame field of view image, the absolute rotation angle corresponding to the previous frame field of view image and the initial smoothing angle, the Kalman filtering algorithm is used to obtain the correction angle corresponding to the current frame field of view image, and the correction angle is determined as the target rotation angle of the current frame field of view image compared with the initial frame field of view image.
[0134] The angular velocity can be determined based on the absolute rotation angle of the previous frame of view image and the absolute rotation angle of the previous frame of view image:
[0135] ;
[0136] in, is the angular velocity, For the preset frame rate, is the absolute rotation angle of the previous frame field of view image, is the absolute rotation angle of the previous frame of field of view image.
[0137] The predicted angle can be determined based on the absolute rotation angle and angular velocity of the previous frame of the field of view image:
[0138] ;
[0139] in, is the prediction angle.
[0140] The forecast covariance can be determined:
[0141] ;
[0142] in, is the process noise, is the covariance of the previous frame, is the prediction covariance. The value of can be 0.1. Initial covariance Can be 1.
[0143] The Kalman gain can be determined:
[0144] ;
[0145] in, is the Kalman gain, is the observation noise. The value can be 0.05.
[0146] The correction angle can be determined based on the initial smoothing angle and the predicted angle:
[0147] ;
[0148] in, The target rotation angle of the current frame field of view image compared to the initial frame field of view image.
[0149] Finally, the covariance can be updated:
[0150] ;
[0151] in, is the current frame covariance.
[0152] It's important to note that determining the initial smoothing angle suppresses short-term noise, such as camera shake. The Kalman filter performs global correction, filtering out sudden angle changes and enhancing the continuity of angle changes. The rotation angle sequence provides denoised observations, which the Kalman filter further optimizes through model predictions. The two complement each other to enhance robustness. For unusual transitions, the Kalman gain automatically reduces confidence in the observations, relying instead on the model predictions for a smooth transition.
[0153] The problem of cumulative errors and sudden changes in the environment during long-term navigation can be solved by periodically or conditionally resetting the reference frame using keyframe field of view images. Since the user needs to know the total rotation angle of the endoscope's current orientation relative to the initial position, it is necessary to determine the target rotation angle of the current frame field of view image compared to the initial frame field of view image. In addition, if only the rotation angle difference between adjacent frames is relied upon, long-term integration will lead to error accumulation. By periodically resetting the reference frame using keyframe field of view images and limiting the error to the keyframe field of view image interval, the accumulation of angle errors caused by infinite accumulation can be avoided.
[0154] In one embodiment, the process of obtaining the initial angle value corresponding to the current keyframe field of view image may include: determining the current keyframe field of view image as the previous keyframe field of view image when the number of image frames after the current keyframe field of view image equals a preset number of frames, or when the cumulative rotation angle of each field of view image after the current keyframe field of view image is greater than a preset angle, or when the feature point matching rate between the current frame field of view image and the current keyframe field of view image is less than a first preset ratio; this step preliminarily determines whether the keyframe field of view image needs to be updated. Subsequently, feature matching can be performed on the ORB feature points and low-texture region feature points of the current frame field of view image and the current keyframe field of view image to determine feature point pairs in the two frames of view image, and RANSAC verification is performed based on the feature point pairs to determine the current inlier ratio. When the current inlier ratio is greater than or equal to the preset inlier ratio, the current frame field of view image is determined as the current keyframe field of view image; when the current inlier ratio is less than the preset inlier ratio, the previous keyframe field of view image is determined as the current keyframe field of view image. The preset inlier ratio can be 30%. Perform feature point matching on the current frame field of view image and the key frame field of view image to verify the geometric consistency of the current frame field of view image and the key frame field of view image, and ensure the reliability of motion parameter calculation.
[0155] It should be noted that the cumulative rotation angle can be the difference between the absolute rotation angle of the current frame field of view image and the absolute rotation angle of the current key frame field of view image. The preset angle can be 15 degrees, the first preset ratio can be 30%, and the preset number of frames can be 30. Feature point matching rate = number of matching points / number of key frame feature points.
[0156] By continuously updating the key frame field of view image, the accumulation of errors can be avoided and the accuracy of the rotation angle determination can be improved.
[0157] Figure 2 FIG. 1 is a block diagram of an endoscope rotation angle determination device shown in an exemplary embodiment of the present application. Figure 2 As shown, the exemplary endoscope rotation angle determination device 200 includes:
[0158] The data acquisition module 210 is used to acquire a current frame field of view image, where the current frame field of view image is the image acquired by the endoscope lens at the current moment.
[0159] The point set determination module 220 is used to determine a high-texture feature point set and a low-texture feature point set in the current frame field of view image based on the current frame field of view image; the high-texture feature point set includes multiple ORB feature points, and the low-texture feature point set includes multiple low-texture area feature points; the visual significance of the ORB feature points in the current frame field of view image is greater than the visual significance of the low-texture area feature points.
[0160] The weight determination module 230 is used to determine the feature weight corresponding to each feature point based on the high texture feature point set and the low texture feature point set; the feature points include ROB feature points and low texture area feature points.
[0161] The data processing module 240 is configured to determine an optical flow field by using a pyramid-layered Lucas-Kanad optical flow method based on the high-texture feature point set, the low-texture feature point set, and the feature weights.
[0162] The first angle determination module 250 is used to determine the rotation angle difference corresponding to the current frame field of view image based on the previous frame field of view image, the optical flow field and the endoscope motion model. The corresponding rotation angle difference of the current frame field of view image is the rotation angle difference between the current frame field of view image and the previous frame field of view image.
[0163] The second angle determination module 260 is configured to determine a target rotation angle of the current frame of the field of view image compared to the initial frame of the field of view image based on the rotation angle difference corresponding to each frame of the field of view image.
[0164] It should be noted that the endoscope rotation angle determination device provided in the above embodiment and the endoscope rotation angle determination method provided in the above embodiment are based on the same concept, and the specific manner in which each module and unit performs operations has been described in detail in the method embodiment and will not be repeated here. In actual application, the endoscope rotation angle determination device provided in the above embodiment can allocate the above functions to different functional modules as needed, that is, divide the internal structure of the system into different functional modules to complete all or part of the functions described above, and this is not limited here.
[0165] An embodiment of the present application also provides an endoscope system, including an endoscope host, the endoscope host including: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by one or more processors, the endoscope system implements the endoscope rotation angle determination method provided in the above-mentioned embodiments.
[0166] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon. When executed by a computer processor, the computer program causes the computer to perform the methods for determining the rotation angle of an endoscope provided in the above-described embodiments. The computer-readable storage medium may be included in the electronic device described in the above-described embodiments, or may exist independently and not be incorporated into the electronic device.
[0167] Another aspect of the present application further provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the endoscope rotation angle determination method provided in each of the above embodiments.
[0168] In the embodiments of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance. Throughout the specification and claims, the terms "including" and "comprising" are open-ended terms and should be interpreted as "including but not limited to."
[0169] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, any equivalent modifications or alterations accomplished by a person of ordinary skill in the art without departing from the spirit and technical concepts disclosed in this application shall be covered by the claims of this application.
Claims
1. A method for determining the rotation angle of an endoscope, characterized in that: include: Acquire a current frame field of view image, where the current frame field of view image is the image acquired by the endoscope lens at the current moment; Based on the current frame field of view image, determining a high texture feature point set and a low texture feature point set in the current frame field of view image; The high-texture feature point set includes a plurality of ORB feature points, and the low-texture feature point set includes a plurality of low-texture region feature points; the visual significance of the ORB feature points in the current frame field of view image is greater than the visual significance of the low-texture region feature points; Based on the high texture feature point set and the low texture feature point set, the feature weight corresponding to each feature point is determined; the feature points include ROB feature points and low texture area feature points; Based on the high-texture feature point set, low-texture feature point set and feature weights, the pyramid-layered Lucas-Kanad optical flow method is used to determine the optical flow field; Determine the rotation angle difference corresponding to the current frame field of view image based on the previous frame field of view image, the optical flow field and the endoscope motion model. The rotation angle difference corresponding to the current frame field of view image is the rotation angle difference between the current frame field of view image and the previous frame field of view image. The target rotation angle of the current frame of the field of view image compared with the initial frame of the field of view image is determined based on the rotation angle difference corresponding to each frame of the field of view image.
2. The method for determining the endoscope rotation angle according to claim 1, wherein: Determining, based on the current frame field of view image, a high texture feature point set and a low texture feature point set in the current frame field of view image, comprising: Downsampling the current frame field of view image to obtain a multi-scale pyramid image; the multi-scale pyramid image includes a plurality of pyramid layer images with different resolutions; The ORB algorithm is used to extract features from the multi-scale pyramid image and determine the high-texture feature point set in the current frame field of view image; Based on the multi-scale pyramid image and the pre-trained convolutional neural network model, the category probability map corresponding to the current frame field of view image is determined, and the low-texture feature point set is determined based on the pixels in the category probability map whose confidence is greater than a preset confidence threshold; the pre-trained convolutional neural network model is used to obtain the output category probability map based on the input current frame field of view image, and the category probability map includes the confidence that each pixel in the current frame field of view image belongs to the low-texture category.
3. The method for determining the endoscope rotation angle according to claim 2, wherein: Based on the high texture feature point set and the low texture feature point set, the feature weight corresponding to each feature point is determined, including: Convert the current frame field of view image into a single-channel grayscale image; For each ORB feature point: taking a first image to be processed within a first preset range with the ORB feature point as the center in the single-channel grayscale image; calculating the variance of the pixel grayscale values within the sliding window in the first image to be processed based on a preset sliding window; Determine the first normalized value corresponding to each ORB feature point based on the variance corresponding to each ORB feature point, and obtain the average texture score based on all the first normalized values; For each low-texture region feature point: taking the low-texture region feature point as the center, taking a second image to be processed within a second preset range around the feature point in the category probability map; determining an average confidence level of the low-texture category in the second image to be processed; Determine the average semantic score based on the average confidence corresponding to all feature points in the low-texture area; Determine the global modality weight based on the average texture score and the average semantic score; Based on the probability category map and the first normalized value of each ORB feature point, the individual weights of the ORB feature points and the individual weights of the low-texture region feature points are determined respectively; based on the individual weights of the ORB feature points, the individual weights of the low-texture region feature points and the global modal weight, the first weight corresponding to each ORB feature point and the second weight corresponding to each low-texture region feature point are determined; The first weight corresponding to the ORB feature point is determined as the feature weight of the ORB feature point, and the difference between the preset value and the second weight corresponding to the low texture area feature point is determined as the feature weight of the low texture area feature point.
4. The method for determining the endoscope rotation angle according to claim 1, wherein: Based on the high-texture feature point set, low-texture feature point set and feature weights, the pyramid-layered Lucas-Kanad optical flow method is used to determine the optical flow field, including: Based on the high-texture feature point set and the low-texture feature point set corresponding to the previous frame field of view image and the current frame field of view image, feature point matching is performed to determine the target feature points belonging to the same physical point in the previous frame field of view image and the current frame field of view image; Based on the target feature points and the feature weights corresponding to the target feature points, the pyramid-layered Lucas-Kanad optical flow method is used to determine the optical flow field.
5. The method for determining the endoscope rotation angle according to claim 1, wherein: The rotation angle difference between the current frame field of view image and the previous frame field of view image is determined based on the optical flow field and the endoscope motion model, including: Based on the previous frame of field of view image and the endoscope motion model, the predicted coordinates of the pixel points in the current frame of field of view image are predicted; Based on the previous frame of field of view image and the optical flow field, the actual coordinates of the pixel points in the current frame of field of view image are determined; Based on the predicted and actual coordinates of the pixel points, the nonlinear least squares method is used to determine the rotation angle difference.
6. The method for determining the rotation angle of an endoscope according to any one of claims 1 to 5, characterized in that: Determining a target rotation angle of the current frame field of view image compared to the initial frame field of view image based on the rotation angle difference corresponding to each frame field of view image includes: Get the initial angle value corresponding to the current key frame field of view image; the initial angle value is the difference between the absolute rotation angle of the current key frame field of view image and the previous key frame field of view image; Based on the rotation angle corresponding to each frame image and the initial value of the angle, determine the absolute rotation angle of each frame image compared to the current key frame field of view image; Based on each absolute rotation angle, the initial smoothing angle corresponding to the current frame field of view image is obtained; Based on the absolute rotation angles corresponding to the previous frame of view image and the previous previous frame of view image, as well as the initial smoothing angle, the Kalman filtering algorithm is used to obtain the correction angle corresponding to the current frame of view image, and the correction angle is determined as the target rotation angle of the current frame of view image compared to the initial frame of view image.
7. The method for determining the endoscope rotation angle according to claim 6, wherein: Get the initial angle value corresponding to the current keyframe field of view image, including: When the number of image frames after the current key frame field of view image is equal to the preset number of frames, or the cumulative rotation angle of each frame field of view image after the current key frame field of view image is greater than the preset angle, or the feature point matching rate between the current frame field of view image and the current key frame field of view image is less than a first preset ratio, the current key frame field of view image is determined as the previous key frame field of view image; Perform feature matching on the feature points of the previous key frame field of view image and the current frame field of view image to determine the feature point pairs in the two frames of field of view images; Perform RANSAC verification based on feature point pairs to determine the current internal point ratio; When the current inlier ratio is greater than or equal to the preset inlier ratio, the current frame field of view image is determined as the current key frame field of view image; When the current inlier ratio is less than the preset inlier ratio, the previous key frame field of view image is determined as the current key frame field of view image.
8. An endoscope rotation angle determination device, characterized in that: include: A data acquisition module is used to acquire a current frame field of view image, where the current frame field of view image is the image acquired by the endoscope lens at the current moment; A point set determination module, configured to determine a high texture feature point set and a low texture feature point set in the current frame field of view image based on the current frame field of view image; The high-texture feature point set includes a plurality of ORB feature points, and the low-texture feature point set includes a plurality of low-texture region feature points; the visual significance of the ORB feature points in the current frame field of view image is greater than the visual significance of the low-texture region feature points; A weight determination module is used to determine the feature weight corresponding to each feature point based on a high-texture feature point set and a low-texture feature point set; the feature points include ROB feature points and low-texture area feature points; A data processing module is used to determine the optical flow field using a pyramid-layered Lucas-Kanad optical flow method based on a high-texture feature point set, a low-texture feature point set, and feature weights; A first angle determination module is used to determine a rotation angle difference corresponding to a current frame of view image based on a previous frame of view image, an optical flow field, and an endoscope motion model, where the rotation angle difference corresponding to the current frame of view image is the rotation angle difference between the current frame of view image and the previous frame of view image; The second angle determination module is configured to determine a target rotation angle of the current frame of the field of view image compared to the initial frame of the field of view image based on the rotation angle difference corresponding to each frame of the field of view image.
9. An endoscope system, characterized in that: The endoscope host comprises: one or more processors; a memory for storing program code executable by the processor; The processor is configured to execute the program code to implement the method for determining the endoscope rotation angle according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the method for determining the endoscope rotation angle according to any one of claims 1 to 7.
Citation Information
Patent Citations
Item position change detection method and device, storage medium, and electronic device
CN109472824A
Single target tracking method based on sparse optical flow motion enhancement
CN116543017A