Camera positioning methods, devices, systems, and non-volatile storage media

By determining the three-dimensional coordinates and pose transformation relationship of the tracking point during bronchoscopy, the problem of camera tracking failure was solved, the success rate and accuracy of camera positioning were improved, and the accuracy and safety of bronchoscopy were enhanced.

CN119863517BActive Publication Date: 2026-01-06HANGZHOU BRONCUS MEDICAL CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411941630.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2026-01-06
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

In existing technologies, camera tracking failures in bronchoscopes lead to discrepancies between virtual bronchoscope images and real bronchial images. This requires pose registration adjustments, which consume significant computing resources and have a low success rate.

Method used

By determining the three-dimensional coordinates of the tracking point in continuously captured physiological channel images, the target pose of the camera is estimated using depth maps and pose transformation relationships. Combined with the sparse pyramid Lucas-Kanade optical flow algorithm and similarity optimization, the accuracy and success rate of pose registration are improved.

Benefits of technology

It reduces the consumption of computing resources, improves the success rate and accuracy of camera repositioning, and enhances the precision and safety of bronchoscopy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119863517B_ABST
    Figure CN119863517B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a camera positioning method, device, system and nonvolatile storage medium. The camera positioning method comprises: determining a plurality of first tracking points in a first image, and determining a corresponding second tracking point of each first tracking point in the plurality of first tracking points in a second image; determining a depth map corresponding to the first image, and determining a three-dimensional point corresponding to each first tracking point according to the depth map and a first tracking point coordinate of each first tracking point in a preset plane rectangular coordinate system, wherein the depth map is used to indicate a depth value of each pixel point in the first image; determining a target pose transformation relationship of a pose of the camera when shooting the second image relative to a pose when shooting the first image according to the three-dimensional point corresponding to each first tracking point and the second tracking point, and determining pose estimation data of the camera when shooting the second image according to the target pose transformation relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image data processing, and more particularly to a camera positioning method, apparatus, system, and non-volatile storage medium. Background Technology

[0002] In related technologies, when a bronchoscope moves within the physiological passageway, camera tracking of the bronchoscope can easily fail. This means that the virtual bronchoscope image displayed in the window during real-time tracking deviates significantly from the actual bronchial image. When this occurs, relying solely on pose registration to adjust the virtual bronchoscope's pose to achieve repositioning requires substantial computational resources and is rarely successful. Summary of the Invention

[0003] This application provides a camera positioning method, apparatus, system, and non-volatile storage medium to at least solve the technical problem in the related art that relying solely on pose registration to adjust the pose of a virtual bronchoscope to achieve repositioning of the bronchoscope requires a large amount of computing resources and has a low positioning success rate.

[0004] This application provides a camera positioning method, including: determining a plurality of first tracking points in a first image, and determining a second tracking point corresponding to each of the plurality of first tracking points in a second image, wherein the first image and the second image are physiological channel images continuously captured by the camera in the physiological channel; determining a depth map corresponding to the first image, and determining a three-dimensional point corresponding to each first tracking point based on the depth map and the coordinates of each first tracking point in a preset Cartesian coordinate system, wherein the depth map is used to indicate the depth value of each pixel in the first image; determining a target pose transformation relationship between the camera's pose when capturing the second image and its pose when capturing the first image based on the three-dimensional point corresponding to the first tracking point and the second tracking point, and determining target pose estimation data of the camera when capturing the second image based on the pose transformation relationship.

[0005] Optionally, determining the target pose transformation relationship of the camera when capturing the second image relative to the pose when capturing the first image, based on the 3D point corresponding to the first tracking point and the second tracking point, includes: determining multiple point sets, wherein each point set contains a preset number of 3D points and corresponding second tracking points, wherein the 3D point and its corresponding second tracking point correspond to the same first tracking point; for each point set, determining candidate pose transformation relationships based on the point set, wherein the candidate pose transformation relationship includes an estimated change matrix between the camera's pose when capturing the second image and the pose when capturing the first image, determined based on the point set; for each candidate pose transformation relationship, determining the projection point of the 3D point corresponding to each first tracking point in the second image based on the candidate pose transformation relationship, and determining the position error value between the projection point and the second tracking point corresponding to each 3D point; for each candidate pose transformation relationship, determining the number of 3D points whose position error value is less than a preset value, and determining the candidate pose transformation relationship with the largest number of 3D points whose position error value is less than the preset value as the target pose transformation relationship.

[0006] Optionally, determining the projection point of the 3D point corresponding to each first tracking point in the second image based on the candidate pose transformation relationship includes: determining the first 3D point coordinates of the 3D point in the first camera coordinate system, and determining the second 3D point coordinates of the 3D point in the second camera coordinate system based on the candidate pose transformation relationship and the first 3D point coordinates, wherein the first camera coordinate system is the camera coordinate system when the camera captures the first image, and the second camera coordinate system is the camera coordinate system when the camera captures the second image; and determining the projection point coordinates of the projection point of each 3D point in the second image in a preset Cartesian coordinate system based on the second 3D point coordinates of each 3D point.

[0007] Optionally, determining the coordinates of the first three-dimensional point in the first camera coordinate system includes: determining the camera intrinsics of the camera, wherein the camera intrinsics include the focal length and the principal point position; determining the depth value of each first tracking point based on the depth map; and determining the coordinates of the first three-dimensional point corresponding to each first tracking point based on the depth value of each first tracking point, the camera intrinsics, and the coordinates of the first tracking point of each first tracking point.

[0008] Optionally, determining the second tracking point corresponding to each of the multiple first tracking points in the second image includes: determining the coordinates of the second tracking point in a preset Cartesian coordinate system; determining the distance between each first tracking point and its corresponding second tracking point based on the coordinates of the first tracking point and the coordinates of the second tracking point corresponding to each first tracking point; calculating the average distance between each first tracking point and its corresponding second tracking point, and calculating the standard deviation of the distance between each first tracking point and its corresponding second tracking point based on the average distance; calculating a distance threshold based on the average distance and the standard deviation, and removing the first tracking points and their corresponding second tracking points whose distances are greater than the distance threshold.

[0009] Optionally, the camera localization method further includes: acquiring virtual images of the physiological channel model captured by the virtual camera corresponding to the camera under the pose indicated by each preset pose data, and filtering optimized pose data of the camera and the virtual camera from each preset pose data according to the similarity between the virtual image and the second image, wherein each preset pose data is pose data obtained by adding a random pose offset to the pose estimation data as the initial value; and updating the pose estimation data to the optimized pose data.

[0010] Optionally, selecting optimized pose data for the camera and the virtual camera from various preset pose data based on the similarity between the virtual image and the second image includes: determining the root mean square error (RMSE) value between the virtual image and the second image, wherein the RMSE value is used to characterize the similarity between the virtual image and the second image, and the RMSE value and similarity are negatively correlated; and determining the preset pose data corresponding to the virtual image with the smallest RMSE value as the optimized pose data.

[0011] Optionally, determining the root mean square error between the virtual image and the second image includes: dividing the virtual image into multiple partially overlapping virtual image blocks, and dividing the second image into multiple partially overlapping real image blocks; determining the virtual feature values ​​of the virtual image blocks and the real feature values ​​of the real image blocks, wherein both the virtual feature values ​​and the real feature values ​​include the standard deviation of pixel intensity values ​​within the image block, and a whiteness value used to represent the proportion of specified pixels in the image block; deleting virtual image blocks and real image blocks with whiteness values ​​greater than a preset whiteness value; and for the retained virtual image blocks, adjusting the pixel intensity value standard of each virtual image block. The differences are arranged in descending order to obtain the first virtual image block sequence. For each pair of retained real image blocks, they are arranged in descending order of the standard deviation of pixel intensity values ​​to obtain the first real image block sequence. A predetermined number of virtual image blocks are selected from the virtual image sequence in a forward-to-back order to obtain the second virtual image block sequence. A predetermined number of real image blocks are also selected from the real image sequence to obtain the second real image block sequence. The root mean square error between the second virtual image block sequence and the second real image block sequence is determined as the root mean square error between the virtual image and the second image.

[0012] Optionally, a specified pixel in an image block is determined by: determining the saturation and brightness values ​​of the pixels in the image block; and determining the pixel as the specified pixel when the saturation value of the pixel is less than a preset saturation value and the brightness value of the pixel is greater than a preset brightness value.

[0013] This application embodiment also provides a camera positioning device, including: a matching module, used to determine a plurality of first tracking points in a first image, and to determine a second tracking point corresponding to each of the plurality of first tracking points in a second image, wherein the first image and the second image are physiological channel images continuously captured by the camera in the physiological channel; a first pose estimation module, used to determine a depth map corresponding to the first image, and to determine a three-dimensional point corresponding to each first tracking point based on the depth map and the coordinates of each first tracking point in a preset Cartesian coordinate system, wherein the depth map is used to indicate the depth value of each pixel in the first image; and a second pose estimation module, used to determine a target pose transformation relationship between the camera's pose when capturing the second image and its pose when capturing the first image based on the three-dimensional point corresponding to each first tracking point and the second tracking point, and to determine pose estimation data of the camera when capturing the second image based on the target pose transformation relationship.

[0014] This application embodiment also provides a camera positioning system, which includes a camera and an electronic device. The electronic device is used to determine a plurality of first tracking points in a first image and to determine a second tracking point corresponding to each of the plurality of first tracking points in a second image. The first image and the second image are physiological channel images continuously captured by the camera in the physiological channel. The electronic device is used to determine a depth map corresponding to the first image and to determine a three-dimensional point corresponding to each first tracking point based on the depth map and the coordinates of each first tracking point in a preset Cartesian coordinate system. The depth map is used to indicate the depth value of each pixel in the first image. The electronic device is used to determine a target pose transformation relationship between the pose of the camera when capturing the second image and the pose when capturing the first image based on the three-dimensional point and the second tracking point corresponding to each first tracking point, and to determine pose estimation data of the camera when capturing the second image based on the target pose transformation relationship.

[0015] This application embodiment also provides a non-volatile storage medium storing a computer program, which implements a camera positioning method when executed by a processor.

[0016] This application also provides a computer program product, including a computer program that implements a camera positioning method when executed by a processor.

[0017] Based on the above scheme, this application determines the target pose transformation relationship when the camera captures the second image relative to when it captures the first image by continuously acquiring the first and second images. Based on the target pose transformation relationship, it determines the pose estimation data when the camera captures the second image, thereby providing an initial pose estimation value that is closer to the real pose for the subsequent pose registration process, reducing the computing resources required for the registration process, and improving the success rate of camera repositioning. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of the structure of a camera positioning system in an exemplary embodiment of this application;

[0020] Figure 2 This is a schematic flowchart of the camera tracking and positioning process in an exemplary embodiment of this application;

[0021] Figure 3 This is a schematic flowchart of a camera positioning method in an exemplary embodiment of this application;

[0022] Figure 4 This is a schematic diagram of physiological channel images captured by a camera in an exemplary embodiment of this application;

[0023] Figure 5 This is a schematic diagram of image tracking points in an exemplary embodiment of this application;

[0024] Figure 6 This is a schematic diagram of a matching point pair in an exemplary embodiment of this application;

[0025] Figure 7 This is a schematic diagram of the matching point pairs retained after filtering in an exemplary embodiment of this application;

[0026] Figure 8 This is a schematic diagram of a physiological channel image and its corresponding depth map in an exemplary embodiment of this application;

[0027] Figure 9 This is a schematic diagram of the camera pose in an exemplary embodiment of this application;

[0028] Figure 10 This is a schematic diagram of the physiological channel model in an exemplary embodiment of this application;

[0029] Figure 11 This is a schematic flowchart of an image preprocessing process in an exemplary embodiment of this application;

[0030] Figure 12 This is a schematic diagram of an image sub-block in an exemplary embodiment of this application;

[0031] Figure 13 This is a schematic diagram of a real physiological channel image and a virtual physiological channel image in an exemplary embodiment of this application;

[0032] Figure 14 This is a schematic diagram of the camera positioning device in an exemplary embodiment of this application;

[0033] Figure 15 This is a schematic diagram of the structure of an electronic device in an exemplary embodiment of this application. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0035] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0036] The technical solutions of this application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0037] Bronchoscopy navigation is an auxiliary technique used in bronchoscopy. It aims to guide the bronchoscope's movement through the physiological passage by tracking and positioning a camera within the bronchoscope, thereby improving the accuracy and safety of the procedure. Bronchoscopy is a medical procedure used to examine the trachea and bronchi, commonly used to diagnose and treat respiratory diseases such as lung cancer, bronchitis, and emphysema. The following requirements must be met when performing bronchoscopy navigation:

[0038] Improving examination accuracy: By introducing navigation technology, the anatomical structures of physiological passages such as the bronchi can be displayed in real-time images, helping operators to guide the bronchoscope more accurately, thereby improving the accuracy of the examination. This is crucial for the early detection and diagnosis of lesions.

[0039] Reducing patient discomfort: Traditional bronchoscopy may require repeated insertion of the bronchoscope, which can lead to patient discomfort and the risk of complications. Bronchoscopic navigation can help doctors locate the target position more quickly, reducing the number of insertions and thus alleviating patient discomfort.

[0040] Overall, the background of bronchoscopy navigation is based on the pursuit of improving the effectiveness of bronchoscopy and the patient experience. It uses modern technology to provide doctors with more information and tools to perform bronchoscopy more accurately and safely.

[0041] In related technologies, the common process for tracking and locating cameras within cameras is offline calculation + pose registration + real-time tracking. Offline calculation involves reconstructing a 3D bronchial tree based on the patient's CT data, calculating the centerline, and planning the route to the target point. Pose registration involves adjusting the pose of the virtual camera within the reconstructed 3D bronchial tree, ensuring that the virtual camera and the real camera are in the same position and posture within the bronchial tree and the actual lungs, respectively. Real-time tracking involves synchronously updating the images captured by the virtual camera and the real camera in real time if tracking is successful.

[0042] However, real-time camera navigation algorithms in related technologies are prone to tracking failures. This means that the virtual camera image displayed in the window deviates significantly from the actual bronchial image during real-time tracking. This is especially true when the camera moves too quickly through physiological channels such as the bronchi, encounters obstacles, or the captured images lack sufficient physiological channel features (such as airway openings, bifurcations, and folds). In these situations, relying solely on the matching between the virtual camera image and the real image captured by the real camera to adjust the virtual camera's pose is unlikely to successfully achieve re-matching and tracking of the real and virtual cameras.

[0043] To address the aforementioned issues, this application provides a camera positioning method that provides high-precision initial pose estimation data for the pose registration process, thereby improving the success rate of re-registration and tracking positioning of real and virtual cameras and reducing the required computing resources.

[0044] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0045] Figure 1 This is a schematic diagram of a camera positioning system provided in an embodiment of this application. This system can be used to execute the camera positioning method provided in the embodiment of this application. Figure 1 The camera positioning system shown includes a camera 10 and electronic equipment 12, wherein,

[0046] Electronic device 12 is configured to determine a plurality of first tracking points in a first image, and to determine a second tracking point corresponding to each of the plurality of first tracking points in a second image, wherein the first image and the second image are physiological channel images continuously captured by camera 10 in the physiological channel; determine a depth map corresponding to the first image, and determine a three-dimensional point corresponding to each first tracking point based on the depth map and the coordinates of each first tracking point in a preset Cartesian coordinate system, wherein the depth map is used to indicate the depth value of each pixel in the first image; determine the target pose transformation relationship between the pose of camera 10 when capturing the second image and the pose when capturing the first image based on the three-dimensional point and the second tracking point corresponding to each first tracking point, and determine the target pose estimation data of camera 10 when capturing the second image based on the pose transformation relationship.

[0047] In some embodiments of this application, the camera positioning system further includes a bronchoscope, wherein the camera 10 can be integrated into the bronchoscope.

[0048] As an optional approach, the complete process of the aforementioned camera positioning system during camera positioning is as follows: Figure 2 As shown, the process includes four steps: offline data computation, pose estimation data determination, pose registration, and real-time tracking. The offline data computation step primarily processes CT data to reconstruct a 3D model of the physiological pathway, such as a 3D lung airway tree. The pose estimation data determination step determines the camera's pose estimation data based on two consecutive frames of images captured by the camera, which involves executing... Figure 3 The camera localization method shown is as follows. In the pose registration step, the initial pose of the virtual camera is adjusted and optimized based on pose estimation data, ensuring that the similarity between the virtual image corresponding to the virtual camera and the image captured by the real camera 10 meets preset requirements. These preset requirements can be a similarity greater than a preset similarity threshold, or the highest similarity among multiple selectable initial poses. The real-time tracking step updates the pose of the virtual camera in the 3D model in real time and synchronously displays the image corresponding to the current virtual camera and the image captured by the real camera 10, providing navigation information for the bronchoscopy operator.

[0049] Figure 3 A flowchart of a camera positioning method provided in this application embodiment specifically includes the following steps:

[0050] Step 302: Determine multiple first tracking points in the first image, and determine the second tracking point corresponding to each of the multiple first tracking points in the second image, wherein the first image and the second image are physiological channel images continuously captured by the camera in the physiological channel;

[0051] Optionally, the first image captured by the camera is as follows: Figure 4 As shown, the first tracking point determined in the first image is as follows: Figure 5 As shown. From Figure 5 As can be seen, when determining the first tracking point, the first image can be divided into multiple image blocks of the same size and shape, and the geometric center point of each image block can be selected as the first tracking point.

[0052] Optionally, after determining the first tracking point, when determining the coordinates of the second tracking point, the sparse pyramid Lucas-Kanade optical flow algorithm and the first tracking point can be used to determine the position of the second tracking point corresponding to each first tracking point in the second image, and then determine the coordinates of the second tracking point.

[0053] The sparse pyramid Lucas-Kanade optical flow algorithm specifically includes the following steps:

[0054] The first step is to initialize the parameters, including the number of layers in the image pyramid, the pyramid's scale factor, and the iteration termination condition.

[0055] The second step involves constructing a pyramid from the input image based on the given number of pyramid layers and scale factor, thereby obtaining a series of images at different scales.

[0056] The third step involves using the sparse Lucas-Kanade optical flow algorithm to track feature points at each layer of the image pyramid. This algorithm utilizes the grayscale changes in local image regions to calculate the displacement of feature points, thereby obtaining the optical flow field.

[0057] The fourth step involves iteratively optimizing the optical flow estimation results, such as the optical flow field, until a preset termination condition is met, such as reaching the maximum number of iterations or achieving the specified accuracy requirements.

[0058] The fifth step is to output the final optical flow field, which is used to represent the displacement information of the first tracking point in the image, that is, the offset information of the position of the second tracking point in the image relative to the position of the first tracking point in the image.

[0059] The second tracking point obtained through the aforementioned optical flow field is as follows: Figure 6 As shown. Among them. Figure 6 The left side shows the first tracking point in the first image. Figure 6 The right side shows the second tracking point in the second image. From Figure 6 As can be seen, the first and second tracking points, which have a corresponding relationship, will be connected.

[0060] In the technical solution provided in step 302, after determining the second tracking point, due to the possibility of incorrect associations during the determination process, the first and second tracking points, which are considered to be related, may not be projections of the same object. Therefore, it is necessary to filter the first and second tracking points. The filtered first and second tracking points are as follows: Figure 7 As shown, where, Figure 7 The left side represents the first tracking point, and the right side represents the second tracking point. Matched first and second tracking points are connected. Optionally, the step of determining the second tracking point corresponding to each of the multiple first tracking points in the second image includes: determining the coordinates of the second tracking point in a preset Cartesian coordinate system; determining the distance between each first tracking point and its corresponding second tracking point based on the coordinates of each first tracking point and the coordinates of the second tracking point corresponding to each first tracking point; calculating the average distance between each first tracking point and its corresponding second tracking point, and calculating the standard deviation of the distance between each first tracking point and its corresponding second tracking point based on the average distance; calculating a distance threshold based on the average distance and the standard deviation, and removing first tracking points and their corresponding second tracking points whose distances are greater than the distance threshold.

[0061] Optionally, assume that set A contains all first tracking points and set B contains all second tracking points. Furthermore, for ease of subsequent calculation and description, we define the second tracking point matched by the i-th first tracking point in set A as the i-th second tracking point in set B. Then the coordinates of the i-th first tracking point in set A are (x... Ai ,y Ai ), the coordinates of the i-th second tracking point that matches the i-th first tracking point in set B are (x Bi ,y Bi The formula for calculating its distance is as follows:

[0062]

[0063] The distance between the first and second tracking points in each matching point pair can then be determined using the formula above. The average distance MeanDist for all matching point pairs can then be determined using the following formula:

[0064]

[0065] The standard deviation StdDevDist for all matching pairs can then be determined using the following formula:

[0066]

[0067] Then, for each matching point pair, it is retained if the distance between the matching points meets the following condition, otherwise it is discarded:

[0068] Dist<=MeanDist+Thresh*StdDevDist

[0069] Thresh is a preset adjustable parameter in the above conditions, for example, it can be set to 1.

[0070] As an optional implementation, when determining the coordinates of the tracking points in each frame of the image in the preset Cartesian coordinate system, the positions of each frame of the image in the preset Cartesian coordinate system are also kept consistent. For example, the lower left vertex of each image coincides with the origin of the preset Cartesian coordinate system.

[0071] Step 304: Determine the depth map corresponding to the first image, and determine the three-dimensional point corresponding to each first tracking point based on the depth map and the coordinates of each first tracking point in the preset Cartesian coordinate system. The depth map is used to indicate the depth value of each pixel in the first image.

[0072] In the technical solution provided in step 304, such as Figure 8 As shown, the depth value corresponding to each pixel in the first image captured by the camera can be determined by a depth estimation model, thus obtaining the depth value as shown in the figure. Figure 8 The depth map shown on the right. The grayscale value of each point in the depth map is the depth value of the corresponding pixel in the first image.

[0073] Step 306: Determine the target pose transformation relationship between the camera's pose when capturing the second image and its pose when capturing the first image based on the three-dimensional point corresponding to the first tracking point and the second tracking point, and determine the camera's pose estimation data when capturing the second image based on the target pose transformation relationship.

[0074] It should be noted that the camera pose data mentioned in this application refers to the camera pose in the computed tomography coordinate system (CT coordinate system). It can also refer to the camera pose in the world coordinate system.

[0075] In the technical solution provided in step 306, determining the pose transformation relationship between the camera's pose when capturing the second image and the target pose when capturing the first image, based on the 3D point and the second tracking point corresponding to each first tracking point, includes: determining multiple point sets, wherein each point set contains a preset number of 3D points and corresponding second tracking points, wherein the 3D point and the corresponding second tracking point correspond to the same first tracking point, and the preset number can be an integer greater than or equal to 4; for each point set, determining the candidate pose transformation relationship corresponding to the point set, wherein the candidate pose transformation relationship includes determining the pose transformation relationship based on the point set. The estimated change matrix between the pose of the camera when capturing the second image and the pose when capturing the first image is determined. For each candidate pose transformation relationship, the projection point of the 3D point corresponding to each first tracking point in the second image is determined based on the candidate pose transformation relationship, and the position error value between the projection point and the second tracking point corresponding to each 3D point is determined. For each candidate pose transformation relationship, the number of 3D points whose position error value is less than a preset value is determined, and the candidate pose transformation relationship with the largest number of 3D points whose position error value is less than the preset value is determined as the target pose transformation relationship.

[0076] As an optional implementation, the step of determining the projection point of the three-dimensional point corresponding to each first tracking point in the second image based on the alternative pose transformation relationship includes: determining the first three-dimensional point coordinates of the three-dimensional point in the first camera coordinate system, and determining the second three-dimensional point coordinates of the three-dimensional point in the second camera coordinate system based on the alternative pose transformation relationship and the first three-dimensional point coordinates, wherein the first camera coordinate system is the camera coordinate system when the camera captures the first image, and the second camera coordinate system is the camera coordinate system when the camera captures the second image; and determining the projection point coordinates of the projection point of each three-dimensional point in the second image in a preset Cartesian coordinate system based on the second three-dimensional point coordinates of each three-dimensional point.

[0077] In some embodiments of this application, the following pose estimation data determination process can also be used. This process includes the following stages:

[0078] Data preparation stage: In this stage, the input data is prepared, which includes the first three-dimensional point coordinates of each three-dimensional point, and the second tracking point coordinates of each three-dimensional point.

[0079] Random sampling stage: A set of points is randomly selected from the input data. This selected set of points is used to estimate the candidate pose estimation data of the camera when capturing the second image. The set of points includes a preset number of randomly selected point pairs. Each point pair includes a 3D point and its corresponding second tracking point. The preset number can be the minimum number of point pairs required to estimate the camera pose, typically three or four, but can be set according to actual needs. This embodiment does not impose any limitations on this.

[0080] Camera pose determination stage: Based on the point set selected in the random sampling stage, the pose transformation relationship is obtained by solving the PnP (Perspective-n-Point) problem;

[0081] Projection error calculation stage: Based on the alternative pose transformation relationship, determine the projection point coordinates of each 3D point in the second image, and determine the error between the projection point coordinates and the corresponding second tracking point coordinates.

[0082] Specifically, assume that after filtering the first and second tracking points, n point pairs are retained, where each point pair includes n first tracking points and n second tracking points. The n first tracking points in the first image correspond to n 3D points in the first camera coordinate system, where the coordinates of one of the first tracking points is p. i =(u i ,v i The corresponding 3D point has the coordinates P in the first camera coordinate system. i =(x i ,y i ,z i The coordinates of the second tracking point p′ in the second image are: i =(u′ i ,v′ i And setting the transformation matrix in the candidate pose transformation relationship as T, calculate the second three-dimensional point coordinates P′ of each three-dimensional point in the second camera coordinate system. i =(x′) i ,y′ i ,z′ i If the value is P', then the calculation formula is as follows: i =TP i

[0083] Suppose the camera's intrinsic parameters are represented by a matrix. K Then the projection coordinates p″ of the three-dimensional point in the second image are... i =(u″ i ,v″ i The calculation formula for ) is as follows:

[0084]

[0085] The formula for calculating the error between the coordinates of the projection point and the corresponding coordinates of the second tracking point is as follows:

[0086] Projection error

[0087] Interior point selection stage: Point pairs with projection errors less than the preset error threshold are marked as interior points, and point pairs with projection errors not less than the preset error threshold are marked as exterior points.

[0088] Iteration Phase: The steps between the random sampling phase and the interior point selection phase are executed iteratively. After meeting the preset iteration termination condition, the candidate pose transformation relationship with the highest number of interior points is selected as the final target pose transformation relationship. Then, as follows... Figure 9 As shown, the pose estimation data for the camera when capturing the second image is determined based on the final determined target pose transformation relationship and the known pose of the camera when capturing the first image.

[0089] Specifically, Figure 9 In this context, the subscript CT indicates the world coordinate system, the subscript c indicates the camera coordinate system based on the camera, and Q... (i-1) The pose of the camera when it takes the first image, Q (i) Let ΔQ(i) be the pose of the camera when capturing the second image, and let ΔQ(i) be the offset of the camera's pose when capturing the second image relative to its pose when capturing the first image.

[0090] As an optional implementation, the step of determining the coordinates of the first three-dimensional point in the first camera coordinate system includes: determining the camera intrinsic parameters of the camera, wherein the camera intrinsic parameters include the focal length and the principal point position, the focal length being the distance from the focal point of the camera optical system to the imaging plane, and the principal point position being the intersection of the camera's optical axis and the imaging plane, i.e., the center of the image; determining the depth value of each first tracking point based on the depth map; and determining the coordinates of the first three-dimensional point corresponding to each first tracking point based on the depth value of each first tracking point, the camera intrinsic parameters, and the coordinates of the first tracking point of each first tracking point.

[0091] Optionally, the 3D points (X, Y, Z) corresponding to each tracking point (u, v) retained after filtering in the first image can be determined based on the depth map. The value of Z can be obtained by accessing the corresponding depth map, and the calculation formulas for X and Y are as follows:

[0092]

[0093] c in the above formula x c y f x f y Here, c is the camera's intrinsic parameter. xc y f represents the pixel coordinates of the principal point positions in the horizontal and vertical directions, respectively. x f y These are the focal lengths in the horizontal and vertical directions, respectively.

[0094] As an optional implementation, during continuous tracking and positioning, the camera continuously captures multiple frames of images, and the camera's pose typically differs when capturing different images. To facilitate continuous tracking and positioning, the camera's pose can be initialized when the first frame is captured. Specifically, when the bronchoscope containing the camera reaches some prominent features in the physiological passage (such as the carina), the transformation matrix from the camera to the CT coordinate system can be calculated. T The subsequent update process for each frame involves calculating the pose increment of adjacent image frames, i.e., the camera's transformation matrix from the previous moment to the present moment, using the steps described above. I Finally, solve for the current pose C = TI of the camera in the CT coordinate system (i.e., the world coordinate system) when acquiring the current frame.

[0095] In some embodiments of this application, after determining the pose estimation data, the camera positioning method further includes: acquiring virtual images of the physiological channel model captured by the virtual camera corresponding to the camera under the pose indicated by each preset pose data, and filtering optimized pose data of the camera and the virtual camera from each preset pose data based on the similarity between the virtual image and the second image, wherein each preset pose data is pose data obtained by adding a random pose offset to the pose estimation data as the initial value; updating the pose estimation data to the optimized pose data. The physiological channel model is as follows: Figure 10 As shown.

[0096] Specifically, during the tracking and positioning of the camera, the acquired image data may contain some noise, and the incorrectly matched first and second tracking points removed during the selection of the second tracking points may not be sufficient. This can lead to a certain deviation between the determined pose estimation data and the actual pose data of the camera when capturing the second image. Therefore, it is necessary to register the virtual image acquired by the virtual camera with the second image to accurately determine the pose data of the second camera when capturing the second image, and to synchronize the poses of the virtual camera and the real camera. That is, the relative position and orientation of the virtual camera in the physiological channel model are the same as the relative position and orientation of the real camera in the physiological channel.

[0097] When registering the virtual image and the second image, multiple preset pose data can be obtained by adding random pose offsets to the pose estimation data as the initial value, and the virtual image captured by the virtual camera with each preset pose data can be determined. This step can be implemented using rendering technology to obtain a 2D VB image. Then, the similarity between the rendered 2D image data and the second image captured by the real camera can be compared, and the preset pose data with the highest similarity can be determined as the optimized pose data of the virtual camera in the physiological channel model, which is also the optimized pose data of the camera in the physiological channel.

[0098] As an optional implementation, the step of selecting optimized pose data of the camera and the virtual camera from various preset pose data based on the similarity between the virtual image and the second image includes: determining the root mean square error value between the virtual image and the second image, wherein the root mean square error value is used to characterize the similarity between the virtual image and the second image, and the root mean square error value and similarity are negatively correlated; and determining the preset pose data corresponding to the virtual image with the smallest root mean square error value as the optimized pose data.

[0099] Specifically, the optimization formula for optimizing pose data is determined based on the root mean square error value as follows:

[0100]

[0101] In the above formula, ΔQ (i) Let C represent the pose of the camera when capturing the second image, relative to the pose when capturing the first image. Let C represent the pose estimation data, and ΔQ represent the random pose offset. Indicates the second image, I V (CΔQ) represents the virtual image obtained after adding a random pose offset. This represents the minimum root mean square error between the second image and the virtual image. The specific meaning of the above optimization formula is to use C as the initial estimate, and then use the BOBYQA algorithm to calculate the update amount ΔQ, so that the rendered image... Images captured by a real camera Most similar.

[0102] ΔQ (i) Let be the camera's pose when capturing the second image, relative to its pose when capturing the first image. arg is an abbreviation for minimizer, indicating that the above formula seeks a ΔQ that minimizes the root mean square error.

[0103] Among them, the BOBYQA (Bound Optimization by Quadratic Approximation) algorithm is a numerical optimization algorithm, particularly suitable for solving nonlinear constraint optimization problems without gradient information. The BOBYQA algorithm stops iterating when a stopping condition is met, such as the change in the objective function being less than a preset threshold or the maximum number of iterations being reached.

[0104] The BobyQA algorithm searches for possible local optima through an iterative optimization process, without relying on the gradient of the objective function. This makes the algorithm widely applicable when solving optimization problems for which gradient information cannot be directly obtained.

[0105] As an alternative implementation method, the registration process can also be achieved by manually adjusting the pose corresponding to the virtual image.

[0106] In some embodiments of this application, such as Figure 11 As shown, the process for determining the root mean square error between the virtual image and the second image can be divided into an image segmentation stage, a feature value calculation stage, a sub-block selection stage, and a similarity calculation stage. Specifically, it includes the following steps: dividing the virtual image into multiple partially overlapping virtual image blocks, and dividing the second image into multiple partially overlapping real image blocks; determining the virtual feature values ​​of the virtual image blocks and the real feature values ​​of the real image blocks, wherein both the virtual and real feature values ​​include the standard deviation of pixel intensity values ​​within the image block, and a whiteness value used to represent the proportion of specified pixels in the image block; deleting virtual image blocks and real image blocks with whiteness values ​​greater than a preset whiteness value; and processing the retained virtual images... The virtual image blocks are arranged in descending order of the standard deviation of pixel intensity values ​​to obtain a first virtual image block sequence. Similarly, the real image blocks are arranged in descending order of the standard deviation of pixel intensity values ​​to obtain a first real image block sequence. A predetermined number of virtual image blocks are selected from the virtual image sequences in a forward-to-back order to obtain a second virtual image block sequence. A predetermined number of real image blocks are also selected from the real image sequences to obtain a second real image block sequence. The root mean square error between the second virtual image block sequence and the second real image block sequence is determined as the root mean square error between the virtual image and the second image.

[0107] Specifically, in Figure 11 In the image segmentation stage, the acquired k-th frame image B can be divided into... (k) (Images can be captured by any camera, such as the first image or the second image) divided into, for example... Figure 12 The diagram shows M×N overlapping sub-blocks. Overlapping sub-blocks prevents a small structure from being divided into two parts. Sub-block Dm,n Defined as:

[0108]

[0109] Where m and n satisfy the following constraints:

[0110] 2≤m≤M-1 and 2≤n≤N-1.

[0111] In the above constraints, W and H These represent the width and height of the input image, respectively. M N and p represent the number of width and height divisions of the image, respectively, and p and q represent the large sub-blocks D. m,n The range of values ​​for the horizontal axis (width) and vertical axis (height), where m and n represent the nth small block on the horizontal and vertical axes, respectively.

[0112] When calculating the eigenvalues, for each sub-block D m,n Its feature values ​​include the standard deviation of the intensity value of the sub-block and the whiteness of the sub-block. The intensity value can be a grayscale value or a pixel value. The whiteness of the sub-block is the number of specified pixels in the sub-block, where specified pixels can be pixels considered white.

[0113] Sub-block D m,n Standard deviation of strength value The calculation formula is as follows:

[0114]

[0115] Where |D m,n | is sub-block D m,n The number of pixels inside, It is the average intensity of all pixels within that sub-block.

[0116] As an optional implementation, a designated pixel in an image block is determined by: determining the saturation value and brightness value of a pixel in the image block; and determining the pixel as a designated pixel when the saturation value of the pixel is less than a preset saturation value and the brightness value of the pixel is greater than a preset brightness value.

[0117] Specifically, sub-block D m,n whiteness The calculation formula is as follows:

[0118]

[0119] Where W is a function used to determine whether the pixel is white, and the function is described as follows:

[0120]

[0121] In the formula, S and B correspond to the saturation and brightness values ​​of the pixel in the HSV space. The formula means that the saturation value S(v) of the pixel is less than or equal to a certain threshold T. s And the brightness value B(v) is greater than a certain threshold T B When it is white, its value is 1.

[0122] In some embodiments of this application, the selection of a sub-block in the (k)th image can be achieved using the following steps:

[0123] The first step is to add all sub-blocks of the (k)th image to the candidate list A. (k) middle.

[0124] The second step is to remove candidates with whiteness values ​​greater than a certain threshold T from the candidate list. w The sub-block, i.e.

[0125] The third step is to determine the standard deviation of the strength values. Arrange the remaining sub-blocks in descending order.

[0126] Step 4, in list A (k) Keep the first α·M·N and remove the remaining sub-blocks. (α is a scaling factor with a value range of 0≤α≤1, and M·N is the total number of sub-blocks into which the entire image is divided).

[0127] The above steps allow for the selection of only sub-blocks with characteristic structures, such as airway bifurcations or folds, for similarity calculation. Image sub-blocks with these characteristic structures exhibit a high standard deviation. However, while sub-blocks containing alveoli also have a high standard deviation, they can interfere with the similarity assessment results. Therefore, the influence of alveoli is eliminated by comprehensively considering the saturation and brightness values ​​of the sub-blocks; that is, the image sub-blocks containing alveoli are excluded if their whiteness value exceeds a certain threshold.

[0128] The virtual image I obtained by rendering the virtual camera can then be rendered using the methods described above. V The same processing is performed to obtain a set of real image sub-blocks and a set of virtual image sub-blocks arranged in descending order of intensity standard deviation. The similarity between the real and virtual images can then be calculated using the following formula:

[0129]

[0130] Where |A (k) | is list A (k) The number of remaining sub-block pairs, where |D| represents the number of pixels within the sub-block. and The average intensity of sub-block D corresponding to the real image and the virtual image, respectively. and This corresponds to the intensity value of each pixel in the real image and the virtual image sub-block D. This is the root mean square error between the virtual image and the real image, which is the similarity between the real image and the virtual image.

[0131] This application determines the pose transformation relationship of the camera when capturing the second image relative to when capturing the first image by continuously acquiring the first and second images by the camera, and determines the pose estimation data of the camera when capturing the second image based on the pose transformation relationship. This provides an initial pose estimation value that is closer to the real pose for the subsequent pose registration process, reduces the computing resources required for the registration process, and improves the success rate of camera repositioning.

[0132] After using the camera positioning method provided in this application embodiment, the images captured by the camera and its corresponding virtual camera are as follows: Figure 13 As shown, where Figure 13 The left side shows an image of a real physiological channel, and the right side shows an image of a virtual physiological channel.

[0133] Figure 14 A structural block diagram of a camera positioning device provided in this application embodiment, the device comprising:

[0134] The matching module 140 is used to determine a plurality of first tracking points in a first image and to determine a second tracking point corresponding to each of the plurality of first tracking points in a second image, wherein the first image and the second image are physiological channel images continuously captured by a camera in the physiological channel; the first pose estimation module 142 is used to determine a depth map corresponding to the first image and to determine a three-dimensional point corresponding to each first tracking point based on the depth map and the coordinates of each first tracking point in a preset Cartesian coordinate system, wherein the depth map is used to indicate the depth value of each pixel in the first image; the second pose estimation module 144 is used to determine the target pose transformation relationship between the pose of the camera when capturing the second image and the pose when capturing the first image based on the three-dimensional point corresponding to each first tracking point and the second tracking point, and to determine the target pose estimation data of the camera when capturing the second image based on the pose transformation relationship.

[0135] In some embodiments of this application, the step of the matching module 140 in determining the second tracking point corresponding to each of the plurality of first tracking points in the second image includes: determining the coordinates of the second tracking point in a preset Cartesian coordinate system; determining the distance between each first tracking point and its corresponding second tracking point based on the first tracking point coordinates of each first tracking point and the second tracking point coordinates of the second tracking point corresponding to each first tracking point; calculating the average distance between each first tracking point and its corresponding second tracking point, and calculating the standard deviation of the distance between each first tracking point and its corresponding second tracking point based on the average distance; calculating a distance threshold based on the average distance and the standard deviation, and removing the first tracking point and its corresponding second tracking point whose distance is greater than the distance threshold.

[0136] In some embodiments of this application, the second pose estimation module 144 determines the target pose transformation relationship of the camera when capturing the second image relative to the pose when capturing the first image based on the three-dimensional point and the second tracking point corresponding to each first tracking point. This includes: determining multiple point sets, wherein each point set contains a preset number of three-dimensional points and the second tracking points corresponding to the three-dimensional points, wherein the three-dimensional points and the second tracking points corresponding to the three-dimensional points correspond to the same first tracking point; for each point set, determining the candidate pose transformation relationship corresponding to the point set, wherein the candidate pose transformation relationship includes an estimated change matrix between the camera pose when capturing the second image and the pose when capturing the first image, determined based on the point set; for each candidate pose transformation relationship, determining the projection point of the three-dimensional point corresponding to each first tracking point in the second image based on the candidate pose transformation relationship, and determining the position error value between the projection point and the second tracking point corresponding to each three-dimensional point; for each candidate pose transformation relationship, determining the number of three-dimensional points whose position error value is less than a preset value corresponding to the candidate pose transformation relationship, and determining the candidate pose transformation relationship with the largest number of three-dimensional points whose position error value is less than the preset value as the target pose transformation relationship.

[0137] In some embodiments of this application, the step of the second pose estimation module 144 determining the projection point of the three-dimensional point corresponding to each first tracking point in the second image based on the candidate pose transformation relationship includes: determining the first three-dimensional point coordinates of the three-dimensional point in the first camera coordinate system, and determining the second three-dimensional point coordinates of the three-dimensional point in the second camera coordinate system based on the candidate pose transformation relationship and the first three-dimensional point coordinates, wherein the first camera coordinate system is the camera coordinate system when the camera captures the first image, and the second camera coordinate system is the camera coordinate system when the camera captures the second image; and determining the projection point coordinates of the projection point of each three-dimensional point in the second image in a preset Cartesian coordinate system based on the second three-dimensional point coordinates of each three-dimensional point.

[0138] In some embodiments of this application, the step of the second pose estimation module 144 determining the coordinates of the first three-dimensional point in the first camera coordinate system includes: determining the camera intrinsic parameters of the camera, wherein the camera intrinsic parameters include the focal length and the principal point position; determining the depth value of each first tracking point based on the depth map; and determining the coordinates of the first three-dimensional point corresponding to each first tracking point based on the depth value of each first tracking point, the camera intrinsic parameters, and the coordinates of the first tracking point of each first tracking point.

[0139] In some embodiments of this application, the camera positioning device is further configured to: acquire virtual images of physiological channel models collected by the virtual camera corresponding to the camera under the poses indicated by various preset pose data, and filter out optimized pose data of the camera and the virtual camera from the various preset pose data based on the similarity between the virtual image and the second image, wherein each preset pose data is pose data obtained by adding a random pose offset to the pose estimation data as the initial value; and update the pose estimation data to the optimized pose data.

[0140] In some embodiments of this application, the step of the camera positioning device selecting optimized pose data of the camera and the virtual camera from various preset pose data based on the similarity between the virtual image and the second image includes: determining the root mean square error value between the virtual image and the second image, wherein the root mean square error value is used to characterize the similarity between the virtual image and the second image, and the root mean square error value and the similarity are negatively correlated; and determining the preset pose data corresponding to the virtual image with the smallest root mean square error value as the optimized pose data.

[0141] In some embodiments of this application, the step of the camera positioning device determining the root mean square error value between the virtual image and the second image includes: dividing the virtual image into multiple partially overlapping virtual image blocks, and dividing the second image into multiple partially overlapping real image blocks; determining the virtual feature values ​​of the virtual image blocks and the real feature values ​​of the real image blocks, wherein both the virtual feature values ​​and the real feature values ​​include the standard deviation of pixel intensity values ​​within the image block, and a whiteness value used to represent the proportion of specified pixels in the image block; deleting virtual image blocks and real image blocks with whiteness values ​​greater than a preset whiteness value; and, for the retained virtual image blocks, processing them according to their respective characteristics. The pixel intensity values ​​are arranged in descending order to obtain a first virtual image block sequence. Similarly, the reserved real image blocks are arranged in descending order of their pixel intensity value standard deviations to obtain a first real image block sequence. A predetermined number of virtual image blocks are selected from the virtual image sequence in a forward-to-back order to obtain a second virtual image block sequence, and a predetermined number of real image blocks are selected from the real image sequence to obtain a second real image block sequence. The root mean square error between the second virtual image block sequence and the second real image block sequence is determined as the root mean square error between the virtual image and the second image.

[0142] In some embodiments of this application, a designated pixel in an image block is determined by: determining the saturation value and brightness value of a pixel in the image block; and determining the pixel as a designated pixel when the saturation value of the pixel is less than a preset saturation value and the brightness value of the pixel is greater than a preset brightness value.

[0143] See Figure 15 The present application provides a schematic diagram of the structure of an electronic device according to an embodiment. Figure 15 As shown, the electronic device includes a memory 1501 and a processor 1502.

[0144] The memory 1501 stores an executable computer program 1503. The processor 1502, coupled to the memory 1501, calls the executable computer program 1503 stored in the memory to execute the camera positioning method provided in the above embodiment.

[0145] For example, the computer program 1503 can be divided into one or more modules / units, which are stored in the memory 1501 and executed by the processor 1502 to complete this application. The one or more modules / units may include various modules in the camera positioning device in the above embodiments, such as: a matching module, a first pose estimation module, and a second pose estimation module.

[0146] Furthermore, the device also includes at least one input device and at least one output device.

[0147] The processor 1502, memory 1501, input devices, and output devices mentioned above can be connected via a bus.

[0148] The input device can be a camera, touch panel, physical buttons, or mouse, etc. The output device can be a display screen.

[0149] Furthermore, the device may include more components than illustrated, or combine certain components, or different components, such as network access devices, sensors, etc.

[0150] The processor 1502 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0151] The memory 1501 can be, for example, a hard disk drive, non-volatile memory (such as flash memory or other electronically programmable erasure-restricted memory used to form a solid-state drive), volatile memory (such as static or dynamic random access memory), etc., and this application embodiment is not limited thereto. Specifically, the memory 1501 can be an internal storage unit of the electronic device, such as the hard disk or RAM of the electronic device. The memory 1501 can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device. Further, the memory 1501 can also include both internal storage units and external storage devices of the electronic device. The memory 1501 is used to store computer programs and other programs and data required by the terminal. The memory 1501 can also be used to temporarily store data that has been output or will be output.

[0152] This application also provides a computer-readable non-volatile storage medium, which may be disposed in the electronic device described in the above embodiments. Figure 15The memory 1501 in the illustrated embodiment stores a computer program on the computer-readable storage medium. When executed by a processor, the program implements the image-based target position adjustment method described in the foregoing embodiments. Furthermore, the computer-readable storage medium can also be various media capable of storing program code, such as a USB flash drive, external hard drive, read-only memory (ROM), RAM, magnetic disk, or optical disk.

[0153] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the following camera positioning method: determining a plurality of first tracking points in a first image, and determining a second tracking point corresponding to each of the plurality of first tracking points in a second image, wherein the first image and the second image are physiological channel images continuously captured by the camera in the physiological channel; determining a depth map corresponding to the first image, and determining a three-dimensional point corresponding to each first tracking point based on the depth map and the coordinates of each first tracking point in a preset Cartesian coordinate system, wherein the depth map is used to indicate the depth value of each pixel in the first image; determining a target pose transformation relationship between the camera's pose when capturing the second image and its pose when capturing the first image based on the three-dimensional point and the second tracking point corresponding to each first tracking point, and determining pose estimation data of the camera when capturing the second image based on the target pose transformation relationship.

[0154] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0155] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0156] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0157] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned readable storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0158] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0159] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0160] In the description of this application, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this application, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this application, as well as the features of different embodiments or examples.

[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A camera positioning method characterized by, The method comprises the following steps: determining a plurality of first tracking points in a first image, and determining a corresponding second tracking point in a second image for each of the plurality of first tracking points, wherein the first image and the second image are physiological channel images successively captured by a camera in a physiological channel; determining a depth map corresponding to the first image, and determining a three-dimensional point corresponding to each of the first tracking points according to the depth map and a first tracking point coordinate of each of the first tracking points in a preset planar rectangular coordinate system, wherein the depth map is used to indicate a depth value of each pixel point in the first image; determining a target pose transformation relationship of a pose of the camera when capturing the second image relative to a pose when capturing the first image according to the three-dimensional point corresponding to each of the first tracking points and the second tracking point, and determining pose estimation data of the camera when capturing the second image according to the target pose transformation relationship; acquiring a virtual image of a physiological channel model captured by a virtual camera corresponding to the camera at a pose indicated by each of preset pose data; dividing the virtual image into a plurality of partially overlapped virtual image blocks, and dividing the second image into a plurality of partially overlapped real image blocks; determining a virtual feature value of the virtual image block and a real feature value of the real image block, wherein the virtual feature value and the real feature value each include a pixel point intensity value standard deviation within the image block and a whiteness value reflecting a specified pixel point proportion in the image block; deleting the virtual image block and the real image block with a whiteness value greater than a preset whiteness value; arranging the retained virtual image blocks in a descending order of pixel point intensity value standard deviation of each of the virtual image blocks to obtain a first virtual image block sequence, and arranging the retained real image blocks in a descending order of pixel point intensity value standard deviation of each of the real image blocks to obtain a first real image block sequence; selecting a preset number of virtual image blocks from the first virtual image block sequence in a front-to-back order to obtain a second virtual image block sequence, and selecting the preset number of real image blocks from the first real image block sequence to obtain a second real image block sequence; determining a root mean square error value between the second virtual image block sequence and the second real image block sequence as a root mean square error value between the virtual image and the second image; determining the preset pose data corresponding to the virtual image with the minimum root mean square error value as optimized pose data; updating the pose estimation data to the optimized pose data.

2. The camera positioning method of claim 1, wherein, The method comprises the following steps: determining a plurality of point sets, wherein each of the point sets contains a preset number of three-dimensional points and second tracking points corresponding to the three-dimensional points, wherein the three-dimensional points and the second tracking points corresponding to the three-dimensional points correspond to the same first tracking point; For each of the point sets, a candidate pose transformation relationship corresponding to the point set is determined according to the point set, wherein the candidate pose transformation relationship comprises an estimated change matrix between a pose of the camera when capturing the second image and a pose of the camera when capturing the first image, which is determined according to the point set; For each of the candidate pose transformation relationships, a projection point of the three-dimensional point corresponding to each of the first tracking points in the second image is determined according to the candidate pose transformation relationship, and a position error value between the projection point corresponding to each of the three-dimensional points and the second tracking point is determined; For each of the candidate pose transformation relationships, a number of the three-dimensional points corresponding to the candidate pose transformation relationship and having a position error value less than a preset value is determined, and the candidate pose transformation relationship corresponding to the largest number of the three-dimensional points having a position error value less than the preset value is determined as the target pose transformation relationship.

3. The camera positioning method of claim 2, wherein, Determining the projection point of the three-dimensional point corresponding to each of the first tracking points in the second image according to the candidate pose transformation relationship comprises: determining a first three-dimensional point coordinate of the three-dimensional point in a first camera coordinate system, and determining a second three-dimensional point coordinate of the three-dimensional point in a second camera coordinate system according to the candidate pose transformation relationship and the first three-dimensional point coordinate, wherein the first camera coordinate system is a camera coordinate system when the camera captures the first image, and the second camera coordinate system is a camera coordinate system when the camera captures the second image; determining a projection point coordinate of the projection point of each of the three-dimensional points in the second image in the preset plane rectangular coordinate system according to the second three-dimensional point coordinate of each of the three-dimensional points.

4. The camera positioning method of claim 3, wherein, Determining the first three-dimensional point coordinate of the three-dimensional point in the first camera coordinate system comprises: determining camera intrinsic parameters of the camera, wherein the camera intrinsic parameters comprise a focal length and a principal point position; determining a depth value of each of the first tracking points according to the depth map; determining a first three-dimensional point coordinate of the three-dimensional point corresponding to each of the first tracking points according to the depth value of each of the first tracking points, the camera intrinsic parameters and the first tracking point coordinate of each of the first tracking points.

5. The camera positioning method of claim 1, wherein, Determining the second tracking point corresponding to each of the first tracking points in the second image comprises: determining a second tracking point coordinate of the second tracking point in the preset plane rectangular coordinate system; determining a distance between each of the first tracking points and the second tracking point corresponding thereto according to the first tracking point coordinate of each of the first tracking points and the second tracking point coordinate of the second tracking point corresponding thereto; calculating an average distance between each of the first tracking points and the second tracking point corresponding thereto, and calculating a standard deviation of the distance between each of the first tracking points and the second tracking point corresponding thereto according to the average distance; calculating a distance threshold value according to the average distance and the standard deviation, and eliminating the first tracking point and the second tracking point corresponding thereto having a distance greater than the distance threshold value.

6. The camera positioning method of claim 1, wherein, Further comprising: The optimization pose data of the camera and the virtual camera is screened from the preset pose data according to the similarity between the virtual image and the second image, wherein the preset pose data is obtained by adding a random pose offset to the pose estimation data.

7. The camera positioning method of claim 6, wherein, Screening the optimization pose data of the camera and the virtual camera from the preset pose data according to the similarity between the virtual image and the second image comprises: determining a root mean square error value between the virtual image and the second image, wherein the root mean square error value is used to represent the similarity between the virtual image and the second image, and the root mean square error value and the similarity are negatively correlated.

8. The camera positioning method of claim 1, wherein, The specified pixel point in the image block is determined by: determining the saturation value and the brightness value of the pixel point in the image block; when the saturation value of the pixel point is less than a preset saturation value and the brightness value of the pixel point is greater than a preset brightness value, determining the pixel point as the specified pixel point.

9. A camera positioning device, characterized by comprises: a matching module, configured to determine a plurality of first tracking points in a first image, and determine a corresponding second tracking point of each first tracking point in the plurality of first tracking points in a second image, wherein the first image and the second image are physiological channel images continuously captured by a camera in a physiological channel; a first pose estimation module, configured to determine a depth map corresponding to the first image, and determine a three-dimensional point corresponding to each first tracking point according to the depth map and a first tracking point coordinate of each first tracking point in a preset plane rectangular coordinate system, wherein the depth map is used to indicate a depth value of each pixel point in the first image; The second pose estimation module is configured to determine a target pose conversion relationship of the pose of the camera when capturing the second image relative to the pose when capturing the first image according to the three-dimensional point corresponding to the first tracking point and the second tracking point, and determine pose estimation data of the camera when capturing the second image according to the target pose conversion relationship; acquire a virtual image of a physiological channel model captured by a virtual camera corresponding to the camera at each preset pose data indicated pose; divide the virtual image into a plurality of partially overlapped virtual image blocks, and divide the second image into a plurality of partially overlapped real image blocks; determine virtual feature values of the virtual image blocks, and real feature values of the real image blocks, wherein the virtual feature values and the real feature values each include a pixel point intensity value standard deviation in the image block, and a whiteness value for reflecting a specified pixel point proportion in the image block; delete the virtual image blocks and the real image blocks with a whiteness value greater than a preset whiteness value; arrange the retained virtual image blocks in descending order of pixel point intensity value standard deviation of each virtual image block to obtain a first virtual image block sequence, and arrange the retained real image blocks in descending order of pixel point intensity value standard deviation of each real image block to obtain a first real image block sequence; select a preset number of virtual image blocks from the first virtual image block sequence in the order from front to back to obtain a second virtual image block sequence, and select the preset number of real image blocks from the first real image block sequence to obtain a second real image block sequence; determine a root mean square error value between the second virtual image block sequence and the second real image block sequence as a root mean square error value between the virtual image and the second image; determine the preset pose data corresponding to the virtual image with the minimum root mean square error value as optimized pose data; and update the pose estimation data to the optimized pose data.

10. A camera positioning system, characterized by The camera positioning system comprises a camera, an electronic device, wherein The electronic device is configured to determine a plurality of first tracking points in a first image, and determine a corresponding second tracking point of each of the plurality of first tracking points in a second image, wherein the first image and the second image are physiological channel images successively captured by a camera in a physiological channel; determine a depth map corresponding to the first image, and determine a three-dimensional point corresponding to each of the first tracking points according to the depth map and a first tracking point coordinate of each of the first tracking points in a preset planar rectangular coordinate system, wherein the depth map is used to indicate a depth value of each pixel point in the first image; determine a target pose transformation relationship of a pose of the camera when capturing the second image relative to a pose when capturing the first image according to the three-dimensional point corresponding to the first tracking point and the second tracking point, and determine pose estimation data of the camera when capturing the second image according to the target pose transformation relationship; acquire a virtual image of a physiological channel model captured by a virtual camera corresponding to the camera at a pose indicated by each preset pose data; divide the virtual image into a plurality of partially overlapped virtual image blocks, and divide the second image into a plurality of partially overlapped real image blocks; determine a virtual feature value of the virtual image block and a real feature value of the real image block, wherein the virtual feature value and the real feature value each include a pixel point intensity value standard deviation in the image block and a whiteness value used to reflect a specified pixel point proportion in the image block; delete the virtual image block and the real image block with a whiteness value greater than a preset whiteness value; arrange the retained virtual image blocks in a descending order of pixel point intensity value standard deviation of each virtual image block to obtain a first virtual image block sequence, and arrange the retained real image blocks in a descending order of pixel point intensity value standard deviation of each real image block to obtain a first real image block sequence; select a preset number of virtual image blocks from the first virtual image block sequence in a front-to-back order to obtain a second virtual image block sequence, and select the preset number of real image blocks from the first real image block sequence to obtain a second real image block sequence; determine a root mean square error value between the second virtual image block sequence and the second real image block sequence as a root mean square error value between the virtual image and the second image; determine the preset pose data corresponding to the virtual image with the minimum root mean square error value as optimized pose data; and update the pose estimation data to the optimized pose data. 11.A non-volatile storage medium storing a computer program, the computer program being executed by a processor to implement the method of any one of claims 1 to 8. 12.A computer program product comprising a computer program, the computer program being executed by a processor to implement the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Endoscope positioning method, electronic equipment and non-transient computer readable storage medium

    CN117710279A

  • Guidance method based on 3D-2D pose estimation and 3D-CT registration with application to live bronchoscopy

    US20070015997A1