Monocular vision 6D pose measurement method and device for segmented varifocal scene
By using the object detection network and 3D sparse key point set in monocular vision 6D pose measurement, combined with the joint optimization method of perspective reprojection and unidirection reprojection residual function, the pose measurement accuracy and robustness problems in the zoom camera and complex environment are solved, and efficient pose measurement is achieved.
Patent Information
- Application Number
- CN202510628763.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
When dealing with zoom cameras and complex environments, existing monocular position measurement methods have problems such as insufficient accuracy, poor robustness and cumbersome calibration process.
A monocular visual 6D pose measurement method for segmented zoom distance scenes is proposed. By acquiring the image sequence, the object detection network and the 3D sparse key point set are used to perform pose estimation, and the perspective reprojection and unidirection reprojection residual functions are combined to obtain the camera internal parameters and pose initial values.
Improves the robustness and accuracy of position measurement, simplifies the camera internal parameter calibration process, and is suitable for zoom cameras and complex environments.
Smart Images

Figure CN120147429A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vision technology, and particularly to a monocular vision 6D pose measurement method and device for segmented variable focal length scenarios. Background Art
[0002] Real-time, robust, and high-precision inter-platform monocular pose measurement is one of the core key technologies in fields such as aircraft automatic landing, autonomous driving, high-precision navigation of robots, and augmented reality and virtual reality. Due to complex environmental interference and application conditions, monocular pose measurement remains a challenging task.
[0003] Existing monocular pose measurement methods can be divided into cooperative methods and non-cooperative methods according to whether cooperative markers are used. Cooperative methods measure the pose by extracting cooperative markers deployed on the target. Such methods are simple and efficient, but their application scope is relatively limited. Non-cooperative methods only use the characteristics of the target itself to achieve pose estimation, and have advantages such as simplicity, low cost, and wide application scenarios. Non-cooperative methods can be further divided into traditional methods and deep learning-related methods. Traditional methods are based on feature points, templates, etc. for pose estimation, and it is difficult to handle problems such as textureless objects and cluttered backgrounds. Benefiting from the powerful feature extraction and expression ability of neural networks, deep learning-based methods have achieved excellent performance in monocular pose estimation. Such methods can be further divided into direct methods and indirect methods. The direct method regresses the 6D pose end-to-end through the network, with a simple process, but the pose estimation accuracy is insufficient. The indirect method outputs the detected intermediate representation through the network, and combines the high-precision three-dimensional information of the target, and uses methods such as Iterative Closest Point (ICP) and Perspective-n-Point (PnP) to solve the pose.
[0004] Currently, visual pose measurement is mostly based on fixed focal length cameras, which have simple camera internal parameter calibration and convenient application. However, when a fixed focal length camera images a distant target, the short focal length results in low imaging resolution and large pose measurement errors. In contrast, variable focal length cameras can keep the target in high-resolution imaging all the time, improving the pose measurement accuracy, but their internal parameter calibration is complex. Some scholars use a look-up table to model the internal parameters of the zoom camera. Due to the lack of consideration of the interaction between parameters, the calibration accuracy is poor, the process is cumbersome, and each camera needs to be processed one by one. Self-calibration methods based on absolute dual quadric surfaces, etc., determine the camera internal parameters by means of multiple uncalibrated images. Although flexible, their accuracy and robustness are insufficient, and they mostly focus on focal length calibration, with high requirements for the accuracy of the principal point. The active vision calibration method requires the camera to perform specific movements, which limits the application scope. In addition, some studies expand the PnP problem and solve the camera focal length and distortion coefficient in a single frame, but such methods assume that the camera principal point is at the center of the image, which does not conform to the actual situation. Summary of the Invention
[0005] Based on this, it is necessary to provide a monocular vision 6D pose measurement method and device for segmented variable focal length scenarios that can improve robustness and accuracy in view of the above technical problems.
[0006] A monocular vision 6D pose measurement method for segmented variable focal length scenarios, the method is applied to a non-cooperative system, the system includes a tracking target and a tracking party that perform relative motion, and the method is implemented in the tracking party: Obtain an image sequence set of the tracking target in a segmented variable focal length scenario, where the image sequence set includes multiple target images sorted by time; Use a target detection network to identify the region where the target is located in each frame of the target image to obtain a target bounding box, and use a preset set of 3D sparse key points of the target to perform key point identification within the target bounding box to obtain the target key points in each frame of the image; For each frame of the target image, if it is the first frame of the target image, obtain the camera internal parameters and the initial pose value of the first frame based on the initial value of the optical center. If it is a non-first frame of the target image, obtain the initial pose value of the current frame according to the optimized camera internal parameters of the previous frame of the image; Use the set of 3D sparse key points of the target and planar feature points to perform reprojection respectively to obtain 3D key point reprojection and planar feature point reprojection. For each frame of the target image, use the corresponding target key points and the 3D key point reprojection to construct a perspective reprojection residual function. At the same time, use the corresponding planar feature points in the current frame of the target image and the previous frame of the target image to construct a homography reprojection residual function; Jointly optimize the camera internal parameters and the initial pose value of the current frame by minimizing the perspective reprojection residual function and the homography reprojection residual function to obtain the measurement value of the 6D pose in each frame of the target image.
[0007] In one embodiment, when constructing the set of 3D sparse key points of the target: Generate a target three-dimensional dense model based on the target entity; Use a human pose estimation method to extract multiple target key points with significant structural features on the target three-dimensional dense model; Generate the set of 3D sparse key points of the target according to the multiple target key points.
[0008] In one embodiment, when using a target detection network to identify the region where the target is located in each frame of the target image to obtain a target bounding box: Use the target detection network to generate a first-frame target bounding box according to the first-frame target image; Based on the first-frame target bounding box, use the STAPLE algorithm to detect the target bounding box in the subsequent frames of the target image.
[0009] In one embodiment, the object detection network employs the YOLOX neural network.
[0010] In one embodiment, when acquiring each frame of target image in the image sequence set: The imaging size of the tracking target in each frame of target image is monitored in real time. When it is detected that the imaging size exceeds the preset threshold, the focal length of the imaging system is adjusted.
[0011] In one embodiment, the perspective reprojection residual function is expressed as: ; In the above formula, represents the two-dimensional projection of the detected th predefined three-dimensional key point on the th image, represents the camera internal parameters to be optimized, represents the pose to be optimized of the th frame of target image, represents the homogeneous coordinates of the th predefined three-dimensional key point, and respectively represent the number of selected images and key points.
[0012] In one embodiment, the homography reprojection residual function is constructed by reprojecting the planar feature points of the previous frame of target image to the current frame of target image through a homography relationship.
[0013] In one embodiment, the homography reprojection residual function is expressed as: ; In the above formula, represents the number of feature point matches between the th frame and the th frame of target image, and respectively represent the th feature point on the th frame and the th frame of target image, and respectively represent the camera internal parameters of the th frame and the th frame of target image, and respectively represent the rotation values of the poses of the th frame and the th frame of target image, and respectively represent the th frame and the The translation component of the frame target image, and respectively represent the plane normal vector and the plane equation intercept in the object coordinate system.
[0014] In one embodiment, when jointly optimizing the camera internal parameters and the initial pose values of the current frame by using the perspective reprojection residual function and the homography reprojection residual function: In the homography reprojection residual function, if the camera internal parameters of the current frame and the target image of the previous frame are the same, then fix the camera internal parameters in the homography reprojection residual function and only optimize the pose; If the camera internal parameters of the current frame and the target image of the previous frame are different, then fix the camera internal parameters of the target image of the previous frame in the homography reprojection residual function and only optimize the camera internal parameters and the pose of the target image of the current frame.
[0015] The present application also provides a monocular vision 6D pose measurement device for a segmented variable focal length scene. The device includes: An image sequence set acquisition module, configured to acquire an image sequence set of a tracking target in a segmented variable focal length scene, where the image sequence set includes multiple frames of target images sorted by time; A target key point extraction module, configured to use a target detection network to identify the area where the target is located in each frame of the target image to obtain a target bounding box, and perform key point identification within the target bounding box by using a preset target 3D sparse key point set to obtain the target key points in each frame of the image; A camera internal parameter and pose initial value obtaining module, configured to, for each frame of the target image, if it is the first frame of the target image, obtain the camera internal parameters and the initial pose value of the first frame based on the initial value of the optical center, and if it is a non-first frame of the target image, obtain the current frame pose initial value according to the optimized camera internal parameters of the previous frame of the image; A joint optimization module, configured to respectively perform reprojection by using the target 3D sparse key point set and the plane feature points to obtain the three-dimensional key point reprojection and the plane feature point reprojection. For each frame of the target image, use the corresponding target key points and the three-dimensional key point reprojection to construct a perspective reprojection residual function, and at the same time, use the corresponding plane feature points in the current frame of the target image and the previous frame of the target image to construct a homography reprojection residual function; A pose measurement value obtaining module, configured to jointly optimize the camera internal parameters and the pose initial values of the current frame by minimizing the perspective reprojection residual function and the homography reprojection residual function to obtain the measurement values of the 6D pose in each frame of the target image.
[0016] A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented: Obtain an image sequence set of the tracking target in a segmented variable focal length scenario, where the image sequence set includes multiple frames of target images sorted by time; Use a target detection network to identify the region where the target is located in each frame of the target image, obtain a target bounding box, and perform key point identification within the target bounding box using a preset set of 3D sparse key points of the target to obtain the target key points in each frame of the image; For each frame of the target image, if it is the first frame of the target image, obtain the camera internal parameters and the initial pose of the first frame based on the initial value of the optical center. If it is a non-first frame of the target image, obtain the initial pose of the current frame according to the optimized camera internal parameters of the previous frame; Use the set of 3D sparse key points of the target and planar feature points to perform reprojection respectively to obtain the 3D key point reprojection and the planar feature point reprojection. For each frame of the target image, use the corresponding target key points and the 3D key point reprojection to construct a perspective reprojection residual function. At the same time, use the corresponding planar feature points in the current frame of the target image and the previous frame of the target image to construct a homography reprojection residual function; Jointly optimize the camera internal parameters and the initial pose of the current frame by minimizing the perspective reprojection residual function and the homography reprojection residual function to obtain the measurement values of the 6D pose in each frame of the target image.
[0017] A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: Obtain an image sequence set of the tracking target in a segmented variable focal length scenario, where the image sequence set includes multiple frames of target images sorted by time; Use a target detection network to identify the region where the target is located in each frame of the target image, obtain a target bounding box, and perform key point identification within the target bounding box using a preset set of 3D sparse key points of the target to obtain the target key points in each frame of the image; For each frame of the target image, if it is the first frame of the target image, obtain the camera internal parameters and the initial pose of the first frame based on the initial value of the optical center. If it is a non-first frame of the target image, obtain the initial pose of the current frame according to the optimized camera internal parameters of the previous frame; Use the set of 3D sparse key points of the target and planar feature points to perform reprojection respectively to obtain the 3D key point reprojection and the planar feature point reprojection. For each frame of the target image, use the corresponding target key points and the 3D key point reprojection to construct a perspective reprojection residual function. At the same time, use the corresponding planar feature points in the current frame of the target image and the previous frame of the target image to construct a homography reprojection residual function; Jointly optimize the camera internal parameters and the initial pose of the current frame by minimizing the perspective reprojection residual function and the homography reprojection residual function to obtain the measurement values of the 6D pose in each frame of the target image.
[0018] The above monocular vision 6D pose measurement method and device for the segmented variable focal length scenario first use a target detection network to identify the target area in each frame of the target image to obtain the target bounding box, and use a preset set of 3D sparse key points of the target to identify key points within the target bounding box to obtain the target key points in each frame of the image. For each frame of the target image, the initial pose value of the current frame is obtained according to the optimized camera internal parameters of the previous frame of the image. The 3D sparse key point set of the target is used for perspective reprojection on the image to obtain the 3D key point reprojection. For each frame of the target image, a perspective reprojection residual function is constructed using the corresponding target key points and the 3D key point reprojection. At the same time, a homography reprojection residual function is constructed using the matching feature points in the current frame of the target image and the previous frame of the target image. Finally, the camera internal parameters and the initial pose value of the current frame are jointly optimized by minimizing the perspective reprojection residual function and the homography reprojection residual function to obtain the measurement values of the 6D pose in each frame of the target image. Using this method can effectively improve the robustness and pose measurement accuracy. Description of the Drawings
[0019] Figure 1 It is a schematic flowchart of a monocular vision 6D pose measurement method for the segmented variable focal length scenario in an embodiment; Figure 2 It is a schematic diagram of a target 3D sparse model in an embodiment, where Figure 2 (a) represents the actual target model, Figure 2 (b) represents the 3D model obtained by scanning the target in the left figure; Figure 3 It is a schematic diagram of the relative pose measurement of the double-actuating platform rendezvous vision guidance in an embodiment; Figure 4 It is a schematic framework diagram of a monocular vision 6D pose measurement method for the segmented variable focal length scenario in an embodiment; Figure 5 It is a schematic flowchart of a monocular vision 6D pose measurement method for the segmented variable focal length scenario in another embodiment; Figure 6 It is a structural block diagram of a monocular vision 6D pose measurement device for the segmented variable focal length scenario in an embodiment; Figure 7 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments
[0020] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0021] There are many problems in existing monocular pose measurement and visual pose measurement technologies. For example, in monocular pose measurement, the application scope of cooperative methods is limited, non-cooperative traditional methods are difficult to handle scenes with no texture and cluttered backgrounds, the accuracy of deep learning direct methods is poor, and indirect methods rely on specific conditions. In the use of cameras for visual pose measurement, fixed-focus cameras have low imaging resolution and large measurement errors for distant targets, and the internal parameter calibration of zoom cameras is complex. Existing methods such as lookup table modeling, self-calibration, active vision calibration, and extended PnP problem solving have problems such as low calibration accuracy, poor robustness, limited application, and non-conformity to reality respectively.
[0022] To address the above problems, in this application, as Figure 1 shown, a monocular vision 6D pose measurement method for segmented variable-focus scenes is provided. This method is applied to a non-cooperative system, which includes a tracking target and a tracking party that perform relative motion. The method is implemented in the tracking party and specifically includes the following steps: Step S100, obtain an image sequence set of the tracking target in a segmented variable-focus scene, where the image sequence set includes multiple frames of target images sorted by time.
[0023] Step S110, use a target detection network to identify the region where the target is located in each frame of the target image to obtain a target bounding box, and use a preset set of 3D sparse key points of the target to perform key point identification within the target bounding box to obtain the target key points in each frame of the image.
[0024] Step S120, for each frame of the target image, if it is the first frame of the target image, obtain the camera internal parameters and the initial pose value of the first frame based on the initial value of the optical center. If it is a non-first frame of the target image, obtain the initial pose value of the current frame according to the optimized camera internal parameters of the previous frame.
[0025] Step S130, use the set of 3D sparse key points of the target and planar feature points to perform reprojection respectively to obtain the three-dimensional key point reprojection and the planar feature point reprojection. For each frame of the target image, use the corresponding target key points and the three-dimensional key point reprojection to construct a perspective reprojection residual function. At the same time, use the corresponding planar feature points in the current frame of the target image and the previous frame of the target image to construct a homography reprojection residual function.
[0026] Step S140, jointly optimize the camera internal parameters and the initial pose value of the current frame by minimizing the perspective reprojection residual function and the homography reprojection residual function to obtain the measurement values of the 6D pose in each frame of the target image.
[0027] In this embodiment, the non - cooperative system includes a tracking target and a tracking party, and both of them are in motion. This system can refer to a system with a ship as the tracking target and a drone as the tracking party, and its task is for the drone to land accurately at a preset position based on the relative pose measured. It can also be in a city, with a moving vehicle as the tracking target and a drone as the tracking party, and the task is also for the drone to land on the top of the vehicle. This method is implemented on the tracking party. The tracking target is photographed by a camera system set on the tracking party, and the relative pose is measured through the obtained images. In this article, a ship is taken as the tracking target and a drone as the tracking party as an example for illustration.
[0028] In step S100, first, the tracking target is continuously photographed by a camera set on the drone, so as to obtain multiple target images sorted by time. Among them, when each frame of the target image is photographed, the focal length of the camera changes, so that when the tracking target moves in the tracking direction, its target always maintains a high resolution on the image.
[0029] Specifically, when obtaining each frame of the target image in the image sequence set: the imaging size of the tracking target in each frame of the target image is monitored in real time. When it is detected that the imaging size exceeds the preset threshold, the focal length of the imaging system is adjusted. This makes the target imaging stable within the optimal size range, thus ensuring the accuracy and stability of visual measurement.
[0030] In step S110, when constructing the target 3D sparse key - point set: first, a target three - dimensional dense model is generated based on the target entity, and a human - pose estimation method is used to extract multiple target key points with significant structural features on the target three - dimensional dense model. Finally, a target 3D sparse key - point set is generated according to the multiple target key points.
[0031] Considering that the method for measuring the target pose based on the three - dimensional dense model is usually slow and it is difficult to directly apply it to subsequent optimization steps. In this embodiment, referring to the idea of setting key points of human body parts in the human - pose estimation method, first, a reasonable object three - dimensional coordinate system is established in the target three - dimensional dense model , and a set of sparse 3D key - point sets with significant structural features is proposed to represent the 3D sparse model of the target, as Figure 2 shown. Use to represent the predefined 3D sparse key - point set, which is expressed as: (1) In formula (1), represents the number of three - dimensional key points used. In order to simulate the actual scene, the 3D sparse key - point coordinates of the target are scaled to the actual size.
[0032] In step S110, when using the target detection network to identify the target area in each frame of the target image and obtain the target bounding box: The target detection network generates the first-frame target bounding box based on the first-frame target image, and then, based on the first-frame target bounding box, the STAPLE algorithm is used to track the target bounding boxes in the subsequent frames of the target image.
[0033] In this embodiment, the target detection network uses the YOLOX neural network.
[0034] Specifically, after selecting the predefined 3D key points the YOLOX algorithm is used to detect the target area. YOLOX outputs the bounding box of the target in the image, denoted as . Since YOLOX takes a long time for target detection on a single-frame image, in order to improve efficiency, after obtaining the target area in the first frame, the STAPLE algorithm is used to track the target in the subsequent input images. To keep the symbols unified, the target bounding box output by the STAPLE algorithm is also denoted as .
[0035] In this embodiment, to ensure the reliability of target tracking, the intersection ratio between the target detection box and the pose reprojection box is further calculated, expressed as: (2) In formula (2), if , it is considered that , and at this time the target tracking result is unreliable, and YOLOX is used to detect the target again.
[0036] In one embodiment, when using the preset target 3D sparse key point set to perform key point recognition within the target bounding box, the RTMPose algorithm is used to achieve efficient and accurate detection of 2D key points in the input image. The RTMPose algorithm performs key point detection as a classification task and has the characteristics of simplicity, high efficiency, and easy deployment. In fact, in this method, other various efficient and accurate key point detection methods can also be used, not limited to RTMPose.
[0037] Specifically, after the key point detection in the target image, the 3D-2D key point matching can actually be obtained, which can be applied to the subsequent pose initial value solution.
[0038] In step S120, when obtaining the camera internal parameters and the initial pose values of each frame of the target image, considering that the zoom camera is usually set to a relatively small focal length before working, and the internal parameters of the camera at this focal length Zhang Zhengyou calibration method can be used for precise calibration in advance. When the system starts to work, the camera first takes an image at this focal length and calculates the current pose based on solving the PnP problem . At this time, the coordinates of the camera optical center in the object coordinate system are . Subsequently, the zoom controller is driven to zoom in and a high-resolution image of the target is taken.
[0039] Furthermore, for the first-frame target image, the calculated camera optical center is substituted into the following formula and further transformed to obtain: (3) In formula (3), is the depth of the 3D key point, is in homogeneous form, and are the initial values of the camera internal parameters and the rotation amount of the first-frame target image after zooming to be solved. is called the infinite homography. Furthermore, formula (3) is written in the form of a homogeneous linear equation system as: (4) In formula (4), is the column vector formed by arranging by rows. The elements of the infinite homography are obtained based on the SVD decomposition method under scale uncertainty. Then, the RQ decomposition is used to decompose into and , and the initial translation value .
[0040] Furthermore, for non-first-frame images, the optimized camera internal parameters of the previous frame are used as the initial values of the current frame internal parameters, and the initial pose values of the current frame are obtained by solving the Perspective-n-Point (PnP) problem.
[0041] Next, in steps S130 and S140, the initial values of the camera internal parameters and the pose of the current frame image are optimized.
[0042] In this embodiment, a feature point detection algorithm is used to match the feature points in the adjacent image plane regions, and the perspective transformation reprojection relationship of the 3D key points and the homography transformation reprojection relationship between the plane feature points are established through the camera internal parameters and the pose. Using the multi-view geometric constraint information in the sequence images, the focal length, principal point, and relative pose that constitute the camera internal parameters are used as the parameters to be optimized, and an objective function is established to minimize the 3D key point reprojection residual and the inter-frame homography transformation reprojection residual. By solving this optimization problem, the joint estimation of the camera internal parameters and the relative 6D pose is realized.
[0043] As Figure 3 shown, since the pose relationship between the monocular zoom camera and the tracking platform can be pre-calibrated and remain fixed, therefore, in this application, the pose relationship between the dual moving platforms can be substantially equivalent to the relative pose relationship between the camera and the target platform. Figure 3 In is the camera coordinate system, and its origin is located at the camera optical center. is the object coordinate system, and its coordinate origin is located on the target platform.
[0044] As Figure 3 shown, in this embodiment, use to represent the rigid body transformation from the object coordinate system to the camera coordinate system, that is, the relative pose between the two coordinate systems, namely: (5) In formula (5), and respectively represent the special Euclidean group and the special orthogonal group. and are respectively 's rotation and translation components.
[0045] Furthermore, the internal parameter matrix of the zoom camera is expressed as: (6) In formula (6), and are respectively and 's equivalent focal lengths in the directions, and
[0046] is the camera principal point coordinate. Since the three-dimensional model of the tracking target is known, a vertex of the tracking target is represented in homogeneous form (7) In formula (7), is the two-dimensional image point, .
[0047] Furthermore, the homography matrix describes the mapping relationship between two planes. If the feature points in the scene all fall on the same plane, then motion estimation can be performed through homography. Assume that the target plane in the object coordinate system is , and its plane equation is: (8) In formula (8), is the plane normal vector. Let a pair of adjacent frame images captured by the zoom camera be and , and there is a pair of matching feature points and on the plane. Based on the camera internal parameters and the 6D pose representation , the homography reprojection relationship can be further established: (9) In this embodiment, when performing joint optimization, a sliding window is constructed by backtracking frames from the current frame. The 3D key points are reprojected and compared with the 2D key point detection values through the camera internal parameters and the initial 6D pose values of the images within the sliding window. The established perspective reprojection residual function is expressed as: (10) In formula (10), represents the 2D projection of the detected th predefined 3D key point on the th image, represents the camera internal parameters to be optimized, represents the pose to be optimized of the target image in the th frame, represents the homogeneous coordinates of the th predefined 3D key point, , respectively represent the number of selected images and key points.
[0048] Since the residual function represented by formula (10) is constructed only based on the relative pose between the object system and the camera system and does not consider the relative pose constraints between images. Therefore, in this embodiment, the target plane area in the image is located and ORB (Oriented FAST and Rotated BRIEF) feature points are extracted on the plane area, and further the ORB feature points between adjacent frame images are matched.
[0049] In this embodiment, the homography reprojection residual function is constructed by reprojecting the plane feature points of the previous frame target image to the current frame target image through the homography relationship.
[0050] Specifically, within the above constructed sliding window, the plane feature points of the th frame image are reprojected to the th frame image through the homography relationship shown in formula (9). The established homography reprojection residual function is expressed as: (11) In formula (11), Indicates the number of feature point matches between the frame and the target image of the -th frame. and respectively represent the -th feature point on the target images of the frame and the -th frame. and respectively represent the internal camera parameters of the target images of the frame and the -th frame. and respectively represent the rotation values of the poses of the target images of the frame and the -th frame. and respectively represent the translation components of the target images of the frame and the -th frame. and
[0051] respectively represent the plane normal vector and the plane equation intercept in the object coordinate system. (12) In Equation (12), only the relative pose between frames is constrained, which is an underdetermined equation. There is redundancy in the degrees of freedom in the optimization problem regarding the internal camera parameters and the absolute pose, resulting in an infinite number of solution spaces. To limit the randomness of the solution, when jointly optimizing the internal camera parameters and the initial pose values of the current frame using the perspective reprojection residual function and the homography reprojection residual function: If in the homography reprojection residual function, the internal camera parameters of the target images of the current frame and the previous frame are the same, then the internal camera parameters in the homography reprojection residual function are fixed, and only the pose is optimized, that is, if , then in is fixed, and only the pose is optimized. If the internal camera parameters of the target images of the current frame and the previous frame are different, then the internal camera parameters of the target image of the previous frame in the homography reprojection residual function are fixed, and the internal camera parameters and the pose of the target image of the current frame are optimized, that is, if , then in is fixed, and only and the 6D pose are optimized.
[0052] For exampleFigure 4 and Figure 5 As shown in Figure 5 , it is the implementation method framework diagram and algorithm flow chart of this method.
[0053] In the above monocular vision 6D pose measurement method for segmented variable focal length scenarios, a target detection model is used to locate the imaging position of the target platform. Then, a key point detection algorithm commonly used in human pose estimation is used to detect the key points of the target platform. Secondly, the infinite homography matrix is solved through the estimated optical center value and decomposed to obtain the initial camera internal parameters and pose initial values of the initial frame. Using the multi-view geometric constraint information in the sequence images, the focal length, principal point, and relative pose that constitute the camera internal parameters are used as parameters to be optimized. An optimization objective function is established by minimizing the perspective reprojection error of three-dimensional key points and the homography transformation reprojection residual of inter-frame feature points. By solving this optimization problem, the joint estimation of camera internal parameters and relative 6D poses is achieved. This method realizes robust, efficient, and high-precision relative pose measurement between platforms in monocular vision guidance. At the same time, the camera internal parameters and 6D poses can be optimized in real time. In this method, the absolute constraint based on the perspective reprojection residual of three-dimensional key points and the relative constraint based on the homography reprojection residual of two-dimensional feature points are combined to establish an objective function for minimizing the reprojection residual, and the camera internal parameters and 6D poses are jointly optimized.
[0054] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0055] In one embodiment, as Figure 6 shown in Figure 6 , a monocular vision 6D pose measurement device for segmented variable focal length scenarios is provided, including: an image sequence set acquisition module 200, a target key point extraction module 210, a camera internal parameter and pose initial value obtaining module 220, a joint optimization module 230, and a pose measurement value obtaining module 240, where: The image sequence set acquisition module 200 is configured to acquire an image sequence set of the tracking target in a segmented variable focal length scenario, and the image sequence set includes multiple frames of target images sorted by time; The target key point extraction module 210 is used to identify the target area in each frame of the target image by using a target detection network, obtain the target bounding box, and perform key point identification within the target bounding box by using a preset 3D sparse key point set of the target, so as to obtain the target key points in each frame of the image; The camera internal parameter and pose initial value obtaining module 220 is used to, for each frame of the target image, if it is the first frame of the target image, obtain the camera internal parameter and the first frame pose initial value based on the initial value of the optical center, and if it is a non-first frame of the target image, obtain the current frame pose initial value according to the optimized camera internal parameter of the previous frame of the image; The joint optimization module 230 is used to respectively perform reprojection by using the 3D sparse key point set of the target and the planar feature points to obtain the 3D key point reprojection and the planar feature point reprojection. For each frame of the target image, a perspective reprojection residual function is constructed by using the corresponding target key points and the 3D key point reprojection. At the same time, a homography reprojection residual function is constructed by using the corresponding planar feature points in the current frame of the target image and the previous frame of the target image; The pose measurement value obtaining module 240 is used to jointly optimize the camera internal parameter and the pose initial value of the current frame by minimizing the perspective reprojection residual function and the homography reprojection residual function, so as to obtain the measurement value of the 6D pose in each frame of the target image.
[0056] For the specific limitations of the monocular vision 6D pose measurement device for the segmented variable focal length scene, reference can be made to the limitations of the monocular vision 6D pose measurement method for the segmented variable focal length scene in the above text, which will not be elaborated here. Each module in the above monocular vision 6D pose measurement device for the segmented variable focal length scene can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in the form of hardware or be independent of the processor, or can be stored in the memory in the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above modules.
[0057] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 7As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a monocular vision 6D pose measurement method for a segmented variable focal length scenario. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0058] Those skilled in the art can understand that Figure 7 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0059] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented: Obtain an image sequence set of the tracking target in a segmented variable focal length scenario. The image sequence set includes multiple frames of target images sorted by time; Use a target detection network to identify the region where the target is located in each frame of the target image to obtain a target bounding box, and use a preset set of 3D sparse key points of the target to perform key point recognition within the target bounding box to obtain the target key points in each frame of the image; For each frame of the target image, if it is the first frame of the target image, obtain the camera internal parameters and the initial pose value of the first frame based on the initial value of the optical center. If it is a non-first frame of the target image, obtain the initial pose value of the current frame according to the optimized camera internal parameters of the previous frame of the image; Use the set of 3D sparse key points of the target and planar feature points to perform reprojection respectively to obtain 3D key point reprojection and planar feature point reprojection. For each frame of the target image, use the corresponding target key points and the 3D key point reprojection to construct a perspective reprojection residual function. At the same time, use the corresponding planar feature points in the current frame of the target image and the previous frame of the target image to construct a homography reprojection residual function; Jointly optimize the camera internal parameters and the initial pose values of the current frame by minimizing the perspective reprojection residual function and the homography reprojection residual function to obtain the measurement values of the 6D pose in each frame of the target image.
[0060] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: Obtain an image sequence set of the tracking target in a segmented variable focal length scenario, where the image sequence set includes multiple frames of target images sorted by time; Use a target detection network to identify the region where the target is located in each frame of the target image to obtain a target bounding box, and use a preset set of 3D sparse key points of the target to perform key point identification within the target bounding box to obtain the target key points in each frame of the image; For each frame of the target image, if it is the first frame of the target image, obtain the camera internal parameters and the initial pose value of the first frame based on the initial optical center value. If it is a non-first frame of the target image, obtain the current frame pose initial value according to the optimized camera internal parameters of the previous frame image; Use the set of 3D sparse key points of the target and the planar feature points to perform reprojection respectively to obtain the 3D key point reprojection and the planar feature point reprojection. For each frame of the target image, use the corresponding target key points and the 3D key point reprojection to construct a perspective reprojection residual function. At the same time, use the corresponding planar feature points in the current frame target image and the previous frame target image to construct a homography reprojection residual function; Jointly optimize the camera internal parameters and the pose initial value of the current frame by minimizing the perspective reprojection residual function and the homography reprojection residual function to obtain the measurement values of the 6D pose in each frame of the target image.
[0061] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0062] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0063] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A monocular vision 6D pose measurement method for segmented variable focal length scenes, characterized in that: The method is applied to a non-cooperative system, the system comprising a tracking target performing relative motion and a tracking party, and the method is implemented in the tracking party: Acquire an image sequence set of the tracking target in a segmented variable focal length scenario, wherein the image sequence set includes multiple frames of target images sorted in time; The target detection network is used to identify the target area in each frame of the target image to obtain a target bounding box, and a preset target 3D sparse key point set is used to perform key point recognition in the target bounding box to obtain the target key points in each frame of the image; For each frame target image, if it is the first frame target image, the camera intrinsic parameters and the first frame pose initial value are obtained based on the initial value of the optical center. If it is not the first frame target image, the current frame pose initial value is obtained based on the camera intrinsic parameters optimized for the previous frame image. Using the target 3D sparse key point set and the plane feature points, reprojection is performed to obtain three-dimensional key point reprojection and plane feature point reprojection respectively. For each frame of the target image, a perspective reprojection residual function is constructed using the corresponding target key points and the three-dimensional key point reprojection. At the same time, a homography reprojection residual function is constructed using the corresponding plane feature points in the current frame target image and the previous frame target image; The camera intrinsic parameters and initial pose values of the current frame are jointly optimized by minimizing the perspective reprojection residual function and the homography reprojection residual function to obtain the measurement value of the 6D pose in each frame target image.
2. The monocular vision 6D pose measurement method for segmented variable focal length scenes according to claim 1, characterized in that: When constructing the target 3D sparse key point set: Generate a target three-dimensional dense model based on the target entity; A human body posture estimation method is used to extract a plurality of target key points with significant structural features on the target three-dimensional dense model; The target 3D sparse key point set is generated according to the multiple target key points.
3. The monocular vision 6D pose measurement method for segmented variable focal length scenes according to claim 2, characterized in that: When using the target detection network to identify the target area in each frame of the target image and obtain the target bounding box: Generate a first-frame target bounding box based on the first-frame target image using the target detection network; Based on the first frame target bounding box, the STAPLE algorithm is used to detect the target bounding boxes in the subsequent frame target images.
4. The monocular vision 6D pose measurement method for segmented variable focal length scenes according to claim 3, characterized in that: The target detection network adopts the YOLOX neural network.
5. The monocular vision 6D pose measurement method for segmented variable focal length scenes according to any one of claims 1 to 4, characterized in that: When acquiring each frame of the target image in the image sequence set: The imaging size of the tracked target in each frame of the target image is monitored in real time, and when it is detected that the imaging size exceeds a preset threshold, the focal length of the imaging system is adjusted.
6. The monocular vision 6D pose measurement method for segmented variable focal length scenes according to claim 4, characterized in that: The perspective reprojection residual function is expressed as: In the above formula, Indicates the detected The predefined 3D key points are The two-dimensional projection on the image, Indicates the camera internal parameters to be optimized. Indicates The pose to be optimized of the frame target image, Indicates homogeneous coordinates of predefined 3D key points, , Respectively represent the number of selected images and key points.
7. The monocular vision 6D pose measurement method for segmented variable focal length scenes according to claim 5, characterized in that: The homography reprojection residual function is constructed by reprojecting the planar feature points of the previous frame target image onto the current frame target image through the homography relationship.
8. The monocular vision 6D pose measurement method for segmented variable focal length scenes according to claim 6, characterized in that: The homography reprojection residual function is expressed as: In the above formula, Indicates Frame and The number of feature point matches between frame target images, , Respectively represent Frame and Frame target image feature points, , Respectively represent Frame and The camera intrinsic parameters of the frame target image, , Respectively represent Frame and The rotation value of the frame target image pose, , Respectively represent Frame and The translation component of the frame target image, and They represent the plane normal vector and plane equation intercept in the object coordinate system respectively.
9. The monocular vision 6D pose measurement method for segmented variable focal length scenes according to claim 7, characterized in that: When the perspective reprojection residual function and the homography reprojection residual function are used to jointly optimize the camera intrinsic parameters and the initial pose value of the current frame: If in the homography reprojection residual function, if the camera intrinsic parameters of the current frame and the previous frame target image are the same, the camera intrinsic parameters in the homography reprojection residual function are fixed, and only the posture is optimized; If the camera intrinsic parameters of the target image of the current frame are different from those of the previous frame, the camera intrinsic parameters of the target image of the previous frame in the homography reprojection residual function are fixed, and only the camera intrinsic parameters and posture of the target image of the current frame are optimized.
10. A monocular vision 6D pose measurement device for segmented variable focal length scenes, characterized in that: The device comprises: An image sequence set acquisition module is used to acquire an image sequence set of a tracking target in a segmented variable focal length scenario, wherein the image sequence set includes multiple frames of target images sorted in time; The target key point extraction module is used to use the target detection network to identify the target area in each frame target image to obtain the target bounding box, and use the preset target 3D sparse key point set to perform key point recognition in the target bounding box to obtain the target key points in each frame image; The module for obtaining the camera intrinsic parameters and initial pose values is used to obtain the camera intrinsic parameters and the initial pose values of the first frame based on the initial value of the optical center for each frame target image. If it is a non-first frame target image, the initial pose value of the current frame is obtained based on the optimized camera intrinsic parameters of the previous frame image. A joint optimization module is used to use the target 3D sparse key point set and the plane feature points to reproject respectively to obtain three-dimensional key point reprojection and plane feature point reprojection, and for each frame of the target image, a perspective reprojection residual function is constructed using the corresponding target key points and the three-dimensional key point reprojection, and at the same time, a homography reprojection residual function is constructed using the corresponding plane feature points in the current frame target image and the previous frame target image; The pose measurement value obtaining module is used to jointly optimize the camera intrinsic parameters and the pose initial value of the current frame by minimizing the perspective reprojection residual function and the homography reprojection residual function to obtain the measurement value of the 6D pose in each frame target image.
Citation Information
Patent Citations
Monocular vision SLAM (Simultaneous Localization and Mapping) method for dynamic environment
CN115471748A
Visual positioning method and system based on single-purpose indoor office scene
CN117522971A
Double-acting platform interaction method and system based on non-cooperative mode
CN117830906A
Panorama camera external parameter verification method and system and vehicle
CN118196210A
Power transmission line unmanned aerial vehicle point cloud map construction method and device, computer equipment and readable storage medium
CN118864729A