Monocular vision-based flight target orbit reconstruction method and system
By initializing the target and camera positions in the three-dimensional engine and determining the rotation matrix and position vector through multiple iterative rotation sampling, the problems of high equipment dependence and complexity of orbit reconstruction in the prior art are solved, and a highly authentic monocular visual flight target orbit reconstruction is achieved.
Patent Information
- Application Number
- CN202510179105.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The prior art relies on multi-camera calibration in the reconstruction of flight target tracks, resulting in high equipment dependence and increased maintenance costs, and recalibration is required when the camera moves or vibrates, affecting the splicing effect.
A method of orbit reconstruction of flight targets based on monocular vision is proposed. By initializing the target and camera positions in a three-dimensional engine, the three degrees of freedom of the split target are rotated to 2+1, and the rotation matrix and position vector are determined through multiple iterative rotation sampling and comparison with the original image.
Reduces device dependence, reduces sampling complexity, improves the authenticity and computing efficiency of track reconstruction, and simplifies the target position calculation process.
Smart Images

Figure CN120107475A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of aerospace technology, and further relates to a method and system for reconstructing a flight target trajectory, which can be used for three-dimensional engine scene simulation. Through pre-experimental shooting data, the flight trajectory of the target is restored in the simulation scene for simulation. Background Art
[0002] In the process of reconstructing the trajectory of a flying target, the core task is model matching, which aims to infer the rotation matrix and position vector of the target in three-dimensional space by comparing the similarity between the existing three-dimensional model of the flying target and the image data obtained from the monocular camera. This process usually involves extracting feature points from the image and matching these feature points to determine the relative position and posture between the camera and the target. Since the monocular vision system can only provide two-dimensional image information, how to effectively restore the three-dimensional trajectory of the target from these two-dimensional data is the main challenge faced by this field.
[0003] Model matching is not only widely used in the aerospace field, but also shows important value in many other fields. In the field of unmanned driving, by using monocular vision for target tracking and trajectory prediction, real-time perception and decision support of the surrounding environment can be achieved, improving the safety and reliability of the autonomous driving system. In addition, in robotics technology, monocular vision is widely used in navigation and obstacle avoidance, helping robots to locate and plan paths in real time through accurate track reconstruction. In addition, in other high-precision positioning scenarios such as object detection and recognition, drone monitoring, etc., monocular vision track reconstruction technology also shows great application potential.
[0004] Existing technical solutions for calculating rotation matrices use a single round of rough sampling plus local interpolation or multiple rounds of progressive optimization. Both methods extract images of the target at multiple angles and match them with the target image. The former is a rough estimate through a single round of sampling plus local interpolation, and the latter is calculated through multiple rounds of sampling and gradual approximation. Existing technical solutions for calculating position vectors include triangulation, trilateration, etc. Triangulation is a positioning method based on angle observation. It measures the angle between the target and two or more known observation points, and then uses geometric relationships to calculate the position of the target. The trilateration method measures the distance between the target and multiple known reference points and combines geometric methods to determine the three-dimensional coordinates of the target.
[0005] The patent application document with application number CN 201810911836.5 discloses a panoramic stitching method and panoramic stitching system based on multi-camera calibration. The main steps of the method are: internal parameter calibration and distortion coefficient solution for each camera; each adjacent two cameras are grouped together, and external parameter calibration data is collected respectively; corner point detection is performed on paired checkerboard images after distortion correction; the RANSAC algorithm is used to screen internal points and estimate the initial homography matrix between two adjacent cameras, and the initial rotation matrix is calculated based on the homography matrix and internal parameters; the LM algorithm is used to globally optimize the rotation matrix between all cameras to obtain the final optimization result. In this system, the external parameters of all cameras are optimized by the global optimization method to eliminate the cumulative error of the stitching parameters in the traditional method. At the same time, this method does not rely on the feature point information in the scene, but realizes stitching by calibrating the geometric relationship between cameras, avoiding the process of feature point extraction and matching for each stitching, thereby improving the stitching speed. Moreover, when there are few feature points or the overlap rate is low, this method can obtain a good stitching effect. However, this method relies heavily on the precise calibration of the intrinsic and extrinsic parameters of each camera, and the intrinsic parameters and distortion coefficients of the camera may change due to environmental changes such as temperature, humidity, or mechanical changes in the camera itself, such as loose lenses. Therefore, the calibration results need to be updated regularly, increasing maintenance costs. At the same time, since this method assumes that the position of the camera array is fixed and the relative position relationship between the cameras does not change during the stitching process, if the camera moves or vibrates during use, it needs to be recalibrated, so the requirements for hardware equipment are relatively high. Summary of the invention
[0006] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and to propose a method and system for reconstructing the trajectory of a flying target based on monocular vision, so as to reduce the dependence on equipment, reduce the sampling complexity, and improve the authenticity of trajectory reconstruction.
[0007] In order to achieve the above object, the technical solution of the present invention includes:
[0008] 1. A method for reconstructing a flying target trajectory based on monocular vision, characterized in that:
[0009] (1) Initialize the 3D engine and adjust the position relationship between the target and the camera to ensure that the target is fully projected within the camera's field of view;
[0010] (2) Split the three degrees of freedom of target rotation into 2+1;
[0011] (3) Read a frame of the original image in sequence, perform multiple rounds of iterative rotation sampling on the first two degrees of freedom of the target around the direction perpendicular to the camera's viewing direction in the 3D engine, and compare them with the original image to determine the precise values of the two degrees of freedom;
[0012] (4) Perform multiple rounds of iterative rotation sampling of the target around the camera's viewing direction and compare it with the original image to determine the precise value of the third degree of freedom;
[0013] (5) Calculate the rotation matrix of the target based on the numerical results of the three rotational degrees of freedom:
[0014] (6) According to the camera imaging principle, the distance from the camera to the plane where the target is located is calculated by using the proportional relationship between the ratio of the distance from the camera to the plane where the target is located and the focal length of the camera and the ratio of the actual size of the target to the imaged size; the spatial vector of the target in the camera coordinate system is determined by using the imaging position of the target in the original image and the distance from the camera to the plane where the target is located; the spatial vector of the target in the camera coordinate system is converted into the spatial vector in the world coordinate system by using the coordinate transformation in graphics;
[0015] (7) Use the space vector and rotation matrix to obtain the precise position and posture of the target in the orbit of the frame;
[0016] (8) Repeat steps (3) to (7) to calculate the orbital pose of all frames of the target and complete the flight target orbit reconstruction.
[0017] Preferably, the three degrees of freedom of the target rotation are split into 2+1, which is to split the three degrees of freedom of the target rotation into two degrees of freedom perpendicular to the camera viewing direction and one degree of freedom parallel to the target viewing direction.
[0018] The two degrees of freedom perpendicular to the camera's viewing direction are the x-axis and y-axis of the world coordinate system;
[0019] The degree of freedom parallel to the target's viewing direction is the z-axis of the world coordinate system.
[0020] 2. A flying target trajectory reconstruction system based on monocular vision, characterized by comprising:
[0021] System initialization module, used to initialize the 3D engine and provide a basic environment for subsequent sampling simulation;
[0022] Track frame extraction module, used to extract the image and camera parameters of each frame in track frame order;
[0023] The target posture matching module is used to iteratively sample and compare the target to determine the target's precise posture matrix;
[0024] The target position calculation module is used to calculate the actual position of the target in space and determine the precise position vector of the target.
[0025] Compared with the prior art, the present invention has the following advantages:
[0026] First, since the raw data of the present invention is collected from a single camera, no complex equipment is required, thus reducing the dependence on equipment.
[0027] Secondly, the present invention splits the three degrees of freedom of target rotation into 2+1 and reduces the complexity of sampling through multiple rounds of progressive optimization, thereby reducing the matching time while ensuring the authenticity of the reconstruction result.
[0028] Thirdly, since the present invention calculates the position of the target from the perspective of the camera imaging principle, it not only has the advantage of a simple and efficient calculation process, but also uses this calculation based on physical laws to ensure the authenticity and credibility of the target position calculation results. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a flow chart of the implementation of the method for reconstructing the trajectory of a flying target based on monocular vision of the present invention;
[0030] Figure 2 It is a schematic diagram of the longitude and latitude sphere used as a reference for the target rotation in the method of the present invention;
[0031] Figure 3 It is a schematic diagram of the comparison between the target image after rotation and matching around the x-axis and the y-axis in the method of the present invention and the original image;
[0032] Figure 4 It is a schematic diagram of comparison between the target posture matching in the method of the present invention and the original image;
[0033] Figure 5 It is a schematic diagram of the camera imaging principle in the method of the present invention;
[0034] Figure 6 is a schematic diagram of comparison between the final image after the pose is determined and the original image in the method of the present invention;
[0035] Figure 7 It is a block diagram of the flying target trajectory reconstruction system based on monocular vision of the present invention. DETAILED DESCRIPTION
[0036] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, other embodiments obtained by ordinary technicians in this field without creative work should all fall within the scope of protection of the present invention.
[0037] Example 1: A flying target trajectory reconstruction method based on monocular vision.
[0038] This example adds a target model and a camera to be detected to the 3D engine, adjusts the camera parameters to be consistent with the actual shooting camera parameters, adjusts the target's rotation posture through frame-by-frame rotation sampling, and obtains the target's rotation angle by comparing the camera screenshot with the original image; calculates the target's position vector through the imaging position and size of the target in the graphic using the camera imaging principle and graphics methods.
[0039] Reference Figure 1 The implementation steps of this example include the following:
[0040] Step 1: Initialize the 3D engine and adjust the target and camera position relationship
[0041] In this example, the target posture is matched in the 3D engine. To obtain accurate target posture results, the accurate model of the target needs to be loaded, and the camera parameters in the 3D engine need to be adjusted to be consistent with the parameters of the actual shooting camera. In addition, to ensure that the target presents a complete image in the camera field of view, the distance between the target and the camera needs to be appropriately adjusted.
[0042] The three-dimensional engine refers to the process of abstracting real-world objects into polygons or various curves, constructing a virtual three-dimensional world in the computer, and displaying it on the screen. The specific implementation of this step includes the following:
[0043] 1.1) Adjust the camera field of view and imaging pixel value in the 3D engine to make them consistent with the real camera;
[0044] 1.2) Place the target at the origin of the coordinate system, place the camera in the positive direction of the z-axis, and look toward the origin to facilitate the rotation of the target;
[0045] 1.3) Adjust the distance between the target and the camera so that the target forms a complete image within the camera's field of view:
[0046] 1.3.1) Calculate the camera's horizontal field of view angle fovx based on the vertical field of view angle fovy value provided by the camera:
[0047]
[0048] where width c and height c are the width and height of the camera’s field of view, respectively;
[0049] 1.3.2) According to the length width of the target in the x-axis direction t , calculate the camera distance d when the target occupies half of the camera's lateral field of view camx :
[0050]
[0051] 1.3.3) According to the length of the target in the y-axis direction height t , calculate the camera distance d when the target occupies half of the camera's longitudinal field of view camy :
[0052]
[0053] 1.3.4) Take d camx With d camy The maximum value of the camera is adjusted to the distance d between the camera and the origin. cam , ensuring that the target forms a complete image within the camera's field of view:
[0054] d cam =max(d camx ,d camy ).
[0055] Step 2: Get the original frame data of the track.
[0056] Read a frame of original image sequentially;
[0057] Get the camera orientation f when the image was taken c , the upward direction u, and the camera position p.
[0058] Step 3: Adjust the rotation of the target around the x-axis and y-axis so that each side of the target faces the camera. Perform multiple rounds of iterative rotation sampling and compare it with the original image to determine the precise values of the two degrees of freedom.
[0059] In order to reasonably operate the target rotation and range compression, the concept of longitude and latitude sphere is borrowed to rotate the intersection of longitude and latitude to the positive direction of z axis, and the longitude and latitude are converted into Euler angles for rotation sampling. SIFT is used to extract feature points for image comparison, and the number and geometric consistency are used as evaluation indicators. The longitude and latitude sphere is as follows: Figure 2 shown.
[0060] The specific implementation of this step includes the following:
[0061] 3.1) The target is rotated around the x-axis and y-axis for sampling:
[0062] 3.1.1) Assume that the first sampling latitude range is [0,180], the longitude range is [0,360], the step size is step, the longitude of the intersection of the longitude and latitude is θ, the latitude is φ, the radius r = 1, and calculate the Cartesian coordinate system:
[0063]
[0064] Convert Cartesian coordinates to Euler angles:
[0065]
[0066] Among them, roll, pitch, and yaw are the angles of rotation around the x, y, and z axes of the target coordinate system respectively;
[0067] 3.1.2) The target is rotated according to the Euler angle, and all intersection points are rotated around the x-axis and y-axis of the world coordinate system to the positive direction of the z-axis:
[0068] 3.1.3) After each rotation, operate the 3D engine to take a screenshot to obtain a sampled image;
[0069] 3.2) Use SIFT algorithm to extract feature points from all sampled images and original images;
[0070] 3.3) Match the feature points of all sampled images and original images respectively by Euclidean distance, and enable cross matching in the matching process to obtain multiple sets of matching point pairs;
[0071] 3.4) Use the matching point ratio and the inlier point ratio to compare the sampled image with the original image and select the best comparison result:
[0072] 3.4.1) Calculate the matching point ratio R match :
[0073]
[0074] 3.4.2) Screen the matching points through geometric consistency and affine transformation, remove the matching points that do not meet the geometric constraints, and obtain the number of inliers of the matching points;
[0075] 3.4.3) Calculate the inlier ratio measure R inlier :
[0076]
[0077] 3.4.4) According to R match With R inlier Calculate the comparison index R XY :
[0078] R XY =0.7·R match +0.3·R inlier ;
[0079] 3.4.5) Calculate R for all sampled images and original images XY , select R XY The longitude and latitude (lon, lat) corresponding to the largest image is taken as the result of this round of matching;
[0080] 3.5) With (lon, lat) as the center, change the longitude range to (lon-step, lon+step), the latitude range to (lan-step, lan+step), and reduce the step size by half, step = step / 2;
[0081] 3.6) Repeat 3.1)-3.5) to gradually reduce the step until the step is less than 1°, and the final longitude and latitude result (lon last ,lat last ) is converted to Euler angles through the Cartesian coordinate system (row last ,pitch last ), the Euler angle is the precise value of the target's rotation around the x-axis and y-axis.
[0082] At this time, the image of the rotated target in the camera has the same shape as the original image, but the rotation angle, size, and position in the two-dimensional plane of the camera's perspective are different, such as Figure 3 As shown, among which Figure 3 (a) is the original image, Figure 3 (b) is the image of the target from the camera’s perspective after determining the first two degrees of freedom.
[0083] Step 4: Perform multiple rounds of iterative rotation sampling on the target around the z-axis and compare it with the original image to determine the exact value of the third degree of freedom.
[0084] Since the image of the target in the camera field of view obtained in step 3 has the same shape as the original image, but the rotation angle in the two-dimensional plane of the camera's view is different, it is necessary to rotate the target around the z-axis of the world coordinate system to realize rotation sampling in the two-dimensional plane of the camera's view and compare it with the original image to determine the exact value of the third degree of freedom. Its implementation includes the following:
[0085] 4.1) Keep the rotation angle of the target around the x-axis and y-axis unchanged, set the first round sampling step size as step, the rotation range around the z-axis as [0,360], rotate around the z-axis of the world coordinate system in the rotation range according to the step size, and take a screenshot to obtain the sampled image;
[0086] 4.2) Use SIFT algorithm to extract feature points from the sampled image and the original image, each feature point has its own unique descriptor;
[0087] 4.3) Match the feature points of all sampled images and original images respectively through Euclidean distance. Cross matching is enabled in the matching process to obtain multiple sets of matching point pairs:
[0088] In the process of image feature point matching, the Euclidean distance is used as the metric to match the feature points of all sampled images with the original image, that is, the Euclidean distance between each feature point descriptor in the sampled image and all feature point descriptors in the original image is calculated. The smaller the Euclidean distance, the closer the two feature points are in the feature space, and the greater the possibility of matching.
[0089] During the matching process, the cross-matching mechanism is enabled, that is, not only matching is performed from the sampled image to the original image, but also reverse matching is performed from the original image to the sampled image. Only when the matching results in the two directions are consistent, the pair of feature points is considered to be a valid match. In this way, some false matches are effectively filtered out, and multiple groups of matching point pairs are obtained, thereby improving the accuracy and reliability of matching.
[0090] 4.4) Use the intersection points between the matching point pairs to compare the images:
[0091] Since the imaging target of this round of sampling rotates in the camera imaging plane, the results obtained by comparing the number and geometric consistency are not much different. Therefore, this step first screens the matching points by geometric consistency and affine transformation to eliminate the matching points that do not meet the geometric constraints; then calculate the number of intersections between the line segments connected by the matching point pairs, so that in an ideal case the matching points completely overlap or the extensions of these line segments intersect at one point. The specific implementation includes the following:
[0092] 4.4.1) Using the key point coordinate pairs, an optimal affine transformation matrix A is estimated using the affine transformation model:
[0093]
[0094] where a 11 and a 22 Represents the scaling factors of the x-axis and y-axis respectively, a 12 and a 21 They represent the shear factor of the y-axis to the x-axis and the shear factor of the x-axis to the y-axis, respectively. x ,t y Represents the translation of the x-axis and y-axis respectively; the matrix describes the geometric relationship between the first set of key points and the second set of key points;
[0095] 4.4.2) For each matching point pair, use the affine transformation matrix A to transform the point p of the first image 1 Transform the coordinates p' to the coordinate system of the second image 2 :
[0096] p' 2 =A·p 1 ;
[0097] 4.4.3) According to p 1The matching point p in the second image 2 And the affine transformation point p' 2 Calculate the Euclidean distance dis between the actual corresponding point and the transformed point:
[0098] dis=||p' 2 -p 2 ||;
[0099] 4.4.4) Set the threshold T according to the calculation accuracy and compare it with the Euclidean distance dis to determine whether the point meets the geometric consistency:
[0100] If dis<T, the point is considered to be geometrically consistent and step 4.4.5 is executed;
[0101] Otherwise, the matching pair corresponding to the point is eliminated and the point will not be considered in the future;
[0102] This example sets T=3 but is not limited thereto.
[0103] 4.4.5) Use the coordinates of the selected matching point pairs to form line segments, and combine all line segments in pairs to determine whether they intersect:
[0104] 4.4.5a) Define the two line segments as [p 1 ,p 2 ] and [q 1 ,q 2 ], the vector form is:
[0105]
[0106] 4.4.5b) Calculate the cross product d of two line segment vectors 1 ×d 2 , determine whether the result is zero:
[0107] If the result is zero, the line segments are parallel or collinear;
[0108] Otherwise, calculate the relative position t of the intersection point on the two line segments 1 ,t 2 :
[0109]
[0110] 4.4.5c) Determine t 1 ,t 2 Is it in the range [0,1]?
[0111] If so, the two line segments intersect;
[0112] Otherwise, the two line segments do not intersect;
[0113] 4.4.5d) Using the method of steps 4.4.5a)-4.4.5c), calculate the number of intersections between the lines connecting all matching point pairs of the single sample image and the original image;
[0114] 4.5) Using the method in step 4.4), calculate the number of intersections between the lines connecting all the matching point pairs of the sampled image and the original image, and select the rotation angle yaw corresponding to the sampled image with the least intersections as the result of this round of matching;
[0115] 4.6) With yaw as the center, change the rotation range to (yaw-step, yaw+step), and reduce the step length by half
[0116] step = step / 2;
[0117] 4.7) Repeat steps 4.1)-4.6) to gradually reduce step until step is less than 1°. The final rotation result yaw last This is the exact value of the rotation around the z-axis.
[0118] After the rotation, the image of the target in the camera is basically consistent with the shape of the target in the original image, such as Figure 4 As shown, Figure 4 (a) is the original image, Figure 4 (b) is the image of the target in the camera's field of view after determining the three degrees of freedom.
[0119] Step 5, calculate the rotation matrix of the target according to the numerical results of the three rotational degrees of freedom.
[0120] Since steps 3 and 4 are the target postures obtained when the target is at the origin of the coordinate system and the camera is in the positive direction of the z-axis, it is necessary to rotate the camera to the posture when the frame was taken so that the target rotates synchronously with the camera and obtain the rotation matrix corresponding to the original image. The implementation includes the following:
[0121] 5.1) According to the rotation order of rotating around the Z axis first, then around the Y axis, and finally around the X axis, multiply the rotation matrices of the three axes to calculate the rotation matrix R of the target simulation result:
[0122] R=Rx(row last )·Ry(pitch last )·Rz(yaw last )
[0123] Where Rx(row last )、Rx(row last )、Rx(row last ) are the rotation matrices of the x-axis, y-axis, and z-axis respectively;
[0124] 5.2) Rotate the camera and target:
[0125] 5.2.1) According to the camera parameters of the simulation state, the camera orientation f1, the upward direction u1 and the camera position p1, the simulation camera transformation matrix M1 is constructed:
[0126]
[0127] Among them, r1=f1×u1 is the right vector of the camera in the simulation state, and u1=r1×f1 is the upward vector of the camera after adjustment in the simulation state;
[0128] 5.2.2) According to the initial camera parameters when taking the picture, the camera orientation f2, the upward direction u2 and the camera position p2, the camera transformation matrix M2 when taking the picture is constructed:
[0129]
[0130] Among them, r2=f2×u2 is the right vector of the camera when taking the picture, and u2=r2×f2 is the upward vector of the camera after adjustment when taking the picture;
[0131] 5.2.3) Rotate the camera from the simulated posture to the posture during real shooting, that is, transform from the M1 matrix to the M2 matrix. The target rotates synchronously with the camera to obtain the rotation matrix R2 of the target in this frame:
[0132]
[0133] Among them, M2 rot is the upper left 3×3 matrix of M2, is the transposed matrix of the upper left 3×3 matrix of M1.
[0134] Step 6, calculate the spatial vector of the target.
[0135] 6.1) According to the imaging position and size of the target in the figure, the spatial vector of the target center in the camera coordinate system is accurately calculated using the camera imaging principle:
[0136] 6.1.1) Obtain the geometric relationship between the camera and the target:
[0137] From Figure 5 From the imaging principle diagram of the camera lens shown in , we can see that since the object forms an inverted, reversed, and reduced image on the imaging sensor after passing through the camera lens, and the camera processes the image formed by the sensor into a positively reduced image as an image output, the image seen by the eye can be equivalently understood as the equivalent imaging plane in the figure, thereby obtaining the following geometric relationship between the camera and the target:
[0138]
[0139] Where D is the distance between the target plane and the camera, f is the focal length of the camera, l is the length of any position of the target in the image, and L is the actual length of the target at the position corresponding to l;
[0140] 6.1.2) Using the geometric relationship between the camera and the target, we can get the geometric relationship between the target and the camera at different distances, and calculate the distance D from the camera to the plane where the target is located during the actual shooting. 1
[0141]
[0142] Among them, L 1 is the minimum external matrix length of the original image target, L 2 is the minimum external matrix length of the target in the final pose downsampled image, D 2 is the distance between the final camera and the target plane;
[0143] 6.1.3) Let the center of the image be the origin of the coordinate system, the right side be the positive direction of the x-axis, and the upward direction be the positive direction of the y-axis. Determine the imaging center position of the target in the original image as (x p ,y p ), determine the spatial vector ct of the target imaging center in the camera coordinate system according to the target imaging position and the camera focal length:
[0144] ct=(x p ,-y p ,-f)
[0145] Where, f is the focal length of the camera, and the camera uses a right-handed coordinate system;
[0146] 6.1.4) Calculate the distance d between the camera and the target imaging center p :
[0147]
[0148] 6.1.5) Calculate the distance d between the target center and the camera:
[0149]
[0150] 6.1.6) Normalize the spatial vector ct of the target imaging center in the camera coordinate system to the unit direction vector e, and calculate the spatial vector α of the target in the camera coordinate system:
[0151] α = de;
[0152] 6.2) Use graphics methods to calculate the position vector of the target in the world coordinate system:
[0153] 6.2.1) Construct the homogeneous vector P of the target space position based on the target space vector camera:
[0154] P camera =(α x ,α y ,α z ,1)
[0155] Among them, α x , α y , α z are the sizes of the target in the x, y and z directions in the camera coordinate system respectively;
[0156] 6.2.2) According to the transformation matrix M2 and homogeneous vector P in step 5.2.2) camera , calculate the position vector of the target in the world coordinate system as P world :
[0157] P world =M2*P camera .
[0158] Step 7, save the rotation matrix calculated in step 5 and the position vector calculated in step 6 as the reconstructed pose parameters of the target in the current frame, and apply these parameters as the pose parameters of the target in the 3D engine to obtain the image from the camera's perspective.
[0159] Compare the image from the camera's perspective with the original image, and the result is as follows Figure 6 As shown. Figure 6 (a) is the original image, Figure 6 (b) is the image from the camera perspective in the 3D engine. Figure 6 It can be seen that the imaging of the target in the camera's perspective is basically consistent with the shape and position of the target in the original image, indicating that the reconstruction result of the current frame has a high degree of authenticity.
[0160] Step 8: Repeat steps 2 to 7 to calculate the target pose of all frames of the track and complete the flight target track reconstruction based on monocular vision.
[0161] Example 2: Flying target trajectory reconstruction system based on monocular vision.
[0162] Reference Figure 7 ,This example system includes an initialization module, a track frame extraction module, a target posture matching module, a target position calculation module and a judgment module, among which:
[0163] The initialization module initializes the three-dimensional engine to provide a basic environment for subsequent sampling;
[0164] The track frame extraction module sequentially extracts the image and camera parameters of a track frame, and transmits the original image and camera parameters of the extracted frame to the target posture matching module and the target position calculation module in sequence;
[0165] The target posture matching module first splits the three degrees of freedom of target rotation into 2+1, then performs iterative rotation sampling on the first two degrees of freedom and the last degree of freedom respectively, and uses different comparison indicators to compare with the original picture of the frame transmitted by the track frame extraction module, and gradually narrows the range of iterative sampling according to the comparison results. When the range is narrowed to the set range, the precise values of the three degrees of freedom can be determined, and finally the precise rotation matrix of the target is determined according to the precise values of the three degrees of freedom and the camera parameters transmitted by the track frame extraction module;
[0166] The target position calculation module calculates the distance from the camera to the plane where the target is located by using the proportional relationship between the ratio of the distance from the camera to the plane where the target is located and the focal length of the camera and the ratio of the actual size of the target to the imaged size, determines the spatial vector of the target in the camera coordinate system by using the distance from the camera to the plane where the target is located and the imaging position of the target in the original picture of the frame transmitted by the track frame extraction module, and converts the spatial vector of the target in the camera coordinate system into the spatial vector in the world coordinate system by using the camera parameters transmitted by the track frame extraction module;
[0167] The judgment module judges whether the current frame is the last frame. If so, the system ends the operation. Otherwise, the system runs the track frame extraction module, extracts the original data of the next frame and runs the target posture matching module and the target position calculation module in sequence to perform calculations until the last frame is calculated.
[0168] It should be noted that the step numbers in the specification and claims of the present invention are only for a clear description of the implementation scheme of the present invention to facilitate understanding, and the sequence of the step numbers is not limited.
Claims
1. A method and system for reconstructing a flying target trajectory based on monocular vision, characterized in that: include: (1) Initialize the 3D engine and adjust the position relationship between the target and the camera to ensure that the target is fully projected within the camera's field of view; (2) Split the three degrees of freedom of target rotation into 2+1; (3) Read a frame of the original image in sequence, perform multiple rounds of iterative rotation sampling on the first two degrees of freedom of the target around the direction perpendicular to the camera's viewing direction in the 3D engine, and compare them with the original image to determine the precise values of the two degrees of freedom; (4) Perform multiple rounds of iterative rotation sampling of the target around the camera's viewing direction and compare it with the original image to determine the precise value of the third degree of freedom; (5) Calculate the rotation matrix of the target based on the numerical results of the three rotational degrees of freedom: (6) According to the camera imaging principle, the distance from the camera to the plane where the target is located is calculated by using the proportional relationship between the ratio of the distance from the camera to the plane where the target is located and the focal length of the camera and the ratio of the actual size of the target to the imaged size; the spatial vector of the target in the camera coordinate system is determined by using the imaging position of the target in the original image and the distance from the camera to the plane where the target is located; the spatial vector of the target in the camera coordinate system is converted into the spatial vector in the world coordinate system through the coordinate transformation relationship in graphics; (7) Use the space vector and rotation matrix to obtain the precise position and posture of the target in the orbit of the frame; (8) Repeat steps (3) to (7) to calculate the target pose for all frames of the trajectory and complete the flight target trajectory reconstruction.
2. The method according to claim 1, characterized in that: In step (1), the 3D engine is initialized and the position relationship between the target and the camera is adjusted. The implementation includes the following steps: (1a) Adjust the camera field of view and imaging pixel value in the engine to make them consistent with the real camera; (1b) Place the target at the origin of the coordinate system, place the camera in the positive direction of the z-axis, and look toward the origin to facilitate the rotation of the target; (1c1) Calculate the camera distance d when the target occupies half of the camera's lateral field of view camx : where width t is the length of the target in the x-axis direction, and fovx is the horizontal field of view of the camera; (1c2) Calculate the camera distance d when the target occupies half of the camera's longitudinal field of view camy : where height t is the length of the target in the y-axis direction, and fovy is the vertical field of view of the camera; (1c3) Take d camx With d camy The maximum value of the camera is adjusted to the distance d between the camera and the origin. cam , ensuring that the target occupies half of the camera's field of view: d cam =max(d camx ,d camy )。 3. The method according to claim 1, characterized in that In step (2), the three degrees of freedom of target rotation are split into 2+1, which means that the three degrees of freedom of target rotation are split into two degrees of freedom perpendicular to the camera's viewing direction and one degree of freedom parallel to the target's viewing direction. The two degrees of freedom perpendicular to the camera's viewing direction are the x-axis and y-axis of the world coordinate system; The one degree of freedom parallel to the target's viewing direction is the z-axis of the world coordinate system.
4. The method according to claim 1, characterized in that: In step (3), multiple rounds of iterative rotation sampling are performed on the first two degrees of freedom of the target around the direction perpendicular to the camera's viewing direction in the three-dimensional engine to determine the precise values of the two degrees of freedom. The implementation includes the following: (3a) Based on the concept of longitude and latitude sphere, the latitude range is set to [0,180] and the longitude range is set to [0,360] for the first sampling, and the step length is set to step. The operation target is to rotate the intersection point around the x-axis and y-axis of the world coordinate system to the positive direction of the z-axis, and take a screenshot to obtain the sampled image; (3b) Extract feature points from all sampled images and original images using the SIFT algorithm; (3c) matching the feature points of all sampled images and original images respectively by Euclidean distance to obtain multiple sets of matching point pairs; (3d) Use the matching point ratio and the inlier point ratio to compare the sampled image with the original image and select the most comparable result: (3d1) Calculate the matching point ratio R match : (3d2) Screen the matching points through geometric consistency and affine transformation, remove the matching points that do not meet the geometric constraints, and obtain the number of inner points of the matching points; (3d3) Calculate the interior point ratio measure R inlier : (3d4) According to R match With R inlier Calculate the comparison index R XY : R XY =0.7·R match +0.3·R inlier ; (3d5) Calculate R for all sampled images and the original image XY , select R XY The longitude and latitude (lon, lat) corresponding to the largest image is taken as the result of this round of matching; (3e) With (lon, lat) as the center, the longitude range is modified to (lon-step, lon+step), the latitude range is modified to (lan-step, lan+step), and the step length is reduced by half, step = step / 2; (3f) Repeat (3a)-(3e) to gradually reduce the step until the step is less than 1°. The final longitude and latitude result (lon last ,lat last ) is converted to Euler angles through the Cartesian coordinate system (row last ,pitch last ), the Euler angle is the precise value of the target's rotation around the x-axis and y-axis.
5. The method according to claim 1, characterized in that In step (4), the target is subjected to multiple rounds of iterative rotation sampling around the camera viewing direction and compared with the original image to determine the precise value of the third degree of freedom. The implementation includes the following: (4a) Assume that the sampling step is step, the rotation range is [0,360], the operation target is rotated around the z-axis of the world coordinate system in the rotation range according to the step, and a screenshot is taken to obtain the sampled image; (4b) Extract feature points from the sampled image and the original image using the SIFT algorithm; (4c) matching the feature points of the sampled image and the original image by Euclidean distance to obtain multiple sets of matching point pairs; (4d) Screening the matching points through geometric consistency and affine transformation, and eliminating the matching points that do not meet the geometric constraints; (4e) Compare all sampled images with the original image, calculate the number of intersections between the lines connecting the matching point pairs, and select the rotation angle yaw corresponding to the sampled image with the least intersections as the result of this round of matching; (4f) With yaw as the center, the rotation range is modified to (yaw-step, yaw+step), and the step length is reduced by half step=step / 2; (4g) Repeat (4a)-(4f) to gradually reduce the step until the step is less than 1°. The final rotation result yaw last This is the exact value of the rotation around the z-axis.
6. The method according to claim 1, characterized in that In step (5), the rotation matrix of the target is calculated according to the numerical results of the three rotational degrees of freedom, and its implementation includes the following: (5a) According to the rotation order of rotating around the Z axis first, then around the Y axis, and finally around the X axis, the rotation matrices of the three axes are multiplied to calculate the rotation matrix R of the target simulation result: R=Rx(row last )·Ry(pitch last )·Rz(yaw last ) Where Rx(row last )、Rx(row last )、Rx(row last ) are the rotation matrices of the x-axis, y-axis, and z-axis respectively; (5b) According to the camera transformation matrix M1 during simulation and the camera transformation matrix M2 during picture taking, the camera is transformed from the M1 matrix to the M2 matrix to obtain the rotation matrix R2 of the target in this frame: Among them, M2 rot is the upper left 3×3 matrix of M2, is the transposed matrix of the upper left 3×3 matrix of M1.
7. The method according to claim 1, characterized in that In step (6), the distance from the camera to the target plane is calculated, and its implementation includes the following: (6a) According to the camera imaging principle, the following geometric relationship between the camera and the target is obtained: Where D is the distance between the target and the camera, f is the focal length of the camera, l is the length of any position of the target in the image, and L is the actual length of the target at the position corresponding to l; (6b) Using the geometric relationship between the camera and the target, we can obtain the following geometric relationship when the target is at different distances from the camera, and calculate the distance D1 from the camera to the plane where the target is located during the actual shooting. Where L1 is the minimum external matrix length of the target in the original image, L2 is the minimum external matrix length of the target in the final pose downsampled image, and D2 is the distance between the final camera and the plane where the target is located.
8. The method according to claim 1, characterized in that: In step (6), the spatial vector of the target in the camera coordinate system is determined by using the imaging position of the target in the original image and the distance from the camera to the plane where the target is located. The implementation includes the following steps: (6c) In the image, let the center of the image be the origin of the coordinate system, the right side be the positive direction of the x-axis, and the upward direction be the positive direction of the y-axis. The imaging center position of the target in the original image is determined as (x p ,y p ), determine the spatial vector ct of the target imaging center in the camera coordinate system according to the target imaging position and the camera focal length: ct=(x p ,-y p ,-f) Where, f is the focal length of the camera, and the camera uses a right-handed coordinate system; (6d) Calculate the distance d between the camera and the target imaging center p : (6e) Calculate the distance d between the target center and the camera: (6f) Normalize the spatial vector ct of the target imaging center in the camera coordinate system to the unit direction vector e, and calculate the spatial vector α of the target in the camera coordinate system: α=de.
9. The method according to claim 1, characterized in that: In step (6), the spatial vector of the target in the camera coordinate system is converted into the spatial vector in the world coordinate system through the coordinate transformation relationship in graphics, and its implementation includes the following: (6g) According to the camera orientation f c , the upward direction u and the camera position p construct the transformation matrix M2 of the camera in the original state: Where r = f c ×u is the right vector, u=r×f c is the adjusted upward vector; (6h) Construct the homogeneous vector P of the target space position according to the target space vector camera : P camera =(α x ,α y ,α z ,1) Among them, α is the spatial vector of the target in the camera coordinate system; (6i) According to the transformation matrix M2 and the homogeneous vector P camera , using the graphics coordinate transformation formula, calculate the position vector of the target in the world coordinate system as P world : P world =M2*P camera 。 10. A flying target trajectory reconstruction system based on monocular vision, characterized in that: include: System initialization module, used to initialize the 3D engine and provide a basic environment for subsequent sampling simulation; Track frame extraction module, used to extract the image and camera parameters of each frame in track frame order; The target posture matching module is used to iteratively sample and compare the target to determine the target's precise posture matrix; The target position calculation module is used to calculate the actual position of the target in space and determine the precise position vector of the target; The judgment module is used to judge whether the current frame is the last frame of the track. If it is, the system operation ends, otherwise the track frame extraction module is executed.
Citation Information
Patent Citations
Panoramic splicing method based on multi-camera calibration, panoramic splicing system
CN109064404A
Target attitude error measuring method for underwater binocular vision
CN111256732A
Monocular camera-based planar moving target visual navigation method and system
CN112179357A
Space non-cooperative target relative pose estimation method based on UPF
CN113175929A
On-orbit calibration method of vision and laser fused space target pose measurement system
CN119146857A