Multi-camera calibration method and positioning method

Through the multi-camera calibration method, the neural network model is used to automatically extract features and calculate the transformation matrix, which solves the cumbersome problems of manual calibration in the existing technology, and realizes efficient and automatic multi-camera calibration and positioning.

CN119991829BActive Publication Date: 2025-06-24CHENGDU CHANGHONG NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510465923.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-24
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The existing multi-view visual positioning method requires manual calibration of the camera's internal and external parameters, which is cumbersome.

Method used

Through the multi-camera calibration method, the neural network model is used to extract the pixel coordinates of semantic features and feature points from the video, calculate the transformation matrix between the cameras, and automatically complete the calibration of the camera.

Benefits of technology

It realizes automatic calibration of multiple cameras without manual calibration, improves positioning accuracy and robustness, and simplifies the operation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991829B_ABST
    Figure CN119991829B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-camera calibration method and a positioning method, which relate to image processing technology. By determining a calibration object, multiple cameras are used to respectively collect videos containing the object, and a picture sequence corresponding to each camera is obtained. Features of the calibration object are extracted from the pictures, and the features include semantic features and pixel coordinates of multiple feature points. Using a neural network model, similar pictures between cameras are obtained from the feature map sequences through the semantic features. The pixel coordinates of the multiple feature points in each camera are used in the similar pictures, and based on the fact that the true distances between the feature points are equal, the three-dimensional coordinates of the feature points relative to each camera are calculated. Through the three-dimensional coordinates of the feature points relative to each camera, the transformation matrix between the cameras is obtained, and the calibration of the cameras is completed, solving the problem that existing multi-view positioning requires manual calibration. The present invention is applicable to multi-camera automatic calibration and positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image processing technology, and particularly to a multi-camera calibration method and a positioning method. Background Art

[0002] Currently, because the position service has higher and higher requirements for the accuracy and real-time performance of the position, visual positioning is an important development direction in the field of computer vision.

[0003] Two cameras with different angles in multi-view visual positioning can complement each other in the dimension information perpendicular to the camera plane lost during the shooting process, and can correct the calculation results for each other, providing support for calculating the three-dimensional size of an object, and also improving the measurement accuracy and robustness.

[0004] The cost of home surveillance cameras is getting lower and lower. Due to considerations of security, anti-theft, etc., more and more families install multiple surveillance cameras and connect them to the home Internet of Things terminal. Therefore, under the condition of the home Internet of Things, it becomes possible to use home surveillance cameras for indoor object recognition and positioning.

[0005] However, existing multi-view visual positioning often requires manual camera internal and external parameter calibration steps according to the calibration algorithm with the help of auxiliary tools before use, and the calibration process is relatively cumbersome. Summary of the Invention

[0006] The technical problem to be solved by the present invention: Provide a multi-camera calibration method and a positioning method to solve the problem that existing multi-view positioning requires manual calibration.

[0007] The technical solution adopted by the present invention to solve the above technical problem: The multi-camera calibration method includes the following steps:

[0008] S1. Determine the calibration object;

[0009] S2. Use multiple cameras to respectively collect videos containing the calibration object, extract the corresponding picture sequences of the videos, and obtain the picture sequences corresponding to each camera;

[0010] S3. Extract the features of the calibration object from the pictures, and the features include semantic features and pixel coordinates of multiple feature points;

[0011] S4. Use a neural network model to obtain similar pictures between cameras from the feature map sequences through semantic features, and the similar pictures include the corresponding pictures of the calibration object in each camera at the same moment;

[0012] S5. Use the pixel coordinates of multiple feature points in similar images in each camera, and based on the fact that the actual distances between the feature points are equal, calculate the three-dimensional coordinates of the feature points relative to each camera;

[0013] S6. Obtain the transformation matrix between the cameras through the three-dimensional coordinates of the feature points relative to each camera, and complete the calibration of the cameras.

[0014] Furthermore, the semantic features include object size, object pose, object local information, object global information, and the similarity between the object in adjacent frames.

[0015] Furthermore, in S3, extracting the features of the calibration object from the image includes: determining the position of the calibration object in the image through the detection module, and obtaining the features of the calibration object in the image through the feature extraction model.

[0016] Furthermore, in S4, the similar images are the images with the highest similarity or the images with a similarity higher than the threshold.

[0017] Furthermore, the calibration object is a human body, the multiple feature points are multiple of the human body feature points, and the human body feature points include nose, neck, right shoulder, left shoulder, right elbow, left elbow, right wrist, left wrist, right hip, left hip, right knee, left knee, right ankle, left ankle, right eye, left eye, right ear, left ear, right toe tip, and left toe tip.

[0018] Furthermore, in S5, it includes the following steps:

[0019] S51. Denote the th feature point as , and the pixel coordinates in the th camera as: , then in the th camera's imaging plane, the two-dimensional coordinates are denoted as: , where represents the ratio of the actual size to the pixel size of the image taken by the th camera;

[0020] S52. Expand the two-dimensional coordinates of in the th camera's imaging plane into three-dimensional coordinates to obtain 's three-dimensional coordinates relative to the th camera, denoted as: , where represents 's depth relative to the th camera, and thus obtain the three-dimensional coordinates of each feature point relative to each camera;

[0021] S53. Establish an equal - distance relationship equation based on the equal true distances between feature points: , where represents the three - dimensional coordinates of the th feature point relative to the th camera, represents the three - dimensional coordinates of the th feature point relative to the th camera, represents the three - dimensional coordinates of the th feature point relative to the th camera, represents the distance calculation symbol;

[0022] S54. Use the equal true distances between multiple feature points to obtain corresponding equations, and solve for to obtain the three - dimensional coordinates of each feature point.

[0023] Furthermore, in S6, the transformation matrix between cameras is obtained through the three - dimensional coordinates of feature points relative to each camera, including the following steps:

[0024] S61. Subtract the three - dimensional coordinates of relative to the th camera from the three - dimensional coordinates of relative to the th camera to obtain the three - dimensional coordinates of the th camera relative to the th camera;

[0025] S62. Combine the rotation matrix of the th camera and the th camera to obtain the homogeneous transformation matrix of the th camera and the th camera. The homogeneous transformation matrix is: , where represents the homogeneous transformation matrix of the th camera and the th camera, represents the rotation matrix of the th camera and the th camera, , , , , represents the attitude angle, represents the rotation angle around the axis, represents the rotation angle around the The angle of rotation of the axis represents the angle of rotation about the axis, represents the matrix transpose of the three-dimensional coordinates of the th camera relative to the th camera.

[0026] Furthermore, in S6, it also includes obtaining the internal parameter matrix of each camera through the three-dimensional coordinates of the feature points relative to each camera. The internal parameter matrix is: , where represents the internal parameter matrix, represents the pixels of the th camera, represents the focal length. The calculation formula of the internal parameter matrix is:

[0027] , where represents the three-dimensional coordinates of the th feature point relative to the th camera, represents the ratio of the real size to the pixel size of the image captured by the th camera.

[0028] The present invention also provides a multi-camera positioning method, which is applied to the calibrated cameras obtained by using the multi-camera calibration method as described above. The positioning method includes the following steps:

[0029] S01. Select two cameras, denoted as the first camera and the second camera respectively, and obtain the internal parameter matrix of the first camera and the homogeneous transformation matrix between the first camera and the second camera;

[0030] S02. Use the two cameras to capture the object to be positioned, and obtain the pixel coordinates of the feature points of the object to be positioned in the first camera and the second camera;

[0031] S03. The three-dimensional coordinates of the th feature point relative to the first camera are . Through the formula , calculate , where represents the pixel coordinates of the feature point in the first camera, represents the internal parameter matrix of the first camera; the three-dimensional coordinates of the th feature point relative to the second camera are . Through the formula , calculate , where represents the pixel coordinates of the feature point in the second camera, represents the internal parameter matrix of the second camera;

[0032] S04. Substitute and into , and solve to obtain the three-dimensional coordinates of the th feature point relative to the first camera or the three-dimensional coordinates of the th feature point relative to the second camera , where represents the homogeneous transformation matrix between the first camera and the second camera.

[0033] Further, the multi-camera positioning method further includes the following steps:

[0034] S05. Establish a world coordinate system with a certain camera as the reference;

[0035] S06. Use the homogeneous transformation matrix between the first camera and the certain camera to convert the three-dimensional coordinates of the feature point relative to the first camera to the coordinates of the feature point relative to the world coordinate system, and obtain the world coordinates of the object to be positioned; or, use the homogeneous transformation matrix between the second camera and the certain camera to convert the three-dimensional coordinates of the feature point relative to the second camera to the coordinates of the feature point relative to the world coordinate system, and obtain the world coordinates of the object to be positioned.

[0036] Advantages of the present invention: The present invention provides a multi-camera calibration method and a positioning method. By determining the calibration object, using multiple cameras to respectively collect videos containing the object, obtaining the corresponding picture sequences of each camera, extracting the features of the calibration object from the pictures, the features including semantic features and the pixel coordinates of multiple feature points, using a neural network model, obtaining similar pictures between cameras from the feature map sequences through the semantic features, the similar pictures including the corresponding pictures of the calibration object in each camera at the same moment, using the pixel coordinates of the multiple feature points in the similar pictures in each camera, and based on the fact that the actual distances between the feature points are equal, calculating the three-dimensional coordinates of the feature points relative to each camera, obtaining the transformation matrix between the cameras through the three-dimensional coordinates of the feature points relative to each camera, completing the calibration of the cameras, solving the problem that existing multi-view positioning requires manual calibration, and using the calibrated cameras to position the object to be positioned. Description of the Drawings

[0037] Figure 1 is a schematic flowchart of a multi-camera calibration method provided by the present invention. Detailed Embodiments

[0038] In view of the problem that existing multi-view positioning requires manual calibration, the present invention provides a multi-camera calibration method, asFigure 1 As shown, it includes the following steps:

[0039] S1. Determine the calibration object.

[0040] Specifically, the calibration object is a human body or an object.

[0041] S2. Use multiple cameras to respectively collect videos containing the calibration object, extract the corresponding picture sequences of the videos, and obtain the picture sequences corresponding to each camera;

[0042] Specifically, since a video is composed of consecutive pictures, the video can be split into a picture sequence frame by frame.

[0043] S3. Extract the features of the calibration object from the pictures. The features include semantic features and the pixel coordinates of multiple feature points.

[0044] Specifically, extracting the features of the calibration object from the pictures includes: judging the position of the calibration object in the picture through a detection module, and obtaining the features of the calibration object in the picture through a feature extraction model. The detection module is yolo5. The semantic features refer to high-level features that can represent the content of the image and are transformed into a high-dimensional space, such as: object size, object pose, object local information, object global information, and the similarity between the object in adjacent frames, etc.

[0045] If the object is a human body, the multiple feature points are multiple of the human body feature points. The human body feature points include nose, neck, right shoulder, left shoulder, right elbow, left elbow, right wrist, left wrist, right hip, left hip, right knee, left knee, right ankle, left ankle, right eye, left eye, right ear, left ear, right toe tip, and left toe tip. The feature extraction model is openpose.

[0046] S4. Use a neural network model to obtain similar pictures between cameras from the sequence of feature maps through semantic features. The similar pictures include the corresponding pictures of the calibration object in each camera at the same moment.

[0047] Specifically, due to the different shooting speeds and shooting environments of multiple cameras during the shooting process, and there may also be phenomena such as frame dropping and picture occlusion caused by unstable factors such as the environment and network. Even if the starting time of shooting is the same, the obtained video picture sequences are often not in a one-to-one correspondence relationship. Therefore, it is necessary to perform time alignment of the video picture sequences, that is, to obtain similar pictures between cameras from the sequence of feature maps. The similar pictures are feature pairs with the highest similarity or feature pairs with a similarity higher than the threshold to ensure time alignment, that is, the corresponding pictures of the calibration object in each camera at the same moment.

[0048] S5. Use the pixel coordinates of multiple feature points in similar images in each camera, and based on the fact that the true distances between the feature points are equal, calculate the three-dimensional coordinates of the feature points relative to each camera.

[0049] Specifically, in S5, the following steps are included:

[0050] S51. Denote the th feature point as , and the pixel coordinates in the th camera as: . Then the two-dimensional coordinates in the imaging plane of the th camera are denoted as: , where represents the ratio of the true size to the pixel size of the image captured by the th camera.

[0051] S52. Expand the two-dimensional coordinates of in the imaging plane of the th camera into three-dimensional coordinates to obtain the three-dimensional coordinates of relative to the th camera, denoted as: , where represents the depth of relative to the

[0052] th camera, thereby obtaining the three-dimensional coordinates of each feature point relative to each camera. S53. Based on the fact that the true distances between the feature points are equal, establish an equal-distance relationship formula: , where represents the three-dimensional coordinates of the th feature point relative to the th camera, represents the three-dimensional coordinates of the th feature point relative to the th camera, represents the three-dimensional coordinates of the th feature point relative to the th camera, and

[0053] S54. Use the fact that the true distances between multiple feature points are equal to obtain the corresponding equations, and solve for to obtain the three-dimensional coordinates of each feature point.

[0054] For example, taking five feature points and two cameras as an example, after expanding to three-dimensional coordinates, 10 third-dimension parameters are added. The true distances between any two of the five feature points remain unchanged, and 10 distance-equality formulas can be obtained. Therefore, the values of the 10 third-dimension parameters can be solved, and thus the three-dimensional coordinates of each feature point can be obtained.

[0055] S6. Obtain the transformation matrix between cameras through the three-dimensional coordinates of the feature points relative to each camera, and complete the calibration of the cameras.

[0056] Specifically, in S6, obtaining the transformation matrix between cameras through the three-dimensional coordinates of the feature points relative to each camera includes the following steps:

[0057] S61. Subtract the three-dimensional coordinates relative to the th camera from the three-dimensional coordinates relative to the th camera to obtain the three-dimensional coordinates of the th camera relative to the th camera; th camera relative to the th camera;

[0058] S62. Combine the rotation matrix of the th camera and the th camera to obtain the homogeneous transformation matrix of the th camera and the th camera. The homogeneous transformation matrix is: , where represents the rotation matrix of the th camera and the th camera, , , , , represents the attitude angle, represents the rotation angle around the axis, represents the rotation angle around the axis, represents the rotation angle around the axis, represents the matrix transpose of the three-dimensional coordinates of the th camera relative to the th camera. Specifically, taking the human head as an object, the attitude angle can be calculated through hopenet.

[0059] Based on this, the transformation matrices between all cameras can be obtained, thereby automatically completing the calibration of the cameras without the need for manual calibration.

[0060] Specifically, the internal parameter matrix of each camera can also be obtained through the three-dimensional coordinates of the feature points relative to each camera. The internal parameter matrix is as follows: , where represents the internal parameter matrix, represents the th pixel of the camera, represents the focal length. The calculation formula of the internal parameter matrix is: , where represents the three-dimensional coordinates of the th feature point relative to the th camera, represents the ratio of the physical true value to the pixel of the th camera, and the focal length in the internal parameter matrix is calculated accordingly.

[0061] For the multi-cameras after the above positioning, the present invention also provides a multi-camera positioning method for supporting use, including the following steps:

[0062] S01. Select two cameras, denoted as the first camera and the second camera respectively, and obtain the internal parameter matrix of the first camera and the homogeneous transformation matrix between the first camera and the second camera;

[0063] S02. Use the two cameras to photograph the object to be positioned, and obtain the pixel coordinates of the feature points of the object to be positioned in the first camera and the second camera;

[0064] S03. The three-dimensional coordinates of the th feature point relative to the first camera are . Through the formula , calculate , where represents the pixel coordinates of the feature point in the first camera, represents the internal parameter matrix of the first camera; the three-dimensional coordinates of the th feature point relative to the second camera are . Through the formula , calculate , where represents the pixel coordinates of the feature point in the second camera, represents the internal parameter matrix of the second camera;

[0065] S04. Substitute and into , and solve to obtain the three-dimensional coordinates of the th feature point relative to the first camera or the three-dimensional coordinates of the Represents the homogeneous transformation matrix between the first camera and the second camera.

[0066] Specifically, if the world coordinate system is established with the first camera or the second camera as the coordinate origin, then the world coordinates of each feature point can be determined therefrom.

[0067] If the world coordinate system is established with a coordinate origin other than the first camera or the second camera, then coordinate transformation needs to be performed using the corresponding homogeneous transformation matrix, which specifically includes the following steps:

[0068] S05. Establish the world coordinate system with a certain camera as the reference;

[0069] S06. Use the homogeneous transformation matrix between the first camera and the certain camera to convert the three-dimensional coordinates of the feature point relative to the first camera to the coordinates of the feature point relative to the world coordinate system, and obtain the world coordinates of the object to be located. The coordinate transformation formula is , where represents the homogeneous transformation matrix between the first camera and the certain camera;

[0070] Alternatively, use the homogeneous transformation matrix between the second camera and the certain camera to convert the three-dimensional coordinates of the feature point relative to the second camera to the coordinates of the feature point relative to the world coordinate system, and obtain the world coordinates of the object to be located. The coordinate transformation formula is , where represents the homogeneous transformation matrix between the second camera and the certain camera.

Claims

1. A multi-camera calibration method, characterized in that: The following steps are involved: S1. Determine the calibration object; S2, using multiple cameras to respectively capture videos containing the calibration object, and extracting image sequences corresponding to the videos to obtain image sequences corresponding to each camera; S3, extracting features of the calibration object from the image, wherein the features include semantic features and pixel coordinates of multiple feature points; S4, using a neural network model to obtain similar images between cameras from the feature map sequence pair through semantic features, wherein the similar images include corresponding images of the calibrated object in each camera at the same time; S5, using the pixel coordinates of multiple feature points in similar images in each camera, and based on the fact that the real distances between the feature points are equal, calculating the three-dimensional coordinates of the feature points relative to each camera, including the following steps: S51, record the nth feature point as P n , P n The pixel coordinates in the i-th camera are denoted by: (h ni , w ni ), then P n The two-dimensional coordinates in the shooting plane of the i-th camera are: (h ni k i , w ni k i ), where k i Indicates the ratio of the real size and pixel size captured by the i-th camera; S52, P n The two-dimensional coordinates in the shooting plane of the i-th camera are expanded into three-dimensional coordinates to obtain P n The three-dimensional coordinates relative to the i-th camera are recorded as: P ni =(h ni k i , w ni k i , Z ni ), where Z ni Indicates P n The depth relative to the i-th camera, so as to obtain the three-dimensional coordinates of each feature point relative to each camera; S53, based on the equal real distances between feature points, establish a distance equality relationship: ||P ni -P mi ||=||P nj -P mj ||, where P mi represents the 3D coordinates of the mth feature point relative to the i-th camera, P nj represents the 3D coordinates of the nth feature point relative to the jth camera, P mj represents the 3D coordinates of the mth feature point relative to the jth camera, and || || represents the distance calculation symbol; S54, using the true distances between multiple feature points to be equal, obtain the corresponding equation group, and solve Z ni , obtain the three-dimensional coordinates of each feature point; S6. Obtain the transformation matrix between cameras through the three-dimensional coordinates of the feature points relative to each camera, and complete the camera calibration.

2. The multi-camera calibration method according to claim 1, characterized in that: The semantic features include object size, object posture, object local information, object global information and object similarity between adjacent frames.

3. The multi-camera calibration method according to claim 1, characterized in that: In S3, the features of the calibration object are extracted from the image, including: determining the position of the calibration object in the image through a detection module, and obtaining the features of the calibration object in the image through a feature extraction model.

4. The multi-camera calibration method according to claim 1, characterized in that: In S4, the similar pictures are pictures with the highest similarity or pictures with a similarity higher than a threshold.

5. The multi-camera calibration method according to claim 1, characterized in that: The calibration object is a human body, and the multiple feature points are multiple of human body feature points, and the human body feature points include nose, neck, right shoulder, left shoulder, right elbow, left elbow, right wrist, left wrist, right hip, left hip, right knee, left knee, right ankle, left ankle, right eye, left eye, right ear, left ear, right toe and left toe.

6. The multi-camera calibration method according to claim 1, characterized in that: In S6, the transformation matrix between the cameras is obtained by the three-dimensional coordinates of the feature points relative to each camera, including the following steps: S61, P n Subtract P from the 3D coordinates of the i-th camera n Relative to the three-dimensional coordinates of the j-th camera, obtain the three-dimensional coordinates of the ith camera relative to the j-th camera; S62. Combining the rotation matrices of the ith camera and the jth camera, obtain a homogeneous transformation matrix of the ith camera and the jth camera, wherein the homogeneous transformation matrix is: Among them, T ij represents the homogeneous transformation matrix between the i-th camera and the j-th camera, R ij Represents the rotation matrix between the i-th camera and the j-th camera, R ij =R z (γ)× R y (β)×R x (α), (γ, β, α) represents the attitude angle, γ represents the rotation angle around the Z axis, β represents the rotation angle around the Y axis, and α represents the rotation angle around the X axis. The matrix transpose representing the 3D coordinates of the i-th camera relative to the j-th camera.

7. The multi-camera calibration method according to claim 1, characterized in that: S6 also includes obtaining the intrinsic parameter matrix of each camera through the three-dimensional coordinates of the feature points relative to each camera. The intrinsic parameter matrix is: in, represents the internal parameter matrix, (h i , w i ) represents the pixel of the i-th camera, f represents the focal length, and the calculation formula of the intrinsic parameter matrix is: Among them, (h ni k i , w ni k i , Z ni ) represents the 3D coordinates of the nth feature point relative to the ith camera, k i It represents the ratio of the real size to the pixel size captured by the i-th camera.

8. A multi-camera positioning method, characterized in that: Applied to the calibrated camera obtained by the multi-camera calibration method according to claim 1, the positioning method comprises the following steps: S01, select two cameras, respectively denoted as a first camera and a second camera, and obtain an intrinsic parameter matrix of the first camera and a homogeneous transformation matrix between the first camera and the second camera; S02, using two cameras to shoot the object to be located, and obtaining the pixel coordinates of the feature points of the object to be located in the first camera and the second camera; S03. The three-dimensional coordinates of the nth feature point relative to the first camera are (X n1 , Y n1 , Z n1 ), through the formula calculate Among them, (h n1 , w n1 ) represents the pixel coordinates of the feature point in the first camera, represents the intrinsic parameter matrix of the first camera; the three-dimensional coordinates of the nth feature point relative to the second camera are (X n2 , Y n2 , Z n2 ), through the formula calculate Among them, (h n2 , w n2 ) represents the pixel coordinates of the feature point in the second camera, represents the intrinsic parameter matrix of the second camera; S04, will and Substitute into the formula Solve to obtain the three-dimensional coordinates (X n1 , Y n1 , Z n1 ) or the three-dimensional coordinates (X n2 , Y n2 , Z n2 ), where T 12 Represents the homogeneous transformation matrix between the first camera and the second camera.

9. The multi-camera positioning method according to claim 8, characterized in that: The following steps are also included: S05. Establish a world coordinate system based on a certain camera; S06. Using the homogeneous transformation matrix between the first camera and the one of the cameras, the three-dimensional coordinates of the feature points relative to the first camera are converted into the coordinates of the feature points relative to the world coordinate system to obtain the world coordinates of the object to be located; or, using the homogeneous transformation matrix between the second camera and the one of the cameras, the three-dimensional coordinates of the feature points relative to the second camera are converted into the coordinates of the feature points relative to the world coordinate system to obtain the world coordinates of the object to be located.

Citation Information

Patent Citations

  • Autonomous mobile platform environment perception and mapping method based on bionics

    CN112509051A

  • Human factor information processing method and device based on target object recognition and electronic equipment

    CN118053093A