Multi-camera calibration method and positioning method

Through the multi-camera calibration method, the transformation matrix calculation between cameras is automatically completed by using neural network model and feature points, which solves the cumbersome problem of manual calibration of multi-view visual positioning in the existing technology, and achieves a more efficient and accurate positioning process.

CN119991829AActive Publication Date: 2025-05-13CHENGDU CHANGHONG NETWORK TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510465923.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The existing multi-view visual positioning method requires manual calibration of the camera's internal and external parameters, which is cumbersome.

Method used

Through the multi-camera calibration method, the transformation matrix between the cameras is automatically calculated using the neural network model and the three-dimensional coordinate calculation of feature points to complete the camera calibration.

Benefits of technology

Automatic calibration of multiple cameras is realized, reducing the complexity and time of manual operation, and improving positioning accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991829A_ABST
    Figure CN119991829A_ABST
Patent Text Reader

Abstract

The invention provides a multi-camera calibration method and a multi-camera positioning method, and relates to an image processing technology, and the method comprises the steps: determining a calibration object, collecting videos containing the object through a plurality of cameras, obtaining a picture sequence corresponding to each camera, and extracting the features of the calibration object from the pictures, the features comprise semantic features and pixel coordinates of a plurality of feature points, similar pictures between the cameras are obtained from the feature map sequence pairs through the semantic features by utilizing a neural network model, and the feature points in the similar pictures are utilized to obtain the feature map sequence pairs based on the fact that the real distances between the feature points are equal. The three-dimensional coordinates of the feature points relative to each camera are calculated, the transformation matrix between the cameras is obtained through the three-dimensional coordinates of the feature points relative to each camera, calibration of the cameras is completed, the problem that manual calibration is needed in existing multi-view-angle positioning is solved, and the method is suitable for multi-camera automatic calibration and positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to image processing technology, and in particular to a multi-camera calibration method and a positioning method. Background Art

[0002] At present, because location services have increasingly higher requirements for location accuracy and real-time performance, visual positioning is an important development direction in the field of computer vision.

[0003] The two cameras at different angles in multi-view visual positioning can complement each other's dimensional information perpendicular to the camera plane that is lost during the shooting process, and can correct each other's calculation results, providing support for calculating the three-dimensional size of the object and improving the accuracy and robustness of the measurement.

[0004] The cost of home surveillance cameras is getting lower and lower, and for security and anti-theft reasons, more and more families are installing multiple surveillance cameras and connecting them to home IoT terminals. Therefore, under the conditions of home IoT, it is possible to use home surveillance cameras to identify and locate indoor objects.

[0005] However, the existing multi-view visual positioning often requires manual calibration of camera internal and external parameters based on the calibration algorithm with auxiliary tools before use, and the calibration process is relatively cumbersome. Summary of the invention

[0006] The technical problem solved by the present invention is to provide a multi-camera calibration method and a positioning method to solve the problem that the existing multi-view positioning requires manual calibration.

[0007] The present invention solves the above technical problems by adopting a technical solution: a multi-camera calibration method, comprising the following steps: S1. Determine the calibration object; S2, using multiple cameras to respectively capture videos containing the calibration object, and extracting image sequences corresponding to the videos to obtain image sequences corresponding to each camera; S3, extracting features of the calibration object from the image, wherein the features include semantic features and pixel coordinates of multiple feature points; S4, using a neural network model to obtain similar images between cameras from the feature map sequence pair through semantic features, wherein the similar images include corresponding images of the calibrated object in each camera at the same time; S5, using the pixel coordinates of multiple feature points in similar images in each camera, and based on the fact that the actual distances between the feature points are equal, calculating the three-dimensional coordinates of the feature points relative to each camera; S6. Obtain the transformation matrix between cameras through the three-dimensional coordinates of the feature points relative to each camera, and complete the camera calibration.

[0008] Furthermore, the semantic features include object size, object posture, object local information, object global information and similarity of the object between adjacent frames.

[0009] Furthermore, in S3, features of the calibration object are extracted from the image, including: determining the position of the calibration object in the image through a detection module, and obtaining the features of the calibration object in the image through a feature extraction model.

[0010] Further, in S4, the similar pictures are pictures with the highest similarity or pictures with a similarity higher than a threshold.

[0011] Furthermore, the calibration object is a human body, and the multiple feature points are multiple of human body feature points, and the human body feature points include nose, neck, right shoulder, left shoulder, right elbow, left elbow, right wrist, left wrist, right hip, left hip, right knee, left knee, right ankle, left ankle, right eye, left eye, right ear, left ear, right toe and left toe.

[0012] Furthermore, S5 includes the following steps: S51, The feature points are recorded as , In the The pixel coordinates in each camera are: ,but In the The two-dimensional coordinates of the camera in the shooting plane are: ,in, Indicates The ratio of the real size to the pixel size captured by each camera; S52, will In the The two-dimensional coordinates in the shooting plane of the camera are expanded into three-dimensional coordinates to obtain Relative to the The three-dimensional coordinates of a camera are recorded as: ,in express Relative to the The depth of each camera is used to obtain the three-dimensional coordinates of each feature point relative to each camera; S53, based on the equal real distances between feature points, establish a distance equality relationship: ,in, Indicates The feature points are relative to the The 3D coordinates of the cameras, Indicates The feature points are relative to the The 3D coordinates of the cameras, Indicates The feature points are relative to the The 3D coordinates of the cameras, Indicates the distance calculation symbol; S54, using the true distances between multiple feature points to obtain the corresponding equations, and solving , obtain the three-dimensional coordinates of each feature point.

[0013] Furthermore, in S6, the transformation matrix between the cameras is obtained through the three-dimensional coordinates of the feature points relative to each camera, including the following steps: S61, will Relative to the The 3D coordinates of the cameras are subtracted Relative to the The three-dimensional coordinates of the camera are obtained The camera is relative to the The 3D coordinates of the cameras; S62, combined with The camera and The rotation matrix of the camera is obtained The camera and The homogeneous transformation matrix of the camera is: ,in, Indicates The camera and The homogeneous transformation matrix of the camera is Indicates The camera and The rotation matrix of the camera, , , , , represents the attitude angle, Indicates winding The rotation angle of the axis, Indicates winding The angle of rotation of the axis, Indicates winding The rotation angle of the axis, Indicates The camera is relative to the The matrix transpose of the 3D coordinates of the cameras.

[0014] Furthermore, S6 also includes obtaining the intrinsic parameter matrix of each camera through the three-dimensional coordinates of the feature points relative to each camera, and the intrinsic parameter matrix is: ,in, represents the internal parameter matrix, Indicates The pixels of the camera, Represents the focal length, and the calculation formula of the internal parameter matrix is: ,in, Indicates The feature points are relative to the The 3D coordinates of the cameras, Indicates The ratio of the actual size to the pixel size captured by a camera.

[0015] The present invention also provides a multi-camera positioning method, which is applied to the calibrated cameras obtained by the multi-camera calibration method as described above, and the positioning method comprises the following steps: S01, select two cameras, respectively denoted as a first camera and a second camera, and obtain an intrinsic parameter matrix of the first camera and a homogeneous transformation matrix between the first camera and the second camera; S02, using two cameras to shoot the object to be located, and obtaining the pixel coordinates of the feature points of the object to be located in the first camera and the second camera; S03, No. The three-dimensional coordinates of the feature points relative to the first camera are , through the formula ,calculate ,in, represents the pixel coordinates of the feature point in the first camera, represents the intrinsic parameter matrix of the first camera; The three-dimensional coordinates of the feature points relative to the second camera are , through the formula ,calculate ,in, represents the pixel coordinates of the feature point in the second camera, represents the intrinsic parameter matrix of the second camera; S04, will and Substitution , and solve for The 3D coordinates of the feature points relative to the first camera Or The 3D coordinates of the feature points relative to the second camera ,in, Represents the homogeneous transformation matrix between the first camera and the second camera.

[0016] Furthermore, the multi-camera positioning method further includes the following steps: S05. Establish a world coordinate system based on a certain camera; S06. Using the homogeneous transformation matrix between the first camera and the one of the cameras, the three-dimensional coordinates of the feature points relative to the first camera are converted into the coordinates of the feature points relative to the world coordinate system to obtain the world coordinates of the object to be located; or, using the homogeneous transformation matrix between the second camera and the one of the cameras, the three-dimensional coordinates of the feature points relative to the second camera are converted into the coordinates of the feature points relative to the world coordinate system to obtain the world coordinates of the object to be located.

[0017] Beneficial effects of the present invention: The present invention provides a multi-camera calibration method and a positioning method, which determines a calibration object, uses multiple cameras to respectively capture videos containing the object, obtains a picture sequence corresponding to each camera, extracts features of the calibration object from the picture, and the features include semantic features and pixel coordinates of multiple feature points. A neural network model is used to obtain similar pictures between cameras from feature picture sequence pairs through semantic features, and the similar pictures include corresponding pictures of the calibration object in each camera at the same time. The pixel coordinates of multiple feature points in the similar pictures in each camera are used, and based on the equal real distances between the feature points, the three-dimensional coordinates of the feature points relative to each camera are calculated. The transformation matrix between the cameras is obtained through the three-dimensional coordinates of the feature points relative to each camera, and the camera calibration is completed, which solves the problem that the existing multi-view positioning needs to be manually calibrated, and the calibrated camera is used to locate the object to be located. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flow chart of a multi-camera calibration method provided by the present invention. DETAILED DESCRIPTION

[0019] The present invention aims at the problem that multi-view positioning needs manual calibration, and provides a multi-camera calibration method. Figure 1 As shown, the following steps are included: S1. Determine the calibration object.

[0020] Specifically, the calibration object is a human body or an object.

[0021] S2, using multiple cameras to respectively capture videos containing the calibration object, and extracting image sequences corresponding to the videos to obtain image sequences corresponding to each camera; Specifically, since a video is composed of continuous pictures, the video can be split into picture sequences by frame.

[0022] S3. Extracting features of the calibration object from the image, wherein the features include semantic features and pixel coordinates of multiple feature points.

[0023] Specifically, extracting the features of the calibration object from the image includes: determining the position of the calibration object in the image through a detection module, and obtaining the features of the calibration object in the image through a feature extraction model. The detection module is yolo5. The semantic features refer to high-level features extracted from the image that can represent the image content and are converted into a high-dimensional space, such as object size, object posture, object local information, object global information, and the similarity of the object between adjacent frames.

[0024] If the object is a human body, the plurality of feature points are a plurality of human body feature points, and the human body feature points include nose, neck, right shoulder, left shoulder, right elbow, left elbow, right wrist, left wrist, right hip, left hip, right knee, left knee, right ankle, left ankle, right eye, left eye, right ear, left ear, right toe and left toe. The feature extraction model is openpose.

[0025] S4. Using a neural network model, similar images between cameras are obtained from a pair of feature image sequences through semantic features, wherein the similar images include corresponding images of the calibrated object in each camera at the same time.

[0026] Specifically, due to the different shooting speeds and shooting environments of multiple cameras during the shooting process, there may also be frame drops, image occlusions and other phenomena caused by unstable factors such as the environment and network. Even if the starting time of the shooting is the same, the obtained video image sequence is often not a one-to-one correspondence. Therefore, it is necessary to perform time alignment of the video image sequence, that is, to obtain similar images between cameras from the feature image sequence pairs. The similar images are the feature pairs with the highest similarity or the feature pairs with a similarity higher than a threshold, so as to ensure time alignment, that is, the corresponding images of the calibrated object in each camera at the same time.

[0027] S5. Using the pixel coordinates of multiple feature points in similar images in each camera, and based on the fact that the actual distances between the feature points are equal, calculate the three-dimensional coordinates of the feature points relative to each camera.

[0028] Specifically, S5 includes the following steps: S51, The feature points are recorded as , In the The pixel coordinates in each camera are: ,but In the The two-dimensional coordinates of the camera in the shooting plane are: ,in, Indicates The ratio of the actual size to the pixel size captured by a camera.

[0029] S52, will In the The two-dimensional coordinates in the shooting plane of the camera are expanded into three-dimensional coordinates to obtain Relative to the The three-dimensional coordinates of a camera are recorded as: ,in express Relative to the The depth of each camera is used to obtain the three-dimensional coordinates of each feature point relative to each camera.

[0030] S53, based on the equal real distances between feature points, establish a distance equality relationship: ,in, Indicates The feature points are relative to the The 3D coordinates of the cameras, Indicates The feature points are relative to the The 3D coordinates of the cameras, Indicates The feature points are relative to the The 3D coordinates of the cameras, Represents the distance calculation symbol.

[0031] S54, using the true distances between multiple feature points to obtain the corresponding equations, and solving , obtain the three-dimensional coordinates of each feature point.

[0032] For example, taking five feature points and two cameras as an example, after expanding to three-dimensional coordinates, 10 third-dimensional parameters are added. The actual distance between the five feature points remains unchanged, and the corresponding 10 distance equality formulas can be obtained. Therefore, the values ​​of the 10 third-dimensional parameters can be solved to obtain the three-dimensional coordinates of each feature point.

[0033] S6. Obtain the transformation matrix between cameras through the three-dimensional coordinates of the feature points relative to each camera, and complete the camera calibration.

[0034] Specifically, in S6, the transformation matrix between the cameras is obtained through the three-dimensional coordinates of the feature points relative to each camera, including the following steps: S61, will Relative to the The 3D coordinates of the cameras are subtracted Relative to the The three-dimensional coordinates of the camera are obtained The camera is relative to the The 3D coordinates of the cameras; S62, combined with The camera and The rotation matrix of the camera is obtained The camera and The homogeneous transformation matrix of the camera is: ,in, Indicates The camera and The rotation matrix of the camera, , , , , represents the attitude angle, Indicates winding The rotation angle of the axis, Indicates winding The angle of rotation of the axis, Indicates winding The rotation angle of the axis, Indicates The camera is relative to the The matrix transpose of the three-dimensional coordinates of each camera. In particular, taking the human head as the object, the posture angle can be calculated by Hopenet.

[0035] Based on this, the transformation matrix between all cameras can be obtained, so that the camera calibration is automatically completed, and manual calibration is no longer required.

[0036] In particular, the intrinsic parameter matrix of each camera can be obtained through the three-dimensional coordinates of the feature points relative to each camera. The intrinsic parameter matrix is: ,in, represents the internal parameter matrix, Indicates The pixels of the camera, Represents the focal length, and the calculation formula of the internal parameter matrix is: ,in, Indicates The feature points are relative to the The 3D coordinates of the cameras, Indicates The ratio of the physical real value of each camera to the pixel is used to calculate the focal length in the intrinsic parameter matrix.

[0037] The present invention further provides a multi-camera positioning method for use with the above-mentioned positioned multi-cameras, comprising the following steps: S01, select two cameras, respectively denoted as a first camera and a second camera, and obtain an intrinsic parameter matrix of the first camera and a homogeneous transformation matrix between the first camera and the second camera; S02, using two cameras to shoot the object to be located, and obtaining the pixel coordinates of the feature points of the object to be located in the first camera and the second camera; S03, No. The three-dimensional coordinates of the feature points relative to the first camera are , through the formula ,calculate ,in, represents the pixel coordinates of the feature point in the first camera, represents the intrinsic parameter matrix of the first camera; The three-dimensional coordinates of the feature points relative to the second camera are , through the formula ,calculate ,in, represents the pixel coordinates of the feature point in the second camera, represents the intrinsic parameter matrix of the second camera; S04, will and Substitution , and solve for The 3D coordinates of the feature points relative to the first camera Or The 3D coordinates of the feature points relative to the second camera ,in, Represents the homogeneous transformation matrix between the first camera and the second camera.

[0038] Specifically, if a world coordinate system is established with the first camera or the second camera as the coordinate origin, the world coordinates of each feature point can be determined accordingly.

[0039] If the world coordinate system is established with a non-first camera or the second camera as the coordinate origin, it is also necessary to use the corresponding homogeneous transformation matrix to perform coordinate conversion, which specifically includes the following steps: S05. Establish a world coordinate system based on a certain camera; S06. Using the homogeneous transformation matrix between the first camera and the certain camera, the three-dimensional coordinates of the feature point relative to the first camera are converted into the coordinates of the feature point relative to the world coordinate system to obtain the world coordinates of the object to be located. The coordinate conversion formula is: ,in, represents a homogeneous transformation matrix between the first camera and the certain camera; Alternatively, the homogeneous transformation matrix between the second camera and the one of the cameras is used to transform the three-dimensional coordinates of the feature point relative to the second camera into the coordinates of the feature point relative to the world coordinate system to obtain the world coordinates of the object to be located. The coordinate transformation formula is: ,in, represents the homogeneous transformation matrix between the second camera and the certain camera.

Claims

1. A multi-camera calibration method, characterized in that: The following steps are involved: S1. Determine the calibration object; S2, using multiple cameras to respectively capture videos containing the calibration object, and extracting image sequences corresponding to the videos to obtain image sequences corresponding to each camera; S3, extracting features of the calibration object from the image, wherein the features include semantic features and pixel coordinates of multiple feature points; S4, using a neural network model to obtain similar images between cameras from the feature map sequence pair through semantic features, wherein the similar images include corresponding images of the calibrated object in each camera at the same time; S5, using the pixel coordinates of multiple feature points in similar images in each camera, and based on the fact that the actual distances between the feature points are equal, calculating the three-dimensional coordinates of the feature points relative to each camera; S6. Obtain the transformation matrix between cameras through the three-dimensional coordinates of the feature points relative to each camera, and complete the camera calibration.

2. The multi-camera calibration method according to claim 1, characterized in that: The semantic features include object size, object posture, object local information, object global information and object similarity between adjacent frames.

3. The multi-camera calibration method according to claim 1, characterized in that: In S3, the features of the calibration object are extracted from the image, including: determining the position of the calibration object in the image through a detection module, and obtaining the features of the calibration object in the image through a feature extraction model.

4. The multi-camera calibration method according to claim 1, characterized in that: In S4, the similar pictures are pictures with the highest similarity or pictures with a similarity higher than a threshold.

5. The multi-camera calibration method according to claim 1, characterized in that: The calibration object is a human body, and the multiple feature points are multiple of human body feature points, and the human body feature points include nose, neck, right shoulder, left shoulder, right elbow, left elbow, right wrist, left wrist, right hip, left hip, right knee, left knee, right ankle, left ankle, right eye, left eye, right ear, left ear, right toe and left toe.

6. The multi-camera calibration method according to claim 5, characterized in that: S5 includes the following steps: S51, The feature points are recorded as , In the The pixel coordinates in each camera are: ,but In the The two-dimensional coordinates of the camera in the shooting plane are: ,in, Indicates The ratio of the real size to the pixel size captured by each camera; S52, will In the The two-dimensional coordinates in the shooting plane of the camera are expanded into three-dimensional coordinates to obtain Relative to the The three-dimensional coordinates of a camera are recorded as: ,in express Relative to the The depth of each camera is used to obtain the three-dimensional coordinates of each feature point relative to each camera; S53, based on the equal real distances between feature points, establish a distance equality relationship: ,in, Indicates The feature points are relative to the The 3D coordinates of the cameras, Indicates The feature points are relative to the The 3D coordinates of the cameras, Indicates The feature points are relative to the The 3D coordinates of the cameras, Indicates the distance calculation symbol; S54, using the true distances between multiple feature points to obtain the corresponding equations, and solving , obtain the three-dimensional coordinates of each feature point.

7. The multi-camera calibration method according to claim 6, characterized in that: In S6, the transformation matrix between the cameras is obtained by the three-dimensional coordinates of the feature points relative to each camera, including the following steps: S61, will Relative to the The 3D coordinates of the cameras are subtracted Relative to the The three-dimensional coordinates of the camera are obtained The camera is relative to the The 3D coordinates of the cameras; S62, combined with The camera and The rotation matrix of the camera is obtained The camera and The homogeneous transformation matrix of the camera is: ,in, Indicates The camera and The homogeneous transformation matrix of the camera is Indicates The camera and The rotation matrix of the camera, , , , , represents the attitude angle, Indicates winding The rotation angle of the axis, Indicates winding The angle of rotation of the axis, Indicates winding The rotation angle of the axis, Indicates The camera is relative to the The matrix transpose of the 3D coordinates of the cameras.

8. The multi-camera calibration method according to claim 6, characterized in that: S6 also includes obtaining the intrinsic parameter matrix of each camera through the three-dimensional coordinates of the feature points relative to each camera. The intrinsic parameter matrix is: ,in, represents the internal parameter matrix, Indicates The pixels of the camera, Represents the focal length, and the calculation formula of the internal parameter matrix is: ,in, Indicates The feature points are relative to the The 3D coordinates of the cameras, Indicates The ratio of the actual size to the pixel size captured by a camera.

9. A multi-camera positioning method, characterized in that: Applied to the calibrated camera obtained by the multi-camera calibration method according to claim 1, the positioning method comprises the following steps: S01, select two cameras, respectively denoted as a first camera and a second camera, and obtain an intrinsic parameter matrix of the first camera and a homogeneous transformation matrix between the first camera and the second camera; S02, using two cameras to shoot the object to be located, and obtaining the pixel coordinates of the feature points of the object to be located in the first camera and the second camera; S03, No. The three-dimensional coordinates of the feature points relative to the first camera are , through the formula ,calculate ,in, represents the pixel coordinates of the feature point in the first camera, represents the intrinsic parameter matrix of the first camera; The three-dimensional coordinates of the feature points relative to the second camera are , through the formula ,calculate ,in, represents the pixel coordinates of the feature point in the second camera, represents the intrinsic parameter matrix of the second camera; S04, will and Substitution , and solve for The 3D coordinates of the feature points relative to the first camera Or The 3D coordinates of the feature points relative to the second camera ,in, Represents the homogeneous transformation matrix between the first camera and the second camera.

10. The multi-camera positioning method according to claim 9, characterized in that: The following steps are also included: S05. Establish a world coordinate system based on a certain camera; S06. Using the homogeneous transformation matrix between the first camera and the one of the cameras, the three-dimensional coordinates of the feature points relative to the first camera are converted into the coordinates of the feature points relative to the world coordinate system to obtain the world coordinates of the object to be located; or, using the homogeneous transformation matrix between the second camera and the one of the cameras, the three-dimensional coordinates of the feature points relative to the second camera are converted into the coordinates of the feature points relative to the world coordinate system to obtain the world coordinates of the object to be located.

Citation Information

Patent Citations

  • Convolution neural network-based approach to target pose recognition for robot control

    CN109344882A

  • System and method for efficiently controlling robot

    CN109910010A

  • Cow weight prediction system

    CN110415282A

  • Three-dimensional distance measurement method for adaptively adjusting baseline of binocular camera

    CN111402315A

  • Autonomous mobile platform environment perception and mapping method based on bionics

    CN112509051A