Camera calibration methods, devices and related equipment

By collecting video data to generate point cloud model maps and performing semantic segmentation and feature extraction, the problem of poor camera calibration flexibility is solved, enabling real-time monitoring and automatic calibration of camera status in large scenes, applicable to various camera positions and heights.

CN120525966BActive Publication Date: 2025-10-28北京数原数字化城市研究中心
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510524857.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-10-28
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

Existing technologies have poor camera calibration flexibility, especially in large scenes with a large number of cameras and possible changes in their positions, making it difficult to effectively adjust calibration parameters.

Method used

By collecting video data of the target scene, a point cloud model is generated, semantic segmentation and feature extraction are performed, wall and ground areas are identified, initial parameters are generated, and the standard model is aligned through mapping relationships to correct camera parameters to adapt to various camera positions and heights.

Benefits of technology

It improves the flexibility of camera calibration, is applicable to various camera positions and heights, and enables real-time monitoring and automatic calibration of camera status in large scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120525966B_ABST
    Figure CN120525966B_ABST
Patent Text Reader

Abstract

This application provides a camera calibration method, apparatus, and related equipment, relating to the field of camera calibration. The method includes: acquiring video data of a target scene; generating a point cloud model of the target scene; aligning the point cloud model with a standard model of the target scene to obtain a scene point cloud model; generating initial parameters for the target camera; performing coordinate calculations on the captured image and the scene point cloud model to obtain a mapping relationship; and correcting the initial parameters according to the mapping relationship to obtain calibration parameters for the target camera. This application, after acquiring an initial point cloud map of the target scene, extracts the ground and wall areas to generate a point cloud model, aligns the point cloud model with a standard model, thereby obtaining a mapping relationship, and corrects the camera parameters according to the mapping relationship. This method is applicable to various camera positions and heights, improving the flexibility of camera calibration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of camera calibration, and more particularly to a camera calibration method, apparatus, and related equipment. Background Technology

[0002] For urban sensing, cameras are indispensable front-end sensing devices. Specifically, in scenarios such as subways, train stations, and shopping malls, camera calibration is unavoidable when performing comprehensive pedestrian and vehicle positioning and counting analyses. However, these large-scale scenarios involve a large number of cameras, and their status may need to be adjusted at any time, requiring recalibration. Current technologies generally rely on calibration objects with known spatial information for camera calibration. This method is limited by the camera's installation location and height, easily leading to poor flexibility in camera calibration. Summary of the Invention

[0003] This application provides a camera calibration method, apparatus, and related equipment to solve the problem of poor camera calibration flexibility in the prior art.

[0004] To solve the above problems, this application is implemented as follows:

[0005] In a first aspect, embodiments of this application provide a camera calibration method, the method comprising:

[0006] Collect video data of the target scene, wherein the video data is video of the walls, ground and other objects in the target scene;

[0007] The video data is used to perform 3D modeling and point cloud estimation to obtain an initial point cloud map of the target scene;

[0008] The video data is semantically segmented and feature extracted to identify wall and ground areas in the video data. Based on the wall and ground areas, a point cloud model of the target scene is generated, including wall point cloud and ground point cloud.

[0009] Align the point cloud model with the standard model of the target scene to obtain the scene point cloud model, wherein the standard model is the architectural standard image of the interior walls and ground of the target scene;

[0010] Camera parameters are calculated from the images captured by the target camera to generate initial parameters for the target camera. The initial parameters include the intrinsic and extrinsic parameters of the target camera, and the target camera is the camera that captured the target scene.

[0011] Coordinate calculations are performed on the captured image and the scene point cloud model to obtain a mapping relationship, which is the coordinate transformation relationship between the pixel points of the captured image and the coordinate points of the scene point cloud model.

[0012] The initial parameters are corrected according to the mapping relationship to obtain the calibration parameters of the target parameters.

[0013] Optionally, aligning the point cloud model diagram with the standard model diagram of the target scene to obtain the scene point cloud model diagram includes:

[0014] The ground equation is determined based on the ground point cloud, and the ground equation is used to represent the coordinate position of the ground point.

[0015] The Z-axis coordinate in the ground equation is adjusted to 0 by rotation, translation, and / or scaling transformation;

[0016] The point cloud on the wall is projected onto the first plane to obtain a two-dimensional point cloud projection map;

[0017] The two-dimensional point cloud projection map is aligned with the standard model map of the target scene by rotation, translation and / or scaling transformation to obtain the scene point cloud model map.

[0018] Optionally, before estimating camera parameters from the images captured by the target camera to generate the initial parameters of the target camera, the method further includes:

[0019] Obtain multiple standard background images corresponding to multiple cameras within the target scene, wherein each camera corresponds one-to-one with the multiple standard background images, and the standard background images are images of the target scene taken by the corresponding camera;

[0020] Multiple real-time images are acquired, wherein the multiple real-time images are pictures taken by the multiple cameras, and the multiple real-time images correspond one-to-one with the multiple cameras;

[0021] For each camera, the similarity between the standard background image and the real-time image corresponding to each camera is calculated to obtain multiple similarity values;

[0022] The camera corresponding to the target similarity value is determined as the target camera, and the target similarity value is the similarity value that is greater than a preset threshold among the plurality of similarity values.

[0023] Optionally, the step of calculating camera parameters from the images captured by the target camera to generate initial parameters for the target camera includes:

[0024] The camera parameters of the target camera are calculated from the images captured by the target camera to obtain the intrinsic parameters of the target camera;

[0025] Feature extraction is performed on the image captured by the target camera to obtain key points and feature descriptors, wherein the feature descriptors are used to describe the key points;

[0026] Based on the key points and the feature descriptors, coordinate transformation is performed to obtain the extrinsic parameters of the target camera;

[0027] Based on the intrinsic and extrinsic parameters, the initial parameters of the target camera are generated.

[0028] Optionally, the step of calculating the coordinates of the captured image and the scene point cloud model to obtain the mapping relationship includes:

[0029] The optical center image coordinates, pitch angle, and heading angle of the target camera are determined based on the intrinsic and extrinsic parameters.

[0030] The first coordinate of the target coordinate point in the captured image is determined based on the intrinsic and extrinsic parameters, where the target coordinate point is the point in the captured image that is in contact with the ground.

[0031] Calculate the second coordinates of the target point in the scene point cloud model based on the camera imaging principle, the optical center image coordinates, the pitch angle, and the heading angle;

[0032] The first coordinate and the second coordinate are used to calculate the coordinates to determine the mapping relationship.

[0033] Optionally, the step of calculating the second coordinates of the target point in the scene point cloud model based on the camera imaging principle, the optical center image coordinates, the pitch angle, and the heading angle includes:

[0034] The Y-axis distance is calculated using Wy = H × tan(β + γ), where β = atan((cy - Q1y) ÷ fy), γ is the pitch angle, (cx, cy) are the image coordinates of the optical center, fx / fy is the camera's intrinsic focal length, and the camera's extrinsic parameters include a 3x3 rotation matrix R and a 3x1 translation matrix t, where t is the camera's physical world coordinates (Xc, Yc, Zc). T Then Zc is the camera height H. The 3x3 rotation matrix R is converted into rotation angles Ax / Ay / Az around the X / Y / Z axes. The pitch angle between the camera optical center and the vertical line is 180+Ax, and the yaw angle is -Ay.

[0035] according to Calculate the distance along the X-axis, where

[0036] The second coordinate (Wx', Wy') is obtained based on the Y-axis coordinate distance Wy and the X-axis coordinate distance Wx, where,

[0037] Secondly, embodiments of this application also provide a camera calibration device, comprising:

[0038] The acquisition module is used to acquire video data of the target scene, wherein the video data is video of the walls, ground and other objects in the target scene;

[0039] An estimation module is used to perform 3D modeling and point cloud estimation on the video data to obtain an initial point cloud map of the target scene;

[0040] The recognition module is used to perform semantic segmentation and feature extraction on the video data, identify the wall area and the ground area in the video data, and generate a point cloud model map of the target scene based on the wall area and the ground area. The point cloud model map includes wall point clouds and ground point clouds. The alignment module is used to align the point cloud model map with the standard model map of the target scene to obtain a scene point cloud model map. The standard model map is the architectural standard image of the wall and ground in the target scene.

[0041] The calculation module is used to calculate camera parameters for the images captured by the target camera and generate initial parameters for the target camera. The initial parameters include the intrinsic and extrinsic parameters of the target camera, and the target camera is the camera that captures the target scene.

[0042] The calculation module is used to perform coordinate calculations on the captured image and the scene point cloud model to obtain a mapping relationship, which is the coordinate transformation relationship between the pixel points of the captured image and the coordinate points of the scene point cloud model.

[0043] The correction module is used to correct the initial parameters according to the mapping relationship to obtain the calibration parameters of the target parameters.

[0044] Thirdly, embodiments of this application also provide an electronic device, including: a transceiver, a memory, a processor, and a program stored in the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps in the method described in the first aspect above.

[0045] Fourthly, embodiments of this application also provide a readable storage medium for storing a program, which, when executed by a processor, implements the steps of the method described in the first aspect above.

[0046] Fifthly, embodiments of this application also provide a computer program product, which is stored in a storage medium and executed by at least one processor to implement the steps in the method described in the first aspect above.

[0047] This application provides a camera calibration method, apparatus, and related equipment, relating to the field of camera calibration. The method includes: acquiring video data of a target scene; generating a point cloud model of the target scene; aligning the point cloud model with a standard model of the target scene to obtain a scene point cloud model; generating initial parameters for the target camera; performing coordinate calculations on the captured image and the scene point cloud model to obtain a mapping relationship; and correcting the initial parameters according to the mapping relationship to obtain calibration parameters for the target camera. This application, after acquiring an initial point cloud map of the target scene, extracts the ground and wall areas to generate a point cloud model, aligns the point cloud model with a standard model, thereby obtaining a mapping relationship, and corrects the camera parameters according to the mapping relationship. This method is applicable to various camera positions and heights, improving the flexibility of camera calibration. Attached Figure Description

[0048] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 This is a schematic flowchart of the camera calibration method provided in the embodiments of this application;

[0050] Figure 2 This is one of the coordinate diagrams provided in the embodiments of this application;

[0051] Figure 3 This is the second coordinate diagram provided in the embodiments of this application;

[0052] Figure 4 This is a schematic diagram of the camera calibration device provided in the embodiments of this application;

[0053] Figure 5 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0054] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0055] The terms "first," "second," etc., used in the embodiments of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. Additionally, the use of "and / or" in this application indicates at least one of the connected objects, such as A and / or B and / or C, representing seven possibilities: including A alone, B alone, C alone, and the presence of both A and B, both B and C, both A and C, and the presence of A, B, and C.

[0056] See Figure 1 , Figure 1 This is a schematic flowchart of the camera calibration method provided in the embodiments of this application. Figure 1 The camera calibration method shown can be performed by a terminal or a server.

[0057] like Figure 1 As shown, the camera calibration method may include the following steps:

[0058] Step 101: Collect video data of the target scene, wherein the video data is video of the walls, ground and other objects in the target scene.

[0059] In this embodiment, the target scene can be a street, the interior of a building, or a room, etc. For example, in a large surveillance scene (such as a train station or shopping mall), there are dozens or even hundreds of fixed surveillance cameras. Video data within the target scene can be collected using other recording devices. For example, when collecting video data of the entire scene area using a mobile phone or panoramic camera, the focus is on collecting images within the range that each surveillance camera can illuminate, thereby generating video data of the target scene. The collection principle is to cover the scene as completely as possible, focusing on collecting images within the possible shooting range of each camera. Based on the collected scene video data, the scene is reconstructed to generate a 3D point cloud model.

[0060] Step 102: Perform 3D modeling and point cloud estimation on the video data to obtain the initial point cloud map of the target scene.

[0061] In this embodiment, the purpose of establishing the scene point cloud model is mainly for camera calibration. In this embodiment, 3D reconstruction is performed using a structure for motion recovery (SFM) to obtain an initial point cloud map of the scene. It should be noted that the initial point cloud map is typically used to represent objects and structures in the environment. This point cloud data consists of a large number of 3D coordinate points, each of which typically contains position (x, y, z) and possible color information (such as RGB values).

[0062] Step 103: Perform semantic segmentation and feature extraction on the video data to identify the wall area and ground area in the video data, and generate a point cloud model of the target scene based on the wall area and ground area. The point cloud model includes wall point cloud and ground point cloud.

[0063] In this embodiment, OneFormer is used to perform semantic segmentation and feature extraction on the video data to identify the wall area and the ground area in the video data, thereby estimating the ground point cloud and the wall point cloud based on the ground and wall areas and the point cloud model.

[0064] Step 104: Align the point cloud model with the standard model of the target scene to obtain the scene point cloud model. The standard model is the architectural standard image of the walls and ground in the target scene.

[0065] In this embodiment, the standard model diagram can be a CAD image. By combining the scene ground CAD diagram, the ground point cloud is rotated to be horizontally aligned with the ground CAD. After the wall point cloud is projected onto the ground, it is rotated, translated, and scaled to coincide with the room layout in the CAD diagram, thus obtaining the transformation T of rotation, translation, and scale information from the point cloud model to the CAD diagram. After the initial point cloud undergoes transformation T, the scene point cloud model diagram used for camera calibration is obtained.

[0066] Specifically, for a medium to large-scale monitoring scenario, a 3D point cloud model of the scene is first established, which is consistent with the CAD coordinate system. Then, the status of the monitoring cameras in the scene is monitored in real time. When a change in the status of a camera in the scene is detected, the camera's intrinsic and extrinsic parameters are automatically calibrated in real time based on the 3D point cloud model of the scene without manual intervention.

[0067] Step 105: Perform camera parameter calculation on the image captured by the target camera to generate the initial parameters of the target camera. The initial parameters include the intrinsic and extrinsic parameters of the target camera. The target camera is the camera that captures the target scene.

[0068] In this embodiment, camera parameters are calculated for the target camera requiring calibration. Specifically, key points and feature descriptors are extracted from the camera image using the ALIKED method. These descriptors are then matched with the feature descriptors of each point in the scene point cloud model using the RANSAC method. Nonlinear optimization is then performed to obtain the rotation and translation relationships from the camera to the CAD coordinate system. Finally, the camera intrinsic parameters are estimated using the matched 2D-3D point pairs.

[0069] Step 106: Calculate the coordinates of the captured image and the scene point cloud model to obtain a mapping relationship. The mapping relationship is the coordinate transformation relationship between the pixel points of the captured image and the coordinate points of the scene point cloud model.

[0070] In this embodiment, the OneFormer semantic segmentation model is used to identify the ground and wall regions in the image. Based on the 2D-3D correspondence in the SfM sparse map, the ground and wall point clouds in the SFM sparse point cloud are then identified. Finally, the transformation T from the sparse point cloud to the CAD coordinate system, i.e., the mapping relationship, is obtained.

[0071] Step 107: Correct the initial parameters according to the mapping relationship to obtain the calibration parameters of the target parameters.

[0072] In this embodiment, the obtained initial parameters are corrected by using the obtained mapping relationship, that is, the intrinsic and extrinsic parameters are corrected to obtain the calibration parameters of the target parameters, thereby completing the calibration of the target camera.

[0073] This application generates a point cloud model by extracting the ground and wall areas after acquiring the initial point cloud map of the target scene. The point cloud model is then aligned with a standard model to obtain a mapping relationship. Based on this mapping relationship, the camera parameters are corrected, which is applicable to various camera positions and heights, thus improving the flexibility of camera calibration.

[0074] In some feasible implementations, optionally, aligning the point cloud model diagram with the standard model diagram of the target scene to obtain the scene point cloud model diagram includes:

[0075] The ground equation is determined based on the ground point cloud, and the ground equation is used to represent the coordinate position of the ground point.

[0076] The Z-axis coordinate in the ground equation is adjusted to 0 by rotation, translation, and / or scaling transformation;

[0077] The point cloud on the wall is projected onto the first plane to obtain a two-dimensional point cloud projection map;

[0078] The two-dimensional point cloud projection map is aligned with the standard model map of the target scene by rotation, translation and / or scaling transformation to obtain the scene point cloud model map.

[0079] In this embodiment, the equation of the ground is estimated using ground point cloud, and the point cloud is rotated and translated so that the ground equation is Z=0, thereby aligning the coordinate system of the point cloud model with the CAD coordinate system horizontally; the wall point cloud is projected parallel onto the ground plane to obtain a 2D point cloud projection map, and then manually rotated, translated, and scaled to align the 2D point cloud projection map with the content in the CAD.

[0080]

[0081] Where s is the scalar scaling, R is a 3x3 rotation matrix, and t is a 3x1 vector. The reconstructed map from SFM... Figure 1 By applying this transformation once, the SFM coordinate system can be aligned with the CAD coordinate system. The extrinsic parameters obtained during subsequent camera calibration will then be the extrinsic parameters in the CAD coordinate system.

[0082] Optionally, before estimating camera parameters from the images captured by the target camera to generate the initial parameters of the target camera, the method further includes:

[0083] Obtain multiple standard background images corresponding to multiple cameras within the target scene, wherein each camera corresponds one-to-one with the multiple standard background images, and the standard background images are images of the target scene taken by the corresponding camera;

[0084] Multiple real-time images are acquired, wherein the multiple real-time images are pictures taken by the multiple cameras, and the multiple real-time images correspond one-to-one with the multiple cameras;

[0085] For each camera, the similarity between the standard background image and the real-time image corresponding to each camera is calculated to obtain multiple similarity values;

[0086] The camera corresponding to the target similarity value is determined as the target camera, and the target similarity value is the similarity value that is greater than a preset threshold among the plurality of similarity values.

[0087] In this embodiment, calculations are performed on multiple cameras within the target scene to determine whether camera calibration is required, i.e., whether the multiple cameras are target cameras. Specifically, multiple standard background images corresponding to each of the multiple cameras are acquired. Images currently captured by the multiple cameras are acquired within the same time interval, and the similarity between the currently captured images and the standard background images is calculated to determine whether camera calibration is required.

[0088] Specifically, a set of preset background images for a fixed state is initialized for each camera in the scene. OneFormer is then used to perform background segmentation on the real-time camera images to obtain the real-time background images. Feature extraction and similarity calculation are performed on both the preset background images and the real-time background images. When the duration of a state with a similarity below a threshold Th is greater than d, the camera state is considered to have changed and entered a new stable state. At this point, parameter estimation based on 3D point clouds is performed on the camera.

[0089] Optionally, the step of calculating camera parameters from the images captured by the target camera to generate initial parameters for the target camera includes:

[0090] The camera parameters of the target camera are calculated from the images captured by the target camera to obtain the intrinsic parameters of the target camera;

[0091] Feature extraction is performed on the image captured by the target camera to obtain key points and feature descriptors, wherein the feature descriptors are used to describe the key points;

[0092] Based on the key points and the feature descriptors, coordinate transformation is performed to obtain the extrinsic parameters of the target camera;

[0093] Based on the intrinsic and extrinsic parameters, the initial parameters of the target camera are generated.

[0094] In this embodiment, the camera parameters are first calculated. Specifically, key points and feature descriptors are extracted from the camera image using the ALIKED method. These descriptors are then matched with the feature descriptors of each point in the scene point cloud model using the RANSAC method. Nonlinear optimization is then performed to obtain the rotation and translation relationships from the camera to the CAD coordinate system. Finally, the camera intrinsic parameters are estimated using the matched 2D-3D point pairs.

[0095] End-to-end mapping from pixel coordinates to CAD coordinates. Based on the monocular positioning theory, an end-to-end mapping from pixel coordinates to CAD coordinates is constructed using the camera's intrinsic and extrinsic parameters.

[0096] Furthermore, by using the OneFormer semantic segmentation model to identify ground and wall regions in the image, and based on the 2D-3D correspondence in the SfM sparse map, ground and wall point clouds in the SFM sparse point cloud are further identified. The equation of the ground is estimated using the ground point cloud, and the point cloud is rotated and translated until the ground equation is Z=0, thus horizontally aligning the point cloud model coordinate system with the CAD coordinate system. The wall point cloud is then projected parallel onto the ground plane to obtain a 2D point cloud projection image, which is then manually rotated, translated, and scaled to align the 2D point cloud projection image with the content in the CAD. Finally, the transformation T from the sparse point cloud to the CAD coordinate system is obtained.

[0097]

[0098] Where s is the scalar scaling, R is a 3x3 rotation matrix, and t is a 3x1 vector. The reconstructed map from SFM... Figure 1 By applying this transformation once, the SFM coordinate system can be aligned with the CAD coordinate system. The extrinsic parameters obtained during subsequent camera calibration will then be the extrinsic parameters in the CAD coordinate system.

[0099] Optionally, the step of calculating the coordinates of the captured image and the scene point cloud model to obtain the mapping relationship includes:

[0100] The optical center image coordinates, pitch angle, and heading angle of the target camera are determined based on the intrinsic and extrinsic parameters.

[0101] The first coordinate of the target coordinate point in the captured image is determined based on the intrinsic and extrinsic parameters, where the target coordinate point is the point in the captured image that is in contact with the ground.

[0102] Calculate the second coordinates of the target point in the scene point cloud model based on the camera imaging principle, the optical center image coordinates, the pitch angle, and the heading angle;

[0103] The first coordinate and the second coordinate are used to calculate the coordinates to determine the mapping relationship.

[0104] In this embodiment, the optical center image coordinates, pitch angle, and yaw angle of the target camera are determined using the camera's intrinsic and extrinsic parameters. The first coordinates of the target point in the captured image are then calculated, and the second coordinates of the target point in the scene point cloud model are calculated using the camera's imaging principle. The camera imaging principle mainly involves the processes of light capture, focusing, and recording. Coordinate calculations are performed using the first and second coordinates to establish a mapping relationship, which is then used to update and correct the camera's parameters.

[0105] Optionally, the step of calculating the second coordinates of the target point in the scene point cloud model based on the camera imaging principle, the optical center image coordinates, the pitch angle, and the heading angle includes:

[0106] The Y-axis distance is calculated using Wy = H × tan(β + γ), where β = atan((cy - Q1y) ÷ fy), γ is the pitch angle, (cx, cy) are the image coordinates of the optical center, fx / fy is the camera's intrinsic focal length, and the camera's extrinsic parameters include a 3x3 rotation matrix R and a 3x1 translation matrix t, where t is the camera's physical world coordinates (Xc, Yc, Zc). T Then Zc is the camera height H. The 3x3 rotation matrix R is converted into rotation angles Ax / Ay / Az around the X / Y / Z axes. The pitch angle between the camera optical center and the vertical line is 180+Ax, and the yaw angle is -Ay.

[0107] according to Calculate the distance along the X-axis, where

[0108] The second coordinate (Wx', Wy') is obtained based on the Y-axis coordinate distance Wy and the X-axis coordinate distance Wx, where,

[0109] In this embodiment, specifically, the camera intrinsic parameters obtained through camera parameter calculation refer to the focal lengths fx and fy, representing the pixel size per unit dimension in the horizontal and vertical directions, and the optical center image coordinates (cx, cy), respectively. The camera extrinsic parameters include a 3x3 rotation matrix R and a 3x1 translation matrix t, where t is the camera's physical world coordinates (Xc, Yc, Zc)T, and Zc is the camera height H. The 3x3 rotation matrix R is converted into rotation angles Ax / Ay / Az around the X / Y / Z axes. According to the definition in the positioning principle, the pitch angle between the camera's optical center and the vertical line is γ = 180 + Ax, and the yaw angle is θ = -Ay.

[0110] Then, based on the camera imaging principle, the world coordinates of the calibration point on the horizontal ground can be calculated from the pixel coordinates. The specific method is as follows: The two figures below contain three coordinate systems: the image coordinate system UO1V, the camera coordinate system with O2 as the origin, and the world coordinate system XO3Y. For example, a point in the world coordinate system has projections onto the X and Y axes of the world coordinate system of Q / P, corresponding to pixel points Q1 / P1 in the image. Given the pixel coordinates Q1 / P1 of the target point, its world coordinates Q / P are calculated.

[0111] like Figure 2 and Figure 3 As shown, Figure 2 and Figure 3 As shown in the coordinate diagram of this embodiment, firstly, the coordinates of point Q in the Y-axis direction are calculated, which is the distance d from P to O3 in the figure.

[0112] Wy = H × tan(β + γ)

[0113] Where β = atan((cy-Q1y)÷fy), γ is the pitch angle, and (cx,cy) is the image coordinates of the optical center.

[0114] Calculate the coordinates of point Q along the X-axis, which is the distance from point Q to point P in the diagram below. in

[0115] Based on the camera's heading angle θ and the camera's position in planar coordinates (Xc, Yc), rotate and translate (Wx, Wy) to obtain the final world coordinates (Wx', Wy').

[0116] Wx'=Wx×cosθ+Wy×sinθ+Mx

[0117] Wy'=Wy×cosθ-Wx×sinθ+My

[0118] The method provided in this application enables real-time monitoring of camera status in large-scale surveillance scenarios that require analysis of targets such as people and vehicles. When a change in camera status is detected, the method automatically and instantly completes camera calibration to ensure positioning accuracy.

[0119] This application generates a point cloud model by extracting the ground and wall areas after acquiring the initial point cloud map of the target scene. The point cloud model is then aligned with a standard model to obtain a mapping relationship. Based on this mapping relationship, the camera parameters are corrected, which is applicable to various camera positions and heights, thus improving the flexibility of camera calibration.

[0120] See Figure 4 , Figure 4 This is a structural diagram of the camera calibration device provided in an embodiment of this application. Figure 4 As shown, the camera calibration device 400 includes:

[0121] Acquisition module 410 is used to acquire video data of the target scene, wherein the video data is video of the walls, ground and other objects in the target scene;

[0122] The estimation module 420 is used to perform three-dimensional modeling and point cloud estimation on the video data to obtain an initial point cloud map of the target scene;

[0123] The recognition module 430 is used to perform semantic segmentation and feature extraction on the video data, identify the wall area and the ground area in the video data, and generate a point cloud model map of the target scene based on the wall area and the ground area, wherein the point cloud model map includes wall point cloud and ground point cloud;

[0124] Alignment module 440 is used to align the point cloud model map with the standard model map of the target scene to obtain a scene point cloud model map, wherein the standard model map is a standard architectural image of the interior walls and ground of the target scene;

[0125] The calculation module 450 is used to calculate camera parameters for the images captured by the target camera and generate initial parameters for the target camera. The initial parameters include the intrinsic and extrinsic parameters of the target camera, and the target camera is the camera that captures the target scene.

[0126] The calculation module 460 is used to perform coordinate calculation on the captured image and the scene point cloud model to obtain a mapping relationship, wherein the mapping relationship is the coordinate transformation relationship between the pixel points of the captured image and the coordinate points of the scene point cloud model.

[0127] The correction module 470 is used to correct the initial parameters according to the mapping relationship to obtain the calibration parameters of the target parameters.

[0128] Optionally, the alignment module 440 includes:

[0129] The first determining submodule is used to determine the ground equation based on the ground point cloud, wherein the ground equation is used to represent the coordinate position of the ground point;

[0130] An adjustment submodule is used to adjust the Z-axis coordinate in the ground equation to 0 through rotation, translation, and / or scaling transformation;

[0131] The projection submodule is used to project the wall point cloud onto the first plane to obtain a two-dimensional point cloud projection map.

[0132] The alignment submodule is used to align the two-dimensional point cloud projection map with the standard model map of the target scene through rotation, translation and / or scaling transformation to obtain the scene point cloud model map.

[0133] Optionally, also include:

[0134] The first acquisition module is used to acquire multiple standard background images corresponding to multiple cameras in the target scene, wherein the multiple cameras correspond one-to-one with the multiple standard background images, and the standard background images are images of the target scene taken by the corresponding cameras.

[0135] The second acquisition module is used to acquire multiple real-time images, which are pictures taken by the multiple cameras, and the multiple real-time images correspond one-to-one with the multiple cameras;

[0136] The similarity calculation module is used to calculate the similarity between the standard background image and the real-time image corresponding to each camera for each camera, and obtain multiple similarity values;

[0137] The determination module is used to determine the camera corresponding to the target similarity value as the target camera, wherein the target similarity value is the similarity value among the plurality of similarity values ​​that is greater than a preset threshold.

[0138] Optionally, the solver module 450 includes:

[0139] The calculation module is used to calculate the camera parameters of the images captured by the target camera to obtain the intrinsic parameters of the target camera;

[0140] The extraction module is used to extract features from the image captured by the target camera to obtain key points and feature descriptors, wherein the feature descriptors are used to describe the key points;

[0141] The transformation module is used to perform coordinate transformation based on the key points and the feature descriptors to obtain the extrinsic parameters of the target camera;

[0142] The generation module is used to generate the initial parameters of the target camera based on the intrinsic and extrinsic parameters.

[0143] Optionally, the computing module 460 includes:

[0144] The second determining submodule is used to determine the optical center image coordinates, pitch angle, and heading angle of the target camera based on the intrinsic and extrinsic parameters.

[0145] The third determining submodule is used to determine the first coordinate of the target coordinate point in the captured image based on the internal and external parameters, wherein the target coordinate point is the point in the captured image that is in contact with the ground;

[0146] The calculation submodule is used to calculate the second coordinates of the target coordinate point in the scene point cloud model map based on the camera imaging principle, the optical center image coordinates, the pitch angle and the heading angle;

[0147] The fourth determining submodule is used to perform coordinate calculations on the first coordinate and the second coordinate to determine the mapping relationship.

[0148] Optionally, the computation submodule includes:

[0149] The first calculation unit is used to calculate the Y-axis coordinate distance according to Wy=H×tan(β+γ), where β=atan((cy-Q1y)÷fy), γ is the pitch angle, (cx,cy) are the image coordinates of the optical center, fx / fy is the camera intrinsic focal length, and the camera extrinsic parameters include a 3x3 rotation matrix R and a 3x1 translation matrix t, where t is the camera's physical world coordinates (Xc, Yc, Zc). T Then Zc is the camera height H. The 3x3 rotation matrix R is converted into rotation angles Ax / Ay / Az around the X / Y / Z axes. The pitch angle between the camera optical center and the vertical line is 180+Ax, and the yaw angle is -Ay.

[0150] The second calculation unit is used to calculate based on Calculate the distance along the X-axis, where

[0151] The third calculation unit is used to obtain the second coordinate (Wx', Wy') based on the Y-axis coordinate distance Wy and the X-axis coordinate distance Wx, wherein...

[0152] This application generates a point cloud model by extracting the ground and wall areas after acquiring the initial point cloud map of the target scene. The point cloud model is then aligned with a standard model to obtain a mapping relationship. Based on this mapping relationship, the camera parameters are corrected, which is applicable to various camera positions and heights, thus improving the flexibility of camera calibration.

[0153] This application also provides an electronic device. Please refer to [link to relevant documentation]. Figure 5 The electronic device may include a processor 501, a memory 502, and a program 5021 stored in the memory 502 and capable of running on the processor 501.

[0154] When program 5021 is executed by processor 501, it can achieve the following: Figure 1 Any step in the corresponding method embodiment:

[0155] Collect video data of the target scene, wherein the video data is video of the walls, ground and other objects in the target scene;

[0156] The video data is used to perform 3D modeling and point cloud estimation to obtain an initial point cloud map of the target scene;

[0157] The video data is semantically segmented and feature extracted to identify wall and ground areas in the video data. Based on the wall and ground areas, a point cloud model of the target scene is generated, including wall point cloud and ground point cloud.

[0158] Align the point cloud model with the standard model of the target scene to obtain the scene point cloud model, wherein the standard model is the architectural standard image of the interior walls and ground of the target scene;

[0159] Camera parameters are calculated from the images captured by the target camera to generate initial parameters for the target camera. The initial parameters include the intrinsic and extrinsic parameters of the target camera, and the target camera is the camera that captured the target scene.

[0160] Coordinate calculations are performed on the captured image and the scene point cloud model to obtain a mapping relationship, which is the coordinate transformation relationship between the pixel points of the captured image and the coordinate points of the scene point cloud model.

[0161] The initial parameters are corrected according to the mapping relationship to obtain the calibration parameters of the target parameters.

[0162] Optionally, aligning the point cloud model diagram with the standard model diagram of the target scene to obtain the scene point cloud model diagram includes:

[0163] The ground equation is determined based on the ground point cloud, and the ground equation is used to represent the coordinate position of the ground point.

[0164] The Z-axis coordinate in the ground equation is adjusted to 0 by rotation, translation, and / or scaling transformation;

[0165] The point cloud on the wall is projected onto the first plane to obtain a two-dimensional point cloud projection map;

[0166] The two-dimensional point cloud projection map is aligned with the standard model map of the target scene by rotation, translation and / or scaling transformation to obtain the scene point cloud model map.

[0167] Optionally, before estimating camera parameters from the images captured by the target camera to generate the initial parameters of the target camera, the method further includes:

[0168] Obtain multiple standard background images corresponding to multiple cameras within the target scene, wherein each camera corresponds one-to-one with the multiple standard background images, and the standard background images are images of the target scene taken by the corresponding camera;

[0169] Multiple real-time images are acquired, wherein the multiple real-time images are pictures taken by the multiple cameras, and the multiple real-time images correspond one-to-one with the multiple cameras;

[0170] For each camera, the similarity between the standard background image and the real-time image corresponding to each camera is calculated to obtain multiple similarity values;

[0171] The camera corresponding to the target similarity value is determined as the target camera, and the target similarity value is the similarity value that is greater than a preset threshold among the plurality of similarity values.

[0172] Optionally, the step of calculating camera parameters from the images captured by the target camera to generate initial parameters for the target camera includes:

[0173] The camera parameters of the target camera are calculated from the images captured by the target camera to obtain the intrinsic parameters of the target camera;

[0174] Feature extraction is performed on the image captured by the target camera to obtain key points and feature descriptors, wherein the feature descriptors are used to describe the key points;

[0175] Based on the key points and the feature descriptors, coordinate transformation is performed to obtain the extrinsic parameters of the target camera;

[0176] Based on the intrinsic and extrinsic parameters, the initial parameters of the target camera are generated.

[0177] Optionally, the step of calculating the coordinates of the captured image and the scene point cloud model to obtain the mapping relationship includes:

[0178] The optical center image coordinates, pitch angle, and heading angle of the target camera are determined based on the intrinsic and extrinsic parameters.

[0179] The first coordinate of the target coordinate point in the captured image is determined based on the intrinsic and extrinsic parameters, where the target coordinate point is the point in the captured image that is in contact with the ground.

[0180] Calculate the second coordinates of the target point in the scene point cloud model based on the camera imaging principle, the optical center image coordinates, the pitch angle, and the heading angle;

[0181] The first coordinate and the second coordinate are used to calculate the coordinates to determine the mapping relationship.

[0182] Optionally, the step of calculating the second coordinates of the target point in the scene point cloud model based on the camera imaging principle, the optical center image coordinates, the pitch angle, and the heading angle includes:

[0183] The Y-axis distance is calculated using Wy = H × tan(β + γ), where β = atan((cy - Q1y) ÷ fy), γ is the pitch angle, (cx, cy) are the image coordinates of the optical center, fx / fy is the camera's intrinsic focal length, and the camera's extrinsic parameters include a 3x3 rotation matrix R and a 3x1 translation matrix t, where t is the camera's physical world coordinates (Xc, Yc, Zc). T Then Zc is the camera height H. The 3x3 rotation matrix R is converted into rotation angles Ax / Ay / Az around the X / Y / Z axes. The pitch angle between the camera optical center and the vertical line is 180+Ax, and the yaw angle is -Ay.

[0184] according to Calculate the distance along the X-axis, where

[0185] The second coordinate (Wx', Wy') is obtained based on the Y-axis coordinate distance Wy and the X-axis coordinate distance Wx, where,

[0186] This application generates a point cloud model by extracting the ground and wall areas after acquiring the initial point cloud map of the target scene. The point cloud model is then aligned with a standard model to obtain a mapping relationship. Based on this mapping relationship, the camera parameters are corrected, which is applicable to various camera positions and heights, thus improving the flexibility of camera calibration.

[0187] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described camera calibration method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0188] This application also provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described camera calibration method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0189] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0190] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0191] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A camera calibration method, characterized in that, The method includes: Collect video data of the target scene, wherein the video data is video of the walls, ground and other objects in the target scene; The video data is used to perform 3D modeling and point cloud estimation to obtain an initial point cloud map of the target scene; The video data is semantically segmented and feature extracted to identify wall and ground areas in the video data. Based on the wall and ground areas, a point cloud model of the target scene is generated, including wall point cloud and ground point cloud. Align the point cloud model with the standard model of the target scene to obtain the scene point cloud model, wherein the standard model is the architectural standard image of the interior walls and ground of the target scene; Camera parameters are calculated from the images captured by the target camera to generate initial parameters for the target camera. The initial parameters include the intrinsic and extrinsic parameters of the target camera, and the target camera is the camera that captured the target scene. Coordinate calculations are performed on the captured image and the scene point cloud model to obtain a mapping relationship. This mapping relationship represents the coordinate transformation between the pixels of the captured image and the coordinates of the points in the scene point cloud model. The step of calculating the coordinates of the captured image and the scene point cloud model to obtain the mapping relationship includes: The optical center image coordinates, pitch angle, and heading angle of the target camera are determined based on the intrinsic and extrinsic parameters. The first coordinate of the target coordinate point in the captured image is determined based on the intrinsic and extrinsic parameters, where the target coordinate point is the point in the captured image that is in contact with the ground. Calculate the second coordinates of the target point in the scene point cloud model based on the camera imaging principle, the optical center image coordinates, the pitch angle, and the heading angle; Perform coordinate calculations on the first coordinate and the second coordinate to determine the mapping relationship; The step of calculating the second coordinates of the target point in the scene point cloud model based on the camera imaging principle, the optical center image coordinates, the pitch angle, and the heading angle includes: The Y-axis distance is calculated using Wy = H × tan(β + γ), where β = atan((cy - Q1y) ÷ fy), γ is the pitch angle, (cx, cy) are the image coordinates of the optical center, fx / fy is the camera's intrinsic focal length, and the camera's extrinsic parameters include a 3x3 rotation matrix R and a 3x1 translation matrix t, where t is the camera's physical world coordinates (Xc, Yc, Zc). T Then Zc is the camera height H. The 3x3 rotation matrix R is converted into rotation angles Ax / Ay / Az around the X / Y / Z axes. The pitch angle between the camera optical center and the vertical line is 180+Ax, and the yaw angle is -Ay. according to Calculate the distance along the X-axis, where The second coordinate (Wx', Wy') is obtained based on the Y-axis coordinate distance Wy and the X-axis coordinate distance Wx, where, The initial parameters are corrected according to the mapping relationship to obtain the calibration parameters of the target parameters.

2. The method according to claim 1, characterized in that, Aligning the point cloud model with the standard model of the target scene to obtain the scene point cloud model includes: The ground equation is determined based on the ground point cloud, and the ground equation is used to represent the coordinate position of the ground point. The Z-axis coordinate in the ground equation is adjusted to 0 by rotation, translation, and / or scaling transformation; The point cloud on the wall is projected onto the first plane to obtain a two-dimensional point cloud projection map; The two-dimensional point cloud projection map is aligned with the standard model map of the target scene by rotation, translation and / or scaling transformation to obtain the scene point cloud model map.

3. The method according to claim 1, characterized in that, Before estimating camera parameters from images captured by the target camera to generate initial parameters for the target camera, the method further includes: Obtain multiple standard background images corresponding to multiple cameras within the target scene, wherein each camera corresponds one-to-one with the multiple standard background images, and the standard background images are images of the target scene taken by the corresponding camera; Multiple real-time images are acquired, wherein the multiple real-time images are pictures taken by the multiple cameras, and the multiple real-time images correspond one-to-one with the multiple cameras; For each camera, the similarity between the standard background image and the real-time image corresponding to each camera is calculated to obtain multiple similarity values; The camera corresponding to the target similarity value is determined as the target camera, and the target similarity value is the similarity value that is greater than a preset threshold among the plurality of similarity values.

4. The method according to claim 1, characterized in that, The step of calculating camera parameters from images captured by the target camera to generate initial parameters for the target camera includes: The camera parameters of the target camera are calculated from the images captured by the target camera to obtain the intrinsic parameters of the target camera; Feature extraction is performed on the image captured by the target camera to obtain key points and feature descriptors, wherein the feature descriptors are used to describe the key points; Based on the key points and the feature descriptors, coordinate transformation is performed to obtain the extrinsic parameters of the target camera; Based on the intrinsic and extrinsic parameters, the initial parameters of the target camera are generated.

5. A camera calibration device, characterized in that, The device includes: The acquisition module is used to acquire video data of the target scene, wherein the video data is video of the walls, ground and other objects in the target scene; An estimation module is used to perform 3D modeling and point cloud estimation on the video data to obtain an initial point cloud map of the target scene; The recognition module is used to perform semantic segmentation and feature extraction on the video data, identify the wall area and the ground area in the video data, and generate a point cloud model map of the target scene based on the wall area and the ground area. The point cloud model map includes wall point cloud and ground point cloud. The alignment module is used to align the point cloud model with the standard model of the target scene to obtain a scene point cloud model, wherein the standard model is a standard architectural image of the walls and ground of the target scene. The calculation module is used to calculate camera parameters for the images captured by the target camera and generate initial parameters for the target camera. The initial parameters include the intrinsic and extrinsic parameters of the target camera, and the target camera is the camera that captures the target scene. The calculation module is used to perform coordinate calculations on the captured image and the scene point cloud model to obtain a mapping relationship, which is the coordinate transformation relationship between the pixel points of the captured image and the coordinate points of the scene point cloud model. The computing module includes: The second determining submodule is used to determine the optical center image coordinates, pitch angle, and heading angle of the target camera based on the intrinsic and extrinsic parameters. The third determining submodule is used to determine the first coordinate of the target coordinate point in the captured image based on the internal and external parameters, wherein the target coordinate point is the point in the captured image that is in contact with the ground; The calculation submodule is used to calculate the second coordinates of the target coordinate point in the scene point cloud model map based on the camera imaging principle, the optical center image coordinates, the pitch angle and the heading angle; The fourth determining submodule is used to perform coordinate calculations on the first coordinate and the second coordinate to determine the mapping relationship; The computational submodule includes: The first calculation unit is used to calculate the Y-axis coordinate distance according to Wy=H×tan(β+γ), where β=atan((cy-Q1y)÷fy), γ is the pitch angle, (cx,cy) are the image coordinates of the optical center, fx / fy is the camera intrinsic focal length, and the camera extrinsic parameters include a 3x3 rotation matrix R and a 3x1 translation matrix t, where t is the camera's physical world coordinates (Xc, Yc, Zc). T Then Zc is the camera height H. The 3x3 rotation matrix R is converted into rotation angles Ax / Ay / Az around the X / Y / Z axes. The pitch angle between the camera optical center and the vertical line is 180+Ax, and the yaw angle is -Ay. The second calculation unit is used to calculate based on Calculate the distance along the X-axis, where The third calculation unit is used to obtain the second coordinate (Wx', Wy') based on the Y-axis coordinate distance Wy and the X-axis coordinate distance Wx, wherein... The correction module is used to correct the initial parameters according to the mapping relationship to obtain the calibration parameters of the target parameters.

6. An electronic device, comprising: A memory, a processor, and a program stored in the memory and executable on the processor; characterized in that the processor is configured to read the program in the memory to implement the steps in the camera calibration method as claimed in any one of claims 1 to 4.

7. A readable storage medium for storing a program, characterized in that, When the program is executed by the processor, it implements the steps in the camera calibration method as described in any one of claims 1 to 4.

8. A computer program product, characterized in that, The computer program product is stored in a storage medium and is executed by at least one processor to implement the steps in the camera calibration method as claimed in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Internal and external parameter calibration method and device and electronic equipment

    CN115661262A

  • Sensing equipment calibration method and device and computer equipment

    CN119540366A