Camera calibration and camera target detection and localization methods, devices and electronic equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHONGQING CHANGAN AUTOMOBILE CO LTD
- Filing Date
- 2022-12-13
- Publication Date
- 2026-05-26
AI Technical Summary
[0006]本申请提供一种相机标定和相机的目标检测与定位方法、装置及电子设备,以解决相关技术中机器人搭载的标定板距离相机较远,拍摄出的图像中标定板的区域小,导致计算出的相机内参精度低,且仅应用于室内场景,有一定的局限性等问题
[0026] 1. In this embodiment, when the movement of the target device being captured by the camera is detected, the image coordinates of the midpoint of the lower edge of the detection frame and the set of timestamps are recorded. At the same time, the set of world coordinates of the mobile device within the same timestamp is obtained. Based on the set of image coordinates and the corresponding set of world coordinates, the coordinate transformation matrix of the image coordinates and world coordinates of each camera is calculated and optimized to obtain the coordinate transformation matrix of the camera to be calibrated, thereby completing the camera calibration. This can be applied to both indoor and outdoor traffic scenarios. By optimizing the coordinate transformation matrix, the camera calibration error is reduced.
Smart Images

Figure CN115830142B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent video surveillance technology, and in particular to a camera calibration and a method, apparatus and electronic device for camera target detection and positioning. Background Technology
[0002] With social progress and development, the application of camera devices is becoming more and more widespread. Currently, many scenes or areas use a large number of camera devices to monitor the scene or area, and then use computer vision technology to detect and track moving targets in the scene by retrieving the images collected by the camera devices.
[0003] Calibrate camera devices to obtain accurate measurement results. In practical applications, without calibrated camera parameters, it is impossible to obtain the true trajectory of the target in the actual environment, leading to target tracking failure and reduced controllability. Therefore, camera calibration is essential. Currently, monocular camera devices are the most widely used in surveillance scenarios; however, camera calibration requires manual operation, consuming significant manpower and time.
[0004] The relevant technology involves using a robot equipped with a calibration board to perform SLAM indoors, while the camera to be calibrated takes pictures and detects the calibration board. Using multiple frames of images, the intrinsic parameters of the camera are calibrated according to Zhang Zhengyou's calibration algorithm, and the extrinsic parameters of the camera are solved according to the ICP or PNP algorithm, thereby completing the calibration of the intrinsic and extrinsic parameters of the camera.
[0005] However, in related technologies, intrinsic parameter calibration usually requires the calibration board to occupy most of the image area in order to obtain high calibration accuracy. However, in actual monitoring scenarios, the calibration board carried by the robot is far away from the camera, and the calibration board only covers a small area in the captured image. Therefore, the accuracy of the calculated camera intrinsic parameters will be reduced. The solution of extrinsic parameters depends on the intrinsic parameters, which will cause the accumulation of errors. Moreover, it is only applicable to indoor scenarios and has limitations. Summary of the Invention
[0006] This application provides a camera calibration and a method, apparatus and electronic device for camera target detection and localization, to solve the problems in related technologies such as the calibration board carried by the robot being far away from the camera, the calibration board area being small in the captured image, resulting in low accuracy of the calculated camera intrinsic parameters, and the limitation of only being applicable to indoor scenes.
[0007] The first aspect of this application provides a camera calibration method, comprising the following steps: acquiring a first set of image coordinates and timestamps of a target device acquired by one or more cameras, and acquiring a second set of the position and timestamps of the target device itself in a three-dimensional map; identifying a set of world coordinates in the second set corresponding to the set of image coordinates in the first set at any timestamp; calculating a coordinate transformation matrix between the image coordinates and world coordinates of each camera based on the set of image coordinates and the set of world coordinates; determining the position mapping relationship of the target device between the world coordinate system and the image coordinate system using the coordinate transformation matrix; calibrating the parameters of the corresponding camera using the position mapping relationship; and obtaining a camera calibration result.
[0008] Based on the above technical means, the embodiments of this application can record the image coordinates and timestamps of the midpoint of the lower edge of the detection frame when the movement of the target device is detected by the camera. At the same time, the set of world coordinates of the mobile device within the same timestamp is obtained. Based on the set of image coordinates and the corresponding set of world coordinates, the coordinate transformation matrix of the image coordinates and world coordinates of each camera is calculated and optimized to obtain the coordinate transformation matrix of the camera to be calibrated, thereby completing the calibration of the camera. It can be applied to both indoor and outdoor traffic scenarios. By optimizing the coordinate transformation matrix, the camera calibration error is reduced.
[0009] Optionally, in one embodiment of this application, before determining the positional mapping relationship between the target device and the world coordinate system and the image coordinate system using the coordinate transformation matrix, the method further includes: dividing the image in the image coordinate set into multiple regions, with each region containing a preset number of known image coordinate points; calculating one or more cluster centers of the preset number of known image coordinate points, and using each cluster center as a vertex of a preset polygon, removing preset outliers from the vertices, and using the polygon formed by the remaining vertices to divide the image in the image coordinate set into multiple new regions; calculating a coordinate transformation matrix according to each new region, and determining the positional mapping relationship between the world coordinate system and the image coordinate system in that region based on the coordinate transformation matrix.
[0010] Based on the above technical means, the embodiments of this application can divide an image into multiple regions, each region containing a certain number of known image coordinate points. The cluster center of the known image coordinate points is calculated and the coordinates are recorded as vertices of a preset polygon. Outliers in the vertices are removed. The corresponding coordinate transformation matrix for each region is calculated to represent the mapping relationship between the image coordinates and world coordinates of that region. This allows for optimization to reduce calculation errors and improve the accuracy of camera calibration results in the event of camera distortion.
[0011] A second aspect of this application provides a target detection and localization method for a camera, comprising the following steps: acquiring an image of a target region; extracting detection boxes for one or more targets within the image and identifying the image coordinates of the midpoint of the lower edge of the detection box; detecting targets within the image using the detection boxes, matching a coordinate transformation matrix based on the position of the detection boxes in the image, calculating the world coordinates corresponding to the image coordinates using the coordinate transformation matrix, and locating the actual position of the target based on the world coordinates.
[0012] Based on the above technical means, this application can extract the detection boxes of multiple targets in an image, determine the position of the detection box in the image based on the image coordinates of the midpoint of the lower edge of the detection box of multiple targets, and calculate the world coordinates of the target by selecting the corresponding coordinate transformation matrix, thereby achieving accurate positioning of the target in the image and further improving the accuracy of the camera calibration results.
[0013] Optionally, in one embodiment of this application, before matching the coordinate transformation matrix according to the position of the detection box in the image, the method further includes: obtaining a first set of image coordinates and timestamps of one or more cameras acquiring target devices, and obtaining a second set of position and timestamps of the target device itself in a 3D map; identifying the set of world coordinates in the second set corresponding to the set of image coordinates in the first set at any timestamp; calculating the coordinate transformation matrix between the image coordinates and world coordinates of each camera based on the set of image coordinates and the set of world coordinates; using the coordinate transformation matrix to determine the position mapping relationship of the target device between the world coordinate system and the image coordinate system; using the position mapping relationship to calibrate the parameters of the corresponding camera to obtain the camera calibration result.
[0014] Based on the above technical means, this application embodiment can record the image coordinates and timestamps of the midpoint of the lower edge of the detection frame when the movement of the target device is detected by the camera. At the same time, it can obtain the set of world coordinates of the mobile device within the same timestamp. Based on the image coordinate set and the corresponding world coordinate set, the coordinate transformation matrix of the image coordinates and world coordinates of each camera is calculated and optimized to obtain the coordinate transformation matrix of the camera to be calibrated, thereby completing the calibration of a batch of single cameras. It can be applied to both indoor and outdoor traffic scenes. By optimizing the coordinate transformation matrix, the camera calibration error is reduced.
[0015] Optionally, in one embodiment of this application, before determining the positional mapping relationship between the target device and the world coordinate system and the image coordinate system using the coordinate transformation matrix, the method further includes: dividing the image in the image coordinate set into multiple regions, with each region containing a preset number of known image coordinate points; calculating one or more cluster centers of the preset number of known image coordinate points, and using each cluster center as a vertex of a preset polygon, removing preset outliers from the vertices, and using the polygon formed by the remaining vertices to divide the image in the image coordinate set into multiple new regions; calculating a coordinate transformation matrix according to each new region, and determining the positional mapping relationship between the world coordinate system and the image coordinate system in that region based on the coordinate transformation matrix.
[0016] Based on the above technical means, the embodiments of this application can divide an image into multiple regions, each region containing a certain number of known image coordinate points. The cluster center of the known image coordinate points is calculated and the coordinates are recorded as vertices of a preset polygon. Outliers in the vertices are removed. The corresponding coordinate transformation matrix for each region is calculated to represent the mapping relationship between the image coordinates and world coordinates of that region. This allows for optimization to reduce calculation errors and improve the accuracy of camera calibration results in the event of camera distortion.
[0017] A third aspect of this application provides a camera calibration device, comprising: a first acquisition module, configured to acquire a first set of image coordinates and timestamps of one or more cameras acquiring target devices, and acquire a second set of position and timestamps of the target device itself in a three-dimensional map; a first identification module, configured to identify a set of world coordinates in the second set corresponding to the set of image coordinates in the first set at any timestamp, and calculate a coordinate transformation matrix between the image coordinates and world coordinates of each camera based on the set of image coordinates and the set of world coordinates; and a first calibration module, configured to determine the positional mapping relationship of the target device between the world coordinate system and the image coordinate system using the coordinate transformation matrix, calibrate the parameters of the corresponding camera using the positional mapping relationship, and obtain a camera calibration result.
[0018] Optionally, in one embodiment of this application, it further includes: a first partitioning module, configured to divide the image in the image coordinate set into multiple regions before determining the positional mapping relationship between the target device and the world coordinate system using the coordinate transformation matrix, wherein each region contains a preset number of known image coordinate points; a first calculation module, configured to calculate one or more cluster centers of the preset number of known image coordinate points, and use each cluster center as a vertex of a preset polygon, remove preset outliers from the vertices, and divide the image in the image coordinate set into multiple new regions using the polygon formed by the remaining vertices; and a first determining module, configured to calculate a coordinate transformation matrix according to each new region, and determine the positional mapping relationship between the world coordinate system and the image coordinate system in that region based on the coordinate transformation matrix.
[0019] A fourth aspect of this application provides a target detection and localization device for a camera, comprising: an acquisition module for acquiring an image of a target area; an extraction module for extracting detection boxes of one or more targets within the image and identifying the image coordinates of the midpoint of the lower edge of the detection box; and a localization module for detecting targets within the image using the detection boxes, matching a coordinate transformation matrix according to the position of the detection boxes in the image, calculating the world coordinates corresponding to the image coordinates using the coordinate transformation matrix, and locating the actual position of the target based on the world coordinates.
[0020] Optionally, in one embodiment of this application, it further includes: a second acquisition module, configured to acquire a first set of image coordinates and timestamps of one or more camera-acquired target devices and a second set of positions and timestamps of the target device itself in a 3D map before matching the coordinate transformation matrix according to the position of the detection box in the image; a second identification module, configured to identify the set of world coordinates in the second set corresponding to the set of image coordinates in the first set at any timestamp, and calculate the coordinate transformation matrix between the image coordinates and world coordinates of each camera according to the set of image coordinates and the set of world coordinates; and a second calibration module, configured to use the coordinate transformation matrix to determine the position mapping relationship of the target device between the world coordinate system and the image coordinate system, and use the position mapping relationship to calibrate the parameters of the corresponding camera to obtain the camera calibration result.
[0021] Optionally, in one embodiment of this application, it further includes: a second partitioning module, configured to divide the image in the image coordinate set into multiple regions before determining the positional mapping relationship between the target device and the world coordinate system using the coordinate transformation matrix, wherein each region contains a preset number of known image coordinate points; a second calculation module, configured to calculate one or more cluster centers of the preset number of known image coordinate points, and use each cluster center as a vertex of a preset polygon, remove preset outliers from the vertices, and use the polygon formed by the remaining vertices to divide the image in the image coordinate set into multiple new regions; and a second determining module, configured to calculate a coordinate transformation matrix according to each new region, and determine the positional mapping relationship between the world coordinate system and the image coordinate system in that region based on the coordinate transformation matrix.
[0022] A fifth aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the camera calibration method as described in the above embodiments.
[0023] A sixth aspect of this application provides a camera, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the target detection and localization method as described in the above embodiments.
[0024] A seventh aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the camera calibration method or the camera target detection and localization method as described in the above embodiments.
[0025] Therefore, this application has at least the following beneficial effects:
[0026] 1. In this embodiment, when the movement of the target device being captured by the camera is detected, the image coordinates of the midpoint of the lower edge of the detection frame and the set of timestamps are recorded. At the same time, the set of world coordinates of the mobile device within the same timestamp is obtained. Based on the set of image coordinates and the corresponding set of world coordinates, the coordinate transformation matrix of the image coordinates and world coordinates of each camera is calculated and optimized to obtain the coordinate transformation matrix of the camera to be calibrated, thereby completing the camera calibration. This can be applied to both indoor and outdoor traffic scenarios. By optimizing the coordinate transformation matrix, the camera calibration error is reduced.
[0027] 2. In this embodiment, an image can be divided into multiple regions, each containing a certain number of known image coordinate points. Then, the cluster center of the known image coordinate points is calculated and the coordinates are recorded and used as vertices of a preset polygon. Outliers in the vertices are removed, and the corresponding coordinate transformation matrix for each region is calculated to represent the mapping relationship between the image coordinates and world coordinates of that region. This allows for optimization to reduce calculation errors and improve the accuracy of camera calibration results in the event of camera distortion.
[0028] 3. This application can extract detection boxes of multiple targets in an image, determine the position of the detection box in the image based on the image coordinates of the midpoint of the lower edge of the detection box of multiple targets, and calculate the world coordinates of the target by selecting the corresponding coordinate transformation matrix, thereby achieving accurate positioning of the target in the image and further improving the accuracy of camera calibration results.
[0029] 4. In this embodiment, when the movement of the target device being captured by the camera is detected, the image coordinates of the midpoint of the lower edge of the detection frame and the set of timestamps are recorded. At the same time, the set of world coordinates of the mobile device within the same timestamp is obtained. Based on the set of image coordinates and the corresponding set of world coordinates, the coordinate transformation matrix of the image coordinates and world coordinates of each camera is calculated and optimized to obtain the coordinate transformation matrix of the camera to be calibrated, thereby completing the calibration of a batch of single cameras. It can be applied to both indoor and outdoor traffic scenes. By optimizing the coordinate transformation matrix, the camera calibration error is reduced.
[0030] 5. In this embodiment, an image can be divided into multiple regions, each containing a certain number of known image coordinate points. Then, the cluster center of the known image coordinate points is calculated and the coordinates are recorded and used as vertices of a preset polygon. Outliers in the vertices are removed, and the corresponding coordinate transformation matrix for each region is calculated to represent the mapping relationship between the image coordinates and world coordinates of that region. This allows for optimization to reduce calculation errors and improve the accuracy of camera calibration results in the event of camera distortion.
[0031] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0032] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0033] Figure 1 This is a flowchart of a camera calibration method provided according to an embodiment of this application;
[0034] Figure 2This is a schematic diagram illustrating the calculation of the coordinate transformation matrix according to an embodiment of this application;
[0035] Figure 3 This is an example diagram of the Voronoi Diagram provided according to an embodiment of this application;
[0036] Figure 4 This is a diagram illustrating the composition of the positioning and detection method provided according to the embodiments of this application;
[0037] Figure 5 This is a flowchart of a camera target detection and localization method according to an embodiment of this application;
[0038] Figure 6 This is a block diagram of a camera calibration device according to an embodiment of this application;
[0039] Figure 7 This is a block diagram of a camera target detection and positioning device according to an embodiment of this application;
[0040] Figure 8 A schematic diagram of the structure of an electronic device according to an embodiment of this application is provided;
[0041] Figure 9 A schematic diagram of the camera structure provided in the embodiments of this application.
[0042] Explanation of reference numerals in the attached drawings: First acquisition module-100, First identification module-200, First calibration module-300, Acquisition module-400, Extraction module-500, Positioning module-600, Memory-801, Processor-802, Communication interface-803, Memory-901, Processor-902, Communication interface-903. Detailed Implementation
[0043] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0044] The following description, with reference to the accompanying drawings, describes a camera calibration method, apparatus, and electronic device according to embodiments of this application, including a camera target detection and localization method. Addressing the problems mentioned in the background section, this application provides a camera calibration method. In this method, when movement of the target device being captured by the camera is detected, the image coordinates of the midpoint of the lower edge of its detection frame and a set of timestamps are recorded. Simultaneously, a set of world coordinates of the mobile device within the same timestamp is obtained. Based on the image coordinate set and the corresponding world coordinate set, a coordinate transformation matrix between the image coordinates and world coordinates for each camera is calculated and optimized to obtain the coordinate transformation matrix of the camera to be calibrated, thereby completing the camera calibration. This method can be applied to both indoor and outdoor traffic scenarios. By optimizing the coordinate transformation matrix, the camera calibration error is reduced. This solves the problems in related technologies, such as the calibration board on the robot being far from the camera, resulting in a small area of the calibration board in the captured image, leading to low accuracy of the calculated camera intrinsic parameters, and limitations limited to indoor scenarios.
[0045] Specifically, Figure 1 This is a schematic flowchart of a camera calibration method provided in an embodiment of this application.
[0046] like Figure 1 As shown, the camera calibration method includes the following steps:
[0047] In step S101, a first set of image coordinates and timestamps of the target device acquired by one or more cameras is obtained, and a second set of the location and timestamps of the target device itself in the 3D map is obtained.
[0048] It is understood that in this embodiment, multiple cameras to be calibrated can be deployed in the scene, and the timestamps of all cameras and the calibration device (robot or vehicle) are pre-synchronized. A target detection algorithm capable of detecting the target device (robot or vehicle) is pre-deployed in all cameras to be calibrated. When a target device is detected, the image coordinates of the midpoint of the lower edge of its detection box and the set of timestamps are recorded. Simultaneously, the device (robot / vehicle) equipped with a special identifier for automatic calibration can be controlled to move within the scene, and data collected by its onboard sensors can be used to construct a 3D map of the scene, obtaining the position and timestamp set of the mobile device in the 3D map. The specific process is as follows... Figure 2 As shown.
[0049] In practical implementation, the embodiments of this application can utilize simultaneous localization and mapping (SLAM) technology to construct a 3D map of the scene, which can be understood as a 3D model of the scene. Simultaneously, the position of the robot / vehicle within the 3D scene map is obtained. The SLAM algorithm can be selected differently depending on the sensors used. For example, one option for SLAM using LiDAR is LOAM (LidarOdometry and Mapping in Real-time), another option for SLAM using a vision camera is ORB-SLAM2, and multi-sensor fusion SLAM can also be used (e.g., the V-LOAM algorithm using LiDAR, vision, and IMU). This can be determined according to the actual situation and is not specifically limited.
[0050] In step S102, the world coordinate set in the second set corresponding to the image coordinate set in the first set at any time stamp is identified, and the coordinate transformation matrix between the image coordinates and the world coordinates of each camera is calculated based on the image coordinate set and the world coordinate set.
[0051] In practical applications, the target that needs to be tracked is usually moving on the ground, such as people or vehicles in outdoor scenes, or pedestrians in indoor scenes. For these cases, it is only necessary to calculate the position of the point of contact between the target and the ground on the ground plane. Therefore, this embodiment can obtain the position by solving the coordinate transformation matrix between image coordinates and world coordinates. This embodiment can obtain the set of world coordinates of the mobile device within the same timestamp based on the set of timestamps when the image detects the mobile device, and solve the coordinate transformation matrix H (homography matrix) between the image coordinates and world coordinates of each camera based on the set of image coordinates and the corresponding set of world coordinates.
[0052] In step S103, the coordinate transformation matrix is used to determine the positional mapping relationship between the target device in the world coordinate system and the image coordinate system, and the parameters of the corresponding camera are calibrated using the positional mapping relationship to obtain the camera calibration result.
[0053] It is understood that the embodiments of this application can utilize the homography matrix to describe the positional mapping relationship of an object between the world coordinate system and the pixel coordinate system, thereby completing the camera calibration. This can be applied to both indoor and outdoor traffic scenes. The positional mapping relationship is defined as follows:
[0054]
[0055] Where M is the camera intrinsic parameter, and r and t are the camera extrinsic parameters.
[0056] Mapping relationship between pixel coordinate system and world coordinate system:
[0057]
[0058] Where s represents the scale factor, u and v represent coordinates in the pixel coordinate system, and Xw and Yw represent coordinates in the world coordinate system.
[0059] It should be noted that the multiple cameras to be calibrated in this application embodiment may have overlapping visible areas or no overlapping areas at all.
[0060] In one embodiment of this application, before determining the positional mapping relationship between the target device and the world coordinate system and the image coordinate system using the coordinate transformation matrix, the method further includes: dividing the image in the image coordinate set into multiple regions, with each region containing a preset number of known image coordinate points; calculating one or more cluster centers of the preset number of known image coordinate points, and using each cluster center as a vertex of a preset polygon, removing preset outliers from the vertices, and using the polygon formed by the remaining vertices to divide the image in the image coordinate set into multiple new regions; calculating a coordinate transformation matrix according to each new region, and determining the positional mapping relationship between the world coordinate system and the image coordinate system in that region based on the coordinate transformation matrix.
[0061] Due to camera distortion, the homography matrix calculated using all world coordinates-image coordinate sets will have significant errors, thus requiring optimization. This application's embodiment utilizes the Voronoi Diagram method to divide the image into N regions, each containing a certain number M (M>=4) known image coordinate points. Then, the cluster centers of these M points are calculated and their coordinates recorded. These cluster centers are then used as vertices of the Voronoi Diagram to calculate the Voronoi Diagram, such as... Figure 3 As shown, the number of vertices is 6.
[0062] Furthermore, in this embodiment, RANSANC (Random Sample Consensus) can be used to remove outliers, and a homography matrix can be calculated separately for each region to represent the mapping relationship between the image coordinates and world coordinates of that region. This allows for optimization to reduce calculation errors and improve the accuracy of camera calibration results in the event of camera distortion.
[0063] According to the camera calibration method proposed in this application, when the movement of the target device being captured by the camera is detected, the image coordinates of the midpoint of the lower edge of its detection frame and the set of timestamps are recorded. Simultaneously, the set of world coordinates of the mobile device within the same timestamp is obtained. Based on the set of image coordinates and the corresponding set of world coordinates, the coordinate transformation matrix between the image coordinates and world coordinates of each camera is calculated and optimized to obtain the coordinate transformation matrix of the camera to be calibrated. This completes the calibration of a batch of single cameras, applicable to both indoor and outdoor traffic scenarios. Furthermore, by optimizing the coordinate transformation matrix, the camera calibration error is reduced. This solves the problems in related technologies, such as the calibration board on the robot being far from the camera, resulting in a small area of the calibration board in the captured image, leading to low accuracy of the calculated camera intrinsic parameters, and limitations limited to indoor scenarios.
[0064] This application also proposes a target detection and localization method, which mainly consists of three parts, such as... Figure 4 As shown, the main functions of calculating the coordinate transformation matrix are: calculating the homography matrix; and optimizing using the Voronoi Diagram to obtain the coordinate transformation matrix of the camera to be calibrated. The main function of multi-object detection is: using a multi-object detection algorithm to extract the detection boxes of multiple objects within the image. The main function of multi-object localization is: based on the image coordinates of the midpoint of the lower edge of the detection box of a multi-object, determining which region in the Voronoi Diagram it is located in, and selecting the corresponding homography matrix to calculate the world coordinates of that object.
[0065] Specifically, such as Figure 5 As shown, the target detection and localization method of this camera includes the following steps:
[0066] In step S501, an image of the target area is acquired.
[0067] The embodiments of this application can use lidar, cameras, and inertial measurement units to acquire images of the target area in order to complete the detection and localization of the target.
[0068] In step S502, detection boxes of one or more targets within the image are extracted, and the image coordinates of the midpoint of the lower edge of the detection box are identified.
[0069] It is understood that the embodiments of this application can utilize multi-target detection and tracking algorithms (adopting deep learning-based related algorithms, one optional detection algorithm is the YOLOv3 algorithm, and one optional multi-target detection algorithm is the FairMOT algorithm) to extract the detection boxes of multiple targets in the image and identify the image coordinates of the midpoint of the lower edge of the detection box. For indoor scenes, the target to be located is usually a pedestrian, and for outdoor traffic scenes, the target to be located is usually a vehicle.
[0070] In step S503, a target within the image is detected using a detection box, and a coordinate transformation matrix is matched based on the position of the detection box in the image. The world coordinates corresponding to the image coordinates are calculated using the coordinate transformation matrix, and the actual position of the target is located based on the world coordinates.
[0071] The embodiments of this application can determine the position of the detection box in the image based on the image coordinates of the midpoint of the lower edge of the detection box of multiple targets, and select the corresponding coordinate transformation matrix to calculate the world coordinates of the target, thereby achieving accurate positioning of the target in the image and further improving the accuracy of the camera calibration results.
[0072] It should be noted that, in the indoor scenarios of this application, the target to be located is usually a pedestrian, and the device used for automatic calibration can be a ground mobile robot equipped with special markers and sensors; in the outdoor traffic scenarios, the target to be located is usually a vehicle, and the device used for automatic calibration can be a car equipped with special markers and sensors.
[0073] In one embodiment of this application, before matching the coordinate transformation matrix according to the position of the detection box in the image, the method further includes: obtaining a first set of image coordinates and timestamps of the target device acquired by one or more cameras, and obtaining a second set of position and timestamps of the target device itself in the 3D map; identifying the world coordinate set in the second set corresponding to the image coordinate set in the first set at any timestamp; calculating the coordinate transformation matrix between the image coordinates and world coordinates of each camera based on the image coordinate set and the world coordinate set; using the coordinate transformation matrix to determine the position mapping relationship of the target device between the world coordinate system and the image coordinate system; using the position mapping relationship to calibrate the parameters of the corresponding camera to obtain the camera calibration result.
[0074] The embodiments of this application use position mapping relationships to calibrate the parameters of the corresponding camera and obtain the camera calibration results. You can refer to steps S101, S102 and S103 of the above camera calibration method. To avoid redundancy, they will not be repeated here.
[0075] In one embodiment of this application, before determining the positional mapping relationship between the target device and the world coordinate system and the image coordinate system using the coordinate transformation matrix, the method further includes: dividing the image in the image coordinate set into multiple regions, with each region containing a preset number of known image coordinate points; calculating one or more cluster centers of the preset number of known image coordinate points, and using each cluster center as a vertex of a preset polygon, removing preset outliers from the vertices, and using the polygon formed by the remaining vertices to divide the image in the image coordinate set into multiple new regions; calculating a coordinate transformation matrix according to each new region, and determining the positional mapping relationship between the world coordinate system and the image coordinate system in that region based on the coordinate transformation matrix.
[0076] It should be noted that the positional mapping relationship between the world coordinate system and the image coordinate system in the above camera calibration method is also applicable to the embodiments of this application, and will not be repeated here.
[0077] The target detection and localization method for cameras proposed in this application extracts detection boxes for multiple targets within an image. Based on the image coordinates of the midpoints of the lower edges of these detection boxes, the position of the boxes in the image is determined. Then, a corresponding coordinate transformation matrix is selected to calculate the world coordinates of the target, achieving accurate localization of targets within the image and further improving the accuracy of camera calibration results. This solves the problems in related technologies, such as the calibration board on the robot being far from the camera, resulting in a small area of the calibration board in the captured image, leading to low accuracy of the calculated camera intrinsic parameters, and limitations applicable only to indoor scenes.
[0078] Next, a camera calibration device according to an embodiment of this application is described with reference to the accompanying drawings.
[0079] Figure 6 This is a block diagram of a camera calibration device according to an embodiment of this application.
[0080] like Figure 6 As shown, the camera calibration device 10 includes: a first acquisition module 100, a first identification module 200, and a first calibration module 300.
[0081] The first acquisition module 100 is used to acquire a first set of image coordinates and timestamps of one or more camera acquisition target devices, and to acquire a second set of the position and timestamps of the target device itself in the 3D map; the first recognition module 200 is used to recognize the world coordinate set in the second set corresponding to the image coordinate set in the first set at any timestamp, and to calculate the coordinate transformation matrix between the image coordinates and world coordinates of each camera based on the image coordinate set and the world coordinate set; the first calibration module 300 is used to determine the position mapping relationship between the target device in the world coordinate system and the image coordinate system using the coordinate transformation matrix, and to calibrate the parameters of the corresponding camera using the position mapping relationship to obtain the camera calibration result.
[0082] In one embodiment of this application, the apparatus 10 further includes: a first partitioning module, a first calculation module, and a first determination module.
[0083] The first partitioning module is used to divide the image in the image coordinate set into multiple regions before determining the positional mapping relationship between the target device and the world coordinate system using the coordinate transformation matrix, and each region contains a preset number of known image coordinate points; the first calculation module is used to calculate one or more cluster centers of the preset number of known image coordinate points, and use each cluster center as a vertex of a preset polygon, remove preset outliers from the vertices, and use the polygon formed by the remaining vertices to divide the image in the image coordinate set into multiple new regions; the first determining module is used to calculate the coordinate transformation matrix according to each new region, and determine the positional mapping relationship between the world coordinate system and the image coordinate system in that region based on the coordinate transformation matrix.
[0084] It should be noted that the foregoing explanation of the camera calibration method embodiment also applies to the camera calibration device of this embodiment, and will not be repeated here.
[0085] According to the camera calibration device proposed in this application, when the movement of the target device being captured by the camera is detected, the device records the image coordinates of the midpoint of the lower edge of its detection frame and a set of timestamps. Simultaneously, it acquires a set of world coordinates of the mobile device within the same timestamp. Based on the image coordinate set and the corresponding world coordinate set, it calculates and optimizes the coordinate transformation matrix of each camera to obtain the coordinate transformation matrix of the camera to be calibrated. This completes the calibration of a batch of single cameras, applicable to both indoor and outdoor traffic scenarios. Furthermore, by optimizing the coordinate transformation matrix, the camera calibration error is reduced. This solves the problems in related technologies, such as the calibration board on the robot being far from the camera, resulting in a small area of the calibration board in the captured image, leading to low accuracy of the calculated camera intrinsic parameters, and limitations limited to indoor scenarios.
[0086] Furthermore, a camera target detection and positioning device according to an embodiment of this application is described with reference to the accompanying drawings.
[0087] Figure 7 This is a block diagram of a camera target detection and positioning device according to an embodiment of this application.
[0088] like Figure 7 As shown, the target detection and positioning device 20 of the camera includes: an acquisition module 400, an extraction module 500, and a positioning module 600.
[0089] The acquisition module 400 is used to acquire images of the target area; the extraction module 500 is used to extract detection boxes of one or more targets in the image and identify the image coordinates of the midpoint of the lower edge of the detection box; the positioning module 600 is used to detect targets in the image with detection boxes, match coordinate transformation matrix according to the position of the detection box in the image, calculate the world coordinates corresponding to the image coordinates using the coordinate transformation matrix, and locate the actual position of the target based on the world coordinates.
[0090] In one embodiment of this application, the apparatus 20 further includes: a second acquisition module, a second identification module, and a second calibration module.
[0091] The second acquisition module is used to acquire a first set of image coordinates and timestamps of one or more camera-acquired target devices before matching the coordinate transformation matrix according to the position of the detection box in the image, and to acquire a second set of position and timestamps of the target device itself in the 3D map; the second recognition module is used to recognize the world coordinate set in the second set corresponding to the image coordinate set in the first set at any timestamp, and to calculate the coordinate transformation matrix between the image coordinates and world coordinates of each camera based on the image coordinate set and the world coordinate set; the second calibration module is used to determine the position mapping relationship between the target device in the world coordinate system and the image coordinate system using the coordinate transformation matrix, and to calibrate the parameters of the corresponding camera using the position mapping relationship to obtain the camera calibration result.
[0092] In one embodiment of this application, the apparatus 20 of this application embodiment further includes: a second division module, a second calculation module, and a second determination module.
[0093] The second partitioning module is used to divide the image in the image coordinate set into multiple regions before determining the positional mapping relationship between the target device and the world coordinate system using the coordinate transformation matrix, and each region contains a preset number of known image coordinate points; the second calculation module is used to calculate one or more cluster centers of the preset number of known image coordinate points, and use each cluster center as a vertex of a preset polygon, remove preset outliers from the vertices, and use the polygon formed by the remaining vertices to divide the image in the image coordinate set into multiple new regions; the second determination module is used to calculate the coordinate transformation matrix according to each new region, and determine the positional mapping relationship between the world coordinate system and the image coordinate system in that region based on the coordinate transformation matrix.
[0094] It should be noted that the foregoing explanation of the camera target detection and localization method embodiment also applies to the camera target detection and localization device of this embodiment, and will not be repeated here.
[0095] The target detection and localization device for a camera proposed in this application extracts detection boxes for multiple targets within an image. Based on the image coordinates of the midpoints of the lower edges of these detection boxes, the device determines the position of the detection boxes in the image and calculates the world coordinates of the targets using a corresponding coordinate transformation matrix. This achieves accurate localization of targets within the image and further improves the accuracy of camera calibration results. This solves the problems in related technologies, such as the calibration board on the robot being far from the camera, resulting in a small area of the calibration board in the captured image, leading to low accuracy of the calculated camera intrinsic parameters, and limitations applicable only to indoor scenes.
[0096] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include:
[0097] The memory 801, the processor 802, and the computer program stored on the memory 801 and capable of running on the processor 802.
[0098] When the processor 802 executes the program, it implements the camera calibration method provided in the above embodiments.
[0099] Furthermore, electronic devices also include:
[0100] Communication interface 803 is used for communication between memory 801 and processor 802.
[0101] The memory 801 is used to store computer programs that can run on the processor 802.
[0102] The memory 801 may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile memory, such as at least one disk storage.
[0103] If the memory 801, processor 802, and communication interface 803 are implemented independently, then the communication interface 803, memory 801, and processor 802 can be interconnected via a bus to complete communication between them. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 8 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0104] Optionally, in a specific implementation, if the memory 801, processor 802, and communication interface 803 are integrated on a single chip, then the memory 801, processor 802, and communication interface 803 can communicate with each other through an internal interface.
[0105] The processor 802 may be a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of this application.
[0106] Figure 9 A schematic diagram of the structure of a camera provided in an embodiment of this application. The camera may include:
[0107] The memory 901, the processor 902, and the computer program stored on the memory 901 and capable of running on the processor 902.
[0108] When the processor 902 executes the program, it implements the target detection and localization method for the camera provided in the above embodiments.
[0109] Furthermore, the camera also includes:
[0110] Communication interface 903 is used for communication between memory 901 and processor 902.
[0111] The memory 901 is used to store computer programs that can run on the processor 902.
[0112] The memory 901 may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile memory, such as at least one disk storage.
[0113] If the memory 901, processor 902, and communication interface 903 are implemented independently, then the communication interface 903, memory 901, and processor 902 can be interconnected via a bus to complete communication between them. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 9 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0114] Optionally, in a specific implementation, if the memory 901, processor 902, and communication interface 903 are integrated on a single chip, then the memory 901, processor 902, and communication interface 903 can communicate with each other through an internal interface.
[0115] The processor 902 may be a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of this application.
[0116] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the camera calibration method or the camera target detection and localization method as described in the above embodiments.
[0117] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0118] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0119] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0120] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (FPGAs), field-programmable gate arrays (FPGAs), etc.
[0121] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0122] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A camera calibration method, characterized in that, Includes the following steps: Obtain a first set of image coordinates and timestamps of a target device acquired by one or more cameras, and obtain a second set of the location and timestamps of the target device itself in a 3D map; Identify the world coordinate set in the second set that corresponds to the image coordinate set in the first set at any timestamp, and calculate the coordinate transformation matrix between the image coordinates and world coordinates of each camera based on the image coordinate set and the world coordinate set; The coordinate transformation matrix is used to determine the positional mapping relationship of the target device between the world coordinate system and the image coordinate system. The positional mapping relationship is then used to calibrate the parameters of the corresponding camera to obtain the camera calibration result. Before using the coordinate transformation matrix to determine the positional mapping relationship of the target device between the world coordinate system and the image coordinate system, the method further includes: The image in the image coordinate set is divided into multiple regions, and each region contains a preset number of known image coordinate points; Calculate one or more cluster centers of the preset number of known image coordinate points, and use each cluster center as a vertex of a preset polygon. Remove the preset outliers in the vertices, and use the polygon formed by the remaining vertices to divide the image in the image coordinate set into multiple new regions. Calculate a coordinate transformation matrix for each new region, and determine the positional mapping relationship between the world coordinate system and the image coordinate system in that region based on the coordinate transformation matrix.
2. A target detection and localization method for a camera, characterized in that, Includes the following steps: Acquire images of the target area; Extract detection boxes for one or more targets within the image, and identify the image coordinates of the midpoint of the lower edge of the detection box; The target within the image is detected using the detection box, and a coordinate transformation matrix is matched based on the position of the detection box in the image. The world coordinates corresponding to the image coordinates are calculated using the coordinate transformation matrix, and the actual position of the target is located based on the world coordinates. Before matching the coordinate transformation matrix according to the position of the detection box in the image, the method further includes: Obtain a first set of image coordinates and timestamps of a target device acquired by one or more cameras, and obtain a second set of the location and timestamps of the target device itself in a 3D map; Identify the world coordinate set in the second set that corresponds to the image coordinate set in the first set at any timestamp, and calculate the coordinate transformation matrix between the image coordinates and world coordinates of each camera based on the image coordinate set and the world coordinate set; The coordinate transformation matrix is used to determine the positional mapping relationship of the target device between the world coordinate system and the image coordinate system. The positional mapping relationship is then used to calibrate the parameters of the corresponding camera to obtain the camera calibration result. Before using the coordinate transformation matrix to determine the positional mapping relationship of the target device between the world coordinate system and the image coordinate system, the method further includes: The image in the image coordinate set is divided into multiple regions, and each region contains a preset number of known image coordinate points; Calculate one or more cluster centers of the preset number of known image coordinate points, and use each cluster center as a vertex of a preset polygon. Remove the preset outliers in the vertices, and use the polygon formed by the remaining vertices to divide the image in the image coordinate set into multiple new regions. Calculate a coordinate transformation matrix for each new region, and determine the positional mapping relationship between the world coordinate system and the image coordinate system in that region based on the coordinate transformation matrix.
3. A camera calibration device, characterized in that, include: The first acquisition module is used to acquire a first set of image coordinates and timestamps of one or more camera-captured target devices, and to acquire a second set of the location and timestamps of the target device itself in a three-dimensional map. The first identification module is used to identify the world coordinate set in the second set corresponding to the image coordinate set in the first set at any time stamp, and to calculate the coordinate transformation matrix between the image coordinates and the world coordinates of each camera based on the image coordinate set and the world coordinate set. The first calibration module is used to determine the positional mapping relationship of the target device between the world coordinate system and the image coordinate system using the coordinate transformation matrix, and to calibrate the parameters of the corresponding camera using the positional mapping relationship to obtain the camera calibration result; The first partitioning module is used to divide the image in the image coordinate set into multiple regions before determining the positional mapping relationship between the target device and the world coordinate system using the coordinate transformation matrix, and each region contains a preset number of known image coordinate points; The first calculation module is used to calculate one or more cluster centers of the preset number of known image coordinate points, and use each cluster center as a vertex of a preset polygon, remove preset outliers from the vertices, and use the polygon formed by the remaining vertices to divide the image in the image coordinate set into multiple new regions. The first determining module is used to calculate a coordinate transformation matrix for each new region and determine the positional mapping relationship between the world coordinate system and the image coordinate system in that region based on the coordinate transformation matrix.
4. A target detection and positioning device for a camera, characterized in that, include: The acquisition module is used to acquire images of the target area; An extraction module is used to extract detection boxes of one or more targets within the image and identify the image coordinates of the midpoint of the lower edge of the detection box; The localization module is used to detect targets within the image using the detection box, match a coordinate transformation matrix based on the position of the detection box in the image, calculate the world coordinates corresponding to the image coordinates using the coordinate transformation matrix, and locate the actual position of the target based on the world coordinates. The second acquisition module is used to acquire a first set of image coordinates and timestamps of one or more camera-acquired target devices before matching the coordinate transformation matrix according to the position of the detection box in the image, and to acquire a second set of position and timestamps of the target device itself in the 3D map. The second recognition module is used to identify the world coordinate set in the second set that corresponds to the image coordinate set in the first set at any time stamp, and to calculate the coordinate transformation matrix between the image coordinates and the world coordinates of each camera based on the image coordinate set and the world coordinate set. The second calibration module is used to determine the positional mapping relationship between the target device and the world coordinate system and the image coordinate system using the coordinate transformation matrix, and to calibrate the parameters of the corresponding camera using the positional mapping relationship to obtain the camera calibration result; The second partitioning module is used to divide the image in the image coordinate set into multiple regions before using the coordinate transformation matrix to determine the positional mapping relationship between the target device in the world coordinate system and the image coordinate system, and each region contains a preset number of known image coordinate points. The second calculation module is used to calculate one or more cluster centers of a preset number of known image coordinate points, and use each cluster center as a vertex of a preset polygon, remove preset outliers in the vertices, and use the polygon formed by the remaining vertices to divide the image in the image coordinate set into multiple new regions. The second determining module is used to calculate the coordinate transformation matrix for each new region and determine the positional mapping relationship between the world coordinate system and the image coordinate system in that region based on the coordinate transformation matrix.
5. An electronic device, characterized in that, include: The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the camera calibration method as described in claim 1.
6. A camera, characterized in that, include: The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the target detection and localization method for a camera as described in claim 2.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the camera calibration method as described in claim 1, or the camera target detection and localization method as described in claim 2.