A calibration method and apparatus

By employing methods such as moving calibration objects and global optimization, the problem of flatness requirements in large-scale scene calibration was solved, achieving high-accuracy calibration on uneven ground, which is suitable for multi-camera systems.

CN117422770BActive Publication Date: 2026-05-26HUAWEI TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2022-07-11
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies have strict requirements for the flatness of the site when used for large-scale site calibration, making them difficult to apply to uneven sites such as football fields, resulting in insufficient calibration accuracy.

Method used

By using a moving calibration object, the camera parameters of the acquisition device are determined through common-view relationships. Combined with global optimization and scale transformation, intrinsic parameters, extrinsic parameters, and distortion coefficients are estimated to improve calibration accuracy.

Benefits of technology

It achieves high-accuracy calibration on uneven terrain, is applicable to multi-camera systems, and improves the accuracy and applicability of calibration results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117422770B_ABST
    Figure CN117422770B_ABST
Patent Text Reader

Abstract

A calibration method and apparatus are disclosed, applicable to the field of computer vision technology. This application is not limited by the flatness of the calibration site, improving the accuracy and applicability of calibration. Multiple acquisition devices are deployed at the calibration site, and calibration is performed by moving the target calibration object. This eliminates the influence of inherent visual features / calibration points of the site, broadening the applicable scenarios, including those with large areas. This application uses a moving calibration object to determine the camera parameters of the acquisition devices through co-view relationships, reducing the requirements for site flatness and improving calibration accuracy in such scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a calibration method and apparatus. Background Technology

[0002] Typically, multiple cameras are deployed in large-scale venues to capture video streams to calculate the behavioral trajectories of athletes or other personnel, or to achieve spatial video effects. Examples include highlight moments of a soccer player kicking the ball, or running postures or trajectories in track and field. Large-scale venues refer to large sports stadiums, such as soccer fields, basketball courts, volleyball courts, ice skating and skiing venues, as well as venues of similar scale, such as plazas, exhibition halls, and conference venues. Calculating the behavioral trajectories of athletes or other personnel, or achieving spatial video effects, requires calibrating the intrinsic and extrinsic parameters of all cameras involved in the shooting. Then, based on these parameters, 3D information within the scene (such as 3D human skeletons, 3D scene point clouds, etc.) is calculated.

[0003] The current calibration scheme involves arranging multiple calibration posts in the shooting area according to a pre-determined pattern. The positional relationships between the posts are physically measured to unify the feature points on all posts into the same world coordinate system. Images, including those of the calibration posts, are then captured by multiple cameras. The coordinates of the calibration points in the images are identified, and a direct linear transformation is used to obtain the camera's intrinsic and extrinsic parameters. However, in actual deployment, accurate measurement of the spatial distance between the calibration posts is crucial, and ensuring that all posts are on the same horizontal plane is essential. This places strict limitations on the flatness of the site, making it unsuitable for uneven terrain, such as in scenarios like football fields. Summary of the Invention

[0004] This application provides a calibration method and apparatus that are not limited by the flatness of the calibration site, thereby improving the accuracy and applicability of the calibration.

[0005] In a first aspect, embodiments of this application provide a calibration method, comprising: acquiring multiple video streams acquired by multiple acquisition devices, wherein the multiple acquisition devices are deployed in a designated space of a sports field, and the multiple video streams are synchronously captured by the multiple acquisition devices during the movement of a target calibration object in the sports field; the movement trajectory of the target calibration object in the sports field at least covers a designated area of ​​the sports field, and the target calibration object includes at least two non-coplanar calibration surfaces, each calibration surface including at least two calibration points; each acquisition device's video stream includes multiple image frames; calibration point detection is performed on each image frame acquired by the multiple acquisition devices to obtain the pixel coordinates of multiple calibration points on the target calibration object in the image frame acquired by each acquisition device; and the pixel coordinates of the multiple calibration points included in the target calibration object in the image frame acquired by each acquisition device, and the calibration object coordinate system of the multiple calibration points on the target calibration object are used as the basis for the calibration method. The intrinsic parameter matrix of each acquisition device is estimated to obtain the first intrinsic parameter estimate of each acquisition device based on the three-dimensional coordinates of the target calibration object, and the first extrinsic parameter estimate of each acquisition device based on the first intrinsic parameter estimates of at least two acquisition devices included in each acquisition device group, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of the target calibration object including multiple calibration points in the calibration object coordinate system. Wherein, at least two acquisition devices included in each acquisition device group have a common viewing area, the matching feature point set includes multiple matching feature point groups, each matching feature point group includes at least two matching pixel coordinates, and the at least two matching pixel coordinates are the pixel coordinates of the same calibration point detected by image frames acquired at the same time by different acquisition devices belonging to the same acquisition device group; the multiple acquisition device groups are obtained by grouping the multiple acquisition devices, and any two acquisition device groups in the multiple acquisition device groups include at least one identical acquisition device.

[0006] Currently, multiple fixed calibration objects are used on the sports field. In actual deployment, accurate measurement of the spatial distance between the calibration objects is required, and multiple calibration objects must be on the same horizontal plane. This places strict limitations on the flatness of the field, making it difficult to apply to uneven surfaces, such as those used in soccer fields. This application uses a moving calibration object method, determining the camera parameters of the acquisition device through common-view relationships, which has lower requirements for field flatness. This can improve the calibration accuracy in such scenarios.

[0007] In one possible design, the method further includes:

[0008] Based on the pixel coordinates of multiple calibration points included in the target calibration object in the image frame acquired by each acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system of the target calibration object, the distortion coefficient of each acquisition device is estimated to obtain the first distortion coefficient estimate of each acquisition device.

[0009] In the above design, the distortion coefficient can also be estimated based on the estimation of intrinsic and extrinsic parameters.

[0010] In one possible design, based on the pixel coordinates of multiple calibration points included in the target calibration object in the image frame acquired by each acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the intrinsic parameter matrix of each acquisition device is estimated to obtain the first intrinsic parameter estimate of each acquisition device, including:

[0011] Based on the pixel coordinates of the calibration points on the target calibration object in the image frames included in the image set acquired by the i-th acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the intrinsic parameter matrix of the i-th acquisition device is estimated to obtain the second intrinsic parameter estimate of the i-th acquisition device; the image set includes M1 image frames of the target calibration object in the video stream acquired by the i-th acquisition device, and the M1 image frames correspond one-to-one with the M1 movement positions of the target calibration object; M1 is a positive integer, and M is an integer greater than M1;

[0012] Based on the second intrinsic parameter estimate of the i-th acquisition device, the pixel coordinates of the calibration points on the target calibration object in the image frames included in the image set acquired by the i-th acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the pose set corresponding to the i-th acquisition device is estimated respectively. The pose set corresponding to the i-th acquisition device includes the pose of the target calibration object relative to the i-th acquisition device at M1 moving positions; i takes the value of a positive integer less than or equal to N, where N is the number of acquisition devices deployed in the set space of the sports field;

[0013] Among them, the range of movement positions corresponding to the image frames acquired by different acquisition devices is different;

[0014] Based on the pixel coordinates of the calibration points on the target calibration object in the image frames of the image sets acquired by N acquisition devices, the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, and the poses of the target calibration object corresponding to the N acquisition devices, and on the basis of the initially set distortion coefficients and second intrinsic parameter estimates corresponding to the N acquisition devices, the intrinsic parameter matrices and distortion coefficients of the N acquisition devices are adjusted globally in multiple rounds to obtain the first intrinsic parameter estimates and first distortion coefficient estimates of the N acquisition devices.

[0015] In the above design, the accuracy of the calibrated intrinsic parameters and distortion coefficients can be improved by optimizing the intrinsic parameters and distortion coefficients through global optimization based on the principle of minimizing projection error.

[0016] In one possible design, the intrinsic parameter matrices and distortion coefficients of the N acquisition devices are adjusted through multiple rounds of global iteration to obtain the first intrinsic parameter estimates and the first distortion coefficient estimates of the N acquisition devices, including:

[0017] Based on the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the pose set of the target calibration object corresponding to each acquisition device, the estimated value of the second intrinsic parameter corresponding to each acquisition device, and the initially set distortion coefficient, the pixel coordinates of the multiple calibration points in the image coordinate system of each acquisition device are estimated.

[0018] The error between the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of each acquisition device and the pixel coordinates of the plurality of calibration points extracted from the image frames acquired from each acquisition device is obtained;

[0019] Based on the error, the pose set of the target calibration object corresponding to each acquisition device, the second intrinsic parameter estimate value corresponding to each acquisition device, and the initially set distortion coefficient are adjusted to obtain the intrinsic parameter estimate value and distortion coefficient corresponding to each acquisition device after the current round of adjustment;

[0020] In this round, the estimated intrinsic parameters and distortion coefficients of each acquisition device after the current round of adjustment are used as the basis for the next round of adjustment, until the C round of adjustment is completed to obtain the first estimated intrinsic parameters and the first estimated distortion coefficients of the N acquisition devices.

[0021] In one possible design, based on the three-dimensional coordinates of the plurality of calibration points in the calibration object coordinate system, the pose set of the target calibration object corresponding to each acquisition device, the estimated value of the second intrinsic parameter corresponding to each acquisition device, and the initially set distortion coefficient, the pixel coordinates of the plurality of calibration points in the image coordinate system of each acquisition device are estimated, including:

[0022] Based on the three-dimensional coordinates of multiple calibration points on the target calibration object at the k-th moving position, and the pose of the target calibration object at the k-th moving position corresponding to the i-th acquisition device, determine the coordinates of the multiple calibration points projected onto the camera coordinate system of the i-th acquisition device;

[0023] The distorted coordinates of the plurality of calibration points projected onto the camera coordinate system are determined based on the coordinates of the i-th acquisition device in the camera coordinate system and the distortion coefficient of the i-th acquisition device initially set.

[0024] The pixel coordinates of the plurality of calibration points projected onto the image coordinate system of the i-th acquisition device are estimated based on the distorted coordinates and the second intrinsic parameter estimate of the i-th acquisition device.

[0025] In one possible design, based on the first intrinsic parameter estimates of at least two acquisition devices in each of multiple acquisition device groups, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of the target calibration object including multiple calibration points in the calibration object coordinate system, the first extrinsic parameter estimate of each acquisition device is estimated, including:

[0026] Based on the first intrinsic parameter estimates of at least two acquisition devices in each of the multiple acquisition device groups, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, the second relative pose of the other acquisition devices among the multiple acquisition devices, excluding the reference acquisition device, relative to the reference acquisition device is obtained; the reference acquisition device is any one of the multiple acquisition devices.

[0027] A scale factor is determined, which is the ratio between a first distance and a second distance. The first distance is the distance between two calibration points on the target calibration object, and the second distance is the distance between the two calibration points in the same image coordinate system. The two calibration points are located on the same calibration surface on the target calibration object.

[0028] The first extrinsic parameter estimate of each acquisition device is obtained based on the second relative pose of each acquisition device and the scale factor.

[0029] In the above design, the external parameters are adjusted by scaling, which is simple.

[0030] In some embodiments, after determining the extrinsic parameters of each acquisition device based on the scale factor, the extrinsic parameters of each acquisition device can be globally optimized by using the principle of minimizing projection error.

[0031] In one possible design, based on the first intrinsic parameter estimates of at least two acquisition devices in each of the multiple acquisition device groups, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, the second relative pose of the other acquisition devices (excluding the reference acquisition device) relative to the reference acquisition device is obtained, including:

[0032] The essential matrix between the first acquisition device and the reference acquisition device is determined based on the matching feature point set corresponding to the first acquisition device group. The first acquisition device and the reference acquisition device belong to the first acquisition device group, and the first acquisition device group is one of the plurality of acquisition device groups.

[0033] Based on the singular value decomposition results of the essential matrix, the second relative pose between the first acquisition device and the reference acquisition device is determined.

[0034] In one possible design, obtaining a first extrinsic parameter estimate for each acquisition device based on the second relative pose of each acquisition device and the scale factor includes:

[0035] After determining the second relative poses of the acquisition devices in the g-th acquisition device group relative to the reference acquisition device, the three-dimensional coordinates of the multiple calibration points in the local coordinate system are determined according to the second relative poses of each acquisition device when the target calibration object moves to M2 moving positions. The local coordinate system is the camera coordinate system of the reference acquisition device. Any of the M2 moving positions is located at least within the common viewing area of ​​two acquisition devices in the g-th acquisition device.

[0036] Based on the three-dimensional coordinates of the multiple calibration points at the M2 moving positions in the local coordinate system, the second relative pose of the acquisition devices included in the g-th acquisition device group, and the first intrinsic parameter estimate, the pixel coordinates of the multiple calibration points at the M2 moving positions are estimated to be projected onto the image coordinate system of the acquisition devices included in the g-th acquisition device group.

[0037] The error between the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of the acquisition devices included in the g-th acquisition device group and the pixel coordinates of the plurality of calibration points extracted from the image frames acquired by the acquisition devices included in the g-th acquisition device group is obtained.

[0038] Based on the error, the second relative pose and the first intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group are adjusted to obtain the relative pose and intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group after the current round of adjustment.

[0039] Among them, the intrinsic parameter estimates and relative poses of the acquisition devices in the g-th acquisition device group after the current round of adjustment are used as the basis for the next round of adjustment, until the D-th round of adjustment is completed to obtain the third relative pose and third intrinsic parameter estimates of the acquisition devices included in the g-th acquisition device group;

[0040] The first extrinsic parameter estimate is obtained by adding the scale factor to the third relative pose of the acquisition devices included in the g-th acquisition device group.

[0041] In some embodiments, after determining the extrinsic parameters of each acquisition device based on the scale factor, the extrinsic parameters of each acquisition device are globally optimized using the principle of minimizing projection error. This can improve the accuracy of the calibrated extrinsic parameters. Furthermore, based on the optimization of the extrinsic parameters, the intrinsic parameters are also optimized, which can further improve the accuracy of the calibrated intrinsic parameters.

[0042] In one possible design, obtaining a first extrinsic parameter estimate for each acquisition device based on the second relative pose of each acquisition device and the scale factor includes:

[0043] After determining the second relative poses of the acquisition devices in the g-th acquisition device group relative to the reference acquisition device, the three-dimensional coordinates of the multiple calibration points in the local coordinate system are determined according to the second relative poses of each acquisition device when the target calibration object moves to M2 moving positions. The local coordinate system is the camera coordinate system of the reference acquisition device. Any of the M2 moving positions is located at least within the common viewing area of ​​two acquisition devices in the g-th acquisition device.

[0044] Based on the three-dimensional coordinates of the multiple calibration points at the M2 moving positions in the local coordinate system, the second relative pose of the acquisition devices included in the g-th acquisition device group, the first intrinsic parameter estimate, and the first distortion coefficient estimate, the pixel coordinates of the multiple calibration points at the M2 moving positions projected onto the image coordinate system of the acquisition devices included in the g-th acquisition device group are estimated.

[0045] The error between the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of the acquisition devices included in the g-th acquisition device group and the pixel coordinates of the plurality of calibration points extracted from the image frames acquired by the acquisition devices included in the g-th acquisition device group is obtained.

[0046] The relative pose, the first intrinsic parameter estimate, and the first distortion coefficient of the acquisition devices included in the g-th acquisition device group are adjusted according to the error to obtain the relative pose and intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group after the current round of adjustment.

[0047] Among them, the intrinsic parameter estimates, relative poses and distortion coefficients of the acquisition devices in the g-th acquisition device group after the current round of adjustment are used as the basis for the next round of adjustment, until the D-th round of adjustment is completed to obtain the third relative pose, third intrinsic parameter estimates and second distortion coefficients of the acquisition devices included in the g-th acquisition device group;

[0048] The first extrinsic parameter estimate is obtained by adding the scale factor to the third relative pose of the acquisition devices included in the g-th acquisition device group.

[0049] In some embodiments, after determining the extrinsic parameters of each acquisition device based on the scale factor, the extrinsic parameters of each acquisition device are globally optimized using the principle of minimizing projection error. This can improve the accuracy of the calibrated extrinsic parameters. Furthermore, based on the optimized extrinsic parameters, the intrinsic parameters and distortion coefficients are also optimized, which can further improve the accuracy of the calibrated intrinsic parameters and distortion coefficients.

[0050] In one possible design, each of the plurality of acquisition device groups includes two acquisition devices. Based on the first intrinsic parameter estimates of at least two acquisition devices in each of the plurality of acquisition device groups, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of the target calibration object including multiple calibration points in the calibration object coordinate system, the first extrinsic parameter estimate of each acquisition device is estimated, including:

[0051] The relative poses of multiple pairs of movement positions among the M movement positions of the target calibration object are determined. A first pair of movement positions includes a first movement position and a second movement position. The first and second movement positions are two of the M movement positions of the target calibration object, and the first and second movement positions are located within the common viewing area of ​​at least one group of acquisition devices. The relative poses of the first pair of movement positions are determined based on the matching feature point set corresponding to at least one group of acquisition devices, the three-dimensional coordinates of multiple calibration points of the target calibration object in the calibration object coordinate system, and the first intrinsic parameter estimates of the two acquisition devices included in the at least one group of acquisition devices.

[0052] The poses of M movement positions in the world coordinate system are determined based on the coordinates of the base movement position in the world coordinate system and the relative poses of multiple movement position pairs, wherein the base movement position is one of the M movement positions;

[0053] Based on the coordinates of M moving positions in the world coordinate system, the coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, and the pixel coordinates of the calibration points on the target calibration object in the image frames acquired by each acquisition device, the camera parameters of each acquisition device are globally optimized. The camera parameters include an intrinsic parameter matrix and an extrinsic parameter matrix, or the camera parameters include an intrinsic parameter matrix, an extrinsic parameter matrix, and distortion coefficients.

[0054] In the global optimization process, the camera parameters of each acquisition device and the pose of the calibration surface where the multiple calibration points are located in the calibration object coordinate system are taken as the quantities to be optimized; the first intrinsic parameter estimate of each acquisition device in the global optimization is taken as the initial value of the intrinsic parameter matrix of each acquisition device.

[0055] In the above design, the three-dimensional coordinates of each calibration point in space are determined by setting a basic moving position (reference point) in space. Then, based on the principle of minimizing projection error, the camera parameters of each acquisition device are determined through global optimization, which can improve the accuracy of calibration.

[0056] In one possible design, the relative poses of the first moving position pair satisfy the following condition:

[0057]

[0058] Among them, T 12 This indicates the relative pose between the first moving position and the second moving position. The at least one group of acquisition devices includes a first group of acquisition devices, which in turn includes a first acquisition device and a second acquisition device. It represents the pose from the first moving position to the second moving position, determined based on the pixel coordinates of the calibration point in the image frame acquired by the first acquisition device when the target calibration object moves to the first moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the second moving position; This indicates the pose from the second moving position to the first moving position, determined based on the pixel coordinates of the calibration point in the image frame acquired by the second acquisition device when the target calibration object moves to the second moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the first moving position.

[0059] In one possible design, at least one group of acquisition devices consists of L devices, and the first group of acquisition devices satisfies:

[0060]

[0061] Where I represents the identity matrix, This represents the pose from the first moving position to the second moving position, determined based on the pixel coordinates of the calibration point in the image frame acquired by the first acquisition device in the l-th acquisition device group when the target calibration object moves to the first moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the second moving position. 11 represents the pose from the second moving position to the first moving position, determined by the pixel coordinates of the calibration point in the image frame acquired by the second acquisition device in the l-th acquisition device group when the target calibration object moves to the second moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the first moving position; 12 represents the first acquisition device in the first acquisition device group and 12 represents the second acquisition device in the second acquisition device group.

[0062] In one possible design, determining the poses of M mobile positions in the world coordinate system based on the coordinates of the base mobile position in the world coordinate system and the relative poses of multiple mobile position pairs includes:

[0063] Determine the confidence weight for each of the multiple movement position pairs, where the confidence weight between the first movement position and the second movement position satisfies the following condition: S 12 This represents the confidence weight between the first and second movement positions.

[0064] The shortest path from the third move position to the base move position is determined based on the credibility weight of each move position pair.

[0065] The shortest path is the path with the smallest confidence weight among all paths from the third moving position to the base moving position; the confidence weight of any path is the sum of the confidence weights of the moving position pairs traversed by the path.

[0066] The pose of the third moving position is determined based on the relative poses of the moving position pairs traversed by the shortest path.

[0067] Secondly, embodiments of this application provide a calibration device, comprising:

[0068] An acquisition unit is used to acquire multiple video streams collected by multiple acquisition devices deployed in a designated space within a sports field. The multiple video streams are synchronously captured by the multiple acquisition devices during the movement of a target calibration object within the sports field. The movement trajectory of the target calibration object within the sports field at least covers a designated area of ​​the sports field. The target calibration object includes at least two non-coplanar calibration surfaces, and each calibration surface includes at least two calibration points. Each video stream acquired by the acquisition device includes multiple image frames.

[0069] The processing unit is configured to perform calibration point detection on the image frames acquired by each of the plurality of acquisition devices to obtain the pixel coordinates of multiple calibration points on the target calibration object in the image frames acquired by each acquisition device; based on the pixel coordinates of the multiple calibration points included in the target calibration object in the image frames acquired by each acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system of the target calibration object, estimate the intrinsic parameter matrix of each acquisition device to obtain a first intrinsic parameter estimate value for each acquisition device; and based on the first intrinsic parameter estimate values ​​of at least two acquisition devices included in each acquisition device group in the plurality of acquisition device groups, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of the multiple calibration points included in the target calibration object in the calibration object coordinate system, determine a first extrinsic parameter estimate value for each acquisition device.

[0070] In this system, each acquisition device group includes at least two acquisition devices that share a common viewing area. The matching feature point set includes multiple matching feature point groups, and each matching feature point group includes at least two matching pixel coordinates. The at least two matching pixel coordinates are the pixel coordinates of the same calibration point detected by different acquisition devices belonging to the same acquisition device group at the same time. The multiple acquisition device groups are obtained by grouping the multiple acquisition devices. Any two acquisition device groups in the multiple acquisition device groups include at least one identical acquisition device.

[0071] In one possible design, the processing unit is further configured to:

[0072] Based on the pixel coordinates of multiple calibration points included in the target calibration object in the image frame acquired by each acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system of the target calibration object, the distortion coefficient of each acquisition device is estimated to obtain the first distortion coefficient estimate of each acquisition device.

[0073] In one possible design, the processing unit is specifically used for:

[0074] Based on the pixel coordinates of the calibration points on the target calibration object in the image frames included in the image set acquired by the i-th acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the intrinsic parameter matrix of the i-th acquisition device is estimated to obtain the second intrinsic parameter estimate of the i-th acquisition device; the image set includes M1 image frames of the target calibration object in the video stream acquired by the i-th acquisition device, and the M1 image frames correspond one-to-one with the M1 movement positions of the target calibration object; M1 is a positive integer, and M is an integer greater than M1;

[0075] Based on the second intrinsic parameter estimate of the i-th acquisition device, the pixel coordinates of the calibration points on the target calibration object in the image frames included in the image set acquired by the i-th acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the pose set corresponding to the i-th acquisition device is estimated respectively. The pose set corresponding to the i-th acquisition device includes the pose of the target calibration object relative to the i-th acquisition device at M1 moving positions; i takes the value of a positive integer less than or equal to N, where N is the number of acquisition devices deployed in the set space of the sports field;

[0076] Among them, the range of movement positions corresponding to the image frames acquired by different acquisition devices is different;

[0077] Based on the pixel coordinates of the calibration points on the target calibration object in the image frames of the image sets acquired by N acquisition devices, the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, and the poses of the target calibration object corresponding to the N acquisition devices, and on the basis of the initially set distortion coefficients and second intrinsic parameter estimates corresponding to the N acquisition devices, the intrinsic parameter matrices and distortion coefficients of the N acquisition devices are adjusted globally in multiple rounds to obtain the first intrinsic parameter estimates and first distortion coefficient estimates of the N acquisition devices.

[0078] In one possible design, the processing unit is specifically used for:

[0079] Based on the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the pose set of the target calibration object corresponding to each acquisition device, the estimated value of the second intrinsic parameter corresponding to each acquisition device, and the initially set distortion coefficient, the pixel coordinates of the multiple calibration points in the image coordinate system of each acquisition device are estimated.

[0080] The error between the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of each acquisition device and the pixel coordinates of the plurality of calibration points extracted from the image frames acquired from each acquisition device is obtained;

[0081] Based on the error, the pose set of the target calibration object corresponding to each acquisition device, the second intrinsic parameter estimate value corresponding to each acquisition device, and the initially set distortion coefficient are adjusted to obtain the intrinsic parameter estimate value and distortion coefficient corresponding to each acquisition device after the current round of adjustment;

[0082] In this round, the estimated intrinsic parameters and distortion coefficients of each acquisition device after the current round of adjustment are used as the basis for the next round of adjustment, until the C round of adjustment is completed to obtain the first estimated intrinsic parameters and the first estimated distortion coefficients of the N acquisition devices.

[0083] In one possible design, the processing unit is specifically used for:

[0084] Based on the three-dimensional coordinates of multiple calibration points on the target calibration object at the k-th moving position, and the pose of the target calibration object at the k-th moving position corresponding to the i-th acquisition device, determine the coordinates of the multiple calibration points projected onto the camera coordinate system of the i-th acquisition device;

[0085] The distorted coordinates of the plurality of calibration points projected onto the camera coordinate system are determined based on the coordinates of the i-th acquisition device in the camera coordinate system and the distortion coefficient of the i-th acquisition device initially set.

[0086] The pixel coordinates of the plurality of calibration points projected onto the image coordinate system of the i-th acquisition device are estimated based on the distorted coordinates and the second intrinsic parameter estimate of the i-th acquisition device.

[0087] In one possible design, the processing unit is specifically used for:

[0088] Based on the first intrinsic parameter estimates of at least two acquisition devices in each of the multiple acquisition device groups, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, the second relative pose of the other acquisition devices among the multiple acquisition devices, excluding the reference acquisition device, relative to the reference acquisition device is obtained; the reference acquisition device is any one of the multiple acquisition devices.

[0089] A scale factor is determined, which is the ratio between a first distance and a second distance. The first distance is the distance between two calibration points on the target calibration object, and the second distance is the distance between the two calibration points in the same image coordinate system. The two calibration points are located on the same calibration surface on the target calibration object.

[0090] The first extrinsic parameter estimate of each acquisition device is obtained based on the second relative pose of each acquisition device and the scale factor.

[0091] In one possible design, the processing unit is specifically used for:

[0092] The essential matrix between the first acquisition device and the reference acquisition device is determined based on the matching feature point set corresponding to the first acquisition device group. The first acquisition device and the reference acquisition device belong to the first acquisition device group, and the first acquisition device group is one of the plurality of acquisition device groups.

[0093] Based on the singular value decomposition results of the essential matrix, the second relative pose between the first acquisition device and the reference acquisition device is determined.

[0094] In one possible design, the processing unit is specifically used for:

[0095] After determining the second relative poses of the acquisition devices in the g-th acquisition device group relative to the reference acquisition device, the three-dimensional coordinates of the multiple calibration points in the local coordinate system are determined according to the second relative poses of each acquisition device when the target calibration object moves to M2 moving positions. The local coordinate system is the camera coordinate system of the reference acquisition device. Any of the M2 moving positions is located at least within the common viewing area of ​​two acquisition devices in the g-th acquisition device.

[0096] Based on the three-dimensional coordinates of the multiple calibration points at the M2 moving positions in the local coordinate system, the second relative pose of the acquisition devices included in the g-th acquisition device group, and the first intrinsic parameter estimate, the pixel coordinates of the multiple calibration points at the M2 moving positions are estimated to be projected onto the image coordinate system of the acquisition devices included in the g-th acquisition device group.

[0097] The error between the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of the acquisition devices included in the g-th acquisition device group and the pixel coordinates of the plurality of calibration points extracted from the image frames acquired by the acquisition devices included in the g-th acquisition device group is obtained.

[0098] Based on the error, the second relative pose and the first intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group are adjusted to obtain the relative pose and intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group after the current round of adjustment.

[0099] Among them, the intrinsic parameter estimates and relative poses of the acquisition devices in the g-th acquisition device group after the current round of adjustment are used as the basis for the next round of adjustment, until the D-th round of adjustment is completed to obtain the third relative pose and third intrinsic parameter estimates of the acquisition devices included in the g-th acquisition device group;

[0100] The first extrinsic parameter estimate is obtained by adding the scale factor to the third relative pose of the acquisition devices included in the g-th acquisition device group.

[0101] In one possible design, the processing unit is specifically used for:

[0102] After determining the second relative poses of the acquisition devices in the g-th acquisition device group relative to the reference acquisition device, the three-dimensional coordinates of the multiple calibration points in the local coordinate system are determined according to the second relative poses of each acquisition device when the target calibration object moves to M2 moving positions. The local coordinate system is the camera coordinate system of the reference acquisition device. Any of the M2 moving positions is located at least within the common viewing area of ​​two acquisition devices in the g-th acquisition device.

[0103] Based on the three-dimensional coordinates of the multiple calibration points at the M2 moving positions in the local coordinate system, the second relative pose of the acquisition devices included in the g-th acquisition device group, the first intrinsic parameter estimate, and the first distortion coefficient estimate, the pixel coordinates of the multiple calibration points at the M2 moving positions projected onto the image coordinate system of the acquisition devices included in the g-th acquisition device group are estimated.

[0104] The error between the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of the acquisition devices included in the g-th acquisition device group and the pixel coordinates of the plurality of calibration points extracted from the image frames acquired by the acquisition devices included in the g-th acquisition device group is obtained.

[0105] Based on the error, the second relative pose and the first intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group are adjusted to obtain the relative pose and intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group after the current round of adjustment.

[0106] Among them, the intrinsic parameter estimates and relative poses of the acquisition devices in the g-th acquisition device group after the current round of adjustment are used as the basis for the next round of adjustment, until the D-th round of adjustment is completed to obtain the third relative pose and third intrinsic parameter estimates of the acquisition devices included in the g-th acquisition device group;

[0107] The first extrinsic parameter estimate is obtained by adding the scale factor to the third relative pose of the acquisition devices included in the g-th acquisition device group.

[0108] In one possible design, each of the plurality of acquisition device groups includes two acquisition devices, and the processing unit is specifically used for:

[0109] The relative poses of multiple pairs of movement positions among the M movement positions of the target calibration object are determined. A first pair of movement positions includes a first movement position and a second movement position. The first and second movement positions are two of the M movement positions of the target calibration object, and the first and second movement positions are located within the common viewing area of ​​at least one group of acquisition devices. The relative poses of the first pair of movement positions are determined based on the matching feature point set corresponding to at least one group of acquisition devices, the three-dimensional coordinates of multiple calibration points of the target calibration object in the calibration object coordinate system, and the first intrinsic parameter estimates of the two acquisition devices included in the at least one group of acquisition devices.

[0110] The poses of M movement positions in the world coordinate system are determined based on the coordinates of the base movement position in the world coordinate system and the relative poses of multiple movement position pairs, wherein the base movement position is one of the M movement positions;

[0111] Based on the coordinates of M moving positions in the world coordinate system, the coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, and the pixel coordinates of the calibration points on the target calibration object in the image frames acquired by each acquisition device, the camera parameters of each acquisition device are globally optimized. The camera parameters include an intrinsic parameter matrix and an extrinsic parameter matrix, or the camera parameters include an intrinsic parameter matrix, an extrinsic parameter matrix, and distortion coefficients.

[0112] In the global optimization process, the camera parameters of each acquisition device and the pose of the calibration surface where the multiple calibration points are located in the calibration object coordinate system are taken as the quantities to be optimized; the first intrinsic parameter estimate of each acquisition device in the global optimization is taken as the initial value of the intrinsic parameter matrix of each acquisition device.

[0113] In one possible design, the relative poses of the first moving position pair satisfy the following condition:

[0114]

[0115] Among them, T 12 This indicates the relative pose between the first moving position and the second moving position. The at least one group of acquisition devices includes a first group of acquisition devices, which in turn includes a first acquisition device and a second acquisition device. It represents the pose from the first moving position to the second moving position, determined based on the pixel coordinates of the calibration point in the image frame acquired by the first acquisition device when the target calibration object moves to the first moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the second moving position; This indicates the pose from the second moving position to the first moving position, determined based on the pixel coordinates of the calibration point in the image frame acquired by the second acquisition device when the target calibration object moves to the second moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the first moving position.

[0116] In one possible design, at least one group of acquisition devices consists of L devices, and the first group of acquisition devices satisfies:

[0117]

[0118] Where I represents the identity matrix, This represents the pose from the first moving position to the second moving position, determined based on the pixel coordinates of the calibration point in the image frame acquired by the first acquisition device in the l-th acquisition device group when the target calibration object moves to the first moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the second moving position. 11 represents the pose from the second moving position to the first moving position, determined by the pixel coordinates of the calibration point in the image frame acquired by the second acquisition device in the l-th acquisition device group when the target calibration object moves to the second moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the first moving position; 12 represents the first acquisition device in the first acquisition device group and 12 represents the second acquisition device in the second acquisition device group.

[0119] In one possible design, determining the poses of M mobile positions in the world coordinate system based on the coordinates of the base mobile position in the world coordinate system and the relative poses of multiple mobile position pairs includes:

[0120] Determine the confidence weight for each of the multiple movement position pairs, where the confidence weight between the first movement position and the second movement position satisfies the following condition: S 12 This represents the confidence weight between the first and second movement positions.

[0121] The shortest path from the third move position to the base move position is determined based on the credibility weight of each move position pair.

[0122] The shortest path is the path with the smallest confidence weight among all paths from the third moving position to the base moving position; the confidence weight of any path is the sum of the confidence weights of the moving position pairs traversed by the path.

[0123] The pose of the third moving position is determined based on the relative poses of the moving position pairs traversed by the shortest path.

[0124] Thirdly, embodiments of this application provide a calibration apparatus, including a memory and a processor. The memory is used to store programs or instructions; the processor is used to invoke the programs or instructions to execute the method described in the first aspect or any design of the first aspect.

[0125] Fourthly, this application provides a computer-readable storage medium storing a computer program or instructions that, when executed by a terminal device, cause the processor to perform the methods described in the first aspect or any possible design of the first aspect.

[0126] Fifthly, this application provides a computer program product comprising a computer program or instructions that, when executed by a processor, implement the method described in the first aspect or any possible implementation thereof.

[0127] The technical effects that can be achieved by any of the second to fifth aspects mentioned above can be referred to the description of the beneficial effects in the first aspect mentioned above, and will not be repeated here.

[0128] Based on the implementations provided in the above aspects, this application can be further combined to provide more implementations. Attached Figure Description

[0129] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below.

[0130] Figure 1 This is a schematic diagram of the image coordinate system provided in the embodiments of this application;

[0131] Figure 2 This is a schematic diagram of the camera coordinate system provided in an embodiment of this application;

[0132] Figure 3 A schematic diagram of an information system architecture provided for an embodiment of this application;

[0133] Figure 4This application provides another information system architecture diagram.

[0134] Figure 5 A schematic diagram illustrating a camera deployment method for an athletics track and field stadium, provided as an embodiment of this application;

[0135] Figure 6 A schematic diagram illustrating another camera deployment method for an athletics stadium, provided as an embodiment of this application;

[0136] Figure 7 A schematic diagram illustrating another camera deployment method for an athletics track and field stadium, provided as an embodiment of this application;

[0137] Figure 8 A schematic diagram illustrating a camera deployment method for a football field, provided as an embodiment of this application;

[0138] Figure 9 This is a schematic diagram of the calibration method provided in the embodiments of this application;

[0139] Figure 10 A schematic diagram of the target calibrator provided in the embodiments of this application;

[0140] Figure 11 A schematic diagram of the calibration tower provided in the embodiments of this application;

[0141] Figure 12A A schematic diagram of the movement trajectory of a target marker provided in an embodiment of this application;

[0142] Figure 12B This is a schematic diagram of another target calibration object movement trajectory provided in an embodiment of this application;

[0143] Figure 12C This is another schematic diagram of the movement trajectory of a target marker provided in the embodiments of this application;

[0144] Figure 12D This is another schematic diagram of the movement trajectory of a target marker provided in the embodiments of this application;

[0145] Figure 13 This is a schematic diagram of feature point filtering provided in an embodiment of this application;

[0146] Figure 14 A schematic flowchart illustrating a first possible method for determining external parameters provided in this application embodiment;

[0147] Figure 15 A flowchart illustrating the optimized intrinsic parameters and relative pose provided in the embodiments of this application;

[0148] Figure 16 A flowchart illustrating the optimization of intrinsic parameters, relative pose, and distortion coefficients provided in the embodiments of this application;

[0149] Figure 17 This is a schematic diagram of the movement position of the target marker represented by the graphical model provided in the embodiments of this application;

[0150] Figure 18 This is a schematic flowchart illustrating a second possible method for determining external parameters provided in an embodiment of this application.

[0151] Figure 19 This is a schematic diagram of a calibration device structure provided in an embodiment of this application;

[0152] Figure 20 This is a schematic diagram of another calibration device structure provided in an embodiment of this application. Detailed Implementation

[0153] The following explanations of some terms used in this application are provided to facilitate understanding by those skilled in the art.

[0154] 1) Camera intrinsic parameters: Distortion parameters (k1, k2, k3, p1, p2) and focal length (f) in the pinhole camera model. x ,f y The center point is (u0, v0). The intrinsic parameter matrix involved in the embodiments of this application refers to the matrix composed of the focal length and the center point.

[0155] Camera distortion refers to the degree of distortion in the image formed by a camera's optical system relative to the object itself. It is an inherent characteristic of optical lenses, and its direct cause is the difference in magnification between the edge and center portions of the lens in a camera. Camera distortion mainly includes radial distortion and tangential distortion.

[0156] Radial distortion: mainly caused by the different magnification of different parts of the camera lens, and is divided into pincushion distortion and barrel distortion.

[0157] Tangential distortion: mainly caused by the camera lens not being perpendicular to the imaging plane, similar to the principle of perspective (objects appear larger when closer and smaller when farther away, circles appear as ellipses, etc.).

[0158] The camera distortion formulas are shown in formulas (1-1) and (1-2). There are three radial distortion coefficients, denoted by k1, k2, and k3. There are two tangential distortion coefficients, denoted by p1 and p2. The distorted pixel coordinates (x′, y′) and the undistorted pixel coordinates (x, y) satisfy the conditions shown in formulas (1-1) and (1-2).

[0159] x′=x(1+k1r 2 +k2r 4 +k3r 6 )+[2p1xy+p2(r2 +2x 2 )] Formula (1-1)

[0160] y′=y(1+k1r 2 +k2r 4 +k3r 6 )+[p1(r 2 +2y 2 Formula (1-2)[)+2p2xy]

[0161] Where x, y are the normalized coordinates of a 3D point projected onto the camera coordinate system, and x′, y′ are the distorted coordinates. 2 =x 2 +y 2 .

[0162] 2) Camera extrinsic parameters: The rotation and translation transformations of the camera in a pinhole camera model relative to a coordinate system (such as the world coordinate system or a reference camera pose), i.e., the 6 degrees of freedom (6DOF) pose of the camera in this coordinate system. This represents translation along three directions and rotation about three axes.

[0163] 3) Image physical coordinate system (referred to as image coordinate system).

[0164] A concrete example of an image physical coordinate system can be found in [link to relevant documentation]. Figure 1 As shown, O i The point is the intersection of the camera's optical axis and its imaging plane, and it is the origin of the image's physical coordinate system. (u, v) represent the number of columns and rows of pixels, where (O... p The pixel plane coordinate system is formed by (u, v). The origin O of the pixel plane coordinate system is... p Located at the upper left corner of the camera's imaging plane, the two coordinate axes (O) p u-axis and O p The v-axis points to the right and downwards respectively. The origin O of the image's physical coordinate system is... i Located at the center of the pixel plane coordinate system, with coordinates (u0, v0), and two coordinate axes (O... i x-axis and O i The y-axis points to the right and down respectively. Let dx and dy represent the physical dimensions of a pixel along the u-axis and v-axis respectively, then the relationship between the pixel plane coordinate system and the image physical coordinate system is shown in formula (2).

[0165]

[0166] Formula (2) above can be further expressed as formula (3).

[0167]

[0168] 4) Camera coordinate system.

[0169] For a concrete example of a camera coordinate system, please refer to Figure 2 As shown, the camera coordinate system is a spatial coordinate system with its origin O. c Located at the optical center of the camera. i The point is the intersection of the camera's optical axis and its imaging plane, i.e., the origin of the image's physical coordinate system. For example... Figure 2 As shown, the O coordinate system of the camera c x c Shaft and O c y c The axes are parallel to the O of the image's physical coordinate system. i x-axis and O i y-axis, O c z c The axis passes through point O i Point O i Point O c The distance is the focal length, denoted by f. According to the imaging projection relationship, the camera coordinate system and the image physical coordinate system satisfy the following relationship as shown in formula (4):

[0170]

[0171] 5) World coordinate system.

[0172] A reference coordinate system is chosen in the environment to describe the positions of the camera and objects; this coordinate system can be called the world coordinate system. The relationship between the camera coordinate system and the world coordinate system can be described by the rotation matrix R and the translation parameter T. Therefore, the homogeneous coordinates of a point P in space in the world coordinate system and the camera coordinate system are respectively (x...). w ,y w ,z w ) and (x c ,y c ,z c ), satisfying the relationship shown in the following formula (5).

[0173]

[0174] Where R is a 3×3 rotation matrix and T is a 3×1 translation parameter.

[0175] Based on the relationship between the pixel plane coordinate system, image physical coordinate system, camera coordinate system and world coordinate system mentioned above, the relationship shown in formula (5) is as follows.

[0176]

[0177] In formula (5), This is the intrinsic parameter matrix. The right side of the equals sign This is the extrinsic parameter matrix. dx and dy represent the length units occupied by one pixel in the x and y directions, respectively, that is, the actual physical value represented by one pixel, which is the key to realizing the transformation between the camera coordinate system and the image coordinate system.

[0178] This application provides a calibration method and apparatus for calibrating camera parameters of a data acquisition device deployed in a designated space within a designated site. The camera parameters include an intrinsic parameter matrix, an extrinsic parameter matrix (and distortion coefficients). The designated site can be a circular area, such as a circular running track or a circular speed skating rink. The sports scene can also be a straight track. The sports field can also be other forms, such as a football field, etc., and this application does not specifically limit these.

[0179] See Figure 3 The diagram shown is a schematic representation of an information system architecture provided in an embodiment of this application. The information system includes multiple data acquisition devices and a data processing server. Figure 3 Taking N acquisition devices as an example, where N is a positive integer, the number of cameras included in the information system can be configured according to the size of the sports field. The acquisition devices can be cameras, video cameras, camcorders, etc. These multiple acquisition devices can be deployed within a designated space where the sports field is located. For example, if the sports field is a football field, located in an open-air stadium or football stadium, multiple acquisition devices can be deployed within the open-air stadium or football stadium. The field of view of each acquisition device includes a portion of the sports field. Different acquisition devices have different field of view, and at least two adjacent acquisition devices must share a common field of view. The common field of view is the area captured by both acquisition devices at the same time.

[0180] A data processing server can include one or more servers. If it includes multiple servers, it can be understood as a server cluster composed of multiple servers. For example, the data processing server can operate in two different modes: calibration mode and competition recording mode. In calibration mode, the data processing server performs calibration processing and stores the calibration results. Calibration processing can include obtaining the intrinsic and extrinsic parameters (and distortion coefficients) of multiple cameras. In competition recording mode, the data processing server can be used to process video streams captured by multiple acquisition devices. It can extract synchronized frames from multiple acquisition devices and then perform visual algorithm processing on the synchronized frames frame by frame based on the calibration results to generate spatial video. This spatial video can then be used for motion analysis, athlete technique review, or to extract highlights, etc., and can be sent to a broadcast van, etc.

[0181] In some possible scenarios, the information system may also include one or more routing devices. These routing devices can be used to transmit images captured by the acquisition devices to the data processing server. Routing devices can be routers, switches, etc. For example, see [link to example]. Figure 4 As shown, an information system can deploy multiple layers of switches. Taking a two-layer system as an example, the first-layer switch can connect one or more data acquisition devices, while the second-layer switch can act as the master switch. One end of the master switch connects to the first-layer switch, and the other end connects to the data processing server. For example, see... Figure 4 As shown.

[0182] In other possible scenarios, the information system also supports sending spatial video data to the broadcast van. The information system also supports acquiring motion analysis data via terminal devices. For example, the information system also includes a mobile front-end. For instance, the mobile front-end includes a web page server. See also Figure 4 As shown, the web page server is connected to the data processing server. The mobile front end may also include a wireless router (or wired router), a broadcast van, or terminal devices. Terminal devices can be desktop computers, laptops, mobile phones, or other electronic devices that support web page access. Terminal devices can operate the data server by accessing the web page server, such as sending synchronous acquisition signals or stop recording signals to multiple acquisition devices. Synchronous acquisition signals instruct acquisition devices to synchronously start video recording. Stop recording signals instruct acquisition devices to stop video recording. Other examples include historical video playback or display of motion information.

[0183] The calibration method provided in this application will be described in detail below with reference to embodiments. The acquisition equipment is deployed within a designated space belonging to a set venue (such as a sports field, conference venue, etc.). In the following description, a sports field is used as an example, and a camera is used as the acquisition equipment. When deploying cameras in the designated space belonging to a sports field, the deployment can be determined based on the permissible installation points within that space, such as the presence of columns, trusses, or ceilings. Each camera can cover a portion of the entire track, such as a section of track. Adjacent cameras share a common viewing area, such as 1 / 2 or 1 / 3 of the image. A truss refers to a planar or spatial structure generally composed of straight bars with triangular units, used for custom-made camera supports.

[0184] As an example, consider the deployment of cameras on an athletics track. Figures 5-8 The diagram illustrates three possible camera deployment methods. See also... Figure 5As shown in (a), this example uses 20 camera positions deployed along the track. A camera position refers to a camera located in a different position. Each camera position is located above and outside the track, taking a top-down view of the track from a high vantage point. Figure 5 (a) Takes the camera deployed on a column as an example. Figure 5 (b) is a top view of the camera deployment. Figure 5 Image (c) is a side view of the cameras deployed on the pillars. Cameras are deployed along the extension of the straight section and along the sides of the curves. Each camera captures a 40-meter area, with adjacent cameras sharing a 20-meter field of view. A total of 20 cameras cover a 400-meter track (5 cameras * 2 on the straight section, 5 cameras * 2 on the curves). In some scenarios, after the cameras are installed in their designated positions, the focus, orientation, or field of view can be adjusted to allow each camera to focus on a specific area of ​​the track, with shared fields of view between adjacent camera positions. The cameras are grouped and connected to two switches: cameras 1-10 are connected to one switch, and cameras 11-20 are connected to the other switch. Video frames captured by cameras 1-20 are sent to the data processing server via the two switches.

[0185] See Figure 6 As shown in (a), this example illustrates the deployment of 20 cameras along the track of an athletics field. The cameras are positioned on a suspended truss. Each camera is located above the track, providing a bird's-eye view. The camera lens axis forms an acute angle with the ground, not perpendicular, to cover a wider shooting area. Figure 6 (b) is a top view of the camera deployment. Figure 6 Image (c) is a side view of the cameras deployed on the ceiling truss. In some scenarios, after the cameras are installed in their designated positions, the focus, orientation, or field of view of each camera can be adjusted to focus on a specific area of ​​the track, with shared viewing areas between adjacent camera positions. The cameras are grouped and connected to two switches: cameras 1-10 are connected to one switch, and cameras 11-20 are connected to the other switch. Video frames captured by cameras 1-20 are sent to the data processing server through the two switches.

[0186] See Figure 7 As shown in (a), this example uses 20 camera positions deployed along the track. The cameras are positioned on pillars. The 20 camera positions are deployed along the track: positions 1-5 film the straightaways outside the lane change zone; positions 6-10 and 16-20 film the two curves; and positions 11-15 film the straightaways outside the lane change zone. Figure 7 Figure (b) shows a side view of cameras 1-5 deployed on the column. The camera groups are connected to two switches. Cameras 1-10 are each connected to one switch, and cameras 11-20 are each connected to the other switch. The video frames captured by cameras 1-20 are sent to the data processing server through the two switches.

[0187] By deploying cameras on track and field, it is possible to analyze athletes participating in competitions such as sprints, middle-distance and long-distance running, hurdles, high jump, and long jump to obtain athletes' sports information or highlights.

[0188] As another example, consider the deployment of cameras in a football stadium. Cameras can be deployed on pillars, trusses, or ceilings, or in designated locations within the stands. See, for example... Figure 8 As shown, this example uses 20 cameras deployed along the track. In some scenarios, after installing the cameras in their designated positions, the focus, orientation, or field of view of each camera can be adjusted to ensure that each camera focuses on a specific area of ​​the stadium, with adjacent cameras sharing a common field of view. The cameras are grouped and connected to two switches: cameras 1-10 are connected to one switch, and cameras 11-20 are connected to the other. Video frames captured by cameras 1-20 are sent to the data processing server through the two switches. By deploying cameras in a football stadium, it's possible to capture exciting moments in front of the goal or instances of fouls, etc.

[0189] It should be noted that the camera deployment described above is merely an example, and the specific deployment can be tailored to the actual scenario. This application does not impose specific limitations on the number of cameras deployed, the grouping of cameras, or the number of switches deployed.

[0190] After the cameras are deployed, the intrinsic parameter matrix, extrinsic parameter matrix, and distortion coefficients of each camera need to be calibrated.

[0191] In this embodiment of the application, the calibration of the camera intrinsic parameter matrix, extrinsic parameter matrix, and distortion coefficient is achieved by having each camera acquire video streams during the movement of the target calibration object.

[0192] The following combination Figure 9 The calibration process for camera parameters is explained in detail. Figure 9 The provided method can be executed by the data processing server, or by the processor or processor system within the data processing server.

[0193] 901. Acquire multiple video streams captured by multiple cameras, wherein the multiple cameras are deployed in a designated space of the sports field, and the multiple video streams are obtained by the multiple cameras simultaneously capturing images during the movement of the target calibration object in the sports field; the movement trajectory of the target calibration object in the sports field at least covers the designated area of ​​the sports field, and the target calibration object includes at least two non-coplanar calibration surfaces, each calibration surface including at least two calibration points.

[0194] In some embodiments, the target calibrator may include one or a group of calibrators. Exemplarily, a group of calibrators may consist of multiple calibrators. Each calibrator may include at least two calibration surfaces, and each calibration surface includes at least one calibration point. The calibration point has stable visual characteristics that do not change over time. In some possible examples, the target calibrator has a specific pattern, and the intersections of the lines in the pattern can serve as calibration points. In other possible examples, the first target calibrator may have a light-emitting screen, and the displayed light-emitting points serve as calibration points. Of course, other methods can also be used to set the calibration points on the calibrator, and this application embodiment does not specifically limit this.

[0195] As an example, consider a target marker with a specific pattern. Figure 10 The diagram shown is a possible target calibration structure. Figure 10 Taking a set of target markers as an example, each marker in the set can be a box or frame. The box or frame has specific patterns on its four sides, with different patterns on different sides. Different boxes have different specific patterns. Figure 10 Taking a QR code as an example, the points at each corner of the QR code can be selected as calibration points, or the corner points of a specific grid in the box can be selected as calibration points, or the two bottom corner points of the box can be selected as calibration points, or the two bottom corner points of the rectangle including the QR code can be selected as calibration points, etc. This application does not specifically limit this.

[0196] In some embodiments, the target calibration object may take the form of a tower structure, and thus the target calibration object may also be called a calibration tower, see [link to documentation]. Figure 11 As shown. For ease of movement, the target can be placed on a wheeled flatbed truck.

[0197] For example, the trajectory of the target marker uniformly covers the shooting area (i.e., the movement field). The trajectory includes, but is not limited to, regular paths (e.g., polygonal, spiral), and random paths, etc. See also Figure 12A The diagram shows a spiral path. See also... Figure 12B and Figure 12C The image shows a broken line path. See also... Figure 12D The image shows a random path. Figures 12A-12D The black dots represent cameras. This application embodiment does not impose specific restrictions on the trajectory direction of the path. The starting point can be located at any point on the field; however, for ease of memorization and use, it is often set at a landmark point on the field. Taking a football field as an example, it could be a corner kick point, a penalty kick point, or a corner point of a field line.

[0198] In one possible implementation, to reduce processing load, synchronous sampling is performed on the video streams acquired by the multiple cameras. Alternatively, this can be understood as extracting frames from the video streams acquired by each camera. Taking the first camera as an example, the target marker is included in all the multiple image frames sampled from the first camera. Then, image frames that do not include the target marker can be removed from the extracted image frames, thus forming an image set corresponding to each acquisition device. It is understood that the target marker's movement position differs for different image frames included in the image set corresponding to each acquisition device.

[0199] It is understandable that, because the target object moves within the playing field, for a given camera, during certain time periods, the target object moves outside the camera's field of view, and therefore, the camera does not capture images including the target object during those time periods. Thus, image frames that do not include the target object can be removed from the extracted image frames.

[0200] In one example, when extracting frames for different cameras, frames can be extracted based on the movement position of the target marker. For instance, multiple positions can be set along the target marker's movement trajectory, such as position 1 to position m, allowing image frames to be sampled at multiple positions. In another example, a frame can be extracted at set intervals. The set interval can be based on the target marker's movement rate within a defined area.

[0201] In some possible scenarios, multiple locations can be marked within a defined area. Each time the target object moves to a new location, multiple cameras acquire an image frame, thus forming an image frame set for each camera. Therefore, the number of images acquired by each camera is the same as the number of marked locations.

[0202] 902, Perform calibration point detection on each of the image frames acquired by the plurality of cameras to obtain the pixel coordinates (or pixel coordinates of marker points) of multiple calibration points on the target calibration object in the image frames acquired by each camera.

[0203] Each of the multiple cameras performs calibration point detection on image frames acquired by each camera to obtain feature points. Different feature points in the same image frame correspond to different calibration points. The detected feature points are further filtered to remove those with lower reliability. These feature points are used to represent the pixel coordinates of the calibration points, thereby obtaining the pixel coordinates of multiple calibration points in each image frame acquired by each camera.

[0204] For example, the distance from a feature point to the image boundary can be used as a filtering condition to remove feature points with low reliability. For instance, if the distance from a feature point to the image boundary is less than a certain set value, that feature point is filtered out. Another example is a calibration surface with a fixed shape, such as a rectangle or a circle. For example, if the calibration surface is rectangular, the angle between the diagonals of the feature surface in the acquired image can be used to remove feature points with low reliability. For example, if the angle between the diagonals is less than a set threshold, feature points belonging to that feature surface are removed. Similarly, if the calibration surface is circular, the curvature of the circular feature surface in the acquired image can be used to determine whether feature points on that surface are considered low-reliability. For example, if the curvature is greater than a set curvature value, the feature points on that surface are determined to be unreliable and removed. For circular calibration surfaces, the ratio of the minimum radius to the maximum radius can also be used to determine whether feature points on the surface are reliable. For example, if the ratio of the minimum radius to the maximum radius is less than a set ratio, the feature points on that surface are determined to be unreliable and removed.

[0205] In some embodiments, when performing calibration point detection on image frames sampled by a camera, global calibration point detection can be performed on the first image frame. In subsequent calibration point detection, a tracking algorithm can be used to narrow down the detection range of the target calibration object in the image frame. This approach can improve the detection speed of feature points.

[0206] 903. Based on the pixel coordinates of the multiple calibration points included in the target calibration object in the image frames acquired by each camera, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system of the target calibration object, the intrinsic parameter matrix of each camera is estimated to obtain the first intrinsic parameter estimate of each camera.

[0207] 904. Based on the first intrinsic parameter estimates of at least two cameras in each of the multiple camera groups, the matching feature point set corresponding to each camera group, and the three-dimensional coordinates of the target calibration object including multiple calibration points in the calibration object coordinate system, determine the first extrinsic parameter estimate of each camera.

[0208] In this system, each camera group includes at least two cameras that share a common field of view. The matching feature point set includes multiple matching feature point groups, and each matching feature point group includes at least two matching pixel coordinates. The at least two matching pixel coordinates are the pixel coordinates of the same calibration point detected by image frames acquired at the same time by different cameras belonging to the same camera group. The multiple camera groups are obtained by grouping the multiple cameras together, and any two camera groups in the multiple camera groups include at least one identical camera.

[0209] In one possible implementation, when estimating the intrinsic parameter matrix of each camera in step 903 to obtain the first intrinsic parameter estimate, the intrinsic parameter matrix of each camera can be estimated by combining the pixel coordinates of the multiple calibration points included in the target calibration object in the image frame acquired by each camera, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, using direct linear transform (DLT).

[0210] Based on the pinhole camera imaging model, the three-dimensional point P in space w The projection onto the camera pixel plane can be represented by the following formula (6). w P represents the homogeneous coordinates of the calibration point in the target calibration object within the calibration object's coordinate system. w =[XYZ 1] T The pixel coordinates (using homogeneous coordinates) of the calibration points in the image coordinate system are obtained through P. uv It means that P uv =[uv 1] T .

[0211] P uv =P proj P w Formula (6)

[0212] P proj This represents the projection matrix. The projection matrix is ​​determined using the camera's intrinsic parameters and its pose (in the calibration object coordinate system). c and [R|t] c These represent the camera's intrinsic parameters and pose (in the calibration object coordinate system), respectively. proj =I c [R|t] c The mapping relationship between the target calibration object's coordinate system and the image coordinate system can be expressed as formula (7).

[0213] P uv =I c [R|t] c P w =P proj P w Formula (7)

[0214] DLT utilizes a large number of known quantities (P) uv P w Substitute into formula (6) above to calculate the projection matrix P. proj By decomposing matrix P proj To get I c [R|t] c .

[0215] For example, P can be calculated by solving a system of linear equations. proj Furthermore, QR decomposition of the matrix is ​​used to obtain I. c Projection matrix P proj It can be represented as a 3x4 matrix, for example, a 3x4 matrix can be represented as Then the following formula (8) holds true.

[0216]

[0217] f x = f / dx, where f represents the focal length. y = f / dy. dx and dy represent the length units occupied by one pixel in the x and y directions, respectively, that is, the actual physical value represented by one pixel, which is the key to realizing the transformation between the camera coordinate system and the image coordinate system. u0 and v0 represent the difference in horizontal and vertical pixels between the center pixel coordinates and the origin pixel coordinates of the image.

[0218] Based on this, P is calculated by solving a system of linear equations. proj Middle l1~l 12 Then on The camera's intrinsic parameter matrix is ​​obtained by performing QR decomposition.

[0219] It should be noted that regardless of the target calibration object's location on the sports field, the coordinates of each calibration point on the target calibration object remain unchanged in the calibration object's coordinate system. Since the camera's intrinsic parameters are used to express the transformation relationship between the camera coordinate system and the image coordinate system, these intrinsic parameters remain constant regardless of the calibration object's movement. Based on this, one or more projection matrices can be calculated for feature points in each image frame acquired by a single camera. Then, QR decomposition is performed on the projection matrices to obtain intrinsic and extrinsic parameter matrices. Since the intrinsic parameter matrices obtained from the decomposition of each calculated projection matrix should be identical, the first intrinsic parameter estimate is determined from all the obtained intrinsic parameter matrices. For example, this can be determined using a Gaussian distribution.

[0220] In another possible implementation, when estimating the intrinsic parameter matrix of each camera in step 903 to obtain the first intrinsic parameter estimate, the second intrinsic parameter matrix can be obtained by combining the DLT algorithm with the pixel coordinates of the multiple calibration points included in the target calibration object in the image frame acquired by each acquisition device and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system. Then, based on the principle of minimum projection error, the intrinsic parameter matrix of each camera is globally optimized.

[0221] The second intrinsic parameter estimate is obtained by combining the DLT algorithm to estimate the intrinsic parameter matrix of each camera. For the relevant descriptions of formulas (6) to (8) above, please refer to them. They will not be repeated here.

[0222] Further optimization is performed based on the second intrinsic parameter estimate for each camera. In some embodiments, redundant data can be filtered out from the feature data of the calibration points before optimization. This can also be understood as removing redundant pixel coordinates from the pixel coordinates of the multiple calibration points identified in the image frames acquired by each camera. For example, for each camera, the distribution of calibration points in the image plane of each image frame in the acquired image set is statistically analyzed, and for any camera, overlapping feature points in the image plane (corresponding to the positions of the calibration points) are deleted. Exemplarily, the image plane can also be meshed. Redundancy removal is performed based on the distribution of feature points in the mesh. See [link to documentation]. Figure 13 As shown. For example, in a certain grid of the image plane, due to the movement of the target marker, the target marker may appear multiple times in that grid. The target marker that appears multiple times in the grid can be retained only once. When the target marker appears again, the pixel coordinates of the target marker's calibration point in the acquired image frame are removed.

[0223] For example, the number of cameras deployed in the designated space of the sports field is N. After obtaining the second intrinsic parameter estimate of camera i, the intrinsic parameter matrix of camera i is optimized, that is, the first intrinsic parameter estimate is obtained by optimizing based on the second intrinsic parameter estimate. For example, based on the second intrinsic parameter estimate of camera i, the pixel coordinates of the calibration points on the target calibration object in the image frames included in the image set acquired by camera i, and the three-dimensional coordinates of multiple calibration points in the calibration object coordinate system, the pose set corresponding to camera i is estimated. The pose set corresponding to camera i includes the pose of the target calibration object relative to camera i at M1 moving positions; i takes the value of a positive integer less than or equal to N.

[0224] It should be noted that the N cameras capture M positions of the target object, while camera i captures M or fewer positions of the target object, referred to here as M1. This is because the field of view of camera i may not cover these M positions.

[0225] Taking camera 1 as an example. After obtaining the second intrinsic parameter estimate of camera 1, the intrinsic parameter matrix of camera 1 is optimized, that is, the first intrinsic parameter estimate is obtained by optimizing based on the second intrinsic parameter estimate. For camera 1, based on the second intrinsic parameter estimate of camera 1, the coordinates of each calibration point in the image frame acquired by camera 1 in the image coordinate system (the pixel coordinates of the calibration points in the image plane of camera 1 after redundancy removal), and the coordinates of the calibration points of the target calibration object in the calibration object coordinate system, the pose of different positions relative to camera 1 can be determined. The pose of the calibration object coordinate system relative to each camera coordinate system at different positions can be determined by the above method. See Table 1, which shows the pose of the target calibration object at different positions relative to different cameras. In Table 1, N cameras are deployed on the sports field as an example. The movement trajectory of the target calibration object passing through position 1 to position M is taken as an example. For example, the target calibration object moves to M positions captured by N cameras.

[0226] The pose of the target object at position 1 relative to camera 1 is [R|t]. 1-1 For example, the pixel coordinates of each calibration point in the image frame captured by camera 1 when the target calibration object moves to position 1, the 3D coordinates of each calibration point of the target calibration object in the calibration object coordinate system, and the first intrinsic parameter estimate of camera 1 can be used to estimate the pose [R|t] of the target calibration object at position 1 relative to camera 1 using the PnP algorithm in combination with the above formula (7). 1-1 It should be noted that the field of view of each camera is not able to cover all positions. Therefore, it is not possible to obtain the pose of the uncovered positions relative to the camera. This is indicated by "none" in Table 1.

[0227] Table 1

[0228] Camera 1 Camera 2 …… Camera N Position 1 <![CDATA[[R|t] 1-1 ]]> <![CDATA[[R|t] 1-2 ]]> <![CDATA[[R|t] 1-i ]]> none Position 2 <![CDATA[[R|t] 2-1 ]]> none <![CDATA[[R|t] 2-i ]]> <![CDATA[[R|t] 2-n ]]> …… <![CDATA[[R|t] j-1 ]]> <![CDATA[[R|t] j-2 ]]> <![CDATA[[R|t] j-i ]]> <![CDATA[[R|t] j-n ]]> Position M none <![CDATA[[R|t] m-2 ]]> none <![CDATA[[R|t] m-n ]]>

[0229] After obtaining the pose of the target calibration object relative to each camera at different positions, the camera parameters can be globally optimized based on a preset nonlinear optimization algorithm to minimize the projection error on the image frames of each camera. For example, the preset nonlinear optimization algorithm is the Leven-Marquardt (LM) algorithm.

[0230] In one example, considering camera distortion, the coordinates of the calibration point on the target calibration object at position j projected onto the camera coordinate system of camera i can be estimated by combining the three-dimensional coordinates of the calibration point on the target calibration object at position j and the pose of position j relative to camera i using formula (9).

[0231]

[0232] in, This represents the coordinates of the calibration point on the target calibration object at position j projected onto camera i in the camera coordinate system. This represents the pose of position j relative to camera i. This represents the three-dimensional coordinates of the calibration point on the target calibration object at position j.

[0233] Furthermore, obtain the coordinates of the calibration point on the target calibration object at position j projected onto camera i in the camera coordinate system. Then, the coordinates of the calibration point on the target calibration object at position j are projected onto the camera coordinate system of camera i. Dividing by z yields the normalized coordinates, which are represented in this embodiment as follows:

[0234] Furthermore, based on the normalized coordinates of the calibration point on the target calibration object at position j projected onto the camera coordinate system of camera i using formulas (1-1) and (1-2), the normalized coordinates of the distorted calibration point on the target calibration object at position j projected onto the camera coordinate system of camera i are estimated. Then, based on the second intrinsic parameter estimate of camera i To calculate the pixel coordinates of the calibration point on the target calibration object at position j projected onto the image coordinate system of camera i. See formula (10).

[0235]

[0236] Then, the error between the estimated pixel coordinates and the pixel coordinates of the calibration point identified in the actual image frame acquired by camera i is calculated. Then, based on the nonlinear optimization algorithm, the intrinsic parameter matrix of camera i, the pose of different positions relative to camera i, and the distortion coefficient of camera i are further optimized according to all the calculated errors, so as to obtain the optimized intrinsic parameter matrix and distortion coefficient of camera i. Then, based on the optimized intrinsic parameter matrix and distortion coefficient of camera i, the next round of optimization is carried out. The coordinates of the calibration point on the target calibration object at position j can be estimated in the camera coordinate system of camera i by combining the three-dimensional coordinates of the calibration point on the target calibration object at position j and the pose of position j relative to camera i after the previous round of optimization with formula (9). Furthermore, based on the normalized coordinates of the calibration point on the target calibration object at position j projected into the camera coordinate system of camera i by formula (1-1) and formula (1-2), the normalized coordinates of the calibration point on the target calibration object at position j projected into the camera coordinate system of camera i after distortion are estimated. Then, based on formula (10) and the optimized intrinsic parameter matrix of camera i, the pixel coordinates of the calibration point on the target calibration object at position j projected onto the image coordinate system of camera i are determined. The error between the estimated pixel coordinates and the pixel coordinates of the calibration point identified in the actual image frame acquired by camera i is then calculated. Finally, based on a nonlinear optimization algorithm and all calculated errors, the intrinsic parameter matrix of camera i, the pose relative to camera i at different positions, and the distortion coefficients of camera i are further optimized to obtain the optimized intrinsic parameter matrix and distortion coefficients of camera i.

[0237] In another example, without considering camera distortion, the pixel coordinates of the calibration point on the target calibration object at position j projected onto the image coordinate system of camera i can be estimated by combining the three-dimensional coordinates of the calibration point on the target calibration object at position j, the pose of position j relative to camera i, and the estimated value of the second intrinsic parameter of camera i with formula (11).

[0238]

[0239] in, This represents the pixel coordinates of the calibration point on the target calibration object at position j projected onto the image coordinate system of camera i. Then, the error between the estimated pixel coordinates and the pixel coordinates of the calibration point identified in the actual image frame acquired by camera i is calculated. Next, based on all calculated errors, a nonlinear optimization algorithm is used to further optimize the intrinsic parameter matrix of camera i and the pose relative to camera i at different positions, thus obtaining the optimized intrinsic parameter matrix of camera i and the pose relative to camera i at different positions.

[0240] In one possible implementation, the extrinsic parameter matrix for each camera can be determined in at least two of the following ways:

[0241] In the first possible approach, the second relative poses of all cameras (excluding the reference camera) relative to the reference camera (any one of the multiple cameras) can be obtained first. Then, a scale factor is added to the relative poses to determine the extrinsic parameter matrix of each camera.

[0242] In the second possible approach, the pose of each movement position in the world coordinate system is determined by combining the base movement position and the co-view relationship between cameras. Based on the pose of each movement position in the world coordinate system, the coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, and the pixel coordinates of the calibration points on the target calibration object in the image frames acquired by each acquisition device, the extrinsic parameter matrix of each camera is determined. The determined extrinsic parameter matrix is ​​used as the initial value for further global optimization. During the optimization of the extrinsic parameter matrix, the intrinsic parameter matrix and distortion coefficients can also be further optimized.

[0243] The first possible implementation is described in detail below. Multiple cameras can be grouped based on their shared field of view (CDR) relationships. For example, they can be divided into multiple camera groups. Each camera group includes at least two cameras. Any two camera groups include at least one identical camera. At least two cameras share a common field of view.

[0244] See Figure 14 As shown.

[0245] 1401. Based on the first intrinsic parameter estimates of at least two cameras included in each camera group, the matching feature point set corresponding to each camera group, and the three-dimensional coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, obtain the second relative pose of the other cameras among the multiple cameras, excluding the reference camera, relative to the reference camera; the reference camera is any one of the multiple cameras.

[0246] The matching feature point set includes multiple matching feature point groups, and each matching feature point group includes at least two matching pixel coordinates. The at least two matching pixel coordinates are the pixel coordinates of the same calibration point detected by image frames acquired at the same time by different cameras belonging to the same camera group.

[0247] For example, each camera pair can be identified, comprising two cameras that share a common viewing relationship (or a common viewing area). Further, the number of times each camera is included in the camera pair can be calculated. The camera with the highest number of inclusions in the camera pair is designated as the reference camera.

[0248] 1402. After determining the second relative pose of each camera relative to the reference camera, a scale factor is further determined. The scale factor is the ratio between a first distance and a second distance. The first distance is the distance between two calibration points on the target calibration object, and the second distance is the distance between the two calibration points in the same image coordinate system. The two calibration points are located on the same calibration surface on the target calibration object.

[0249] The first distance is measured, for example, by measuring the side length of a designated edge on the calibration surface of the target calibration object. The endpoints of this edge can also be understood as two calibration points. The calibration surface in the image frame captured by the reference camera (camera 1) is identified, and the side length of the designated edge of the calibration surface is calculated. Taking a rectangular calibration surface as an example, the designated edge can be one of the edges. The four corner points of the calibration surface in the image frame captured by camera 1 can be identified, and the coordinates of the two corner points in the camera coordinate system of camera 1 are determined based on the pixel coordinates of the two corner points corresponding to the designated edge and the intrinsic parameters of camera 1. The side length is further calculated based on the coordinates of the two corner points in the camera coordinate system of camera 1 to obtain an estimated side length value. Taking a circular calibration surface as an example, the designated edge can be the diameter of the circle. The ratio of the measured side length value to the estimated side length value is the adjustment ratio between the camera coordinate system and the world coordinate system. In this embodiment, this adjustment ratio is called the scale. The scale is represented by S.

[0250] It should be noted that, when determining the second distance, cameras other than the reference camera can also be used, and this application embodiment does not specifically limit this. Similarly, when acquiring the first and second distances, any two calibration points located on the same calibration plane can also be used, and this application embodiment does not specifically limit this.

[0251] 1403, and then obtain the first extrinsic parameter estimate for each camera based on the second relative pose of each camera and the scale factor. It should be noted that the two calibration points can be any two calibration points on the same calibration plane.

[0252] In some embodiments, the second relative pose of other cameras that co-view with the reference camera and the reference camera can be determined. The second relative position of a camera that does not co-view with the reference camera can be calculated using the second relative poses of the cameras that co-view with the reference camera and the relative poses between the cameras that co-view with the reference camera. For example, camera 1 is the reference camera. Camera 2 co-views with camera 1, so the second relative pose of camera 2 relative to camera 1 can be calculated. Camera 3 does not co-view with camera 1 but co-views with camera 2. The relative pose of camera 3 relative to camera 2 can be calculated, and then the second relative pose of camera 3 relative to camera 1 can be determined by combining the relative pose of camera 3 relative to camera 2 and the relative pose of camera 2 relative to camera 1.

[0253] The two cameras share a common field of view. Based on this common field of view, their relative poses can be calculated. Taking camera 1 and camera 2 as an example, they share a common field of view. Camera 1 and camera 2 can be understood as a camera group (or camera pair). The essential matrix between camera 1 and camera 2 is determined based on the matching feature point set corresponding to this camera group. Then, singular value QR decomposition is performed on the essential matrix to obtain the second relative pose between camera 1 and camera 2. The essential matrix can also be called the intrinsic matrix, but this embodiment does not specifically limit its meaning.

[0254] Specifically, the coordinates of the same calibration point within the common viewing area under the camera coordinates of two different cameras satisfy the conditions shown in the following formula (6). Taking camera 1 and camera 2 as examples.

[0255] P x1y1 EP x2y2 =0 Formula (12)

[0256] Among them, P x1y1 P represents the normalized coordinates of the calibration point in the camera coordinate system of camera 1. x2y2 This represents the normalized coordinates of the calibration point in the camera coordinate system of camera 2. E represents the eigenvalue matrix.

[0257] The essential matrix E describes the pose relationship between cameras. Definition: Matrix E contains the rotation and translation information related to the two cameras in physical space.

[0258] The normalized coordinates of the calibration point in the camera coordinate system of camera 1 can be determined based on the pixel coordinates of the calibration point in the image coordinate system of camera 1 (i.e., the coordinates of the calibration points included in the matching feature point set in the image coordinate system of camera 1) and the estimated value of the second intrinsic parameter of camera 1. Similarly, the normalized coordinates of the calibration point in the camera coordinate system of camera 2 can be determined based on the pixel coordinates of the calibration point in the image coordinate system of camera 2 (i.e., the pixel coordinates of the calibration points included in the matching feature point set in the image coordinate system of camera 1) and the estimated value of the second intrinsic parameter of camera 2. For the specific calculation method, please refer to the description of the transformation relationship between the image coordinate system and the camera coordinate system in formula (2).

[0259] After determining the eigenma using the above formula (12), E can be decomposed to obtain the rotation matrix R and the translation parameter t. It should be understood that if one of the cameras is taken as the reference camera, the rotation matrix of the reference camera is set to the identity matrix, and the translation parameter is set to 0, then the R and t obtained by decomposition are the rotation matrix R and translation parameter t of the other camera, which is the second relative pose.

[0260] In some embodiments, after determining the second relative pose of each camera relative to a reference camera, a scaling factor can be added to the second relative pose of each camera to obtain the extrinsic parameter matrix of each camera. For example, the determined second relative pose of camera 1 is [R|t]1. The extrinsic parameter matrix of camera 1 after adding the scaling factor can be expressed as follows: The extrinsic parameters of other cameras can be adjusted in this way.

[0261] In other embodiments, after determining the second relative pose of each camera relative to a reference camera, the estimated intrinsic parameter matrices of each camera, the relative poses (and distortion coefficients) of each camera can be globally optimized by combining the second relative poses of each camera. Then, a scale factor is added to the optimized relative poses to obtain the extrinsic parameter matrices of each camera. The optimized distortion coefficients and intrinsic parameter matrices of each camera are used as the final calibrated distortion coefficients and intrinsic parameter matrices of the cameras.

[0262] See Figure 15 As shown, taking the principle of minimizing projection error to optimize the intrinsic parameter matrix and relative pose as an example, the camera distortion coefficient is not considered.

[0263] 1501, after determining the second relative pose of the cameras in the g-th camera group relative to the reference camera, the three-dimensional coordinates of the multiple calibration points in the local coordinate system are determined according to the second relative pose of each camera when the target calibration object moves to M2 moving positions.

[0264] The local coordinate system is the camera coordinate system of the reference camera. Any one of the M2 movement positions is located at least within the shared field of view of two cameras in the g-th camera.

[0265] 1502, based on the three-dimensional coordinates of the plurality of calibration points at the M2 moving positions in the local coordinate system, the second relative pose of the cameras included in the g-th camera group, and the first intrinsic parameter estimation value, estimate the pixel coordinates of the plurality of calibration points at the M2 moving positions projected onto the image coordinate system of the cameras included in the g-th camera group.

[0266] 1503, obtain the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of the cameras included in the g-th camera group, and the error between the estimated pixel coordinates of the plurality of calibration points extracted from the image frames acquired by the cameras included in the g-th camera group.

[0267] 1504. Based on the error, adjust the second relative pose and the first intrinsic parameter estimate of the cameras included in the g-th camera group to obtain the relative pose and intrinsic parameter estimate of the cameras included in the g-th camera group after the current round of adjustment.

[0268] In this process, the intrinsic parameter estimates and second relative poses of the cameras in the g-th camera group after the current round of adjustment are used as the basis for the next round of adjustment, until the third relative pose and third intrinsic parameter estimates of the cameras included in the g-th camera group are obtained after completing the D round of adjustment.

[0269] See Figure 16 As shown, the principle of minimizing projection error is used to optimize the intrinsic parameter matrix, relative pose, and distortion coefficients as an example.

[0270] 1601, after determining the second relative pose of the acquisition devices in the g-th acquisition device group relative to the reference acquisition device, the three-dimensional coordinates of the multiple calibration points in the local coordinate system are determined according to the second relative pose of each acquisition device when the target calibration object moves to M2 moving positions.

[0271] The local coordinate system is the camera coordinate system of the reference acquisition device; any of the M2 moving positions is located at least within the common viewing area of ​​two acquisition devices in the g-th acquisition device.

[0272] 1602, based on the three-dimensional coordinates of the plurality of calibration points at the M2 moving positions in the local coordinate system, the second relative pose of the acquisition devices included in the g-th acquisition device group, the first intrinsic parameter estimate, and the first distortion coefficient estimate, estimate the pixel coordinates of the plurality of calibration points at the M2 moving positions projected onto the image coordinate system of the acquisition devices included in the g-th acquisition device group.

[0273] 1603, obtain the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of the acquisition devices included in the g-th acquisition device group, and the error between the pixel coordinates of the plurality of calibration points extracted from the image frames acquired by the acquisition devices included in the g-th acquisition device group.

[0274] 1604. Based on the error, adjust the second relative pose and the first intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group to obtain the relative pose, intrinsic parameter estimate, and distortion coefficient of the acquisition devices included in the g-th acquisition device group after the current round of adjustment.

[0275] In this round, the intrinsic parameter estimates and relative poses of the acquisition devices in the g-th acquisition device group after the current round of adjustment are used as the basis for the next round of adjustment, until the D-th round of adjustment is completed to obtain the third relative pose, third intrinsic parameter estimates, and second distortion coefficients of the acquisition devices included in the g-th acquisition device group.

[0276] For example, the number of cameras in each of the aforementioned camera groups varies. When optimizing each camera group, the relative poses and intrinsic parameter matrices of the cameras within the group can be optimized in the order of the number of cameras in the group. It can be understood that the first optimized camera group includes two cameras. The second camera group includes three cameras, including two cameras from the first camera group. It can be understood that the second camera group adds one camera to the first camera group. This added camera shares a common field of view with at least one camera in the first camera group.

[0277] In one possible example, after the second relative poses of the two cameras in the first camera group are determined, the second relative poses and intrinsic parameter matrices of the cameras in the first camera group are optimized. Then, based on the optimized relative poses and intrinsic parameter matrices of the two cameras in the first camera group, the relative poses of the newly added cameras in the second camera group are calculated. Then, the relative poses and intrinsic parameter matrices of the cameras in the second camera group are further optimized. This process continues.

[0278] For example, two cameras are selected from multiple cameras to form camera group 1. One of the selected cameras is chosen as the reference camera. The optical axis angle between the two cameras is within a set range, for example, less than 5 degrees. The optical axis angle can be determined by the estimated second intrinsic parameters of each camera and the calibration points in the captured images. The two cameras share a common field of view. Based on the common field of view relationship between the two cameras, the relative pose of the two cameras is calculated. For example, the camera coordinate system of the reference camera is used as the local coordinate system. The method for determining the relative pose of the two cameras is as described above and will not be repeated here.

[0279] Taking camera group 1, specifically camera 1 and camera 2, as an example, and using camera 1 as the reference camera, we determine the image frames captured by both cameras at the same time. We use 10 time points t as an example. Camera 1 at t1…t2… 10 Image frames fr11...fr1 captured separately 10 Camera 2 at t1...t 10 Image frames fr21...fr2 captured separately 10 Image frames fr11...fr1 10 With image frames fr21……fr2 10 Each image frame is captured at the same time and corresponds to a single image frame. Two image frames captured at the same time constitute an image frame pair.

[0280] P u1v1 P represents the pixel coordinates of the calibration point in the image frame corresponding to camera 1 in the image frame pair. u2v2 P represents the pixel coordinates of the same calibration point in the image frame corresponding to camera 2 in the image frame pair. u1v1=[u1 v1 1] T ;P u2v2 =[u2 v2 1] T The second intrinsic parameter estimate of camera 1 is represented by I1, and the second intrinsic parameter estimate of camera 2 is represented by I2. The first extrinsic parameter estimate of camera 1 is represented by [R|t]1, and the first extrinsic parameter estimate of camera 2 can be represented by [R|t]2. The first extrinsic parameter estimate of camera 1 is the identity matrix. The first extrinsic parameter estimate of camera 2 can be obtained through the above decomposition.

[0281] The pixel coordinates of each calibration point in the local coordinate system of the common viewing area of ​​camera 1 and camera 2 are estimated based on the following formula (13).

[0282]

[0283] in, This indicates that the coordinates of each calibration point within the common viewing area are estimated in the local coordinate system. If camera 1 is the reference camera, then in [R|t]1, R is the identity matrix, and the translation parameter t is a vector of all zeros.

[0284] Furthermore, the coordinates of each calibration point within the estimated common viewing area in the local coordinate system can be used as a basis. The first pixel coordinates of each calibration point in the common viewing area in the image coordinate system of camera 2 are estimated by combining the following formula (14).

[0285]

[0286] This represents the estimated first pixel coordinates of each calibration point within the common viewing area in the image coordinate system of camera 2. Further, according to... With P u2v2 The intrinsic parameters of camera 1, camera 2, and the relative pose of camera 2 are adjusted based on the errors between them. Then, the next round of iterative adjustments is performed based on these adjusted intrinsic and extrinsic parameters. Specifically, the coordinates of each calibration point in the local coordinate system within the common field of view are recalculated based on the intrinsic parameters of camera 1 after the first round of adjustment. Then, the estimated pixel coordinates of each calibration point in the image coordinate system of camera 2 within the common field of view are estimated based on the adjusted intrinsic and extrinsic parameters of camera 2. The error between the estimated pixel coordinates and the actual pixel coordinates is further calculated to adjust the intrinsic parameters of camera 1, camera 2, and camera 2, as well as the relative pose of camera 2. This process is repeated for multiple rounds of adjustments.

[0287] In some embodiments, camera distortion is taken into account. The coordinates of each calibration point in the common field of view in the camera coordinate system of camera 1 can be estimated based on the coordinates of each calibration point in the local coordinate system within the common field of view, combined with the following formula (15).

[0288] P u1v1 =I1P x1y1 Formula (15)

[0289] Obtain the estimated coordinates P of each calibration point in the camera coordinate system of camera 1. x1y1 Then, using formulas (1-1) and (1-2), in P x1y1 Based on this, estimate the coordinates before distortion. Then, based on the second relative pose of camera 1, formula (16) is used to calculate the coordinates of each calibration point in the local coordinate system.

[0290]

[0291] Furthermore, the coordinates of each calibration point within the estimated common viewing area in the local coordinate system can be used as a basis. The coordinate estimates P of each calibration point in the common viewing area in the camera coordinate system of camera 2 are estimated by combining the following formula (17). x2y2

[0292]

[0293] Obtain the estimated coordinates P of each calibration point in the camera coordinate system of camera 2. x2y2 Then, normalization processing can be performed to obtain the normalized coordinate estimates of each calibration point in the camera coordinate system of camera 2.

[0294] Furthermore, using formulas (1-1) and (1-2), the normalized coordinate estimates are obtained. Based on this, estimate the normalized coordinates after distortion. Then, based on the second intrinsic parameter estimate I2 of camera 2, formula (18) is used to calculate the pixel coordinates of each calibration point projected onto the image coordinate system of camera 2.

[0295]

[0296] Furthermore, according to With P u2v2The intrinsic parameters of camera 1, the distortion coefficients of camera 1, the intrinsic parameters of camera 2, the relative pose of camera 2, and the distortion coefficients of camera 2 are adjusted based on the errors between them. Then, the next round of iterative adjustment is performed based on the adjusted intrinsic parameters of camera 1, the distortion coefficients of camera 1, the intrinsic parameters of camera 2, the relative pose of camera 2, and the distortion coefficients of camera 2. Specifically, based on the intrinsic parameters and distortion coefficients of camera 1 adjusted in the first round, the coordinates of each calibration point in the common field of view in the local coordinate system are recalculated. Then, based on the adjusted intrinsic parameters, relative pose, and distortion coefficients of camera 2, the estimated pixel coordinates of each calibration point in the common field of view in the image coordinate system of camera 2 are estimated. The error between the estimated pixel coordinates and the actual pixel coordinates is further calculated to adjust the intrinsic parameters of camera 1, the distortion coefficients of camera 1, the intrinsic parameters of camera 2, the relative pose of camera 2, and the distortion coefficients of camera 2. This process is repeated for multiple rounds of adjustment.

[0297] For example, the number of iterations can be pre-configured, and the iteration stops when the configured number of iterations is reached. An error threshold can also be pre-configured, and the iteration stops when the calculated error during a certain iteration is less than or equal to the error threshold.

[0298] Furthermore, a third camera is added to the existing two cameras, resulting in a shared field of view for all three cameras. These three cameras form camera group 2. Joint optimization is achieved by optimizing the intrinsic and extrinsic parameter matrices (and distortion coefficients) of these three cameras. The image frames obtained by each of the three cameras (camera 1, camera 2, and camera 3) simultaneously capturing their respective fields of view (including the shared field of view) are determined. For example, consider 10 time points t. Camera 1 captures images at time points t1…t3…t4…t5…t6…t7 ...…t7……t7……t7……t7……t7………t7………………………………………………………………………………………………………………………………………………………………………………………………………………………………………………………………………………………… 10 Image frames fr11...fr1 captured separately 10 Camera 2 at t1...t 10 Image frames fr21...fr2 captured separately 10 Camera 3 at t1...t 10 Image frames fr31...fr3 were captured separately. 10 Three image frames captured by three cameras at the same time constitute an image frame group. It can be understood that the coordinates of calibration points within the shared field of view captured at the same time should be the same in the local coordinate system of the reference camera; that is, the coordinates of calibration points within the shared field of view of the three cameras should be identical in their local coordinate systems. Based on this, the intrinsic parameters, relative pose, and distortion coefficients of the three cameras are optimized using the pixel coordinates of each calibration point in each image frame group captured by the three cameras at the same time.

[0299] In some embodiments, some calibration points are not located within the common field of view of the three cameras, but are located within the common field of view of two cameras. These feature points can participate in the adjustment of the three cameras.

[0300] For example, cameras 1, 2, and 3 share a common viewing region 1; cameras 1 and 2 share a common viewing region 2 (excluding region 1); cameras 2 and 3 share a common viewing region 3 (excluding region 1); and cameras 2 and 3 have no other common viewing regions besides region 1. The coordinates of the calibration points within regions 1, 2, and 3 in their local coordinate systems can be calculated. These calibration points are then projected onto the image coordinate system of the camera capable of capturing those points, and the error is calculated. Based on the calculated error, the intrinsic parameters, relative pose, and distortion coefficients of the three cameras are optimized.

[0301] Next, select another camera that shares a common field of view with at least two cameras in camera group 2. Cameras sharing a common field of view form camera group 3. Then, optimize the intrinsic parameter matrix of each camera within camera group 3, and so on. This yields the third intrinsic parameter estimate, third relative pose, and second distortion coefficient for each camera.

[0302] In some embodiments, a scaling factor is added to the optimized third relative pose to obtain the extrinsic parameter matrix of each camera. The optimized second distortion coefficient and third intrinsic parameter estimates of each camera are used as the final calibrated intrinsic parameter matrix and distortion coefficient of each camera.

[0303] In other embodiments, after increasing the scale factor to obtain the extrinsic parameter matrix of each camera, and obtaining the second distortion coefficient and the third intrinsic parameter estimate, further global optimization can be performed on the extrinsic parameter matrix, the second distortion coefficient and the third intrinsic parameter estimate of each camera.

[0304] For example, the coordinates of each calibration point in the local coordinate system can be scaled to the world coordinate system. For instance, the normalized coordinates of a calibration point in the local coordinate system are represented as P. w P w =[XYZ 1] T The normalized coordinates of the calibration point in the world coordinate system are represented as P′. w Then P′ w =[XYZ s] T .

[0305] The pixel coordinates of each calibration point in the image coordinate system are estimated based on the normalized coordinates of each calibration point in the world coordinate system, the intrinsic parameters of camera i, and the extrinsic parameters of camera i in the world coordinate system. For example, it can be calculated based on formula (19).

[0306] P′ uivi =I i [R′|t] iP′ wi Formula (19)

[0307] Among them, I i Represents the intrinsic parameters of camera i, [R′|t] i P′ represents the extrinsic parameters of camera i. wi P′ represents the coordinates of the calibration point within the field of view of camera i in the world coordinate system. uivi This represents the estimated pixel coordinates of the calibration point in the image frame acquired by camera i. Let P... uivi P represents the pixel coordinates of the calibration point in the image frame acquired by camera i. uivi This represents the pixel coordinates obtained by identifying calibration points from image frames acquired by camera i. Further, determine P′. uivi With P uivi The error between them. Based on the above formula (19), the pixel coordinates of the calibration point in the image frame acquired by each camera are estimated, and the error between the estimated pixel coordinates of the calibration point and the pixel coordinates obtained by identifying the calibration point is determined. The intrinsic and extrinsic parameters of each camera are further adjusted according to the error. The intrinsic and extrinsic parameter adjustment of the current round of cameras is completed. Then, the next round of iterative adjustment is carried out based on the adjusted intrinsic and extrinsic parameters of the cameras. Specifically, the pixel coordinates of the calibration point in the image frame acquired by each camera are re-estimated according to the intrinsic parameters of each camera after the first round of adjustment, and the error between the estimated pixel coordinates of the calibration point and the pixel coordinates obtained by identifying the calibration point is determined. The intrinsic and extrinsic parameters of each camera are further adjusted according to the error, and the intrinsic and extrinsic parameter adjustment of the current round of cameras is completed. This process is repeated to perform multiple rounds of adjustment.

[0308] In some embodiments, considering camera distortion, the error between the pixel coordinates of the distorted calibration points and the pixel coordinates obtained from identifying the calibration points is determined using formulas (1-1) and (1-2). The intrinsic parameter matrix, extrinsic parameter matrix, and distortion coefficients of each camera are then adjusted based on this error. Specifically, the second distortion coefficient of each camera determined above is used as the initial distortion coefficient for adjustment. After completing the adjustment of the camera's intrinsic parameter matrix, extrinsic parameter matrix, and distortion coefficients for the current round, the next round of iterative adjustment is performed based on the adjusted intrinsic parameter matrix, extrinsic parameter matrix, and distortion coefficients. Specifically, based on the intrinsic parameter matrix, extrinsic parameter matrix, and distortion coefficients of each camera after the first round of adjustment, the pixel coordinates of the calibration points in the image frames acquired by each distorted camera are re-estimated, and the error between the estimated pixel coordinates of the distorted calibration points and the pixel coordinates obtained from identifying the calibration points is determined. The intrinsic parameters, extrinsic parameters, and distortion coefficients of each camera are then adjusted based on this error. This process is repeated for multiple rounds of adjustment.

[0309] The second possible implementation is described in detail below. In this second possible implementation, a calibration location can be set within the sports field. This calibration location serves as the origin of the world coordinate system. The target calibration object passes through this calibration location during its movement across the sports field; this can also be understood as the target calibration object's basic moving position. For example, the location with the largest number of shared-view cameras can be selected as the calibration location. For further explanation of the moving position, please refer to [link to relevant documentation]. Figure 9 The descriptions in the corresponding embodiments will not be repeated here. The target marker moves through multiple positions, including positions 1 to M, during its movement across the sports field. Each position of the target marker can be represented by a graphical model, as shown below. Figure 17 As shown. The specific implementation process can be found in [the image / document]. Figure 18 As shown.

[0310] 1801, determine the relative pose of multiple pairs of movement positions among the M movement positions of the target calibrator, wherein the first movement position pair includes a first movement position and a second movement position, the first movement position and the second movement position are two movement positions among the M movement positions of the target calibrator, and the first movement position and the second movement position are located within the common field of view of at least one acquisition device group.

[0311] The relative pose of the first moving position pair is determined based on the matching feature point set corresponding to at least one acquisition device group, the three-dimensional coordinates of the target calibration object including multiple calibration points in the calibration object coordinate system, and the first intrinsic parameter estimates of the two acquisition devices included in at least one acquisition device group.

[0312] See Figure 17 As shown, cameras 1 to 9 are nine cameras deployed in a ring around the designated space of the sports field. S0 to S5 represent six positions where the target calibration object moves, with S0 representing the calibration position. The coordinates of the calibration position in the world coordinate system are known, i.e., the origin. The relative pose between the positions can be calculated based on the co-view relationship between each position and different cameras. Figure 17 (As shown by the dashed lines in the image). If two moving positions are located within the shared field of view of two cameras, these two moving positions constitute a moving position pair. For example, S0 and S5 are a moving position pair. S0 and S3 are a moving position pair. S5 and S1 are a moving position pair, and so on.

[0313] It should be noted that a pair of moving positions can be located within the shared field of view of multiple pairs of cameras. For example, the pair of moving positions consisting of S0 and S5 is located within the shared field of view of camera 1 and camera 2, and also within the shared field of view of camera 2 and camera 3.

[0314] 1802, Determine the poses of M movement positions in the world coordinate system based on the coordinates of the base movement position in the world coordinate system and the relative poses of multiple movement position pairs, wherein the base movement position is one of the M movement positions.

[0315] 1803. Based on the coordinates of M moving positions in the world coordinate system, the coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, and the pixel coordinates of the calibration points on the target calibration object in the image frames acquired by each acquisition device, the camera parameters of each acquisition device are globally optimized. The camera parameters include an intrinsic parameter matrix and an extrinsic parameter matrix, or the camera parameters include an intrinsic parameter matrix, an extrinsic parameter matrix, and distortion coefficients.

[0316] In the global optimization process, the camera parameters of each acquisition device and the pose of the calibration surface where the multiple calibration points are located in the calibration object coordinate system are taken as the quantities to be optimized; the first intrinsic parameter estimate of each acquisition device in the global optimization is taken as the initial value of the intrinsic parameter matrix of each acquisition device.

[0317] For example, the relative pose of a pair of moving positions can satisfy the condition shown in formula (20) below. The pair of moving positions is located within the common field of view of l camera pairs (camera groups). The pose is calculated based on each camera pair.

[0318]

[0319] in, This represents the position S corresponding to the l-th camera pair. i With position S j The relative poses between them. This indicates that the position of camera a is based on the position of the l-th camera. i Image frames acquired at location S and at location S j The location S determined by the image frame acquired at that location i To position S j The position. This indicates that the camera b is positioned at location S based on the l-th camera. i Image frames acquired at location S and at location S j The location S determined by the image frame acquired at that location j To position S i The position.

[0320] Determining the field of view includes position S i With position S j There are K cameras. Each pair of cameras forms a camera pair. For the first camera a and the second camera b in a camera pair, determine the position S from camera a. i To position S jThe pose, and the position S determined from camera b. j To position S i The position.

[0321] Determine position S from camera a i To position S j The pose can be determined based on the position S i The pose of the calibration object coordinate system relative to the camera coordinate system of camera a, and at position S j The pose of the calibration object's coordinate system relative to the camera coordinate system of camera a is determined. The position S is determined from camera b. j To position S i The pose can be determined based on the position S i The pose of the calibration object coordinate system relative to the camera coordinate system of camera b, and the position S j The pose of the calibration object coordinate system relative to the camera coordinate system of camera b is determined.

[0322] Specifically, the data is acquired by cameras a and b respectively when the target marker moves to position S. i The acquired image frames, and the positions S obtained by cameras a and b at the target marker. j Image frames were captured at the location. The pixel coordinates of the calibration points were obtained by identifying the image frames.

[0323] Based on the camera a moving to position S of the target marker. i The pixel coordinates of the calibration points in the acquired image frames, the coordinates of the calibration points of the target calibration object in the calibration object coordinate system, and the estimated value of the third intrinsic parameter of camera a can be used to determine the position S using the PnP algorithm. i pose of the calibration object coordinate system relative to the camera coordinate system of camera a Based on the camera a moving to position S of the target marker. j The pixel coordinates of the calibration points in the acquired image frames, the coordinates of the calibration points of the target calibration object in the calibration object coordinate system, and the estimated value of the third intrinsic parameter of camera a can be used to determine the position S using the PnP algorithm. j pose of the calibration object coordinate system relative to the camera coordinate system of camera a

[0324] In some embodiments, during execution Figure 18 Before the corresponding steps, you can first base it on Figure 15 or Figure 16 In some embodiments, the intrinsic parameters of each camera are first optimized to obtain a third intrinsic parameter estimate for each camera. Then, based on the third intrinsic parameter estimate, the pose at each location is further determined. In other embodiments, the following steps are performed: Figure 18 Before the corresponding steps, no longer based on Figure 15 or Figure 16The corresponding implementation first optimizes the intrinsic parameters of each camera. This allows for further determination of the pose at each location based on the second intrinsic parameter estimate.

[0325] Then, from the position S determined by camera a i To position S j position The conditions shown in the following formula (21) are satisfied.

[0326]

[0327] Based on the camera b moving to position S of the target marker. i The pixel coordinates of the calibration points in the acquired image frames, the coordinates of the calibration points of the target calibration object in the calibration object coordinate system, and the estimated value of the third intrinsic parameter of camera b can be used to determine the position S using the PnP algorithm. i pose of the calibration object coordinate system relative to the camera coordinate system of camera b Based on the camera b moving to position S of the target marker. j The pixel coordinates of the calibration points in the acquired image frames, the coordinates of the calibration points of the target calibration object in the calibration object coordinate system, and the estimated value of the third intrinsic parameter of camera b can be used to determine the position S using the PnP algorithm. j pose of the calibration object coordinate system relative to the camera coordinate system of camera b

[0328] Then, from the position S determined by camera b j To position S i position The conditions shown in formula (22) are satisfied.

[0329]

[0330] The above method can be used to determine the pose of the calibration object's coordinate system relative to each camera's coordinate system at different locations. See Table 2 for the poses of the calibration object at different locations relative to different cameras. Table 2 uses an example of deploying b cameras on a sports field. The example shows the trajectory of the calibration object passing through position 1 to position M.

[0331] Table 2

[0332]

[0333] It should be understood that a pair of movement positions can lie within the shared field of view of multiple camera pairs (or camera groups). Therefore, for each camera pair, the relative pose from one movement position to another within that pair is calculated. And the relative pose of one moving position to another moving position. According to Choose one relative pose from the relative poses of the movement position pairs corresponding to multiple camera pairs as the relative pose of the movement position pair. For example, it can be determined according to the following formula (23).

[0334]

[0335] Where I is the identity matrix. For example, if the minimum value corresponds to the 3rd camera pair (k=3), then the relative pose of that moving position pair is: The confidence weight of this move pair is then...

[0336] Furthermore, the poses of the M translation positions in the world coordinate system are determined based on the coordinates of the base translation position in the world coordinate system and the relative poses of multiple translation position pairs. This can be determined in the following way:

[0337] A1, determine the credibility weight of each of the multiple movement location pairs. Referring to the example above, the credibility weight of a movement location pair is... Execute A2. A2 determines the shortest path from the third move position to the base move position based on the confidence weight of each move position pair;

[0338] Wherein, the shortest path is the path with the smallest confidence weight among all paths from the third moving position to the base moving position; the confidence weight of any path is the sum of the confidence weights of the moving position pairs traversed by that path. Execute A3.

[0339] A3, determine the pose of the third moving position based on the relative poses of the moving position pairs traversed by the shortest path.

[0340] A graphical model is established for each movement position. The various movement positions of the target marker are used as vertices in the graphical model. The edges connecting any two vertices are defined by confidence weights. A reference position S0 (where the pose is the identity matrix) is selected as the origin of the global coordinate system. Then, the position S... i The pose relative to the reference position S0 can be simply referred to as position S. i Posture T i =T 0k1 T k1k2 …T kni .

[0341] Among them, from position S i To position S0, passing through position S k1 ~S kn And after passing through position S k1 ~S kn It is position S iThe shortest path to the reference position S0. The path can be determined using Dijkstra's algorithm, for example, by calculating the shortest path based on confidence weights. From position S... i The path corresponding to the minimum sum of weights among all paths to the reference position S0 is taken as the shortest path.

[0342] After determining the pose of each position relative to the reference position S0, the coordinates of each calibration point in the world coordinate system at each movement position can be calculated. Then, based on the coordinates of each calibration point in the world coordinate system (also known as the global coordinate system), the extrinsic parameter matrix of each camera is roughly calculated. Finally, the intrinsic parameter matrix, extrinsic parameter matrix, and distortion coefficients of each camera are globally optimized.

[0343] To obtain the pose T at different positions in global coordinates i Then, the pose T can be determined based on different positions. i The coordinates of each calibration point on the target calibration object in the target calibration object coordinate system are obtained, and the coordinates of each calibration point on the target calibration object in the global coordinate system at different locations are obtained.

[0344] Furthermore, the extrinsic parameters of each camera are determined. Taking camera i as an example, the extrinsic parameters of camera i are determined based on the coordinates of each calibration point on the target calibration object at different locations in the global coordinate system, the intrinsic parameters of camera i, and the pixel coordinates of the calibration points in the image frames acquired by camera i from the target calibration object at multiple different locations. The extrinsic parameters of all cameras are obtained in this way.

[0345] Next, the intrinsic and extrinsic parameters of each camera can be optimized by adjusting the coordinates of each calibration point on the target calibration object at different locations in the global coordinate system.

[0346] The coordinates of each calibration point on the target calibration object at different locations can be estimated in the global coordinate system to project the pixel coordinates of the calibration points on the target calibration object onto each camera at different locations. The specific determination method is shown in formula (24).

[0347] A calibration point on the target calibration object is at position S j Pixel coordinates projected onto camera i The following conditions must be met: (24)

[0348]

[0349] Among them, P wj Indicates position S j The coordinates of the calibration point on the target calibration object in the calibration object coordinate system, B m It is the pose of the calibration plate where the calibration point is located in the calibration object coordinate system, T j Indicates the position S of the target calibration object jpose in the world coordinate system, I i and [R|t] i These represent the intrinsic and extrinsic parameters (pose in the global coordinate system) of camera i, respectively. This indicates that the estimated value at position S is... j The pixel coordinates of the calibration point on the target calibration object in the image frame captured by camera i.

[0350] Among them, the above-mentioned B m It can be obtained through measurement. The target calibration object is assembled according to the design dimensions, such as... Figure 11 As shown. There are four calibration plates with QR codes affixed to them, three plates on each side. The printing size error of each plate is negligible. After determining the coordinate system of the calibration object, the design pose B of each calibration plate is obtained according to the design dimensions. m .

[0351] For example, since the actual position of the calibration point deviates from the design position due to assembly errors and manual pasting errors of the target calibration object, the design pose of each calibration plate can be used as a variable to be optimized, thereby improving the calibration accuracy of the camera's intrinsic parameters, extrinsic parameters, and distortion coefficients.

[0352] The pixel coordinates of the calibration points on the target calibration object at different locations in the image frames acquired by each camera are estimated using the above formula (24). Then, the intrinsic parameters, extrinsic parameters, and B of each camera are adjusted according to the error between the estimated pixel coordinates and the pixel coordinates in the recognized image frames. m Then, based on the adjusted camera's intrinsic and extrinsic parameters, B... m The next round of iterative adjustments will then be conducted. Specifically, this will be based on the adjusted intrinsic and extrinsic parameters of each camera, as well as the B-value. m The pixel coordinates of the calibration points on the target calibration object at different locations in the image frames acquired by each camera are estimated again. Then, based on the error between the estimated pixel coordinates and the pixel coordinates in the recognized image frames, the intrinsic parameters, extrinsic parameters, and B-values ​​of each camera are adjusted. m This process is repeated multiple times.

[0353] In some embodiments, considering camera distortion, formulas (1-1) and (1-2) are used to determine the pixel coordinates of the calibration points on the target calibration object at different locations in the image frames acquired by each camera. Then, based on the error between the estimated distorted pixel coordinates and the pixel coordinates in the identified image frames, the intrinsic parameters, extrinsic parameters, and B-values ​​of each camera are adjusted. m In addition to the distortion coefficient, multiple rounds of adjustments are performed to obtain the intrinsic parameters, extrinsic parameters, and distortion coefficients of each camera.

[0354] For example, the number of iterations can be pre-configured, and the iteration stops when the configured number of iterations is reached. An error threshold can also be pre-configured, and the iteration stops when the calculated error during a certain iteration is less than or equal to the error threshold.

[0355] This application's embodiments can be applied to motion analysis scenarios. Each camera acquires video streams captured during an athlete's movement. Motion information is then calculated based on calibrated camera parameters, such as the athlete's running distance, speed, and number of steps. Deeper information, such as team interface data and technical and tactical analysis, can also be obtained. The athlete's spatial position is accurately reconstructed from image frames captured by each camera and the calibrated camera parameters. Further motion information is then acquired. The three-dimensional spatial coordinates of all skeletal points can be calculated using the calibrated camera parameters based on the pixel coordinates of human skeleton points detected in the image frames. The quality of camera calibration directly affects the accuracy of the three-dimensional position calculation, thus affecting the reliability of subsequent motion analysis. Existing calibration schemes use multiple fixed calibration objects deployed on a designated sports field. Specifically, multiple calibration pillars with markers are placed in the main area of ​​the field, ensuring they are on the same horizontal plane. The physical distance between the pillars is measured, and the marker feature points on all calibration pillars are unified in the same coordinate system. Then, the camera parameters are calculated based on a direct linear transformation. However, this existing scheme with multiple fixed calibration objects results in a deviation between the acquired human skeleton points and the actual human body position. Compared to existing calibration schemes that use multiple fixed-position calibration objects, the calibration scheme proposed in this application involves pre-collecting data by moving the calibration objects on the sports field and running corresponding algorithms to calculate all camera parameters. The reprojected skeletal points in this application have a higher degree of overlap with the real human body. The moving scheme provided in this application is not limited by placement location, can cover a wider area, and the calibration results can more fully reflect the spatial relationships of the field area. Furthermore, by calculating the distortion coefficient, image edges can be corrected, reducing the adverse effects of lens distortion.

[0356] This application's embodiments can also be applied to large-scale sporting events, providing a 6 degrees of freedom (6DoF) experience. Viewers can freely choose their viewing position and angle, and even enter the scene, achieving close-ups and extreme close-ups, resulting in an immersive visual experience. To achieve a complete 6DoF video effect, high-precision camera calibration parameters are first required. Then, by analyzing the correlation between the content and features of each camera's captured images, a 3D reconstruction of the scene is performed. The calibration scheme provided in this application's embodiments yields more accurate calibration results, leading to superior 3D reconstruction that more closely resembles real-world effects.

[0357] It is understood that, in order to implement the functions in the above method embodiments, the data processing server includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the modules and method steps described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.

[0358] As an example, see Figure 19 The diagram shown is a schematic representation of a calibration device according to an embodiment of this application. This device can be applied to a data processing server. The calibration device includes an acquisition unit 1901 and a processing unit 1902.

[0359] Acquisition unit 1901 is used to acquire multiple video streams collected by multiple acquisition devices, which are deployed in a designated space of the sports field. The multiple video streams are synchronously captured by the multiple acquisition devices during the movement of a target calibration object in the sports field. The movement trajectory of the target calibration object in the sports field at least covers a designated area of ​​the sports field. The target calibration object includes at least two non-coplanar calibration surfaces, and each calibration surface includes at least two calibration points. Each acquisition device's video stream includes multiple image frames.

[0360] Processing unit 1902 is configured to perform calibration point detection on image frames acquired by each of the plurality of acquisition devices to obtain pixel coordinates of multiple calibration points on the target calibration object in the image frames acquired by each acquisition device; based on the pixel coordinates of the multiple calibration points included in the target calibration object in the image frames acquired by each acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system of the target calibration object, estimate the intrinsic parameter matrix of each acquisition device to obtain a first intrinsic parameter estimate value for each acquisition device; and based on the first intrinsic parameter estimate values ​​of at least two acquisition devices included in each acquisition device group in the plurality of acquisition device groups, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of the multiple calibration points included in the target calibration object in the calibration object coordinate system, determine a first extrinsic parameter estimate value for each acquisition device.

[0361] In this system, each acquisition device group includes at least two acquisition devices that share a common viewing area. The matching feature point set includes multiple matching feature point groups, and each matching feature point group includes at least two matching pixel coordinates. The at least two matching pixel coordinates are the pixel coordinates of the same calibration point detected by different acquisition devices belonging to the same acquisition device group at the same time. The multiple acquisition device groups are obtained by grouping the multiple acquisition devices. Any two acquisition device groups in the multiple acquisition device groups include at least one identical acquisition device.

[0362] In one possible implementation, the processing unit 1902 is further configured to:

[0363] Based on the pixel coordinates of multiple calibration points included in the target calibration object in the image frame acquired by each acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system of the target calibration object, the distortion coefficient of each acquisition device is estimated to obtain the first distortion coefficient estimate of each acquisition device.

[0364] In one possible implementation, the processing unit 1902 is specifically used for:

[0365] Based on the pixel coordinates of the calibration points on the target calibration object in the image frames included in the image set acquired by the i-th acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the intrinsic parameter matrix of the i-th acquisition device is estimated to obtain the second intrinsic parameter estimate of the i-th acquisition device; the image set includes M1 image frames of the target calibration object in the video stream acquired by the i-th acquisition device, and the M1 image frames correspond one-to-one with the M1 movement positions of the target calibration object; M1 is a positive integer, and M is an integer greater than M1;

[0366] Based on the second intrinsic parameter estimate of the i-th acquisition device, the pixel coordinates of the calibration points on the target calibration object in the image frames included in the image set acquired by the i-th acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the pose set corresponding to the i-th acquisition device is estimated respectively. The pose set corresponding to the i-th acquisition device includes the pose of the target calibration object relative to the i-th acquisition device at M1 moving positions; i takes the value of a positive integer less than or equal to N, where N is the number of acquisition devices deployed in the set space of the sports field;

[0367] Among them, the range of movement positions corresponding to the image frames acquired by different acquisition devices is different;

[0368] Based on the pixel coordinates of the calibration points on the target calibration object in the image frames of the image sets acquired by N acquisition devices, the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, and the poses of the target calibration object corresponding to the N acquisition devices, and on the basis of the initially set distortion coefficients and second intrinsic parameter estimates corresponding to the N acquisition devices, the intrinsic parameter matrices and distortion coefficients of the N acquisition devices are adjusted globally in multiple rounds to obtain the first intrinsic parameter estimates and first distortion coefficient estimates of the N acquisition devices.

[0369] In one possible implementation, the processing unit 1902 is specifically used for:

[0370] Based on the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the pose set of the target calibration object corresponding to each acquisition device, the estimated value of the second intrinsic parameter corresponding to each acquisition device, and the initially set distortion coefficient, the pixel coordinates of the multiple calibration points in the image coordinate system of each acquisition device are estimated.

[0371] The error between the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of each acquisition device and the pixel coordinates of the plurality of calibration points extracted from the image frames acquired from each acquisition device is obtained;

[0372] Based on the error, the pose set of the target calibration object corresponding to each acquisition device, the second intrinsic parameter estimate value corresponding to each acquisition device, and the initially set distortion coefficient are adjusted to obtain the intrinsic parameter estimate value and distortion coefficient corresponding to each acquisition device after the current round of adjustment;

[0373] In this round, the estimated intrinsic parameters and distortion coefficients of each acquisition device after the current round of adjustment are used as the basis for the next round of adjustment, until the C round of adjustment is completed to obtain the first estimated intrinsic parameters and the first estimated distortion coefficients of the N acquisition devices.

[0374] In one possible implementation, the processing unit 1902 is specifically used for:

[0375] Based on the three-dimensional coordinates of multiple calibration points on the target calibration object at the k-th moving position, and the pose of the target calibration object at the k-th moving position corresponding to the i-th acquisition device, determine the coordinates of the multiple calibration points projected onto the camera coordinate system of the i-th acquisition device;

[0376] The distorted coordinates of the plurality of calibration points projected onto the camera coordinate system are determined based on the coordinates of the i-th acquisition device in the camera coordinate system and the distortion coefficient of the i-th acquisition device initially set.

[0377] The pixel coordinates of the plurality of calibration points projected onto the image coordinate system of the i-th acquisition device are estimated based on the distorted coordinates and the second intrinsic parameter estimate of the i-th acquisition device.

[0378] In one possible implementation, the processing unit 1902 is specifically used for:

[0379] Based on the first intrinsic parameter estimates of at least two acquisition devices in each of the multiple acquisition device groups, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, the second relative pose of the other acquisition devices among the multiple acquisition devices, excluding the reference acquisition device, relative to the reference acquisition device is obtained; the reference acquisition device is any one of the multiple acquisition devices.

[0380] A scale factor is determined, which is the ratio between a first distance and a second distance. The first distance is the distance between two calibration points on the target calibration object, and the second distance is the distance between the two calibration points in the same image coordinate system. The two calibration points are located on the same calibration surface on the target calibration object.

[0381] The first extrinsic parameter estimate of each acquisition device is obtained based on the second relative pose of each acquisition device and the scale factor.

[0382] In one possible implementation, the processing unit 1902 is specifically used for:

[0383] The essential matrix between the first acquisition device and the reference acquisition device is determined based on the matching feature point set corresponding to the first acquisition device group. The first acquisition device and the reference acquisition device belong to the first acquisition device group, and the first acquisition device group is one of the plurality of acquisition device groups.

[0384] Based on the singular value decomposition results of the essential matrix, the second relative pose between the first acquisition device and the reference acquisition device is determined.

[0385] In one possible implementation, the processing unit 1902 is specifically used for:

[0386] After determining the second relative poses of the acquisition devices in the g-th acquisition device group relative to the reference acquisition device, the three-dimensional coordinates of the multiple calibration points in the local coordinate system are determined according to the second relative poses of each acquisition device when the target calibration object moves to M2 moving positions. The local coordinate system is the camera coordinate system of the reference acquisition device. Any of the M2 moving positions is located at least within the common viewing area of ​​two acquisition devices in the g-th acquisition device.

[0387] Based on the three-dimensional coordinates of the multiple calibration points at the M2 moving positions in the local coordinate system, the second relative pose of the acquisition devices included in the g-th acquisition device group, and the first intrinsic parameter estimate, the pixel coordinates of the multiple calibration points at the M2 moving positions are estimated to be projected onto the image coordinate system of the acquisition devices included in the g-th acquisition device group.

[0388] The error between the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of the acquisition devices included in the g-th acquisition device group and the pixel coordinates of the plurality of calibration points extracted from the image frames acquired by the acquisition devices included in the g-th acquisition device group is obtained.

[0389] Based on the error, the second relative pose and the first intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group are adjusted to obtain the relative pose and intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group after the current round of adjustment.

[0390] Among them, the intrinsic parameter estimates and relative poses of the acquisition devices in the g-th acquisition device group after the current round of adjustment are used as the basis for the next round of adjustment, until the D-th round of adjustment is completed to obtain the third relative pose and third intrinsic parameter estimates of the acquisition devices included in the g-th acquisition device group;

[0391] The first extrinsic parameter estimate is obtained by adding the scale factor to the third relative pose of the acquisition devices included in the g-th acquisition device group.

[0392] In one possible implementation, the processing unit 1902 is specifically used for:

[0393] After determining the second relative poses of the acquisition devices in the g-th acquisition device group relative to the reference acquisition device, the three-dimensional coordinates of the multiple calibration points in the local coordinate system are determined according to the second relative poses of each acquisition device when the target calibration object moves to M2 moving positions. The local coordinate system is the camera coordinate system of the reference acquisition device. Any of the M2 moving positions is located at least within the common viewing area of ​​two acquisition devices in the g-th acquisition device.

[0394] Based on the three-dimensional coordinates of the multiple calibration points at the M2 moving positions in the local coordinate system, the second relative pose of the acquisition devices included in the g-th acquisition device group, the first intrinsic parameter estimate, and the first distortion coefficient estimate, the pixel coordinates of the multiple calibration points at the M2 moving positions projected onto the image coordinate system of the acquisition devices included in the g-th acquisition device group are estimated.

[0395] The error between the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of the acquisition devices included in the g-th acquisition device group and the pixel coordinates of the plurality of calibration points extracted from the image frames acquired by the acquisition devices included in the g-th acquisition device group is obtained.

[0396] Based on the error, the second relative pose and the first intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group are adjusted to obtain the relative pose and intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group after the current round of adjustment.

[0397] Among them, the intrinsic parameter estimates and relative poses of the acquisition devices in the g-th acquisition device group after the current round of adjustment are used as the basis for the next round of adjustment, until the D-th round of adjustment is completed to obtain the third relative pose and third intrinsic parameter estimates of the acquisition devices included in the g-th acquisition device group;

[0398] The first extrinsic parameter estimate is obtained by adding the scale factor to the third relative pose of the acquisition devices included in the g-th acquisition device group.

[0399] In one possible implementation, each of the plurality of acquisition device groups includes two acquisition devices, and the processing unit 1902 is specifically used for:

[0400] The relative poses of multiple pairs of movement positions among the M movement positions of the target calibration object are determined. A first pair of movement positions includes a first movement position and a second movement position. The first and second movement positions are two of the M movement positions of the target calibration object, and the first and second movement positions are located within the common viewing area of ​​at least one group of acquisition devices. The relative poses of the first pair of movement positions are determined based on the matching feature point set corresponding to at least one group of acquisition devices, the three-dimensional coordinates of multiple calibration points of the target calibration object in the calibration object coordinate system, and the first intrinsic parameter estimates of the two acquisition devices included in the at least one group of acquisition devices.

[0401] The poses of M movement positions in the world coordinate system are determined based on the coordinates of the base movement position in the world coordinate system and the relative poses of multiple movement position pairs, wherein the base movement position is one of the M movement positions;

[0402] Based on the coordinates of M moving positions in the world coordinate system, the coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, and the pixel coordinates of the calibration points on the target calibration object in the image frames acquired by each acquisition device, the camera parameters of each acquisition device are globally optimized. The camera parameters include an intrinsic parameter matrix and an extrinsic parameter matrix, or the camera parameters include an intrinsic parameter matrix, an extrinsic parameter matrix, and distortion coefficients.

[0403] In the global optimization process, the camera parameters of each acquisition device and the pose of the calibration surface where the multiple calibration points are located in the calibration object coordinate system are taken as the quantities to be optimized; the first intrinsic parameter estimate of each acquisition device in the global optimization is taken as the initial value of the intrinsic parameter matrix of each acquisition device.

[0404] In one possible implementation, the relative poses of the first movement position pair satisfy the following condition:

[0405]

[0406] Among them, T 12 This indicates the relative pose between the first moving position and the second moving position. The at least one group of acquisition devices includes a first group of acquisition devices, which in turn includes a first acquisition device and a second acquisition device. It represents the pose from the first moving position to the second moving position, determined based on the pixel coordinates of the calibration point in the image frame acquired by the first acquisition device when the target calibration object moves to the first moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the second moving position; This indicates the pose from the second moving position to the first moving position, determined based on the pixel coordinates of the calibration point in the image frame acquired by the second acquisition device when the target calibration object moves to the second moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the first moving position.

[0407] In one possible implementation, at least one group of acquisition devices is L, and the first group of acquisition devices satisfies:

[0408]

[0409] Where I represents the identity matrix, This represents the pose from the first moving position to the second moving position, determined based on the pixel coordinates of the calibration point in the image frame acquired by the first acquisition device in the l-th acquisition device group when the target calibration object moves to the first moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the second moving position. 11 represents the pose from the second moving position to the first moving position, determined by the pixel coordinates of the calibration point in the image frame acquired by the second acquisition device in the l-th acquisition device group when the target calibration object moves to the second moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the first moving position; 12 represents the first acquisition device in the first acquisition device group and 12 represents the second acquisition device in the second acquisition device group.

[0410] In one possible implementation, determining the poses of the M mobile positions in the world coordinate system based on the coordinates of the base mobile position in the world coordinate system and the relative poses of multiple mobile position pairs includes:

[0411] Determine the confidence weight for each of the multiple movement position pairs, where the confidence weight between the first movement position and the second movement position satisfies the following condition: S 12 This represents the confidence weight between the first and second movement positions.

[0412] The shortest path from the third move position to the base move position is determined based on the credibility weight of each move position pair.

[0413] The shortest path is the path with the smallest confidence weight among all paths from the third moving position to the base moving position; the confidence weight of any path is the sum of the confidence weights of the moving position pairs traversed by the path.

[0414] The pose of the third moving position is determined based on the relative poses of the moving position pairs traversed by the shortest path.

[0415] The unit division in this embodiment is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into a single processor, exist as separate physical units, or be integrated into a single unit. The integrated units described above can be implemented in hardware or as software functional units. Figure 19 One or more of the various units within can be implemented using software, hardware, firmware, or a combination thereof. The software or firmware includes, but is not limited to, computer program instructions or code, and can be executed by a hardware processor. The hardware includes, but is not limited to, various integrated circuits, such as a central processing unit (CPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC).

[0416] Based on the above embodiments and the same concept, this application also provides a calibration device for implementing the calibration method provided in this application. Figure 20As shown, the device may include one or more processors 2001, a memory 2002, and one or more computer programs (not shown). In one implementation, the aforementioned devices may be coupled via one or more communication lines 2003. The memory 2002 stores one or more computer programs, which include instructions; the processor 2001 invokes the instructions stored in the memory 2002, causing the device to execute the calibration method provided in the embodiments of this application.

[0417] In the embodiments of this application, the processor may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0418] In the embodiments of this application, the memory can be volatile memory or non-volatile memory, or it can include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory. The memory in the embodiments of this application may also be circuitry or any other means capable of implementing storage functions.

[0419] In one implementation, the device may further include a communication interface 2004 for communicating with other devices via a transmission medium. For example, the communication interface 2004 can be used to communicate with an acquisition device to receive image frames acquired by the acquisition device. In this embodiment, the communication interface 2004 may be a transceiver, circuit, bus, module, or other type of communication interface. In this embodiment, when the communication interface 2004 is a transceiver, the transceiver may include an independent receiver, an independent transmitter, or a transceiver with integrated transceiver functions, or an interface circuit.

[0420] In some embodiments of this application, the processor 2001, memory 2002, and communication interface 2004 can be interconnected via a communication line 2003. The communication line 2003 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication line 2003 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 20 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0421] In the various embodiments of this application, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of different embodiments are consistent and can be referenced by each other. The technical features of different embodiments can be combined to form new embodiments according to their inherent logical relationship.

[0422] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple. In the textual description of this application, the character " / " generally indicates that the preceding and following related objects are in an "or" relationship. In the formulas of this application, the character " / " indicates that the preceding and following related objects are in a "division" relationship. Additionally, in this application, the word "exemplarily" is used to indicate an example, illustration, or explanation. Any embodiment or design described as "example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Alternatively, it can be understood that the use of the word "example" is intended to present concepts in a specific manner and does not constitute a limitation of this application.

[0423] It is understood that the various numerical designations used in this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application. The order of the process numbers described above does not imply the order of execution; the execution order of each process should be determined by its function and inherent logic. Terms such as "first," "second," and similar expressions are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or device is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.

[0424] One embodiment of this application provides a computer-readable medium for storing a computer program that includes instructions for performing the method steps described in the above method embodiments.

[0425] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, optical storage, etc.) containing computer-usable program code.

[0426] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0427] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and variations.

Claims

1. A calibration method characterized by, include: Multiple video streams are acquired from multiple acquisition devices deployed in a designated space within a sports field. The video streams are captured synchronously by the acquisition devices during the movement of a target calibration object within the sports field. The trajectory of the target calibration object within the sports field at least covers a designated area of ​​the sports field. The target calibration object includes at least two non-coplanar calibration surfaces, each of which includes at least two calibration points. Each video stream acquired by each acquisition device includes multiple image frames. Calibration point detection is performed on the image frames acquired by each of the plurality of acquisition devices to obtain the pixel coordinates of multiple calibration points on the target calibration object in the image frames acquired by each acquisition device; Based on the pixel coordinates of multiple calibration points included in the target calibration object in the image frame acquired by each acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system of the target calibration object, the intrinsic parameter matrix of each acquisition device is estimated to obtain the first intrinsic parameter estimate of each acquisition device; Based on the first intrinsic parameter estimates of at least two acquisition devices in each acquisition device group, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, the first extrinsic parameter estimate of each acquisition device is determined. In this system, each acquisition device group includes at least two acquisition devices that share a common viewing area. The matching feature point set includes multiple matching feature point groups, and each matching feature point group includes at least two matching pixel coordinates. The at least two matching pixel coordinates are the pixel coordinates of the same calibration point detected by different acquisition devices belonging to the same acquisition device group at the same time. The multiple acquisition device groups are obtained by grouping the multiple acquisition devices. Any two acquisition device groups in the multiple acquisition device groups include at least one identical acquisition device.

2. The method of claim 1, wherein, The method further includes: Based on the pixel coordinates of multiple calibration points included in the target calibration object in the image frame acquired by each acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system of the target calibration object, the distortion coefficient of each acquisition device is estimated to obtain the first distortion coefficient estimate of each acquisition device.

3. The method of claim 2, wherein, Based on the pixel coordinates of multiple calibration points included in the target calibration object in the image frames acquired by each acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the intrinsic parameter matrix of each acquisition device is estimated to obtain the first intrinsic parameter estimate value of each acquisition device, including: Based on the pixel coordinates of the calibration points on the target calibration object in the image frames included in the image set acquired by the i-th acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the intrinsic parameter matrix of the i-th acquisition device is estimated to obtain the second intrinsic parameter estimate of the i-th acquisition device; the image set includes M1 image frames of the target calibration object in the video stream acquired by the i-th acquisition device, and the M1 image frames correspond one-to-one with the M1 movement positions of the target calibration object among the M movement positions; M1 is a positive integer, and M is an integer greater than M1; Based on the second intrinsic parameter estimate of the i-th acquisition device, the pixel coordinates of the calibration points on the target calibration object in the image frames included in the image set acquired by the i-th acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the pose set corresponding to the i-th acquisition device is estimated respectively. The pose set corresponding to the i-th acquisition device includes the pose of the target calibration object relative to the i-th acquisition device at M1 moving positions; i takes the value of a positive integer less than or equal to N, where N is the number of acquisition devices deployed in the set space of the sports field; Among them, the range of movement positions corresponding to the image frames acquired by different acquisition devices is different; Based on the pixel coordinates of the calibration points on the target calibration object in the image frames of the image sets acquired by N acquisition devices, the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, and the poses of the target calibration object corresponding to the N acquisition devices, and on the basis of the initially set distortion coefficients and second intrinsic parameter estimates corresponding to the N acquisition devices, the intrinsic parameter matrices and distortion coefficients of the N acquisition devices are adjusted globally in multiple rounds to obtain the first intrinsic parameter estimates and first distortion coefficient estimates of the N acquisition devices.

4. The method of claim 3, wherein, The intrinsic parameter matrices and distortion coefficients of the N acquisition devices are adjusted through multiple rounds of global iteration to obtain the first intrinsic parameter estimates and the first distortion coefficient estimates of the N acquisition devices, including: Based on the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the pose set of the target calibration object corresponding to each acquisition device, the estimated value of the second intrinsic parameter corresponding to each acquisition device, and the initially set distortion coefficient, the pixel coordinates of the multiple calibration points in the image coordinate system of each acquisition device are estimated. The error between the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of each acquisition device and the pixel coordinates of the plurality of calibration points extracted from the image frames acquired from each acquisition device is obtained; Based on the error, the pose set of the target calibration object corresponding to each acquisition device, the second intrinsic parameter estimate value corresponding to each acquisition device, and the initially set distortion coefficient are adjusted to obtain the intrinsic parameter estimate value and distortion coefficient corresponding to each acquisition device after the current round of adjustment; In this round, the estimated intrinsic parameters and distortion coefficients of each acquisition device after the current round of adjustment are used as the basis for the next round of adjustment, until the C round of adjustment is completed to obtain the first estimated intrinsic parameters and the first estimated distortion coefficients of the N acquisition devices.

5. The method of claim 4, wherein, Based on the three-dimensional coordinates of the plurality of calibration points in the calibration object coordinate system, the pose set of the target calibration object corresponding to each acquisition device, the estimated value of the second intrinsic parameter corresponding to each acquisition device, and the initially set distortion coefficient, the pixel coordinates of the plurality of calibration points in the image coordinate system of each acquisition device are estimated, including: Based on the three-dimensional coordinates of multiple calibration points on the target calibration object at the k-th moving position, and the pose of the target calibration object at the k-th moving position corresponding to the i-th acquisition device, determine the coordinates of the multiple calibration points projected onto the camera coordinate system of the i-th acquisition device; The distorted coordinates of the plurality of calibration points projected onto the camera coordinate system are determined based on the coordinates of the i-th acquisition device in the camera coordinate system and the distortion coefficient of the i-th acquisition device initially set. The pixel coordinates of the plurality of calibration points projected onto the image coordinate system of the i-th acquisition device are estimated based on the distorted coordinates and the second intrinsic parameter estimate of the i-th acquisition device.

6. The method according to any one of claims 1 to 5, wherein, Based on the estimated first intrinsic parameters of at least two acquisition devices in each acquisition device group, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of multiple calibration points of the target calibration object in the calibration object coordinate system, the estimated first extrinsic parameter of each acquisition device is estimated, including: Based on the first intrinsic parameter estimates of at least two acquisition devices in each of the multiple acquisition device groups, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, the second relative pose of the other acquisition devices among the multiple acquisition devices, excluding the reference acquisition device, relative to the reference acquisition device is obtained; the reference acquisition device is any one of the multiple acquisition devices. A scale factor is determined, which is the ratio between a first distance and a second distance. The first distance is the distance between two calibration points on the target calibration object, and the second distance is the distance between the two calibration points in the same image coordinate system. The two calibration points are located on the same calibration surface on the target calibration object. The first extrinsic parameter estimate of each acquisition device is obtained based on the second relative pose of each acquisition device and the scale factor.

7. The method of claim 6, wherein, Based on the first intrinsic parameter estimates of at least two acquisition devices in each of the multiple acquisition device groups, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, the second relative pose of the other acquisition devices (excluding the reference acquisition device) relative to the reference acquisition device is obtained, including: The essential matrix between the first acquisition device and the reference acquisition device is determined based on the matching feature point set corresponding to the first acquisition device group. The first acquisition device and the reference acquisition device belong to the first acquisition device group, and the first acquisition device group is one of the plurality of acquisition device groups. Based on the singular value decomposition results of the essential matrix, the second relative pose between the first acquisition device and the reference acquisition device is determined.

8. The method of claim 6, wherein, Based on the second relative pose of each acquisition device and the scale factor, the first extrinsic parameter estimate of each acquisition device is obtained, including: After determining the second relative poses of the acquisition devices in the g-th acquisition device group relative to the reference acquisition device, the three-dimensional coordinates of the multiple calibration points in the local coordinate system are determined based on the second relative poses of each acquisition device when the target calibration object moves to M2 moving positions. The local coordinate system is the camera coordinate system of the reference acquisition device. Any of the M2 moving positions is located at least within the common viewing area of ​​two acquisition devices in the g-th acquisition device group. Based on the three-dimensional coordinates of the multiple calibration points at the M2 moving positions in the local coordinate system, the second relative pose of the acquisition devices included in the g-th acquisition device group, and the first intrinsic parameter estimate, the pixel coordinates of the multiple calibration points at the M2 moving positions are estimated to be projected onto the image coordinate system of the acquisition devices included in the g-th acquisition device group. The error between the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of the acquisition devices included in the g-th acquisition device group and the pixel coordinates of the plurality of calibration points extracted from the image frames acquired by the acquisition devices included in the g-th acquisition device group is obtained. Adjust the second relative pose and the first intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group according to the error, and obtain the relative pose and intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group after the current round of adjustment; Among them, the intrinsic parameter estimates and relative poses of the acquisition devices in the g-th acquisition device group after the current round of adjustment are used as the basis for the next round of adjustment, until the D-th round of adjustment is completed to obtain the third relative pose and third intrinsic parameter estimates of the acquisition devices included in the g-th acquisition device group; The first extrinsic parameter estimate is obtained by adding the scale factor to the third relative pose of the acquisition devices included in the g-th acquisition device group.

9. The method of claim 6, wherein, Based on the second relative pose of each acquisition device and the scale factor, the first extrinsic parameter estimate of each acquisition device is obtained, including: After determining the second relative poses of the acquisition devices in the g-th acquisition device group relative to the reference acquisition device, the three-dimensional coordinates of the multiple calibration points in the local coordinate system are determined based on the second relative poses of each acquisition device when the target calibration object moves to M2 moving positions. The local coordinate system is the camera coordinate system of the reference acquisition device. Any of the M2 moving positions is located at least within the common viewing area of ​​two acquisition devices in the g-th acquisition device group. Based on the three-dimensional coordinates of the multiple calibration points at the M2 moving positions in the local coordinate system, the second relative pose of the acquisition devices included in the g-th acquisition device group, the first intrinsic parameter estimate, and the first distortion coefficient estimate, the pixel coordinates of the multiple calibration points at the M2 moving positions projected onto the image coordinate system of the acquisition devices included in the g-th acquisition device group are estimated. The error between the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of the acquisition devices included in the g-th acquisition device group and the pixel coordinates of the plurality of calibration points extracted from the image frames acquired by the acquisition devices included in the g-th acquisition device group is obtained. Adjust the second relative pose and the first intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group according to the error, and obtain the relative pose and intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group after the current round of adjustment; Among them, the intrinsic parameter estimates and relative poses of the acquisition devices in the g-th acquisition device group after the current round of adjustment are used as the basis for the next round of adjustment, until the D-th round of adjustment is completed to obtain the third relative pose and third intrinsic parameter estimates of the acquisition devices included in the g-th acquisition device group; The first extrinsic parameter estimate is obtained by adding the scale factor to the third relative pose of the acquisition devices included in the g-th acquisition device group.

10. The method of any one of claims 1-5, wherein, Each of the plurality of acquisition device groups includes two acquisition devices. Based on the first intrinsic parameter estimates of at least two acquisition devices in each acquisition device group, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of the multiple calibration points included in the target calibration object in the calibration object coordinate system, the first extrinsic parameter estimates of each acquisition device are estimated, including: The relative poses of multiple pairs of movement positions among the M movement positions of the target calibration object are determined. A first pair of movement positions includes a first movement position and a second movement position. The first and second movement positions are two of the M movement positions of the target calibration object, and the first and second movement positions are located within the common viewing area of ​​at least one group of acquisition devices. The relative poses of the first pair of movement positions are determined based on the matching feature point set corresponding to at least one group of acquisition devices, the three-dimensional coordinates of multiple calibration points of the target calibration object in the calibration object coordinate system, and the first intrinsic parameter estimates of the two acquisition devices included in the at least one group of acquisition devices. The poses of M movement positions in the world coordinate system are determined based on the coordinates of the base movement position in the world coordinate system and the relative poses of multiple movement position pairs, wherein the base movement position is one of the M movement positions; Based on the coordinates of M moving positions in the world coordinate system, the coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, and the pixel coordinates of the calibration points on the target calibration object in the image frames acquired by each acquisition device, the camera parameters of each acquisition device are globally optimized. The camera parameters include an intrinsic parameter matrix and an extrinsic parameter matrix, or the camera parameters include an intrinsic parameter matrix, an extrinsic parameter matrix, and distortion coefficients. In the global optimization process, the camera parameters of each acquisition device and the pose of the calibration surface where the multiple calibration points are located in the calibration object coordinate system are taken as the quantities to be optimized; the first intrinsic parameter estimate of each acquisition device in the global optimization is taken as the initial value of the intrinsic parameter matrix of each acquisition device.

11. The method of claim 10, wherein, The relative poses of the first moving position pair satisfy the following conditions: ; wherein, denotes a relative pose between the first mobile position and the second mobile position, the at least one set of acquisition devices comprises a first set of acquisition devices, the first set of acquisition devices comprises a first acquisition device and a second acquisition device, denotes a pose of the first mobile position to the second mobile position determined based on pixel coordinates of the calibration points in the image frames acquired by the first acquisition device when the target calibration object is moved to the first mobile position and pixel coordinates of the calibration points in the image frames acquired by the first acquisition device when the target calibration object is moved to the second mobile position; denotes a pose of the second mobile position to the first mobile position determined based on pixel coordinates of the calibration points in the image frames acquired by the second acquisition device when the target calibration object is moved to the second mobile position and pixel coordinates of the calibration points in the image frames acquired by the second acquisition device when the target calibration object is moved to the first mobile position.

12. The method as described in claim 11, characterized in that, At least one group of acquisition devices consists of L devices, and the first group of acquisition devices satisfies: ; in, Represents the identity matrix. Indicates based on the first The pose from the first moving position to the second moving position is determined by the pixel coordinates of the calibration point in the image frame acquired by the first acquisition device in the group of acquisition devices when the target calibration object moves to the first moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the second moving position; Indicates based on the first The pose from the second moving position to the first moving position is determined by the pixel coordinates of the calibration point in the image frame acquired by the second acquisition device in the acquisition device group when the target calibration object moves to the second moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the first moving position; Indicates the first The first acquisition device in the acquisition device group. 2 indicates the first The second acquisition device in the acquisition device group.

13. The method as described in claim 12, characterized in that, The process of determining the poses of M mobile positions in the world coordinate system based on the coordinates of the base mobile position in the world coordinate system and the relative poses of multiple mobile position pairs includes: Determine the confidence weight for each of the multiple movement position pairs, where the confidence weight between the first movement position and the second movement position satisfies the following condition: ; This represents the confidence weight between the first and second movement positions. The shortest path from the third move position to the base move position is determined based on the credibility weight of each move position pair. The shortest path is the path with the smallest confidence weight among all paths from the third moving position to the base moving position; the confidence weight of any path is the sum of the confidence weights of the moving position pairs traversed by the path. The pose of the third moving position is determined based on the relative poses of the moving position pairs traversed by the shortest path.

14. A calibration device, characterized in that, include: An acquisition unit is used to acquire multiple video streams collected by multiple acquisition devices deployed in a designated space within a sports field. The multiple video streams are synchronously captured by the multiple acquisition devices during the movement of a target calibration object within the sports field. The movement trajectory of the target calibration object within the sports field at least covers a designated area of ​​the sports field. The target calibration object includes at least two non-coplanar calibration surfaces, and each calibration surface includes at least two calibration points. Each video stream acquired by the acquisition device includes multiple image frames. The processing unit is used to perform calibration point detection on the image frames acquired by each of the plurality of acquisition devices, so as to obtain the pixel coordinates of multiple calibration points on the target calibration object in the image frames acquired by each acquisition device; Based on the pixel coordinates of multiple calibration points included in the target calibration object in the image frame acquired by each acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system of the target calibration object, the intrinsic parameter matrix of each acquisition device is estimated to obtain the first intrinsic parameter estimate of each acquisition device; Based on the first intrinsic parameter estimates of at least two acquisition devices in each acquisition device group, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, the first extrinsic parameter estimate of each acquisition device is determined. In this system, each acquisition device group includes at least two acquisition devices that share a common viewing area. The matching feature point set includes multiple matching feature point groups, and each matching feature point group includes at least two matching pixel coordinates. The at least two matching pixel coordinates are the pixel coordinates of the same calibration point detected by different acquisition devices belonging to the same acquisition device group at the same time. The multiple acquisition device groups are obtained by grouping the multiple acquisition devices. Any two acquisition device groups in the multiple acquisition device groups include at least one identical acquisition device.

15. The apparatus as claimed in claim 14, characterized in that, The processing unit is further configured to: Based on the pixel coordinates of multiple calibration points included in the target calibration object in the image frame acquired by each acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system of the target calibration object, the distortion coefficient of each acquisition device is estimated to obtain the first distortion coefficient estimate of each acquisition device.

16. The apparatus as claimed in claim 15, characterized in that, The processing unit is specifically used for: Based on the pixel coordinates of the calibration points on the target calibration object in the image frames included in the image set acquired by the i-th acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the intrinsic parameter matrix of the i-th acquisition device is estimated to obtain the second intrinsic parameter estimate of the i-th acquisition device; the image set includes M1 image frames of the target calibration object in the video stream acquired by the i-th acquisition device, and the M1 image frames correspond one-to-one with the M1 movement positions of the target calibration object among the M movement positions; M1 is a positive integer, and M is an integer greater than M1; Based on the second intrinsic parameter estimate of the i-th acquisition device, the pixel coordinates of the calibration points on the target calibration object in the image frames included in the image set acquired by the i-th acquisition device, and the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the pose set corresponding to the i-th acquisition device is estimated respectively. The pose set corresponding to the i-th acquisition device includes the pose of the target calibration object relative to the i-th acquisition device at M1 moving positions; i takes the value of a positive integer less than or equal to N, where N is the number of acquisition devices deployed in the set space of the sports field; Among them, the range of movement positions corresponding to the image frames acquired by different acquisition devices is different; Based on the pixel coordinates of the calibration points on the target calibration object in the image frames of the image sets acquired by N acquisition devices, the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, and the poses of the target calibration object corresponding to the N acquisition devices, and on the basis of the initially set distortion coefficients and second intrinsic parameter estimates corresponding to the N acquisition devices, the intrinsic parameter matrices and distortion coefficients of the N acquisition devices are adjusted globally in multiple rounds to obtain the first intrinsic parameter estimates and first distortion coefficient estimates of the N acquisition devices.

17. The apparatus as claimed in claim 16, characterized in that, The processing unit is specifically used for: Based on the three-dimensional coordinates of the multiple calibration points in the calibration object coordinate system, the pose set of the target calibration object corresponding to each acquisition device, the estimated value of the second intrinsic parameter corresponding to each acquisition device, and the initially set distortion coefficient, the pixel coordinates of the multiple calibration points in the image coordinate system of each acquisition device are estimated. The error between the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of each acquisition device and the pixel coordinates of the plurality of calibration points extracted from the image frames acquired from each acquisition device is obtained; Based on the error, the pose set of the target calibration object corresponding to each acquisition device, the second intrinsic parameter estimate value corresponding to each acquisition device, and the initially set distortion coefficient are adjusted to obtain the intrinsic parameter estimate value and distortion coefficient corresponding to each acquisition device after the current round of adjustment; In this round, the estimated intrinsic parameters and distortion coefficients of each acquisition device after the current round of adjustment are used as the basis for the next round of adjustment, until the C round of adjustment is completed to obtain the first estimated intrinsic parameters and the first estimated distortion coefficients of the N acquisition devices.

18. The apparatus as claimed in claim 17, characterized in that, The processing unit is specifically used for: Based on the three-dimensional coordinates of multiple calibration points on the target calibration object at the k-th moving position, and the pose of the target calibration object at the k-th moving position corresponding to the i-th acquisition device, determine the coordinates of the multiple calibration points projected onto the camera coordinate system of the i-th acquisition device; The distorted coordinates of the plurality of calibration points projected onto the camera coordinate system are determined based on the coordinates of the i-th acquisition device in the camera coordinate system and the distortion coefficient of the i-th acquisition device initially set. The pixel coordinates of the plurality of calibration points projected onto the image coordinate system of the i-th acquisition device are estimated based on the distorted coordinates and the second intrinsic parameter estimate of the i-th acquisition device.

19. The apparatus according to any one of claims 14-18, characterized in that, The processing unit is specifically used for: Based on the first intrinsic parameter estimates of at least two acquisition devices in each acquisition device group, the matching feature point set corresponding to each acquisition device group, and the three-dimensional coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, the second relative pose of the other acquisition devices in the multiple acquisition devices, excluding the reference acquisition device, relative to the reference acquisition device is obtained. The reference acquisition device is any one of the plurality of acquisition devices; A scale factor is determined, which is the ratio between a first distance and a second distance. The first distance is the distance between two calibration points on the target calibration object, and the second distance is the distance between the two calibration points in the same image coordinate system. The two calibration points are located on the same calibration surface on the target calibration object. The first extrinsic parameter estimate of each acquisition device is obtained based on the second relative pose of each acquisition device and the scale factor.

20. The apparatus as claimed in claim 19, characterized in that, The processing unit is specifically used for: The essential matrix between the first acquisition device and the reference acquisition device is determined based on the matching feature point set corresponding to the first acquisition device group. The first acquisition device and the reference acquisition device belong to the first acquisition device group, and the first acquisition device group is one of the plurality of acquisition device groups. Based on the singular value decomposition results of the essential matrix, the second relative pose between the first acquisition device and the reference acquisition device is determined.

21. The apparatus as claimed in claim 19, characterized in that, The processing unit is specifically used for: After determining the second relative poses of the acquisition devices in the g-th acquisition device group relative to the reference acquisition device, the three-dimensional coordinates of the multiple calibration points in the local coordinate system are determined based on the second relative poses of each acquisition device when the target calibration object moves to M2 moving positions. The local coordinate system is the camera coordinate system of the reference acquisition device. Any of the M2 moving positions is located at least within the common viewing area of ​​two acquisition devices in the g-th acquisition device group. Based on the three-dimensional coordinates of the multiple calibration points at the M2 moving positions in the local coordinate system, the second relative pose of the acquisition devices included in the g-th acquisition device group, and the first intrinsic parameter estimate, the pixel coordinates of the multiple calibration points at the M2 moving positions are estimated to be projected onto the image coordinate system of the acquisition devices included in the g-th acquisition device group. The error between the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of the acquisition devices included in the g-th acquisition device group and the pixel coordinates of the plurality of calibration points extracted from the image frames acquired by the acquisition devices included in the g-th acquisition device group is obtained. Adjust the second relative pose and the first intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group according to the error, and obtain the relative pose and intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group after the current round of adjustment; Among them, the intrinsic parameter estimates and relative poses of the acquisition devices in the g-th acquisition device group after the current round of adjustment are used as the basis for the next round of adjustment, until the D-th round of adjustment is completed to obtain the third relative pose and third intrinsic parameter estimates of the acquisition devices included in the g-th acquisition device group; The first extrinsic parameter estimate is obtained by adding the scale factor to the third relative pose of the acquisition devices included in the g-th acquisition device group.

22. The apparatus as claimed in claim 19, characterized in that, The processing unit is specifically used for: After determining the second relative poses of the acquisition devices in the g-th acquisition device group relative to the reference acquisition device, the three-dimensional coordinates of the multiple calibration points in the local coordinate system are determined based on the second relative poses of each acquisition device when the target calibration object moves to M2 moving positions. The local coordinate system is the camera coordinate system of the reference acquisition device. Any of the M2 moving positions is located at least within the common viewing area of ​​two acquisition devices in the g-th acquisition device group. Based on the three-dimensional coordinates of the multiple calibration points at the M2 moving positions in the local coordinate system, the second relative pose of the acquisition devices included in the g-th acquisition device group, the first intrinsic parameter estimate, and the first distortion coefficient estimate, the pixel coordinates of the multiple calibration points at the M2 moving positions projected onto the image coordinate system of the acquisition devices included in the g-th acquisition device group are estimated. The error between the estimated pixel coordinates of the plurality of calibration points in the image coordinate system of the acquisition devices included in the g-th acquisition device group and the pixel coordinates of the plurality of calibration points extracted from the image frames acquired by the acquisition devices included in the g-th acquisition device group is obtained. Adjust the second relative pose and the first intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group according to the error, and obtain the relative pose and intrinsic parameter estimate of the acquisition devices included in the g-th acquisition device group after the current round of adjustment; Among them, the intrinsic parameter estimates and relative poses of the acquisition devices in the g-th acquisition device group after the current round of adjustment are used as the basis for the next round of adjustment, until the D-th round of adjustment is completed to obtain the third relative pose and third intrinsic parameter estimates of the acquisition devices included in the g-th acquisition device group; The first extrinsic parameter estimate is obtained by adding the scale factor to the third relative pose of the acquisition devices included in the g-th acquisition device group.

23. The apparatus according to any one of claims 14-18, characterized in that, Each of the plurality of data acquisition device groups includes two data acquisition devices, and the processing unit is specifically used for: The relative poses of multiple pairs of movement positions among the M movement positions of the target calibration object are determined. A first pair of movement positions includes a first movement position and a second movement position. The first and second movement positions are two of the M movement positions of the target calibration object, and the first and second movement positions are located within the common viewing area of ​​at least one group of acquisition devices. The relative poses of the first pair of movement positions are determined based on the matching feature point set corresponding to at least one group of acquisition devices, the three-dimensional coordinates of multiple calibration points of the target calibration object in the calibration object coordinate system, and the first intrinsic parameter estimates of the two acquisition devices included in the at least one group of acquisition devices. The poses of M movement positions in the world coordinate system are determined based on the coordinates of the base movement position in the world coordinate system and the relative poses of multiple movement position pairs, wherein the base movement position is one of the M movement positions; Based on the coordinates of M moving positions in the world coordinate system, the coordinates of multiple calibration points included in the target calibration object in the calibration object coordinate system, and the pixel coordinates of the calibration points on the target calibration object in the image frames acquired by each acquisition device, the camera parameters of each acquisition device are globally optimized. The camera parameters include an intrinsic parameter matrix and an extrinsic parameter matrix, or the camera parameters include an intrinsic parameter matrix, an extrinsic parameter matrix, and distortion coefficients. In the global optimization process, the camera parameters of each acquisition device and the pose of the calibration surface where the multiple calibration points are located in the calibration object coordinate system are taken as the quantities to be optimized; the first intrinsic parameter estimate of each acquisition device in the global optimization is taken as the initial value of the intrinsic parameter matrix of each acquisition device.

24. The apparatus as claimed in claim 23, characterized in that, The relative poses of the first moving position pair satisfy the following conditions: ; in, This indicates the relative pose between the first moving position and the second moving position. The at least one group of acquisition devices includes a first group of acquisition devices, which in turn includes a first acquisition device and a second acquisition device. It represents the pose from the first moving position to the second moving position, determined based on the pixel coordinates of the calibration point in the image frame acquired by the first acquisition device when the target calibration object moves to the first moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the second moving position; This indicates the pose from the second moving position to the first moving position, determined based on the pixel coordinates of the calibration point in the image frame acquired by the second acquisition device when the target calibration object moves to the second moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the first moving position.

25. The apparatus as claimed in claim 24, characterized in that, At least one group of acquisition devices consists of L devices, and the first group of acquisition devices satisfies: ; in, Represents the identity matrix. Indicates based on the first The pose from the first moving position to the second moving position is determined by the pixel coordinates of the calibration point in the image frame acquired by the first acquisition device in the group of acquisition devices when the target calibration object moves to the first moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the second moving position; Indicates based on the first The pose from the second moving position to the first moving position is determined by the pixel coordinates of the calibration point in the image frame acquired by the second acquisition device in the acquisition device group when the target calibration object moves to the second moving position and the pixel coordinates of the calibration point in the image frame acquired when the target calibration object moves to the first moving position; Indicates the first The first acquisition device in the acquisition device group. 2 indicates the first The second acquisition device in the acquisition device group.

26. The apparatus as claimed in claim 25, characterized in that, The process of determining the poses of M mobile positions in the world coordinate system based on the coordinates of the base mobile position in the world coordinate system and the relative poses of multiple mobile position pairs includes: Determine the confidence weight for each of the multiple movement position pairs, where the confidence weight between the first movement position and the second movement position satisfies the following condition: ; This represents the confidence weight between the first and second movement positions. The shortest path from the third move position to the base move position is determined based on the credibility weight of each move position pair. The shortest path is the path with the smallest confidence weight among all paths from the third moving position to the base moving position; the confidence weight of any path is the sum of the confidence weights of the moving position pairs traversed by the path. The pose of the third moving position is determined based on the relative poses of the moving position pairs traversed by the shortest path.

27. A calibration device, characterized in that, Including processor and memory; The memory is used to store computer programs; The processor is used to execute a computer program stored in the memory to implement the method as described in any one of claims 1-13.

28. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when run on a processor, causes the processor to perform the method as described in any one of claims 1-13.