A spatial calibration method and device, electronic equipment and storage medium

By allocating and jointly optimizing lenses for virtual shooting scenes with multiple lenses and cameras, the problems of long processing time or insufficient accuracy in existing technologies are solved, achieving efficient spatial calibration and meeting the high-precision requirements of virtual shooting.

CN122636741APending Publication Date: 2026-08-25BEIJING YOUKU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610589150.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-29
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing spatial calibration methods are time-consuming or lack sufficient calibration accuracy in virtual shooting scenarios with multiple lenses and cameras, making it difficult to meet the requirements of high-precision virtual shooting.

Method used

By allocating multiple lenses, the lens set of each camera is determined, and the spatial calibration data of the main lens and sub-lens of each camera is obtained. Joint optimization is then performed to obtain spatial calibration results, including screen offset information, node offset information, and calibration values ​​of lens intrinsic parameters.

Benefits of technology

While ensuring calibration accuracy, the amount of data and time required for collecting spatial calibration data have been significantly reduced, improving spatial calibration efficiency and meeting the high-precision positioning requirements in the virtual shooting process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122636741A_ABST
    Figure CN122636741A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a spatial calibration method, device, electronic equipment and storage medium, which is applied to a scene of virtual shooting by using multiple cameras, each of which can be adapted to multiple lenses, and each of which is fixedly installed with a positioning device; the method comprises: allocating multiple lenses, determining a lens set of each camera, obtaining a first preset number of spatial calibration data corresponding to a main lens of each camera, and a second preset number of spatial calibration data corresponding to a sub-lens of each camera; wherein the second preset number is less than the first preset number; based on the first preset number of spatial calibration data corresponding to the main lens of each camera and the second preset number of spatial calibration data corresponding to the sub-lens of each camera, joint optimization is performed to obtain a spatial calibration result. Through the present disclosure, for the virtual shooting scene of multiple-lens and multiple-camera, the spatial calibration efficiency is greatly improved under the premise of ensuring the spatial calibration accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of virtual photography technology, and in particular to a spatial calibration method, apparatus, electronic device, and storage medium. Background Technology

[0002] Virtual production refers to projecting scene images rendered by a virtual scene rendering engine onto a screen, then having actors perform using the screen as a background while the camera simultaneously captures both the actors and the screen. This places real actors within a virtual scene, achieving the effect of shooting outdoor scenes or science fiction backgrounds in a studio.

[0003] During virtual shooting, to ensure that the footage captured by the camera is strictly aligned with the virtual scene in the virtual scene rendering engine within the same spatial coordinate system (e.g., correct occlusion relationship between the virtual set and the real foreground, consistent perspective, and parallax matching), it is necessary to obtain the camera's pose in real time. This means the position and orientation of the camera lens within the virtual shooting area (usually including the 3D position of the camera's equivalent optical center and the rotational attitude of the camera lens). Inaccurate camera pose will directly manifest as problems such as virtual element drift, jitter, incorrect perspective, and inaccurate landing, especially noticeable in close-ups, telephoto shots, and fast-moving shots. Since common film and television cameras themselves do not possess reliable spatial positioning capabilities and cannot directly output pose like industrial cameras or mobile devices, optical tracking and other positioning systems are typically used to measure pose.

[0004] Because the pose measured by the positioning system is often the pose of the external positioning device mounted on the camera, rather than the camera's pose, there is usually a fixed but unknown (or imprecise) spatial offset between the positioning device's pose and the camera's pose. If this offset is not calibrated, even if the external positioning device's pose is very accurate, the final camera pose will still have systematic errors, resulting in a persistent misalignment between the virtual and real worlds. Therefore, spatial calibration (also known as spatial calibration or extrinsic parameter calibration) is required. However, existing spatial calibration methods are time-consuming or lack sufficient calibration accuracy in multi-lens, multi-camera virtual shooting scenarios. Summary of the Invention

[0005] In view of this, this disclosure provides a space calibration method, apparatus, electronic device, storage medium, and computer program product.

[0006] According to one aspect of this disclosure, a spatial calibration method is provided, applied to a scene where multiple cameras are used for virtual shooting, wherein each camera can be adapted to multiple lenses, and each camera is fixedly equipped with a positioning device; the method includes:

[0007] The multiple lenses are assigned to determine the lens set for each camera, wherein the lens sets of different cameras contain different lenses;

[0008] Acquire a first preset number of spatial calibration data corresponding to the main lens of each camera, and a second preset number of spatial calibration data corresponding to the sub-lens of each camera; wherein the second preset number is less than the first preset number; the sub-lens of any camera is the lens in the lens set of that camera excluding the main lens of that camera;

[0009] Based on a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera, joint optimization is performed to obtain spatial calibration results. The spatial calibration results include: calibration values ​​of screen offset information for the virtual shooting location, calibration values ​​of node offset information for each camera, and calibration values ​​of intrinsic parameters for each of the multiple lenses. The node offset information represents the offset between the equivalent optical center of the positioning device and the camera; the screen offset information represents the offset between the coordinate system corresponding to the screen and the coordinate system corresponding to the positioning device in the virtual shooting location.

[0010] In one possible implementation, the joint optimization based on a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera to obtain spatial calibration results includes:

[0011] Using a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera, a first joint optimization is performed to obtain intermediate spatial calibration results; the intermediate spatial calibration results include: optimized values ​​of the intrinsic parameters of each lens, optimized values ​​of the node offset information of each camera, and optimized values ​​of the screen offset information corresponding to each camera.

[0012] Using a first preset number of spatial calibration data corresponding to the main lens of each camera, a second preset number of spatial calibration data corresponding to the sub-lenses of each camera, and the intermediate spatial calibration results, a second joint optimization is performed to obtain the spatial calibration result.

[0013] In one possible implementation, the first joint optimization, using a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera, to obtain intermediate spatial calibration results, includes:

[0014] For any camera, the initial values ​​of the intrinsic parameters of the main lens, the initial values ​​of the node offset information of the camera, and the initial values ​​of the screen offset information of the camera are obtained by using the first preset number of spatial calibration data corresponding to the main lens of the camera.

[0015] The initial values ​​of the intrinsic parameters of the camera's sub-lenses are obtained by using the second preset number of spatial calibration data corresponding to the sub-lenses of the camera.

[0016] Using a first preset number of spatial calibration data corresponding to the main lens of the camera and a second preset number of spatial calibration data corresponding to the sub-lenses of the camera, with the goal of minimizing reprojection error, the intrinsic parameters of the main lens of the camera, the intrinsic parameters of the sub-lenses of the camera, the node offset information of the camera, and the screen offset information corresponding to the camera are optimized to obtain optimized values ​​for the intrinsic parameters of the main lens of the camera, the intrinsic parameters of the sub-lenses of the camera, the node offset information of the camera, and the screen offset information corresponding to the camera. The initial values ​​in the optimization process include the initial values ​​for the intrinsic parameters of the main lens of the camera, the intrinsic parameters of the sub-lenses of the camera, the node offset information of the camera, and the screen offset information corresponding to the camera.

[0017] In one possible implementation, the second joint optimization, using a first preset number of spatial calibration data corresponding to the main lens of each camera, a second preset number of spatial calibration data corresponding to the sub-lenses of each camera, and the intermediate spatial calibration results, to obtain the spatial calibration result, includes:

[0018] Using a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera, with the goal of minimizing reprojection error, the intrinsic parameters of the main lens of each camera, the intrinsic parameters of the sub-lenses of each camera, the node offset information of each camera, and the screen offset information of the virtual shooting site are optimized to obtain the spatial calibration result; wherein, the initial values ​​in the optimization process include: the optimized value of the intrinsic parameters of the main lens of each camera, the optimized value of the intrinsic parameters of the sub-lenses of each camera, the optimized value of the node offset information of each camera, and the optimized value of the screen offset information corresponding to any camera.

[0019] In one possible implementation, the spatial calibration data includes: images captured by the camera for the calibration screen in a virtual shooting location, and pose data synchronously collected by the positioning device.

[0020] In one possible implementation, each camera has one main lens.

[0021] In one possible implementation, acquiring a first preset number of spatial calibration data corresponding to the main lens of each camera includes:

[0022] In the virtual shooting location, determine the first preset number of location points and the shooting angle corresponding to each location point to cover the virtual shooting location;

[0023] For any given camera, images captured by the camera's main lens at each position point with corresponding shooting angles for the calibration screen are acquired, along with pose data synchronously collected by the positioning device installed on the camera. The calibration screen is then displayed on a screen in a virtual shooting area.

[0024] In one possible implementation, the first preset quantity is twice the second preset quantity.

[0025] According to another aspect of this disclosure, a spatial calibration device is provided for use in scenarios where multiple cameras are used for virtual photography, wherein each camera can be adapted to multiple lenses, and each camera is fixedly equipped with a positioning device; the device includes:

[0026] The allocation module is used to allocate the multiple lenses and determine the lens set for each camera, wherein the lens sets of different cameras contain different lenses;

[0027] The data acquisition module is used to acquire a first preset number of spatial calibration data corresponding to the main lens of each camera, and a second preset number of spatial calibration data corresponding to the sub-lens of each camera; wherein, the second preset number is less than the first preset number; the sub-lens of any camera is the lens in the lens set of that camera excluding the main lens of that camera;

[0028] A spatial calibration module is used to perform joint optimization based on a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera to obtain spatial calibration results. The spatial calibration results include: calibration values ​​of screen offset information of the virtual shooting location, calibration values ​​of node offset information of each camera, and calibration values ​​of intrinsic parameters of each of the multiple lenses. The node offset information represents the offset between the equivalent optical center of the positioning device and the camera; the screen offset information represents the offset between the coordinate system corresponding to the screen and the coordinate system corresponding to the positioning device in the virtual shooting location.

[0029] According to another aspect of this disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above-described method.

[0030] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the above-described method.

[0031] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0032] According to various aspects of this disclosure, it is applied to scenarios where multiple cameras are used for virtual shooting, wherein each camera can be adapted to multiple lenses, and each camera is fixedly equipped with a positioning device; during the spatial calibration process, the multiple lenses are allocated to determine the lens set of each camera, wherein the lens sets of different cameras contain different lenses; a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera are obtained; wherein the second preset number is less than the first preset number; a sub-lens of any camera is a lens in the lens set of that camera excluding the main lens of that camera; based on the first preset number of spatial calibration data corresponding to the main lens of each camera and the second preset number of spatial calibration data corresponding to the sub-lenses of each camera, joint optimization is performed to obtain a spatial calibration result; wherein the spatial calibration result includes: calibration values ​​of screen offset information of the virtual shooting site, calibration values ​​of node offset information of each camera, and calibration values ​​of intrinsic parameters of each lens among the multiple lenses; wherein the node offset information represents the offset between the equivalent optical center of the positioning device and the camera; the screen offset information represents the offset between the coordinate system corresponding to the screen in the virtual shooting site and the coordinate system corresponding to the positioning device. In this way, for virtual shooting scenarios with multiple lenses and cameras, the amount of spatial calibration data collected and the collection time are significantly reduced while ensuring spatial calibration accuracy, thereby greatly improving spatial calibration efficiency. Specifically, since each lens is assigned to only one camera, for any given lens, only the spatial calibration data corresponding to that lens when installed on its assigned camera needs to be acquired, effectively reducing the amount of spatial calibration data collected during the spatial calibration process and significantly shortening the data collection time. Furthermore, since the second preset number is less than the first preset number, meaning the amount of spatial calibration data for the second preset number of sub-lenses is less than the amount of spatial calibration data for the first preset number of main lenses, the amount of spatial calibration data collected for sub-lenses during the spatial calibration process is further reduced, shortening the data collection time. Moreover, through joint optimization, even with less spatial calibration data, the spatial calibration results obtained through the mutual constraints between multiple lenses and cameras couple the equivalent optical center differences between different lenses, ensuring that the calibration values ​​of the node offset information of each camera and the screen offset information of the virtual shooting site are applicable to each lens, thus guaranteeing spatial calibration accuracy.

[0033] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0034] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.

[0035] Figure 1 A schematic diagram of a virtual shooting scene according to an embodiment of the present disclosure is shown.

[0036] Figure 2 A flowchart of a space calibration method according to an embodiment of the present disclosure is shown.

[0037] Figure 3 A flowchart illustrating joint optimization according to an embodiment of the present disclosure is shown.

[0038] Figure 4 A schematic flowchart of a space calibration method according to an embodiment of the present disclosure is shown.

[0039] Figure 5 A structural diagram of a space calibration device according to an embodiment of the present disclosure is shown.

[0040] Figure 6 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. Detailed Implementation

[0041] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0042] As used herein, the terms “comprising,” “including,” “having,” or variations thereof are open-ended and include one or more of the stated features, integrals, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, integrals, elements, steps, components, functions, or groups thereof.

[0043] When an element is referred to as “connected,” “coupled,” “responding,” or a variation thereof relative to another element, it may be directly connected, coupled, or responding to another element, or there may be an intermediate element present.

[0044] Although the terms first, second, third, etc., may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Therefore, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.

[0045] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0046] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0047] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions.

[0048] The following section provides an example of the application scenarios to which this disclosure can be applied.

[0049] Figure 1 The diagram illustrates a scene using multiple cameras for virtual shooting according to an embodiment of the present disclosure, such as... Figure 1 As shown, a screen and multiple cameras (i.e., camera 1, camera 2... camera N in the figure, where N is a positive integer) are arranged in the virtual shooting location.

[0050] The system includes a screen to display scene images rendered by a virtual scene rendering engine, a camera to capture the images displayed on the screen, and a real scene that can be set up in front of the screen. Actors can stand in appropriate positions in front of the screen and perform against its backdrop. The camera can simultaneously capture both the actual scene in which the actors are located and the images displayed on the screen, thus completing the virtual filming. For example, the type, number, and shape of the screen can be configured according to requirements. For instance, the screen type can be an LED (Light-Emitting Diode) screen or a screen made of other materials. The screen shape can be a flat screen, a curved screen, a tri-fold screen, or a multi-faceted, three-dimensional irregularly shaped screen, etc. The number of screens can be one or more. Figure 1(Only one is shown in the illustration), and this disclosure does not limit the scope of the embodiments.

[0051] The number of cameras is multiple ( Figure 1 The diagram shows N cameras, the specific number of which can be configured according to requirements to meet the needs of multi-view shooting. Each camera can be equipped with multiple lenses, such as infrared lenses, fixed-focus lenses, wide-angle lenses, zoom lenses, etc., to meet different shooting needs. Each camera is fixedly equipped with a positioning device (also known as a tracking device or external positioning device) to capture the camera's real-time pose information (such as position and attitude). The positioning device can be fixedly mounted on the camera body, top handle, hot shoe, or gimbal, etc. For example, the positioning device on the camera can be a tracker in an optical positioning system such as OptiTrack or RedSpy.

[0052] In virtual photography, to achieve alignment between virtual and real spatial positions, it is necessary to acquire the camera's accurate pose in real time (including the 3D position of the camera's equivalent optical center and the camera lens's orientation). Since the camera's pose is indirectly obtained through a positioning device mounted on the camera; considering that the positioning device is fixed to the camera, its installation position may deviate from the camera's equivalent optical center, and its orientation may also deviate from the camera lens's orientation—that is, there may be an offset between the camera's pose and the pose of the positioning device mounted on the camera—spatial calibration is required beforehand to accurately calculate the camera's pose based on the pose data provided by the positioning device during actual virtual photography.

[0053] During the actual virtual shooting phase, the production crew often frequently changes lens focal lengths and specifications (primary focus, zoom, different brand mounts and adapters), and may use multiple cameras simultaneously. Each change of lens or camera body may alter the spatial offset between the camera's equivalent optical center and the positioning device, thus requiring recalibration.

[0054] In one related technology, a spatial calibration method involves performing a complete spatial calibration for each camera and each lens. A single high-precision spatial calibration typically takes around 20 minutes or even longer. However, in actual virtual shooting, dozens of lenses with different focal lengths and specifications, and several different cameras are often used. For example, assuming 10 lenses and 2 cameras are used, a total of 20 complete spatial calibrations would be required, taking approximately 6.7 hours. This method of performing a complete high-precision spatial calibration for each lens and each camera is time-consuming, severely impacting valuable shooting time on set and increasing operational costs and the probability of human error. In another spatial calibration method, each camera is calibrated only once. After changing lenses, only the focal length of the virtual camera is changed, while other parameters remain unchanged. However, due to differences in distortion and the position of the equivalent optical center for each lens, adjusting the focal length only barely aligns the lens's field of view, resulting in low calibration accuracy that cannot meet the high-precision shooting positioning requirements in virtual shooting scenarios.

[0055] To address the aforementioned technical issues, this disclosure provides a spatial calibration method (detailed description below). For virtual shooting scenarios with multiple lenses and cameras, this method significantly reduces the amount of spatial calibration data collected and the collection time while ensuring spatial calibration accuracy, thereby greatly improving spatial calibration efficiency and meeting the high-precision shooting and positioning requirements during virtual shooting.

[0056] For example, the spatial calibration method provided in this disclosure can be executed by an electronic device such as a terminal device or a server, or a part of an electronic device (such as a processor). The terminal device can be a desktop terminal or a mobile terminal, such as a laptop, tablet, desktop computer, smartphone, smart speaker, smartwatch, smart TV, in-vehicle terminal, or other types of electronic devices. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.

[0057] The spatial calibration method provided in the embodiments of this disclosure will be described in detail below.

[0058] Figure 2 The flowchart illustrates a spatial calibration method according to an embodiment of the present disclosure. This method is applied to a scene where multiple cameras are used for virtual photography, wherein each camera can be adapted to multiple lenses, and each camera is fixedly equipped with a positioning device; for example, it can be applied to the above-mentioned... Figure 1 The scene shown. (As shown) Figure 2 As shown, the method may include the following steps:

[0059] Step 201: Allocate the multiple lenses to determine the lens set for each camera, wherein the lens sets of different cameras contain different lenses.

[0060] For example, a scene using multiple cameras for virtual shooting has N cameras and M lenses, where M and N are both positive integers, and M is greater than N. Each camera can be adapted to these M lenses, meaning that each lens can be directly mounted or mounted on each camera via an adapter ring.

[0061] Each camera's lens set contains at least one lens, and at least some cameras' lens sets contain at least two lenses. For example, multiple lenses can be evenly distributed to determine the lens set for each camera; wherein, if the number of lenses is divisible by the number of cameras, the evenly distributed lens sets of different cameras contain the same number of lenses; if the number of lenses is not divisible by the number of cameras, the evenly distributed lens sets of different cameras contain the same number of lenses or differ by one.

[0062] For example, after the allocation of multiple lenses is completed, each camera can be used to synchronously collect data based on the corresponding lens set, thereby further saving data collection time and improving the efficiency of spatial calibration.

[0063] Step 202: Obtain a first preset number of spatial calibration data corresponding to the main lens of each camera, and a second preset number of spatial calibration data corresponding to the sub-lens of each camera; wherein, the second preset number is less than the first preset number; the sub-lens of any camera is the lens in the lens set of that camera other than the main lens of that camera.

[0064] The values ​​of the first and second preset quantities can be set according to requirements and are not limited thereto. In this step, for each camera, the first preset quantity of spatial calibration data corresponding to each main lens of that camera and the second preset quantity of spatial calibration data corresponding to each sub-lens are acquired. Since each lens is only assigned to one camera, for any lens, it is only necessary to acquire the spatial calibration data corresponding to that lens when it is installed on the assigned camera, thereby effectively reducing the amount of spatial calibration data that needs to be collected during the spatial calibration process and significantly shortening the data acquisition time. Furthermore, since the second preset quantity is less than the first preset quantity, that is, the amount of spatial calibration data corresponding to the second preset quantity of sub-lenses is less than the amount of spatial calibration data corresponding to the first preset quantity of main lenses, the amount of spatial calibration data collected for sub-lenses during the spatial calibration process is further reduced, shortening the data acquisition time.

[0065] For example, the first preset quantity can be the quantity required to meet the high-precision spatial calibration needs, i.e., the full spatial calibration data. In this way, during the spatial calibration process, only the spatial calibration data corresponding to each lens installed on its assigned camera needs to be acquired. Furthermore, only each main lens of each camera needs to collect full spatial calibration data, while each sub-lens of each camera does not require full spatial calibration data collection. Compared to the existing technology where each lens of each camera needs to collect full spatial calibration data, this significantly reduces the amount of spatial calibration data and the acquisition time. For example, taking a total of N cameras and M lenses as an example, the existing spatial calibration method requires M... Spatial calibration data is collected under N combinations, but in this embodiment of the disclosure, only M combinations of lens and camera are required to collect spatial calibration data, and only N main lenses need to collect full spatial calibration data, while the remaining MN sub-lenses do not need to collect full spatial calibration data.

[0066] In one possible implementation, each camera has only one main lens. For example, still assuming a total of N cameras and M lenses, one lens can be randomly selected from the M lenses for each of the N cameras as its main lens; then the remaining MN lenses can be evenly distributed among the N cameras as their sub-lenses. Since the second preset number is less than the first preset number, each camera is configured with only one main lens, thereby further reducing the amount of spatial calibration data that needs to be collected and improving spatial calibration efficiency.

[0067] For example, for any camera, a lens can be randomly selected from the camera's lens set as the camera's main lens, and the remaining lenses can be used as the camera's sub-lenses.

[0068] In one possible implementation, the first preset quantity is twice the second preset quantity. For example, the first preset quantity is 20 sets, and the second preset quantity is 10 sets. In this way, the amount of spatial calibration data to be collected for each sub-lens is reduced by half compared to the amount of spatial calibration data for each main lens, thereby significantly reducing the amount of spatial calibration data collected and the collection time during the spatial calibration process, and improving the efficiency of spatial calibration.

[0069] In one possible implementation, the spatial calibration data includes: images captured by the camera for the calibration screen in a virtual shooting location, and pose data synchronously collected by the positioning device.

[0070] For example, the screen in the virtual shooting location is the screen used to display the rendered scene in virtual shooting.

[0071] For example, the calibration screen includes at least one feature point. A feature point refers to a point with significant characteristics in the calibration screen, such as edges, corners, dots, or spots; for example, if the calibration screen is a checkerboard pattern, the corners of the checkerboard pattern are feature points; as another example, if the calibration screen is a dot matrix pattern, the dots in the dot matrix pattern are feature points; and as yet another example, the calibration screen may also include a pre-defined location identifier with known positional information, such as an Aruco code, where the edges of the Aruco code are feature points. The size, type, and number of feature points of the calibration screen can be set according to requirements and are not limited thereto.

[0072] For example, for any camera, during the process of acquiring spatial calibration data corresponding to any lens of the camera, a calibration screen can be displayed on the screen in the virtual shooting location, and after the lens is installed on the camera, a shot is taken of the calibration screen to obtain the captured image; the positioning device installed on the camera can synchronously acquire and obtain real-time pose data, that is, the real-time pose at the time of shooting; in this way, a set of spatial calibration data (including the captured image and pose data) can be obtained through one shot.

[0073] As an example, taking the number of main lenses of each camera as one; the step of obtaining the first preset number of spatial calibration data corresponding to the main lens of each camera includes: determining the first preset number of position points and the shooting angle corresponding to each position point in the virtual shooting area to cover the virtual shooting area; for any camera, obtaining the image taken by the main lens of the camera at each position point with the corresponding shooting angle for the calibration screen, and the pose data synchronously collected by the positioning device installed on the camera, and the calibration screen is displayed on the screen in the virtual shooting area.

[0074] For example, a first preset number of location points and the corresponding shooting angle for each location point can be determined based on the layout of the virtual shooting location and the virtual shooting task objective, so as to cover the virtual shooting location. That is, when shooting at these location points with the corresponding shooting angle, as many scene features as possible in the virtual shooting location can be collected. For example, these location points can be evenly distributed throughout the entire virtual shooting location, and for each location point, the shooting angle of that location point can be set according to the possible shooting angle of the camera during the actual shooting process to cover sufficient pose changes. During the data acquisition process, for any camera, the main lens of the camera takes a picture of the calibration scene at each location point with the corresponding shooting angle to obtain the captured image. At the same time, the positioning device installed on the camera synchronously collects pose data. Among them, the image captured at a location point and the collected position data constitute a set of spatial calibration data.

[0075] In this way, a first preset number of location points covering the virtual shooting area and the shooting angle corresponding to each location point are determined in the virtual shooting area. Then, data is collected using the main lens of each camera and the positioning device installed on the camera to achieve full spatial calibration data collection. The collected full spatial calibration data can meet the requirements of high-precision spatial calibration.

[0076] As another example, taking the case where each camera has multiple sub-lenses, the step of obtaining the second preset number of spatial calibration data corresponding to each camera's sub-lenses includes: for any camera, obtaining images captured by each sub-lens of the camera at a second preset number of positions in the virtual shooting area relative to the calibration screen, and pose data synchronously collected by the positioning device installed on the camera, wherein the calibration screen is displayed on a screen in the virtual shooting area.

[0077] For example, the location points of the second preset data can be arbitrarily set in the virtual shooting scene as needed; for example, the second preset number of location points can be randomly selected from the first preset number of location points, or the second preset number of location points can be randomly determined in the virtual shooting area. During the data acquisition process, for each sub-lens of any camera, the sub-lens of the camera takes a picture of the calibration scene at each location point to obtain the captured image. The shooting angle is not limited, as long as the calibration scene can be captured. At the same time, the positioning device installed on the camera synchronously collects pose data. The image captured at a location point and the collected position data constitute a set of spatial calibration data.

[0078] In this way, a second preset number of location points are determined in the virtual shooting location, and then data is collected using the sub-lens of each camera and the positioning device installed on the camera. The location points and shooting angles are not strictly limited, thereby reducing the data collection requirements and further shortening the data collection time.

[0079] Step 203: Based on the first preset number of spatial calibration data corresponding to the main lens of each camera and the second preset number of spatial calibration data corresponding to the sub-lenses of each camera, perform joint optimization to obtain spatial calibration results; wherein, the spatial calibration results include: calibration values ​​of screen offset information of the virtual shooting site, calibration values ​​of node offset information of each camera, and calibration values ​​of intrinsic parameters of each of the multiple lenses; wherein, the node offset information represents the offset between the equivalent optical center of the positioning device and the camera; the screen offset information represents the offset between the coordinate system corresponding to the screen and the coordinate system corresponding to the positioning device in the virtual shooting site.

[0080] Joint optimization involves merging multiple previously independent optimization problems into a unified optimization problem for simultaneous solution. In this step, joint optimization is performed based on a first preset number of spatial calibration data points corresponding to the main lens of each camera and a second preset number of spatial calibration data points corresponding to the sub-lenses of each camera. When the spatial calibration data is limited, the spatial calibration results obtained through the mutual constraints between multiple lenses and cameras incorporate the equivalent optical center differences of different lenses. This ensures that the calibration values ​​of the node offset information of each camera and the screen offset information of the virtual shooting area are applicable to each lens, thereby guaranteeing the accuracy of spatial calibration.

[0081] The specific process of joint optimization is described below.

[0082] For example, the intrinsic parameters of each lens may include: field of view (FOV), equivalent optical center, focal length, distortion parameters (such as radial distortion, tangential distortion), etc.

[0083] For any camera, the node offset information of that camera represents the offset between the positioning device mounted on the camera and the equivalent optical center of the camera, that is, the spatial offset between the pose of the positioning device and the pose of the camera. Exemplarily, this can be represented by an offset matrix, which includes a rotation transformation matrix and a translation transformation matrix. The rotation transformation matrix represents the deviation between the pose of the positioning device and the pose of the camera lens, and the translation transformation matrix represents the deviation between the position of the positioning device and the position of the equivalent optical center of the camera. For a real lens composed of multiple lenses, in geometric optics, an "equivalent single point" represents the perspective center or ray convergence relationship of its imaging. The equivalent optical center of the camera represents the equivalent single point of the imaging perspective center or ray convergence relationship of the lens mounted on the camera. Considering that the equivalent optical center differs when different lenses are mounted on different cameras, in this embodiment of the disclosure, through the above joint optimization, the calibration value of the node offset information of each camera is coupled with the difference in the equivalent optical center of different lenses. Therefore, for any camera, the calibration value of the node offset information of that camera is applicable when different lenses are mounted.

[0084] Since the pose data collected by the positioning device is the pose of the positioning device itself in the corresponding coordinate system; and positioning devices installed on different cameras in the same virtual shooting location belong to the same positioning system (such as an optical positioning system), that is, the corresponding coordinate systems are the same. Therefore, the final spatial calibration result includes a calibration value of screen offset information; that is, it represents the offset between the screen's corresponding coordinate system and the positioning device's corresponding coordinate system in the virtual shooting location. For example, the screen offset information can be represented by an offset matrix of the positioning device's corresponding coordinate system relative to the screen's corresponding coordinate system in the virtual shooting location; where the origin of the positioning device's corresponding coordinate system can be any point in space, and the direction of the coordinate axes can be configured as needed. For example, the center of the virtual shooting location can be used as the origin to determine the positioning device's corresponding coordinate system; the screen's corresponding coordinate system is the three-dimensional coordinate system set when performing screen modeling, and the definition method of the screen's corresponding coordinate system can be set according to the actual situation. For example, the center of the screen can be used as the origin, or a point at the lower left corner or the lower right corner of the screen can be used as the origin, etc.

[0085] In this embodiment, the method is applied to a scenario where multiple cameras are used for virtual shooting. Each camera can be equipped with multiple lenses, and each camera is fixedly equipped with a positioning device. During spatial calibration, the multiple lenses are allocated to determine the lens set of each camera, wherein the lens sets of different cameras contain different lenses. A first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera are obtained, wherein the second preset number is less than the first preset number. A sub-lens of any camera is a lens in the lens set of that camera other than the main lens of that camera. Based on the first preset number of spatial calibration data corresponding to the main lens of each camera and the second preset number of spatial calibration data corresponding to the sub-lenses of each camera, joint optimization is performed to obtain a spatial calibration result. The spatial calibration result includes: calibration values ​​of screen offset information of the virtual shooting site, calibration values ​​of node offset information of each camera, and calibration values ​​of intrinsic parameters of each lens among the multiple lenses. The node offset information represents the offset between the equivalent optical center of the positioning device and the camera. The screen offset information represents the offset between the coordinate system corresponding to the screen in the virtual shooting site and the coordinate system corresponding to the positioning device.

[0086] For virtual shooting scenarios with multiple lenses and cameras, the amount of spatial calibration data collected and the collection time are significantly reduced while ensuring spatial calibration accuracy, thereby greatly improving spatial calibration efficiency. Specifically, since each lens is assigned to only one camera, for any given lens, only the spatial calibration data corresponding to that lens when installed on its assigned camera needs to be acquired, effectively reducing the amount of spatial calibration data collected during the spatial calibration process and significantly shortening the data collection time. Furthermore, because the second preset number is less than the first preset number, meaning the amount of spatial calibration data for the second preset number of sub-lenses is less than the amount of spatial calibration data for the first preset number of main lenses, the amount of spatial calibration data collected for sub-lenses during the spatial calibration process is further reduced, shortening the data collection time. Moreover, through joint optimization, even with less spatial calibration data, the spatial calibration results obtained through the mutual constraints between multiple lenses and cameras couple the equivalent optical center differences between different lenses, ensuring that the calibration values ​​of the node offset information of each camera and the screen offset information of the virtual shooting site are applicable to each lens, thus guaranteeing spatial calibration accuracy.

[0087] Furthermore, during the actual virtual shooting process using the aforementioned spatial calibration results, when a lens is mounted on a camera, pose data can be collected in real time using the positioning device on the camera. Then, using the calibration values ​​of the camera's node offset information and the screen offset information of the virtual shooting location, the camera's pose in the corresponding screen coordinate system (including the position of the camera's equivalent optical center in the corresponding screen coordinate system and the rotation attitude of the camera lens) can be obtained; that is, the relative pose information between the camera and the screen in the virtual shooting location. The virtual scene rendering engine acquires this relative pose information, wherein the virtual scene rendering engine constructs a three-dimensional virtual... Virtual models such as simulated scenes, screen models, and camera models are used. The virtual scene rendering engine can adjust the rendered content based on the relative pose information and the calibration values ​​of the lens's intrinsic parameters, relying on technologies such as real-time rendering. In this way, the relative pose between the camera model and the screen model can be consistent with the relative pose between the camera (real camera) and the screen (real screen) in the virtual shooting location. This achieves synchronization of the movement and composition of the camera model and the camera in the virtual shooting location, and synchronization of the content displayed on the screen model and the screen in the virtual shooting location, ensuring that the images captured by the camera in the virtual shooting location from the screen in the virtual shooting location have a realistic sense of space.

[0088] Since the calibration values ​​of the node offset information of each camera and the screen offset information of the virtual shooting location are obtained through the above joint optimization, and the node offset information and screen offset information are coupled with the optical center differences between different lenses, this can be applied to the case where any lens is installed on the camera; that is, any lens among multiple lenses installed on any camera, the calibration values ​​of the node offset information of that camera and the screen offset information of the virtual shooting location can be used to determine the relative pose information between the camera and the screen in the virtual shooting location. Thus, during virtual shooting, after each change of lens or camera body, based on the calibration values ​​of the current camera's node offset information and the screen offset information of the virtual shooting location, combined with the pose data collected in real time by the positioning device, the relative pose information between the camera with the current lens and the screen of the virtual shooting scene can be accurately calculated without recalibrating the space.

[0089] The following is about the above. Figure 2 In step 203, the possible implementation methods for jointly optimizing the spatial calibration results based on the first preset number of spatial calibration data corresponding to the main lens of each camera and the second preset number of spatial calibration data corresponding to the sub-lenses of each camera are illustrated by example.

[0090] Figure 3 A flowchart illustrating joint optimization according to an embodiment of the present disclosure is shown, such as... Figure 3 As shown, it includes the following steps:

[0091] Step 301: Using the first preset number of spatial calibration data corresponding to the main lens of each camera and the second preset number of spatial calibration data corresponding to the sub-lens of each camera, perform the first joint optimization to obtain the intermediate spatial calibration results; the intermediate spatial calibration results include: the optimized value of the intrinsic parameters of each lens, the optimized value of the node offset information of each camera, and the optimized value of the screen offset information corresponding to each camera.

[0092] In this step, the first preset number of spatial calibration data corresponding to the main lens of each camera and the second preset number of spatial calibration data corresponding to the sub-lenses of each camera are used to jointly optimize the parameters such as the intrinsic parameters of the main lens of each camera, the intrinsic parameters of the sub-lenses of each camera, the node offset information of each camera, and the screen offset information corresponding to each camera. Through joint optimization, these parameters (i.e., the intermediate results of spatial calibration) are coupled with the deviation of the equivalent optical center when different lenses in the corresponding lens set are installed on each camera, thereby realizing the multi-lens fusion calibration of each camera.

[0093] For example, the first preset number of spatial calibration data corresponding to the main lens of each camera includes multiple sets of spatial calibration data, wherein each set of spatial calibration data includes an image taken by the main lens once for the calibration scene and pose data collected by the positioning device during shooting; the second preset number of spatial calibration data corresponding to the sub-lens of each camera also includes multiple sets of spatial calibration data, wherein each set of spatial calibration data includes an image taken by the sub-lens once for the calibration scene and pose data collected by the positioning device during shooting.

[0094] It should be noted that, for any camera, when the number of sub-lenses is multiple, the second preset number of spatial calibration data corresponding to the sub-lenses represents the second preset number of spatial calibration data corresponding to each sub-lens of the camera. When the number of main lenses is multiple, the first preset number of spatial calibration data corresponding to the main lenses represents the first preset number of spatial calibration data corresponding to each main lens of the camera.

[0095] Step 301, as the first joint optimization, is performed on a per-camera basis. Taking a total of N cameras as an example, step 301 can be performed separately for each of the N cameras to obtain the optimized values ​​of the camera's intrinsic parameters, the optimized values ​​of the camera's node offset information, and the optimized values ​​of the corresponding screen offset information. The first joint optimization for each of the N cameras can be performed in parallel.

[0096] The calibration image includes at least one feature point. For example, for each set of spatial calibration data, feature point detection can be performed on the captured image to determine the two-dimensional coordinates of each extracted feature point in the image coordinate system and the three-dimensional coordinates of each feature point in the corresponding screen coordinate system. The two-dimensional and three-dimensional coordinates of the same feature point constitute a 2D-3D coordinate point pair. Then, based on the 2D-3D coordinate point pairs formed by the two-dimensional and three-dimensional coordinates of each feature point, and the pose data synchronously acquired by the positioning device, a first joint optimization can be performed. Feature point detection on the captured image can be implemented using various algorithms, such as Scale-invariant Feature Transform (SIFT), Directional Fast, and Rotation-invariant BRIEF (ORB) algorithms; no limitation is made to this approach. The image coordinate system is a two-dimensional Cartesian coordinate system. Its origin can be the intersection of the camera's optical axis and the plane containing the image sensor. The X and Y axes of the image coordinate system are parallel to the X and Y axes of the camera coordinate system, respectively. The camera coordinate system can have its origin at the camera's equivalent optical center and its Z-axis at the camera's optical axis. Based on this established image coordinate system, after extracting feature points from the captured image, the coordinates of each feature point in the image coordinate system (i.e., two-dimensional coordinates) can be determined. The coordinates of each feature point in the corresponding screen coordinate system are known, or can be calculated using a pre-defined location identifier. This allows the determination of the coordinates of each feature point in the corresponding screen coordinate system (i.e., three-dimensional coordinates).

[0097] In one possible implementation, the first joint optimization using a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera to obtain intermediate spatial calibration results includes: for any camera, using the first preset number of spatial calibration data corresponding to the main lens of the camera to obtain the initial values ​​of the intrinsic parameters of the main lens of the camera, the initial values ​​of the node offset information of the camera, and the initial values ​​of the screen offset information corresponding to the camera; using the second preset number of spatial calibration data corresponding to the sub-lenses of the camera to obtain the initial values ​​of the intrinsic parameters of the sub-lenses of the camera; using the first preset number of spatial calibration data corresponding to the main lens of the camera to obtain the initial values ​​of the intrinsic parameters of the sub-lenses of the camera; using the first preset number of spatial calibration data corresponding to the main lens of the camera to obtain the initial values ​​of the intrinsic parameters of the sub-lenses of the camera; using the second preset number of spatial calibration data corresponding to the sub-lenses of the camera to obtain the initial values ​​of the intrinsic parameters of the sub-lenses of the camera; using the first preset number of spatial calibration data corresponding to the main lens of the camera to obtain the initial values ​​of the intrinsic parameters of the sub-lenses of the camera; using the second preset number of spatial calibration data corresponding to the sub-lenses of the camera to obtain the initial values ​​of the intrinsic parameters of the sub-lenses of the camera; using the second ... main lens of the camera to obtain the initial values ​​of the intrinsic parameters of the sub-lenses of the camera; using the second preset number of spatial calibration data corresponding to the main lens of the camera to obtain the initial values ​​of the intrinsic parameters Spatial calibration data and a second preset number of spatial calibration data corresponding to the sub-lenses of the camera are used to optimize the intrinsic parameters of the main lens, the intrinsic parameters of the sub-lenses of the camera, the node offset information of the camera, and the screen offset information corresponding to the camera, with the goal of minimizing reprojection error. This yields optimized values ​​for the intrinsic parameters of the main lens, the intrinsic parameters of the sub-lenses of the camera, the node offset information of the camera, and the screen offset information corresponding to the camera. The initial values ​​in the optimization process include the initial values ​​for the intrinsic parameters of the main lens, the intrinsic parameters of the sub-lenses of the camera, the node offset information of the camera, and the screen offset information corresponding to the camera. In this way, for any camera, by using a first preset number of spatial calibration data corresponding to the main lens of the camera and a second preset number of spatial calibration data corresponding to the sub-lenses of the camera, the intrinsic parameters of the main lens of the camera, the intrinsic parameters of the sub-lenses of the camera, the node offset information of the camera, and the screen offset information corresponding to the camera are jointly optimized. Through joint optimization, these parameters are coupled with the deviation of the equivalent optical center when different lenses in the lens set of the camera are installed, thereby realizing the multi-lens fusion calibration of the camera.

[0098] As an example, taking the number of main lenses for each camera as one, for any camera, the initial values ​​of the intrinsic parameters of the main lens, the initial values ​​of the node offset information of the camera, and the initial values ​​of the screen offset information of the camera are obtained by using a first preset number of spatial calibration data corresponding to the main lens of the camera. This can include: using the Zhang Zhengyou calibration method and the hand-eye calibration algorithm to process the first preset number of spatial calibration data corresponding to the main lens of the camera to obtain the initial values ​​of the intrinsic parameters of the main lens, the initial values ​​of the node offset information of the camera, and the initial values ​​of the screen offset information of the camera.

[0099] For example, the initial values ​​of the intrinsic parameters of the camera's main lens can be obtained by processing 2D-3D coordinate point pairs composed of multiple sets of spatial calibration data corresponding to the main lens of the camera using the Zhang Zhengyou calibration method. Then, the initial values ​​of the camera's node offset information and the initial values ​​of the camera's screen offset information can be obtained using a hand-eye calibration algorithm, utilizing the initial values ​​of the camera's main lens intrinsic parameters, the 2D-3D coordinate point pairs composed of multiple sets of spatial calibration data corresponding to the main lens of the camera, and the corresponding pose data. The calibration process using the Zhang Zhengyou calibration method and the hand-eye calibration algorithm can refer to existing technologies. In this way, a complete spatial calibration (i.e., calibration of intrinsic parameters, node offset information, and screen offset information) is performed on the main lens. Since the first preset number of spatial calibration data corresponding to the main lens is full spatial calibration data, the accuracy of the calculated initial values ​​of the main lens's intrinsic parameters, the initial values ​​of the camera's node offset information, and the initial values ​​of the camera's screen offset information is guaranteed.

[0100] Regarding the hand-eye calibration algorithm: In the field of robotics, a camera can be fixed to the robot's robotic arm, and the robotic arm and camera work together to complete a specified task, such as express delivery sorting, parts processing, etc. To ensure the smooth completion of the task, spatial calibration is required, and this calibration process is usually called hand-eye calibration. That is, hand-eye calibration refers to calculating the offset between the camera (eye) and the robotic arm (hand) given the poses of the camera (eye) and the robotic arm (hand). At present, existing hand-eye calibration algorithms can usually be used to solve the hand-eye calibration problem, such as open-source hand-eye calibration algorithms: the Tsai-Lenz algorithm, the algorithm proposed by Horaud, the algorithm proposed by Park, etc. Therefore, the initial value of the camera's node offset information in this embodiment can be abstracted as the above-mentioned hand-eye calibration problem in the field of robotics; specifically, the camera can be understood as the eye, the positioning device as the hand, and the origin of the coordinate system corresponding to the positioning device as the origin of the robotic arm's coordinate system. For example, firstly, based on the 2D-3D coordinate point pairs formed by multiple sets of spatial calibration data corresponding to the main lens of the camera, and the initial values ​​of the intrinsic parameters of the main lens of the camera determined by Zhang Zhengyou's calibration method, the extrinsic parameters of each set of spatial calibration data corresponding to the main lens of the camera are obtained using the SolvePnP algorithm; wherein, the SolvePnP algorithm includes, but is not limited to, P3P, DLT, EPnP, etc.; then, the extrinsic parameters of each set of spatial calibration data corresponding to the main lens of the camera, and the positioning data of each set of spatial calibration data corresponding to the main lens of the camera are input into the hand-eye calibration algorithm to obtain the initial value of the camera's node offset information; furthermore, using the initial value of the camera's node offset information, the initial value of the camera's intrinsic parameters, and the 2D-3D coordinate point pairs formed by multiple sets of spatial calibration data corresponding to the main lens of the camera, the initial value of the screen offset information corresponding to the camera is solved using algorithms such as the least squares method. Specifically, the screen offset information represents the offset between the screen coordinate system (i.e., the coordinate system corresponding to the screen) and the positioning device coordinate system (i.e., the coordinate system corresponding to the positioning device) (which can be represented by an offset matrix, including rotation transformation and translation transformation). Given the 3D coordinates of feature points in the screen coordinate system, they can be transformed to the positioning device coordinate system using screen offset information, then to the camera coordinate system using node offset information, and finally projected to the image coordinate system using lens intrinsic parameters to obtain the predicted 2D coordinates. The initial value of the screen offset information corresponding to the camera can be solved using algorithms such as least squares method with the goal of minimizing the error between the predicted 2D coordinates and the actual observed 2D coordinates.

[0101] As an example, taking a camera with multiple sub-lenses as an example, for any camera, the initial value of the intrinsic parameters of the sub-lenses of that camera is obtained using a second preset number of spatial calibration data corresponding to the sub-lenses. This can include: for any sub-lens of the camera, processing the second preset number of spatial calibration data corresponding to the sub-lens of that camera using the Zhang Zhengyou calibration method to obtain the initial value of the intrinsic parameters of the sub-lens of that camera. For example, for any sub-lens of the camera, the initial value of the intrinsic parameters of the sub-lens of the camera can be obtained by processing a pair of 2D-3D coordinate points based on multiple sets of spatial calibration data corresponding to the sub-lens of that camera using the Zhang Zhengyou calibration method. The calibration process using the Zhang Zhengyou calibration method can refer to existing technologies and will not be elaborated here. In this way, intrinsic parameter calibration is performed on the sub-lenses without performing complete spatial calibration; since the second preset number of spatial calibration data corresponding to the sub-lenses is not the full spatial calibration data, only the intrinsic parameters of the sub-lenses need to be calculated, ensuring the accuracy of the initial values ​​of the calculated intrinsic parameters of the sub-lenses.

[0102] As an example, taking a camera with one main lens and multiple sub-lenses, using a first preset number of spatial calibration data corresponding to the main lens and a second preset number of spatial calibration data corresponding to the sub-lenses, with the goal of minimizing reprojection error, the intrinsic parameters of the main lens, the intrinsic parameters of the sub-lenses, the node offset information of the camera, and the screen offset information corresponding to the camera are optimized to obtain optimized values ​​for the intrinsic parameters of the main lens, the intrinsic parameters of the sub-lenses, the node offset information of the camera, and the screen offset information corresponding to the camera. This can include: constructing multiple 2D-3D coordinate point pairs based on multiple sets of spatial calibration data corresponding to each sub-lens and multiple sets of spatial calibration data corresponding to the main lens; for any 2D-3D coordinate point pair, using the parameters to be optimized corresponding to the 2D-3D coordinate point pair (i.e., the intrinsic parameters of the lens used when capturing the image corresponding to the 2D-3D coordinate point pair, the node offset information of the camera, and the screen offset information corresponding to the camera), The latest value of the information is used to project the 3D coordinate point. For example, the projection transformation process can be as follows: using the latest value of the screen offset information corresponding to the camera, the 3D coordinate value of the feature point in the coordinate system corresponding to the screen is first converted into the 3D coordinate value in the coordinate system corresponding to the positioning device. Then, using the latest value of the node offset information of the camera, the 3D coordinate value in the coordinate system corresponding to the positioning device is converted into the 3D coordinate value in the camera coordinate system. Furthermore, using the latest value of the lens intrinsic parameters used when capturing the image corresponding to the 2D-3D coordinate point pair, the 3D coordinate value corresponding to the camera coordinate system is converted into a 2D coordinate value in the image coordinate system, i.e., a predicted 2D coordinate value. The difference between the predicted 2D coordinate value and the actual 2D coordinate value (i.e., the 2D coordinate point in the 2D-3D coordinate point pair) is taken as the reprojection error. With the goal of minimizing this reprojection error, the values ​​of the parameters to be optimized corresponding to the 2D-3D coordinate point pair are adjusted. For example, gradient descent can be used to adjust the values ​​of the parameters to be optimized. The adjusted values ​​are then used as the latest values ​​for the next iteration. In this way, by constructing multiple sets of spatial calibration data corresponding to the sub-lenses of the camera and multiple sets of spatial calibration data corresponding to the main lens of the camera, the above operation is repeated for different 2D-3D coordinate point pairs in multiple 2D-3D coordinate point pairs, and iteration is performed multiple times until the preset termination condition is met (such as reaching the preset number of iterations, convergence of reprojection error, etc.), thereby obtaining the optimized value of the intrinsic parameters of the main lens of the camera, the optimized value of the intrinsic parameters of each sub-lens of the camera, the optimized value of the node offset information of the camera, and the optimized value of the screen offset information corresponding to the camera.The iterative optimization process is based on the high-precision initial values ​​of the intrinsic parameters of the main lens of the camera, the initial values ​​of the intrinsic parameters of each sub-lens of the camera, the initial values ​​of the node offset information of the camera, and the initial values ​​of the screen offset information corresponding to the camera, thereby ensuring the accuracy of the obtained optimization values.

[0103] For example, if there are N cameras, the above optimization process can be performed in parallel for each camera. For instance, for any one camera, if it corresponds to one main lens and two sub-lenses, the main lens collects 10 sets of spatial calibration data, and the two sub-lenses each collect 5 sets of spatial calibration data, calibrating that there are 5 feature points in the image, then there are a total of 20 sets of spatial calibration data (20...). (5 2D-3D coordinate point pairs). One set of 20 sets of spatial calibration data can be selected in any order, and one 2D-3D coordinate point pair in that set of spatial calibration data can be selected in any order. The above optimization process aimed at minimizing reprojection error is executed. After the iteration of 100 2D-3D coordinate point pairs of 20 sets of spatial calibration data reaches the convergence condition, the optimized values ​​of the intrinsic parameters of the main lens of the camera, the optimized values ​​of the intrinsic parameters of each sub-lens of the camera, the optimized values ​​of the node offset information of the camera, and the optimized values ​​of the screen offset information corresponding to the camera are obtained.

[0104] Step 302: Using the first preset number of spatial calibration data corresponding to the main lens of each camera, the second preset number of spatial calibration data corresponding to the sub-lens of each camera, and the intermediate spatial calibration results, perform a second joint optimization to obtain the spatial calibration results.

[0105] In this step, the first preset number of spatial calibration data corresponding to the main lens of each camera, the second preset number of spatial calibration data corresponding to the sub-lens of each camera, and the intermediate spatial calibration results obtained from the first joint optimization are used to further jointly optimize the parameters such as the intrinsic parameters of the main lens of each camera, the intrinsic parameters of the sub-lens of each camera, the node offset information of each camera, and the screen offset information. Through joint optimization, these parameters (i.e., spatial calibration results) are coupled with the deviation of the equivalent optical center when different lenses are installed in multiple lenses of each camera, thereby realizing the fusion calibration of multiple cameras.

[0106] In one possible implementation, the second joint optimization using a first preset number of spatial calibration data corresponding to the main lens of each camera, a second preset number of spatial calibration data corresponding to the sub-lenses of each camera, and the intermediate spatial calibration results to obtain the spatial calibration result includes: using the first preset number of spatial calibration data corresponding to the main lens of each camera and the second preset number of spatial calibration data corresponding to the sub-lenses of each camera, with the goal of minimizing reprojection error, optimizing the intrinsic parameters of the main lens of each camera, the intrinsic parameters of the sub-lenses of each camera, the node offset information of each camera, and the screen offset information of the virtual shooting site to obtain the spatial calibration result; wherein the initial values ​​in the optimization process include: the optimized values ​​of the intrinsic parameters of the main lens of each camera, the optimized values ​​of the intrinsic parameters of the sub-lenses of each camera, the optimized values ​​of the node offset information of each camera, and the optimized values ​​of the screen offset information corresponding to any camera.

[0107] As an example, taking a camera with one main lens and multiple sub-lenses, multiple 2D-3D coordinate point pairs can be constructed based on multiple sets of spatial calibration data corresponding to each sub-lens and the main lens of each camera. For any 2D-3D coordinate point pair, the latest value of the parameters to be optimized corresponding to the 2D-3D coordinate point pair (i.e., the intrinsic parameters of the lens used when capturing the image corresponding to the 2D-3D coordinate point pair, the node offset information of the camera used when capturing the image corresponding to the 2D-3D coordinate point pair, and the screen offset information corresponding to any camera) is used to perform a projection transformation on the 3D coordinate point to obtain a predicted 2D coordinate value. The difference between the predicted 2D coordinate value and the actual 2D coordinate value (i.e., the 2D coordinate point in the 2D-3D coordinate point pair) is taken as the reprojection error. With the goal of minimizing this reprojection error, the values ​​of the parameters to be optimized corresponding to the 2D-3D coordinate point pair are adjusted. For example, gradient descent can be used to adjust the values ​​of the parameters to be optimized. The adjusted values ​​are then used as the latest values ​​for the next iteration. In this way, multiple 2D-3D coordinate point pairs are constructed based on multiple sets of spatial calibration data corresponding to each sub-lens of each camera and multiple sets of spatial calibration data corresponding to the main lens of each camera. The above operation is repeated for different 2D-3D coordinate point pairs, and multiple iterations are performed until the preset termination conditions are met (such as reaching the preset number of iterations, convergence of reprojection error, etc.). This yields the calibration values ​​of the screen offset information of the virtual shooting site, the calibration values ​​of the node offset information of each camera, and the calibration values ​​of the intrinsic parameters of each lens. The iterative optimization process is based on the above-calculated high-precision optimized values ​​of the intrinsic parameters of each lens, the optimized values ​​of the node offset information of each camera, and the optimized values ​​of the screen offset information corresponding to any camera, thereby ensuring the accuracy of the spatial calibration results.

[0108] For example, if there are N cameras, the above optimization process can be performed on each camera in any order. For simplicity, assume there are two cameras (Camera 1, Camera 2), each with one main lens and two sub-lenses. The main lens collects 10 sets of spatial calibration data, and the two sub-lenses each collect 5 sets of spatial calibration data. The calibration image contains 5 feature points, resulting in a total of 40 sets of spatial calibration data (40...). (5 2D-3D coordinate point pairs). The optimized value of the screen offset information obtained in step 301 for any two cameras (e.g., camera 1) can be used as the parameter to be optimized (the final optimized result of this parameter is the calibration value of the screen offset information of the virtual shooting area). The parameter to be optimized also includes the intrinsic parameters of the lens used when shooting the image corresponding to the 2D-3D coordinate point pair, and the node offset information of the camera used when shooting the image corresponding to the 2D-3D coordinate point pair. Assuming optimization is performed in the order of camera 1 and camera 2, firstly, for camera 1, one set of 20 sets of spatial calibration data corresponding to camera 1 can be selected in any order, and one 2D-3D coordinate point pair from that set can be selected in any order. The above optimization process aimed at minimizing reprojection error is then performed. After completing the optimization process for 100 2D-3D coordinate point pairs of camera 1, the calibration values ​​of the intrinsic parameters of the main lens and sub-lens of camera 1, the calibration values ​​of the node offset information of camera 1, and the intermediate calibration values ​​of the screen offset information of the virtual shooting area are obtained. Next, select one set of spatial calibration data corresponding to camera 2 in any order, and then select one 2D-3D coordinate point pair from that set in any order. Perform the optimization process described above, which aims to minimize reprojection error. When optimizing camera 2, the objects to be optimized include the intermediate calibration values ​​of the screen offset information of the virtual shooting area (obtained from the optimization of camera 1). After completing the optimization process of 100 2D-3D coordinate point pairs of camera 2, the calibration values ​​of the intrinsic parameters of the main lens and sub-lens of camera 2, the calibration values ​​of the node offset information of camera 2, and the calibration values ​​of the screen offset information of the virtual shooting area are obtained.

[0109] In this embodiment, a first joint optimization is performed using a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera to obtain an intermediate spatial calibration result. A second joint optimization is then performed using the first preset number of spatial calibration data corresponding to the main lens of each camera, the second preset number of spatial calibration data corresponding to the sub-lenses of each camera, and the intermediate spatial calibration result to obtain the final spatial calibration result. Thus, multi-lens fusion calibration is performed in the first joint optimization, and then multi-camera fusion calibration is performed in the second joint optimization. This two-stage joint optimization achieves high-precision spatial calibration with a relatively small amount of collected spatial calibration data.

[0110] For example, Figure 4 This diagram illustrates a flow chart of a space calibration method according to an embodiment of the present disclosure, as shown below. Figure 4 As shown, it mainly includes: data acquisition (corresponding to the above). Figure 2 Step 202), multi-lens fusion calibration (corresponding to the above) Figure 3 Step 301), Multi-camera fusion calibration (corresponding to the above) Figure 3 The process involves three stages: step 302). In the data acquisition stage, the main lens of each camera acquires full spatial calibration data, while the sub-lenses of each camera do not require full spatial calibration data acquisition. For each main lens, a complete spatial calibration is performed using the acquired spatial calibration data, while for each sub-lens, intrinsic parameter calibration is performed using the acquired spatial calibration data. In the multi-lens fusion calibration stage, the spatial calibration results of the main lens and the intrinsic parameter calibration results of the sub-lenses within the same camera are jointly optimized to obtain intermediate spatial calibration results. In the multi-camera fusion stage, based on the intermediate spatial calibration results of each camera, further joint optimization is performed to obtain the final spatial calibration result.

[0111] Based on the same inventive concept as the above method embodiments, the present disclosure also provides a space calibration device, which can be used to execute the technical solutions described in the above method embodiments.

[0112] Figure 5 The diagram illustrates a structural diagram of a spatial calibration device according to an embodiment of the present disclosure. This device is applied to a scene employing multiple cameras for virtual photography, wherein each camera can be fitted with multiple lenses, and each camera is fixedly equipped with a positioning device; as shown... Figure 5 As shown, the device may include:

[0113] The allocation module 501 is used to allocate the multiple lenses and determine the lens set of each camera, wherein the lens sets of different cameras contain different lenses.

[0114] The data acquisition module 502 is used to acquire a first preset number of spatial calibration data corresponding to the main lens of each camera, and a second preset number of spatial calibration data corresponding to the sub-lens of each camera; wherein, the second preset number is less than the first preset number; and the sub-lens of any camera is the lens in the lens set of that camera excluding the main lens of that camera.

[0115] The spatial calibration module 503 is used to perform joint optimization based on a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera to obtain spatial calibration results. The spatial calibration results include: calibration values ​​of screen offset information of the virtual shooting location, calibration values ​​of node offset information of each camera, and calibration values ​​of intrinsic parameters of each of the multiple lenses. The node offset information represents the offset between the equivalent optical center of the positioning device and the camera; the screen offset information represents the offset between the coordinate system corresponding to the screen and the coordinate system corresponding to the positioning device in the virtual shooting location.

[0116] In this embodiment, it is applied to a scenario where multiple cameras are used for virtual shooting. Each camera can be adapted to multiple lenses, and each camera is fixedly equipped with a positioning device. During the spatial calibration process, the multiple lenses are allocated to determine the lens set of each camera, wherein the lens sets of different cameras contain different lenses. A first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera are obtained. The second preset number is less than the first preset number. The sub-lenses of any camera are the lenses in the lens set of that camera other than the main lens of that camera. Based on the first preset number of spatial calibration data corresponding to the main lens of each camera and the second preset number of spatial calibration data corresponding to the sub-lenses of each camera, joint optimization is performed to obtain a spatial calibration result. The spatial calibration result includes: the calibration value of the screen offset information of the virtual shooting site, the calibration value of the node offset information of each camera, and the calibration value of the intrinsic parameters of each lens among the multiple lenses. The node offset information represents the offset between the equivalent optical center of the positioning device and the camera. The screen offset information represents the offset between the coordinate system corresponding to the screen in the virtual shooting site and the coordinate system corresponding to the positioning device. In this way, for virtual shooting scenes with multiple lenses and cameras, the amount of spatial calibration data collected and the collection time are significantly reduced while ensuring spatial calibration accuracy, thereby greatly improving spatial calibration efficiency. Specifically, since each lens is assigned to only one camera, for any given lens, only the spatial calibration data corresponding to that lens when installed on its assigned camera needs to be acquired, effectively reducing the amount of spatial calibration data collected during the spatial calibration process and significantly shortening the data collection time. Furthermore, since the second preset number is less than the first preset number, meaning the amount of spatial calibration data for the second preset number of sub-lenses is less than the amount of spatial calibration data for the first preset number of main lenses, the amount of spatial calibration data collected for sub-lenses during the spatial calibration process is further reduced, shortening the data collection time. Moreover, through joint optimization, even with less spatial calibration data, the spatial calibration results obtained through the mutual constraints between multiple lenses and cameras couple the equivalent optical center differences between different lenses, ensuring that the calibration values ​​of the node offset information of each camera and the screen offset information of the virtual shooting site are applicable to each lens, thus guaranteeing spatial calibration accuracy.

[0117] In one possible implementation, the spatial calibration module 503 is further configured to: perform a first joint optimization using a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera to obtain intermediate spatial calibration results; the intermediate spatial calibration results include: optimized values ​​of the intrinsic parameters of each lens, optimized values ​​of the node offset information of each camera, and optimized values ​​of the screen offset information corresponding to each camera; and perform a second joint optimization using the first preset number of spatial calibration data corresponding to the main lens of each camera, the second preset number of spatial calibration data corresponding to the sub-lenses of each camera, and the intermediate spatial calibration results to obtain the spatial calibration results.

[0118] In one possible implementation, the spatial calibration module 503 is further configured to: for any camera, obtain initial values ​​of the intrinsic parameters of the main lens of the camera, initial values ​​of the node offset information of the camera, and initial values ​​of the screen offset information corresponding to the camera using a first preset number of spatial calibration data corresponding to the main lens of the camera; obtain initial values ​​of the intrinsic parameters of the sub-lenses of the camera using a second preset number of spatial calibration data corresponding to the sub-lenses of the camera; and use the first preset number of spatial calibration data corresponding to the main lens of the camera and the second preset number of spatial calibration data corresponding to the sub-lenses of the camera to... With the goal of minimizing reprojection error, the intrinsic parameters of the camera's main lens, the intrinsic parameters of the camera's sub-lenses, the camera's node offset information, and the corresponding screen offset information are optimized to obtain optimized values ​​for the camera's main lens intrinsic parameters, sub-lenses intrinsic parameters, node offset information, and corresponding screen offset information. The initial values ​​used in the optimization process include: initial values ​​for the camera's main lens intrinsic parameters, sub-lenses intrinsic parameters, node offset information, and corresponding screen offset information.

[0119] In one possible implementation, the spatial calibration module 503 is further configured to: utilize a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera, with the goal of minimizing reprojection error, optimize the intrinsic parameters of the main lens of each camera, the intrinsic parameters of the sub-lenses of each camera, the node offset information of each camera, and the screen offset information of the virtual shooting site to obtain the spatial calibration result; wherein the initial values ​​in the optimization process include: the optimized values ​​of the intrinsic parameters of the main lens of each camera, the optimized values ​​of the intrinsic parameters of the sub-lenses of each camera, the optimized values ​​of the node offset information of each camera, and the optimized values ​​of the screen offset information corresponding to any camera.

[0120] In one possible implementation, the spatial calibration data includes: images captured by the camera for the calibration screen in a virtual shooting location, and pose data synchronously collected by the positioning device.

[0121] In one possible implementation, each camera has one main lens.

[0122] In one possible implementation, the data acquisition module 502 is further configured to: determine the first preset number of location points and the shooting angle corresponding to each location point in the virtual shooting area to cover the virtual shooting area; for any camera, acquire the image captured by the main lens of the camera at each location point with the corresponding shooting angle for the calibration screen, and the pose data synchronously collected by the positioning device installed on the camera, wherein the calibration screen is displayed on the screen in the virtual shooting area.

[0123] In one possible implementation, the first preset quantity is twice the second preset quantity.

[0124] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0125] This disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.

[0126] This disclosure also provides a non-volatile computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0127] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0128] Figure 6 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 may be provided as a server or a terminal device. (Refer to...) Figure 6 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.

[0129] Electronic device 1900 may also include a power supply component 1926 configured to perform power management of electronic device 1900, a wired or wireless network interface 1950 configured to connect electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). Electronic device 1900 can operate on an operating system, such as Windows Server, stored in memory 1932. TM Mac OS X TM Unix TM Linux TM FreeBSD TM Or similar.

[0130] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of an electronic device 1900 to perform the above-described method.

[0131] Computer-readable storage media can be tangible devices capable of holding and storing programs / instructions used by instruction execution devices. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0132] The computer program (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage medium in the respective computing / processing device.

[0133] The computer program (or computer program instructions) used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions to implement various aspects of this disclosure.

[0134] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0135] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0136] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0137] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0138] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A space calibration method, characterized in that, This method is applicable to scenarios involving virtual shooting using multiple cameras, where each camera can be fitted with multiple lenses, and each camera is fixedly equipped with a positioning device; the method includes: The multiple lenses are assigned to determine the lens set for each camera, wherein the lens sets of different cameras contain different lenses; Acquire a first preset number of spatial calibration data corresponding to the main lens of each camera, and a second preset number of spatial calibration data corresponding to the sub-lens of each camera; wherein the second preset number is less than the first preset number; the sub-lens of any camera is the lens in the lens set of that camera excluding the main lens of that camera; Based on a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera, joint optimization is performed to obtain spatial calibration results. The spatial calibration results include: calibration values ​​of screen offset information for the virtual shooting location, calibration values ​​of node offset information for each camera, and calibration values ​​of intrinsic parameters for each of the multiple lenses. The node offset information represents the offset between the equivalent optical center of the positioning device and the camera; the screen offset information represents the offset between the coordinate system corresponding to the screen and the coordinate system corresponding to the positioning device in the virtual shooting location.

2. The method according to claim 1, characterized in that, The spatial calibration results are obtained by jointly optimizing the spatial calibration data based on a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera, including: Using a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera, a first joint optimization is performed to obtain intermediate spatial calibration results; the intermediate spatial calibration results include: optimized values ​​of the intrinsic parameters of each lens, optimized values ​​of the node offset information of each camera, and optimized values ​​of the screen offset information corresponding to each camera. Using a first preset number of spatial calibration data corresponding to the main lens of each camera, a second preset number of spatial calibration data corresponding to the sub-lenses of each camera, and the intermediate spatial calibration results, a second joint optimization is performed to obtain the spatial calibration result.

3. The method according to claim 2, characterized in that, The first joint optimization is performed using a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera to obtain intermediate spatial calibration results, including: For any camera, the initial values ​​of the intrinsic parameters of the main lens, the initial values ​​of the node offset information of the camera, and the initial values ​​of the screen offset information of the camera are obtained by using the first preset number of spatial calibration data corresponding to the main lens of the camera. The initial values ​​of the intrinsic parameters of the camera's sub-lenses are obtained by using the second preset number of spatial calibration data corresponding to the sub-lenses of the camera. Using a first preset number of spatial calibration data corresponding to the main lens of the camera and a second preset number of spatial calibration data corresponding to the sub-lenses of the camera, with the goal of minimizing reprojection error, the intrinsic parameters of the main lens of the camera, the intrinsic parameters of the sub-lenses of the camera, the node offset information of the camera, and the screen offset information corresponding to the camera are optimized to obtain optimized values ​​for the intrinsic parameters of the main lens of the camera, the intrinsic parameters of the sub-lenses of the camera, the node offset information of the camera, and the screen offset information corresponding to the camera. The initial values ​​in the optimization process include the initial values ​​for the intrinsic parameters of the main lens of the camera, the intrinsic parameters of the sub-lenses of the camera, the node offset information of the camera, and the screen offset information corresponding to the camera.

4. The method according to claim 2, characterized in that, The second joint optimization is performed using a first preset number of spatial calibration data corresponding to the main lens of each camera, a second preset number of spatial calibration data corresponding to the sub-lenses of each camera, and the intermediate spatial calibration results to obtain the spatial calibration results, including: Using a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera, with the goal of minimizing reprojection error, the intrinsic parameters of the main lens of each camera, the intrinsic parameters of the sub-lenses of each camera, the node offset information of each camera, and the screen offset information of the virtual shooting site are optimized to obtain the spatial calibration result; wherein, the initial values ​​in the optimization process include: the optimized value of the intrinsic parameters of the main lens of each camera, the optimized value of the intrinsic parameters of the sub-lenses of each camera, the optimized value of the node offset information of each camera, and the optimized value of the screen offset information corresponding to any camera.

5. The method according to claim 1, characterized in that, The spatial calibration data includes: images captured by the camera for the calibration screen in a virtual shooting location, and pose data synchronously collected by the positioning device.

6. The method according to claim 1, characterized in that, Each camera has one main lens.

7. The method according to claim 1, characterized in that, The step of acquiring a first preset number of spatial calibration data corresponding to the main lens of each camera includes: In the virtual shooting location, determine the first preset number of location points and the shooting angle corresponding to each location point to cover the virtual shooting location; For any given camera, images captured by the camera's main lens at each position point with corresponding shooting angles for the calibration screen are acquired, along with pose data synchronously collected by the positioning device installed on the camera. The calibration screen is then displayed on a screen in a virtual shooting area.

8. The method according to claim 1, characterized in that, The first preset quantity is twice the second preset quantity.

9. A space calibration device, characterized in that, This device is applicable to scenarios where multiple cameras are used for virtual shooting, wherein each camera can be adapted to multiple lenses, and each camera is fixedly equipped with a positioning device; the device includes: The allocation module is used to allocate the multiple lenses and determine the lens set for each camera, wherein the lens sets of different cameras contain different lenses; The data acquisition module is used to acquire a first preset number of spatial calibration data corresponding to the main lens of each camera, and a second preset number of spatial calibration data corresponding to the sub-lens of each camera; wherein, the second preset number is less than the first preset number; and the sub-lens of any camera is the lens in the lens set of that camera excluding the main lens of that camera. A spatial calibration module is used to perform joint optimization based on a first preset number of spatial calibration data corresponding to the main lens of each camera and a second preset number of spatial calibration data corresponding to the sub-lenses of each camera to obtain spatial calibration results. The spatial calibration results include: calibration values ​​of screen offset information of the virtual shooting location, calibration values ​​of node offset information of each camera, and calibration values ​​of intrinsic parameters of each of the multiple lenses. The node offset information represents the offset between the equivalent optical center of the positioning device and the camera; the screen offset information represents the offset between the coordinate system corresponding to the screen and the coordinate system corresponding to the positioning device in the virtual shooting location.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.

11. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.