Method, device and electronic device for determining delay between camera and capture device
By using the spatial coordinates and pixel coordinates of markers in curtain wall pictures and videos, combined with the capturer posture sequence, and using the PnP algorithm and coordinate conversion relationship, the problem of difficulty in accurately determining the delay between the camera and the capturer is solved, and high-precision delay measurement is achieved.
Patent Information
- Application Number
- CN202211393903.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-11-08
AI Technical Summary
There is a problem of outputs that are not synchronized between the camera and the capture, which makes it difficult to accurately determine the delay.
By obtaining the curtain wall pictures and videos taken by the camera at different locations, as well as the pose sequences collected synchronously by the capturer, the delay between the camera and the capturer is determined using the PnP algorithm and coordinate conversion relationship.
Accurate and rapid determination of the delay between the camera and the capture is achieved, improving the synchronization and stability of the system.
Smart Images

Figure CN116489488B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology. Specifically, the present application relates to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for determining the delay between a camera and a capture device. Background Art
[0002] Virtual production is a broad term referring to various digital workflows and methods that utilize computer-aided production and film visualization production. An important device in virtual production is a camera with an optical camera capture system (hereinafter referred to as a capture device). By installing the capture device on the camera and shooting a reflective marker sticker attached to the capture area, the camera position information and lens information can be collected in real time with high precision.
[0003] Although the capture device can synchronously collect information with the camera, that is, when the camera captures an image or video frame, the capture device can also synchronously collect its own and the camera's postures, due to the difference in processing capabilities of the two devices, the problem of asynchronous output may occur between the two devices. Summary of the Invention
[0004] Embodiments of the present application provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for determining the delay between a camera and a capture device, which can solve the above problems in the prior art. The technical solutions are as follows:
[0005] According to the first aspect of the embodiments of the present application, a method for determining the delay between a camera and a capture device is provided. The capture device is used to synchronously collect the capture device posture of the capture device itself in the first coordinate system when the camera captures an image. The method includes:
[0006] Obtain the curtain wall pictures of the curtain wall taken by the camera at different positions, and the first capture device postures synchronously collected by the capture device. Obtain the curtain wall video of the curtain wall taken by the camera during the movement, and the capture device posture sequence composed of the second capture device postures synchronously collected by the capture device during the movement. The curtain wall is used to display markers, and the spatial coordinates of each marker in the second coordinate system have been determined in advance;
[0007] According to each curtain wall picture and the curtain wall video, obtain a first posture sequence. The first posture sequence includes the first camera postures of the camera in the second coordinate system when the camera captures each video frame in the curtain wall video;
[0008] According to each curtain wall picture, each first capture device posture, and the capture device posture sequence, obtain a second posture sequence. The second posture sequence includes the second camera postures of the camera in the second coordinate system when the capture device captures each second capture device posture;
[0009] Determine the delay between the camera and the capture device according to the first pose sequence and the second pose sequence.
[0010] As an alternative implementation, determining the delay between the camera and the capture device according to the first pose sequence and the second pose sequence includes:
[0011] Take one of the first pose sequence and the second pose sequence as the reference sequence, and the other as the target sequence;
[0012] Take at least one camera pose in the reference sequence as the reference pose, and determine the camera pose with the highest similarity to the reference pose in the target sequence as the target pose of the reference pose;
[0013] For each reference pose, determine the difference between the first time sequence and the second time sequence corresponding to the reference pose, where the first time sequence is the time sequence of the reference pose in the reference sequence, and the second time sequence is the time sequence of the target pose in the target sequence;
[0014] Determine the delay according to the differences corresponding to each reference pose.
[0015] As an alternative implementation, obtaining the first pose sequence according to each curtain wall picture and the curtain wall video includes:
[0016] Obtain the internal parameters of the camera according to the spatial coordinates and pixel coordinates of the markers in each curtain wall picture;
[0017] Obtain the first pose sequence according to the curtain wall video and the internal parameters.
[0018] As an alternative implementation, obtaining the second pose sequence according to each curtain wall picture, each first capture device pose, and the capture device pose sequence includes:
[0019] Obtain the offset of the camera relative to the capture device in the first coordinate system and the coordinate transformation relationship between the first coordinate system and the second coordinate system according to the internal parameters of the camera, each curtain wall picture, and the first capture device pose corresponding to each curtain wall picture; the internal parameters are determined according to the spatial coordinates and pixel coordinates of the markers in each curtain wall picture;
[0020] Obtain the second pose sequence according to the offset, the coordinate transformation relationship, and the capture device pose sequence.
[0021] As an alternative implementation, obtaining the first pose sequence according to the curtain wall video and the internal parameters includes:
[0022] Obtain the first camera pose of the camera in the second coordinate system when shooting each video frame according to the internal parameters, the spatial coordinates and pixel coordinates of the markers in each video frame;
[0023] Obtain a first pose sequence according to the time sequence of each video frame in the curtain wall video and the first camera pose corresponding to each video frame.
[0024] As an alternative implementation, according to the internal parameters of the camera, each curtain wall picture, and the first capturer pose corresponding to each curtain wall picture, obtain the offset of the camera relative to the capturer in the first coordinate system and the coordinate transformation relationship between the first coordinate system and the second coordinate system, including:
[0025] According to the internal parameters, the spatial coordinates and pixel coordinates of the markers in each curtain wall picture, obtain the third camera pose of the camera in the second coordinate system when shooting each curtain wall picture;
[0026] According to the first capturer pose and the third camera pose corresponding to each curtain wall picture, obtain the offset and the coordinate transformation relationship.
[0027] As an alternative implementation, according to the first capturer pose and the third camera pose corresponding to each curtain wall picture, obtain the coordinate transformation relationship, including:
[0028] For each curtain wall picture, according to the offset, the first capturer pose, and the third camera pose corresponding to the curtain wall picture, obtain the initial coordinate transformation relationship corresponding to the curtain wall picture;
[0029] Fit the initial coordinate transformation relationships corresponding to all curtain wall pictures to obtain the coordinate transformation relationship.
[0030] As an alternative implementation, the marker is a two-dimensional code, and the two-dimensional code is used to record the spatial coordinates of the two-dimensional code in the second coordinate system.
[0031] According to the second aspect of the embodiments of the present application, there is provided a device for determining the delay between a camera and a capturer, the device including:
[0032] An image pose acquisition module, configured to obtain curtain wall pictures of the curtain wall taken by the camera at different positions, and the first capturer pose synchronously acquired by the capturer, obtain a curtain wall video of the curtain wall taken by the camera during movement, and a capturer pose sequence composed of the second capturer poses synchronously acquired by the capturer during movement, where the curtain wall is used to display markers, and the spatial coordinates of each marker in the second coordinate system have been determined in advance;
[0033] A first pose sequence determination module, configured to obtain a first pose sequence according to each curtain wall picture and the curtain wall video, where the first pose sequence includes the first camera pose of the camera in the second coordinate system when shooting each video frame in the curtain wall video;
[0034] A second attitude sequence determination module, configured to obtain a second attitude sequence according to each curtain wall picture, each first capturer attitude, and the capturer attitude sequence, where the second attitude sequence includes the second camera attitude of the camera in the second coordinate system when the capturer collects each second capturer attitude;
[0035] A comparison module, configured to determine the delay between the camera and the capturer according to the first attitude sequence and the second attitude sequence.
[0036] According to the third aspect of the embodiments of the present application, an electronic device is provided, which includes a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method in the first aspect.
[0037] According to the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method in the first aspect are implemented.
[0038] According to an aspect of the embodiments of the present application, a computer program product is provided, including a computer program that implements the steps of the method in the first aspect when executed by a processor.
[0039] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are as follows:
[0040] On the basis that there is a rigid body relationship between the capturer and the camera and the frame rate of the capturer collecting its own attitude is the same as the frame rate of the camera collecting images, by obtaining the curtain wall pictures taken by the camera at different positions and the first capturer attitudes synchronously collected by the capturer, as well as the curtain wall videos taken by the camera during movement and the capturer attitude sequence synchronously collected by the capturer during movement, where the first capturer attitude and the capturer attitude sequence are both the attitudes of the capturer itself in the first coordinate system, and for the curtain wall in the curtain wall pictures and curtain wall videos, the spatial coordinates of each marker in the second coordinate system have been determined in advance. The present application can use the spatial coordinates and pixel coordinates of the markers in each curtain wall picture and curtain wall video, and use the PnP algorithm to obtain the first camera attitude of the camera in the second coordinate system when shooting each video frame of the curtain wall video. On the other hand, using the curtain wall pictures, each first capturer attitude, and the capturer attitude sequence, the second camera attitude of the camera in the second coordinate system when the capturer collects a second capturer attitude can be obtained. By comparing the coincidence degree of the first attitude sequence and the second attitude sequence, the delay between the camera and the capturer can be determined, achieving the effect of accurately and quickly determining the delay. Description of the Drawings
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.
[0042] Figure 1 This is a schematic diagram of the system architecture for implementing the method for determining the delay between a camera and a capture device provided by an embodiment of the present application;
[0043] Figure 2 This is a schematic flowchart of the inventive concept provided by an embodiment of the present application;
[0044] Figure 3 This is a schematic flowchart of a method for determining the delay between a camera and a capture device provided by an embodiment of the present application;
[0045] Figure 4 This is a schematic diagram for performing LED curtain wall modeling and lens-related settings provided by an embodiment of the present application;
[0046] Figure 5 This is a schematic flowchart for determining the delay according to the first pose sequence and the second pose sequence provided by an embodiment of the present application;
[0047] Figure 6 This is a schematic diagram of the geometric structure of a PnP problem provided by an embodiment of the present application;
[0048] Figure 7 This is a schematic flowchart of a curtain wall modeling provided by an embodiment of the present application;
[0049] Figure 8 This is a schematic flowchart of another method for determining the delay between a camera and a capture device provided by an embodiment of the present application;
[0050] Figure 9 is a schematic diagram of a scenario for determining the delay between a camera and a capture device provided by an embodiment of the present application;
[0051] Figure 10 This is a schematic diagram of the structure of a device for determining the delay between a camera and a capture device provided by an embodiment of the present application;
[0052] Figure 11 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0053] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.
[0054] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", and "the" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations supported by the technical field of the present application. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein can include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term. For example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".
[0055] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0056] First, several terms related to the present application are introduced and explained:
[0057] Genlock, synchronous locking. When the Genlock function is enabled for the capturer and the camera, the two devices acquire data at the same sampling rate. For example, when the Genlock function is set to 25 FPS, it means that the capturer acquires poses every 40 milliseconds, and the camera acquires one frame of image every 40 milliseconds.
[0058] Zhang Zhengyou camera calibration: Provide a calibration map with corner points, define the origin as needed, and then provide the three-dimensional positions of the corner points corresponding to the origin. The camera takes multiple calibration maps in different poses. According to the pixel coordinates and spatial coordinates of multiple sets of corner points in the taken pictures, the internal parameters are obtained.
[0059] The camera internal parameters are parameters related to the characteristics of the camera itself, such as the focal length and pixel size of the camera. The camera external parameters are parameters in the world coordinate system, such as the position and rotation direction of the camera. Camera calibration is the mapping from world coordinates to pixel coordinates.
[0060] The PnP (Perspective-n-Point) algorithm is a method for solving the correspondence between 3D and 2D points. It describes how to estimate the pose of the camera when the 3D positions of n 3D space points and their positions are known. If the 3D positions of the feature points in one of the two images are known, then at least 3 point pairs (and at least one additional verification point to verify the result) are required to calculate the motion of the camera.
[0061] Support vector machines (SVM) is a binary classification model. Its basic model is a linear classifier with the largest margin defined in the feature space. The largest margin makes it different from the perceptron; SVM also includes the kernel trick, which makes it a substantially non-linear classifier. The learning strategy of SVM is to maximize the margin, which can be formalized as a problem of solving a convex quadratic programming, and is also equivalent to the minimization problem of the regularized hinge loss function. The learning algorithm of SVM is the optimization algorithm for solving convex quadratic programming.
[0062] Hand-eye calibration, where the hand refers to the robotic arm and the eye refers to the camera. In industrial applications, there are two positional relationships between the hand and the eye. One is that the camera (eye) is fixed on the robotic arm (hand), and the eye moves with the hand. The other is that the camera (eye) and the robotic arm (hand) are separated, and the position of the eye is fixed relative to the hand. In this application, the camera and the catcher are rigidly connected, belonging to the first group of positional relationships. For the catcher and the camera, calibration refers to determining the coordinate transformation relationship between the camera and the catcher. What the camera knows is the pixel coordinates, while what the catcher knows is the spatial coordinate system. Therefore, hand-eye calibration is the coordinate transformation relationship between the pixel coordinate system and the spatial robot coordinate system.
[0063] The method, device, electronic device, computer-readable storage medium, and computer program product for determining the delay between a camera and a catcher provided by this application aim to solve the above technical problems of the prior art.
[0064] Next, through the description of several exemplary embodiments, the technical solutions of the embodiments of this application and the technical effects produced by the technical solutions of this application will be described. It should be noted that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.
[0065] Figure 1 It is a schematic diagram of the system architecture for implementing the method for determining the delay between a camera and a catcher provided by the embodiments of this application. The system architecture 100 may include a server 2000 and a cluster of terminal devices. Among them, the cluster of terminal devices may specifically include one or more terminal devices, and the number of terminal devices in the cluster of terminal devices will not be limited here. As Figure 1 shown, the multiple terminal devices may specifically include terminal device 3000a, terminal device 3000b, terminal device 3000c,..., terminal device 3000n; terminal device 3000a, terminal device 3000b, terminal device 3000c,..., terminal device 3000n may be directly or indirectly connected to the server 2000 through wired or wireless communication methods, so that each terminal device can perform data interaction with the server 2000 through this network connection.
[0066] Among them, each terminal device in the terminal device cluster may include: intelligent terminals with data processing functions such as smartphones, tablets, laptops, desktop computers, smart home appliances, wearable devices, vehicle-mounted terminals, intelligent voice interaction devices, cameras, etc. For ease of understanding, embodiments of the present application may be in Figure 1 select one terminal device as the target terminal device from the multiple terminal devices shown. For example, embodiments of the present application may use Figure 1 the terminal device 3000b shown as the target terminal device.
[0067] Among them, the server 2000 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0068] It should be understood that a shooting component for collecting a target image associated with a spatial object may be integrated on the target terminal device. Here, the shooting component may be a camera on the target terminal device for taking photos or videos and a catcher rigidly connected to the camera. Among them, multiple cameras may be integrally installed on one target terminal device, and each camera is rigidly connected to a catcher. Embodiments of the present application take one camera and the corresponding catcher on one target terminal device as one shooting component for illustration. Among them, the spatial object may be a QR code green screen, and the QR code green screen means a green screen printed with QR codes. Optionally, the spatial object may also be a checkerboard green screen, and the checkerboard green screen means a green screen printed with rectangular frames of solid color (for example, black). It should be understood that embodiments of the present application are described by taking the spatial object as a QR code green screen as an example.
[0069] Among them, the QR code green screen may have three faces: the left wall + the right wall + the ground. Optionally, the QR code green screen may also be any one of the left wall + the right wall + the ground, and the QR code green screen may also be any two of the left wall + the right wall + the ground. All the QR codes in the QR code green screen have unique patterns and numbers, and the identification code detection algorithm (QR code detection algorithm) can detect them in the target image and accurately obtain their corner coordinates on the target image. Among them, for a single QR code, the four vertices formed by the QR code border can be called QR code corners.
[0070] It can be understood that in the present application, a two-dimensional code that can be correctly recognized by a two-dimensional code detection algorithm can be referred to as an observable two-dimensional code. It can be understood that when the two-dimensional code is blocked, the two-dimensional code is unclear, or a partial area of the two-dimensional code exceeds the picture boundary of the target image, the two-dimensional code detection algorithm will be unable to detect the two-dimensional code. At this time, such a two-dimensional code is not considered an observable two-dimensional code.
[0071] It should be understood that the above network framework can be applied to the field of virtual-real fusion, for example, video production (virtual production), live broadcast, and virtual-real fusion in post-video special effects. Virtual-real fusion refers to implanting a real shooting subject into a virtual scene. Compared with the traditional method of completely real shooting, virtual-real fusion can conveniently replace the scene, greatly reducing the cost of scene layout (virtual-real fusion only requires a green screen), and can provide very cool environmental effects. In addition, virtual-real fusion is also highly compatible with concepts such as VR (Virtual Reality), the metaverse, and the omnipresent Internet, and can provide the very basic ability to implant real people into virtual scenes for them.
[0072] The virtual-real fusion method of the embodiments of the present application needs to synchronize the camera and the capture device before formal shooting to ensure the correct visual effect of the subsequent synthesized picture.
[0073] Please refer to Figure 2 , which is a schematic flowchart of the inventive concept of the embodiments of the present application. As shown in the figure, first, curtain wall modeling is performed, that is, a curtain wall showing at least one marker is constructed. The position of each marker in the LED display screen can be fixed. Determine the spatial coordinates of each marker in the second coordinate system. The second coordinate system of the embodiments of the present application can be a spatial coordinate system based on the curtain wall construction. The camera takes pictures of the curtain wall from at least 4 positions to obtain curtain wall pictures, and obtains the first capture device attitude synchronously collected by the capture device. The first capture device attitude is the attitude of the capture device itself in the first coordinate system. The first coordinate system can be a spatial coordinate system based on the capture device.
[0074] According to the spatial coordinates and pixel coordinates of the markers in the curtain wall pictures, the optimal fitting internal parameters of the camera can be obtained based on the camera calibration method. According to the internal parameters and the spatial coordinates and pixel coordinates of the markers in each curtain wall picture, based on the 3D-to-2D motion method, the third camera attitude of the camera in the second coordinate system when taking each curtain wall picture can be obtained. Combining the first capture device attitude of the capture device in the first coordinate system when the camera takes each curtain wall picture, the offset of the camera relative to the capture device in the first coordinate system and the coordinate transformation relationship between the first coordinate system and the second coordinate system can be obtained.
[0075] Then, a camera captures the curtain wall in a moving state, obtaining a curtain wall video of the curtain wall captured by the camera during movement and a sequence of capturer postures composed of the second capturer postures synchronously collected by the capturer during movement.
[0076] On the one hand, using the spatial coordinates and pixel coordinates of the markers in each video frame of the curtain wall video, and combining with the internal parameters of the camera, a first posture sequence can be obtained, which includes the first camera postures of the camera in the second coordinate system when the camera captures each video frame of the curtain wall video; on the other hand, using each curtain wall picture, each first capturer posture, and the sequence of capturer postures, a second posture sequence can be obtained, which includes the second camera postures of the camera in the second coordinate system when the capturer collects each second capturer posture.
[0077] By comparing the first camera postures and the second camera postures, and using the difference in the sequence numbers of the first camera posture and the second camera posture with the highest similarity in their respective sequences, the number of frames difference between the camera and the capturer, that is, the delay, can be obtained.
[0078] An embodiment of the present application provides a method for determining the delay between a camera and a capturer, as Figure 3 shown, the method includes:
[0079] S101. Obtain curtain wall pictures of the curtain wall captured by the camera at different positions, and the first capturer postures synchronously collected by the capturer, and obtain a curtain wall video of the curtain wall captured by the camera during movement and a sequence of capturer postures composed of the second capturer postures synchronously collected by the capturer during movement.
[0080] It should be noted that, before executing step S101, an embodiment of the present application can construct a curtain wall for displaying markers. The type of the curtain wall corresponding to the embodiment of the present application is not limited. For example, it can be an ordinary storage rack. In one embodiment, the curtain wall of the embodiment of the present application is a display screen, that is, the marker is an image displayed on the display screen, rather than an entity in the conventional sense. Using the display screen as the curtain wall in the present application can display the marker more quickly, and is also more accurate when determining the spatial coordinates of the marker - by pre - establishing the correspondence between the spatial coordinates and the pixel points in the display screen, after determining the pixel points where the marker is located, the control coordinates of the marker can be determined in combination with the above - mentioned correspondence. It should be understood that the spatial coordinates of the marker in the present application are coordinates in the second coordinate system. The second coordinate system can be a coordinate system constructed by the curtain wall, or a coordinate system constructed by other objects that remain relatively stationary with the curtain wall. The embodiment of the present application does not make specific limitations.
[0081] Please refer to Figure 4, which exemplarily shows a schematic diagram of the LED curtain wall modeling and lens-related settings in the embodiments of the present application. As shown in the figure, the embodiments of the present application need to set information such as the LED screen name, basic screen information, camera name, and camera physical parameters. It should be understood that the parameters shown in the figure are only examples, and more precise parameters can also be set in actual applications, which are not limited in the present application.
[0082] The embodiments of the present application need to obtain curtain wall pictures taken by the camera at different positions. Optionally, when the camera is at different positions, there are also differences in the shooting angles. Such a setting can, in the case of a small number of markers captured, due to the differences in shooting angles, make the pixel coordinates of the same marker different in different curtain wall pictures, and obtain more accurate internal parameters of the camera. It should be understood that the number of curtain wall pictures taken in the embodiments of the present application is at least 4, and more curtain wall pictures can be obtained according to the actual situation to obtain more accurate internal parameters.
[0083] The capturer can synchronously collect the attitude of the capturer itself when the camera captures an image (such as a curtain wall picture). In the embodiments of the present application, the first capturer attitude of the capturer is obtained when the camera captures each curtain wall picture. The first capturer attitude is the attitude of the capturer itself in the first coordinate system, and the first coordinate system can be a coordinate system constructed by the capturer. Generally, the first capturer attitude can be represented by a matrix of size 4*4.
[0084] In actual applications, whenever the camera is at a new shooting position, it can first stabilize at this shooting position for one or two seconds before taking the curtain wall picture, so as to ensure that the first capturer attitude collected by the capturer is the attitude of the camera when taking the curtain wall picture at this position.
[0085] Step S101 of the embodiments of the present application also needs to obtain a capturer attitude sequence composed of the curtain wall video taken by the camera during the movement process and the second capturer attitude synchronously collected by the capturer during the movement process. In actual applications, the duration of the movement process of the camera in the embodiments of the present application can be 1 to 2 seconds, and the number of video frames included in the curtain wall video is about 50 frames. It can be seen that the time cost for the operator to obtain the capturer attitude sequence is very low, and because the duration of the movement process in the embodiments of the present application is very short, the situation of frame loss is basically eliminated. It can be understood that the second capturer attitude can also be represented by a matrix of size 4*4.
[0086] As can be seen from the above description, although the capturer and the camera are synchronously collected, there may be differences when they output. For example, when the camera outputs the third video frame, the capturer may only output the second capturer attitude. Therefore, the capturer attitude sequence obtained in the present application and the video frames in the curtain wall video are not in one-to-one correspondence in order.
[0087] In the embodiments of the present application, there is no specific limitation on the sequence of obtaining curtain wall pictures and curtain wall videos. One can first obtain curtain wall pictures and then obtain curtain wall videos, or first obtain curtain wall videos and then obtain curtain wall pictures.
[0088] S102. Obtain a first pose sequence according to each curtain wall picture and the curtain wall video. The first pose sequence includes the first camera pose of the camera in the second coordinate system when the camera captures each video frame of the curtain wall video.
[0089] Since both the curtain wall pictures and the curtain wall video are captured by the camera, and the curtain wall pictures and the curtain wall video include both the pixel coordinates of the markers and the spatial coordinates of the markers can be correspondingly obtained. Through coordinate system conversion processing (such as the PnP algorithm), the movement trajectory of the camera when capturing the curtain wall video can be obtained. This movement trajectory is specifically characterized by a sequence composed of the poses of the camera in the second coordinate system.
[0090] S103. Obtain a second pose sequence according to each curtain wall picture, each first capture device pose, and the capture device pose sequence. The second pose sequence includes the second camera pose of the camera in the second coordinate system when the capture device captures each second capture device pose.
[0091] On the other hand, by using the pixel coordinates of the markers included in the curtain wall pictures and the spatial coordinates of the markers, the internal parameters of the camera, the offset of the camera relative to the capture device in the first coordinate system, and the conversion relationship between the first coordinate system and the second coordinate system can be obtained. Further, in combination with each second capture device pose in the capture device pose sequence, the second camera pose of the camera in the second coordinate system when the capture device captures each second capture device pose can be obtained.
[0092] S104. Determine the delay between the camera and the capture device according to the first pose sequence and the second pose sequence.
[0093] Both pose sequences describe the poses of the camera during movement. If the capture device and the camera output information synchronously, then the poses in the same order in the two pose sequences should be the same. Based on this, in the embodiments of the present application, the number of frames difference between the camera and the capture device, that is, the delay, can be obtained based on the matching degree of the first pose sequence and the second pose sequence.
[0094] The method for determining the delay between a camera and a capturer according to the embodiments of the present application, on the basis that there is a rigid body relationship between the capturer and the camera and the frame rate of the capturer collecting its own attitude is the same as the frame rate of the camera collecting images, obtains curtain wall pictures taken by the camera at different positions and the first capturer attitude synchronously collected by the capturer, as well as the curtain wall video taken by the camera during movement and the sequence of capturer attitudes synchronously collected by the capturer during movement. The first capturer attitude and the sequence of capturer attitudes are both the attitudes of the capturer itself in the first coordinate system, and for the curtain wall in the curtain wall pictures and the curtain wall video, the spatial coordinates of each marker in the second coordinate system have been determined in advance. The present application can utilize the spatial coordinates and pixel coordinates of the markers in each curtain wall picture and curtain wall video, and use the PnP algorithm to obtain the first camera attitude of the camera in the second coordinate system when shooting each video frame of the curtain wall video. On the other hand, by using the curtain wall pictures, each first capturer attitude and the sequence of capturer attitudes, the second camera attitude of the camera in the second coordinate system when the capturer collects a second capturer attitude can be obtained. By comparing the coincidence degree of the first attitude sequence and the second attitude sequence, the delay between the camera and the capturer can be determined, achieving the effect of accurately and quickly determining the delay.
[0095] On the basis of the above embodiments, as an optional embodiment, determining the delay between the camera and the capturer according to the first attitude sequence and the second attitude sequence includes:
[0096] S201. Take one of the first attitude sequence and the second attitude sequence as the reference sequence, and take the other attitude sequence as the target sequence;
[0097] S202. Take at least one camera attitude in the reference sequence as the reference attitude, and determine the camera attitude with the highest similarity to the reference attitude in the target sequence as the target attitude of the reference attitude;
[0098] S203. For each reference attitude, determine the difference between the first time sequence and the second time sequence corresponding to the reference attitude, where the first time sequence is the time sequence of the reference attitude in the reference sequence and the second time sequence is the time sequence of the target attitude in the target sequence.
[0099] Please refer to Figure 5, which exemplarily shows a schematic flow chart of determining the delay according to the first pose sequence and the second pose sequence in the embodiments of the present application. As shown in the figure, after determining the first pose sequence and the second pose sequence in the embodiments of the present application, any one of the two sequences can be used as a reference sequence, and the other sequence can be used as a target sequence. Then, at least one camera pose is determined from the reference sequence as a reference pose, that is, the number of reference poses determined in the embodiments of the present application can be one or more. The following describes the case where the number of reference poses is 1. The similarity between the reference pose and each pose in the target sequence is compared to determine the similarity comparison result between the reference pose and each pose in the target sequence, and the target pose with the highest similarity is determined therefrom. Further, the time sequence (the first time sequence) of the reference pose in the reference sequence and the time sequence (the second time sequence) of the target pose in the target sequence are determined, and the delay is calculated by calculating the difference between the first time sequence and the second time sequence.
[0100] Based on the above embodiments, as an optional embodiment, obtaining the first pose sequence according to each curtain wall picture and the curtain wall video includes:
[0101] S301. Obtain the internal parameters of the camera according to the spatial coordinates and pixel coordinates of the markers in each curtain wall picture;
[0102] S302. Obtain the first pose sequence according to the curtain wall video and the internal parameters.
[0103] Calibration is the link connecting the world coordinates and the pixel coordinates, and the purpose is to obtain the internal and external parameters of the camera and the projector, which is crucial for 3D imaging. The embodiments of the present application provide a camera calibration method based on a 2D planar target, and use the curtain wall pictures obtained after the camera shoots the curtain wall from multiple angles to obtain the internal parameters of the camera. The basic idea of this method is the nonlinear least squares idea, which is a parameter estimation method for estimating the parameters of a nonlinear static model with the criterion of minimizing the sum of the squares of the errors.
[0104] The spatial coordinates of the markers in the curtain wall and the pixel coordinates in the picture are meaningfully corresponding, so the corresponding relationship between the spatial coordinates and the pixel coordinates can be obtained. It can be understood that the internal parameters are determined by the internal parameters of the camera and are defined as a x and a y are the scale factors of the u-axis and the v-axis in the picture (related to the focal length of the camera), r is the non-perpendicular factor of the u-axis and the v-axis, and (u0, v0) is the principal point coordinate (the intersection of the optical axis and the imaging plane).
[0105] The external parameters are determined by the relative position of the camera and the curtain wall (the camera coordinates coincide with the second coordinates after rotation and translation, and this rotation and translation matrix is the external parameter). It is defined as:
[0106]
[0107] r1, r2, and r3 are the direction vectors of the three coordinate axes of the camera coordinates in the second coordinate system respectively. r1, r2, and r3 are perpendicular to each other, and t is the translation vector from the origin of the second coordinate system to the optical center.
[0108] Let the spatial coordinates of the marker M be M = (x y z) T , and the pixel coordinates of the marker in the curtain wall picture be m = (u v) T , and the corresponding homogeneous coordinates are Through coordinate system transformation, we can obtain s is a constant.
[0109] Assume that the curtain wall is located on the xy plane of the second coordinate system, that is, z = 0. Therefore, there is:
[0110]
[0111] Still use M to represent the spatial coordinates of the marker on the curtain wall, but at this time M = (x y) T , Thus, a one-to-one correspondence is obtained:
[0112]
[0113] H = λ * A * (r1 r2 t) = (h1 h2 h3)
[0114] Next, the solution can be carried out. Since multiple groups of markers can be counted from multiple curtain wall pictures, therefore, through these multiple groups of markers, the least squares method can be used to find H. The calculation of H is a process that minimizes the difference between the actual image coordinates m i and those obtained through M . The objective function is:
[0115]
[0116] After solving for H, the internal and external parameters of the camera can be obtained. Specifically, using the formula H = λ * A * (r1 r2 t) = (h1 h2 h3), and the orthogonality of R we can get:
[0117]
[0118] Equation ① is two basic constraints on the internal parameters of the camera. One transformation matrix H can obtain two constraints on the internal parameters of the camera. Therefore, to find A, multiple transformation matrices are required (each curtain wall picture can obtain a transformation matrix).
[0119] For the convenience of solution, let's set here:
[0120]
[0121] It should be noted that B is a symmetric matrix, so it can be represented as a six-dimensional vector
[0122] b = (B 11 B 12 B 22 B 13 B 23 B 33 ) T .
[0123] The i-th column vector in H is h i = (h i1 h i2 h i3 ) T , and it can be deduced that
[0124] where v ij = (h i1 h j1 h i1 h j2 + h i2 h j1 h i2 h j2 h i3 h j1 + h i1 h j3 h i3 h j2 + h i2 h j3 h i3 h j3 )
[0125] Thus, equation ① is simplified to:
[0126]
[0127] By superimposing the equations of all n curtain wall pictures, we can obtain:
[0128] Vb = 0... ②
[0129] where V is a 2n×6 matrix.
[0130] By solving ②, b can be obtained. The solution can be obtained by solving the eigenvector corresponding to the minimum eigenvalue of the matrix, or by performing a singular value decomposition on the matrix V. After obtaining B, A can be obtained by using the Cholesky matrix decomposition algorithm -1 , and then the internal parameter A can be obtained through the inverse.
[0131] After obtaining the internal parameters of the camera, the pixel coordinates and spatial coordinates of the markers in each video frame of the curtain wall video can be combined to obtain the pose of the camera in the second coordinate system when each video frame is captured, and then the first pose sequence can be obtained.
[0132] Based on the above embodiments, as an alternative embodiment, obtaining a second pose sequence according to each curtain wall picture, each first capturer pose, and the capturer pose sequence includes:
[0133] S401. Obtain the offset of the camera relative to the capturer in the first coordinate system and the coordinate conversion relationship between the first coordinate system and the second coordinate system according to the internal parameters of the camera, each curtain wall picture, and the first capturer pose corresponding to each curtain wall picture.
[0134] S402. Obtain the second pose sequence according to the offset, the coordinate conversion relationship, and the capturer pose sequence.
[0135] As can be seen from the above embodiments, the internal parameters of the camera can be determined according to the spatial coordinates and pixel coordinates of the markers in each curtain wall picture. Further, since the capturer also synchronously collects its own pose in the first coordinate system (i.e., the first capturer pose) when the camera captures the curtain wall pictures, the hand-eye calibration algorithm can be used. Under the steel body relationship, when there are more than 4 curtain wall pictures, according to the pose of the camera in the second coordinate system and the first capturer pose corresponding to each curtain wall picture, the offset of the camera relative to the capturer in the first coordinate system and the coordinate conversion relationship between the first coordinate system and the second coordinate system can be obtained.
[0136] After obtaining the offset and the conversion relationship, for each second capturer pose in the capturer pose sequence, using the offset of the camera relative to the capturer in the first coordinate system, the second camera pose of the camera in the first coordinate system when each second capturer pose is captured can be obtained. Then, using the coordinate conversion relationship between the first coordinate system and the second coordinate system, the second camera pose of the camera in the second coordinate system when each second capturer pose is captured can be obtained, and then a second pose sequence is formed.
[0137] Based on the above embodiments, as an alternative embodiment, obtaining a first pose sequence according to the curtain wall video and the internal parameters includes:
[0138] S601. Obtain the first camera pose of the camera in the second coordinate system when each video frame is captured according to the internal parameters, the spatial coordinates, and the pixel coordinates of the markers in each video frame.
[0139] S602. Obtain the first pose sequence according to the time sequence of each video frame in the curtain wall video and the first camera pose corresponding to each video frame.
[0140] The geometric structure of the Perspective-n-Point (PnP) problem is as follows Figure 6 shown. Given the coordinates of 3D points, the corresponding 2D point coordinates, and the intrinsic matrix, the pose of the camera (the attitude in the second coordinate system) is solved. In the embodiments of the present application, the spatial coordinates P1, P2, …, P i , …, P n in the second coordinate system of n markers are known, and the corresponding coordinates p1, p2, …, p i , …, p n on the curtain wall picture and the intrinsic parameter K of the camera. According to these parameters, the pose of the camera in the camera coordinate system (O c X c Y c Z c ) relative to the second coordinate system (O w X w Y w Z w ) needs to be obtained, that is, [Rt] in the following formula.
[0141]
[0142] The homogeneous coordinates of the spatial coordinates of the markers can be denoted as [X w Y w Z w 1] T , and the homogeneous coordinates of the pixel coordinates can be expressed as [u v 1] T . The perspective projection model is:
[0143]
[0144] Expanded as:
[0145]
[0146] After organizing the above model, it can be obtained:
[0147] f 11 X w +f 12 Y w +f 13 Z w +f 14 -f 31 X w u-f 32 Y w u c -f 33 Z w u c -f 34 uc
[0148] f 21 X w +f 22 Y w +f 23 Z w +f 24 -f 31 X w u-f 32 Y w u c -f 33 Z w u c -f 34 u c
[0149] The spatial coordinates and pixel coordinates of each marker can correspond to the above two equations, with a total of 12 unknowns. Therefore, at least 6 groups of markers are required.
[0150] Suppose there are N groups of markers, then:
[0151]
[0152] The above equation is written in matrix form:
[0153] AF = 0
[0154] When N = 6, the linear equations can be directly solved.
[0155] When N is greater than 6, the least-squares solution under the constraint of |F| = 1 can be obtained and solved by SVD. The last column of the V matrix is the solution sought.
[0156] F = UDV T
[0157] Since F = [KR Kt], the rotation matrix and translation matrix can be expressed as:
[0158]
[0159]
[0160] After obtaining the first camera pose of the camera in the second coordinate system when shooting each video frame, combined with the time sequence of each video frame in the curtain wall video, the first pose sequence can be obtained. For example, if the first camera pose when the camera shoots video frame 1 is a, the first camera pose when the camera shoots video frame 2 is b, and the first camera pose when the camera shoots video frame 3 is c, and the time sequences of video frames 1, 2, and 3 in the curtain wall video are also 1, 2, and 3, then the first pose sequence can be expressed as (a, b, c).
[0161] Based on the above embodiments, as an alternative embodiment, according to the internal parameters of the camera, each curtain wall picture, and the attitude of the first capturer corresponding to each curtain wall picture, obtain the offset of the camera relative to the capturer in the first coordinate system and the coordinate transformation relationship between the first coordinate system and the second coordinate system, including:
[0162] S701. According to the internal parameters, the spatial coordinates and pixel coordinates of the markers in each curtain wall picture, obtain the third camera attitude of the camera in the second coordinate system when shooting each curtain wall picture;
[0163] S702. According to the attitude of the first capturer corresponding to each curtain wall picture and the third camera attitude, obtain the offset and the coordinate transformation relationship.
[0164] It should be understood that when the present application determines the third camera attitude of the camera in the second coordinate system when shooting the curtain wall picture, the mechanism is the same as that of step S601 above, both are obtained based on the PnP algorithm, and the embodiments of the present application will not be elaborated here.
[0165] It should be noted that after executing step S701, the attitude of the first capturer and the third camera attitude corresponding to each curtain wall picture are obtained. Since the camera and the capturer satisfy the rigid body relationship, if the attitude of the first capturer is assumed to be A and the third camera attitude is B, then the relationship between the two can be simplified to AX = XB, where X is the offset of the camera relative to the capturer in the first coordinate system. Specifically, the embodiments of the present application can determine the offset through the hand-eye calibration method. Briefly speaking, the hand-eye calibration method includes two steps:
[0166] The first step: Obtain AX = XB, and simplify it to solve Ra = b, where a and b are actually the rotation axes of matrices A and B.
[0167] The second step: Optimize the solution. By introducing the Rodrigues matrix, the 9 unknowns in the rotation matrix R are changed to 3. Considering that there are errors in the coefficient matrices on both sides of the equation, the singular value decomposition (SVD) algorithm is introduced to obtain the optimal offset.
[0168] Based on the above embodiments, as an alternative embodiment, according to the attitude of the first capturer corresponding to each curtain wall picture and the third camera attitude, obtain the coordinate transformation relationship, including:
[0169] For each curtain wall picture, according to the offset, the attitude of the first capturer corresponding to the curtain wall picture, and the third camera attitude, obtain the initial coordinate transformation relationship corresponding to the curtain wall picture;
[0170] Fit the initial coordinate transformation relationships corresponding to all curtain wall pictures to obtain the coordinate transformation relationship.
[0171] It should be noted that the third camera pose is the camera pose in the second coordinate system when the camera captures the curtain wall picture, the offset is the offset of the camera relative to the catcher in the first coordinate system, and the first catcher pose is the catcher pose of the catcher in the first coordinate system when the camera captures the curtain wall picture. For each curtain wall picture, there is a corresponding relationship:
[0172] Offset × First catcher pose × Initial coordinate transformation relationship = Third camera pose.
[0173] Through the corresponding relationships of N (N is an integer greater than 4) curtain wall pictures and fitting by the singular value decomposition method, the coordinate transformation relationship can be obtained.
[0174] Based on the above embodiments, as an optional embodiment, the marker in the embodiment of the present application is a two-dimensional code, and the two-dimensional code is used to record the spatial coordinates of the two-dimensional code in the second coordinate system. By recording its own spatial coordinates in the two-dimensional code in the embodiment of the present application, when processing the curtain wall picture and each video frame, the two-dimensional code in the curtain wall picture can be directly recognized to obtain the spatial coordinates of the two-dimensional code, improving the processing efficiency.
[0175] Please refer to Figure 7 , which exemplarily shows a schematic flow chart of curtain wall modeling provided by the embodiment of the present application. First, determine the size of each two-dimensional code in the curtain wall, determine the positions of all two-dimensional codes in the curtain wall according to the size of the two-dimensional code, calibrate the spatial coordinates of the two-dimensional code in the second coordinate system based on the position, and then generate a two-dimensional code from the spatial coordinates and display the generated two-dimensional code at the corresponding position of the curtain wall. In one embodiment, the spatial coordinates of the two-dimensional code include the spatial coordinates of the four corner points of the two-dimensional code in the second coordinate system.
[0176] Please refer to Figure 8 , which exemplarily shows a schematic flow chart of the method for determining the delay between the camera and the catcher in the embodiment of the present application. As shown in the figure, it includes:
[0177] S801. Curtain wall modeling. The curtain wall includes multiple two-dimensional codes (i.e., markers), and the two-dimensional codes are used to record the spatial coordinates of the two-dimensional codes in the second coordinate system;
[0178] S802. Obtain the curtain wall pictures of the curtain wall captured by the camera at different positions, and the first catcher pose synchronously collected by the catcher;
[0179] S803. Obtain the internal parameters of the camera according to the spatial coordinates and pixel coordinates of the two-dimensional codes in each curtain wall picture. According to the internal parameters, the spatial coordinates and pixel coordinates of the markers in each curtain wall picture, obtain the third camera pose of the camera in the second coordinate system when capturing each curtain wall picture;
[0180] S804. Obtain the offset of the camera relative to the capture device in the first coordinate system and the coordinate conversion relationship between the first coordinate system and the second coordinate system according to the attitude of the first capture device and the attitude of the third camera corresponding to each curtain wall picture;
[0181] S805. Obtain the curtain wall video of the curtain wall captured by the camera during movement, and the capture device attitude sequence composed of the attitudes of the second capture device synchronously collected by the capture device during movement;
[0182] S806. According to the internal parameters, the spatial coordinates and pixel coordinates of the markers in each video frame, obtain the first camera attitude of the camera in the second coordinate system when shooting each video frame, and obtain the first attitude sequence according to the time sequence of each video frame in the curtain wall video and the first camera attitude corresponding to each video frame;
[0183] S806'. Obtain the second attitude sequence according to the offset, the coordinate conversion relationship and the capture device attitude sequence;
[0184] S807. Use one of the attitude sequences in the first attitude sequence and the second attitude sequence as the reference sequence, and use the other attitude sequence as the target sequence;
[0185] S808. Use at least one camera attitude in the reference sequence as the reference attitude, and determine the camera attitude with the highest similarity to the reference attitude in the target sequence as the target attitude of the reference attitude;
[0186] S809. For each reference attitude, determine the difference between the first time sequence and the second time sequence corresponding to the reference attitude, where the first time sequence is the time sequence of the reference attitude in the reference sequence, and the second time sequence is the time sequence of the target attitude in the target sequence;
[0187] S810. Determine the delay between the camera and the capture device according to the differences corresponding to each reference attitude.
[0188] Please refer to FIG. 9. For ease of understanding, further, please refer to Figure 9a ., Figure 9b and Figure 9c ., Figure 9a is a schematic diagram of a scenario for determining the delay between a camera and a capture device provided by an embodiment of the present application Figure 1 ., Figure 9b is a schematic diagram of a scenario for determining the delay between a camera and a capture device provided by an embodiment of the present application Figure 2 ., Figure 9c is a schematic diagram of a scenario for determining the delay between a camera and a capture device provided by an embodiment of the present application Figure 3 . As Figure 9a shown, the server 20a can be the above-mentioned Figure 1The server 2000 in the corresponding embodiment, such as Figure 9a , Figure 9b and Figure 9c The terminal device 20b shown can be the target terminal device in the above Figure 1 corresponding embodiment. Among them, the user corresponding to the target terminal device can be the object 20c, and the target terminal device is integrated with a shooting component.
[0189] Such as Figure 9a shown, the object 20c can shoot the curtain wall through the shooting component in the terminal device 20b to obtain a target image. Among them, the target image can be a curtain wall picture directly shot by the shooting component at different positions, and also includes a curtain wall video shot by the shooting component in a moving state. The curtain wall can include a left wall 21a, a right wall 21b, a ground 21c, and a shooting subject 22b. The shooting subject 22b can be a puppy standing on the ground 21c. Aluminum alloy brackets can be installed behind the left wall 21a and the right wall 21b, and the left wall 21a and the right wall 21b are supported by the aluminum alloy brackets.
[0190] Among them, the spatial object corresponds to at least two coordinate axes. Each two of the at least two coordinate axes are used to form a spatial plane, and there is a perpendicular relationship between each two coordinate axes. Such as Figure 9a shown, the spatial object can correspond to three coordinate axes. The three coordinate axes can specifically be the x-axis, the y-axis, and the z-axis. Each two of the three coordinate axes can be used to form a spatial plane. The x-axis and the z-axis can be used to form the spatial plane formed by the left wall 21a, the y-axis and the z-axis can be used to form the spatial plane formed by the right wall 21b, and the x-axis and the y-axis can be used to form the spatial plane formed by the ground 21c. Among them, the spatial plane includes two-dimensional codes with the same size. For example, the spatial plane formed by the left wall 21a can include a two-dimensional code 22a. It should be understood that the number of two-dimensional codes in the spatial plane in the embodiments of the present application is not limited.
[0191] Such as Figure 9b shown, when the terminal device 20b shoots the curtain wall picture and the curtain wall video, it also obtains the first capturer attitude and the capturer attitude sequence synchronously collected by the capturer. Based on the spatial coordinates and pixel coordinates of the markers in the curtain wall picture, the optimal fitting internal parameters of the camera can be obtained based on the camera calibration method.
[0192] Such as Figure 9bAs shown, based on the method of 3D-to-2D motion, according to the internal reference and the spatial coordinates and pixel coordinates of the markers in each curtain wall picture, the third camera pose of the camera in the second coordinate system when shooting each curtain wall picture can be obtained. Combining the first capturer pose of the capturer in the first coordinate system when the camera shoots each curtain wall picture, the offset of the camera relative to the capturer in the first coordinate system and the coordinate conversion relationship between the first coordinate system and the second coordinate system can be obtained.
[0193] As Figure 9b shown, on the one hand, using the spatial coordinates and pixel coordinates of the markers in each video frame of the curtain wall video and combining with the internal reference of the camera, a first pose sequence can be obtained, which includes the first camera pose of the camera in the second coordinate system when shooting each video frame in the curtain wall video; on the other hand, using each curtain wall picture, each first capturer pose and the capturer pose sequence, a second pose sequence can be obtained, which includes the second camera pose of the camera in the second coordinate system when the capturer acquires each second capturer pose.
[0194] As Figure 9b shown, by comparing the first camera pose and the second camera pose and using the difference in the sequence numbers of the first camera pose and the second camera pose with the highest similarity in their respective sequences, the number of frames difference between the camera and the capturer, that is, the delay, can be obtained.
[0195] Meanwhile, as Figure 9a shown, the terminal device 20b can also obtain the virtual scene for virtual-reality fusion from the server 20a. The virtual scene here can be the Figure 9c virtual scene 23a shown.
[0196] As Figure 9c shown, after the terminal device 20b determines the delay between the camera and the capturer in the shooting component of the terminal device 20b and obtains the virtual scene 23a from the server 20a, the shooting subject 22b in the target image can be implanted into the virtual scene 23a based on the internal reference of the component to obtain the fused scene 23b with correct perspective relationship.
[0197] It can be seen that the embodiments of the present application can achieve calibration of the delay between the camera and the capturer with a small number of pictures and a 1- to 2-second video, and without using hardware devices to calibrate the internal reference of the components, which will significantly reduce the speed and cost of calibrating the delay between the camera and the capturer.
[0198] Virtual-real fusion requires calibrating the shooting component. The embodiments of the present application can support shooting techniques with real-time optical zoom (e.g., Hitchcock zoom) in cooperation with spatial objects, enabling the production of videos with very cool visual effects, thereby enhancing the viewing experience of virtual-real fusion and attracting more users. In addition, the present application can greatly reduce the hardware cost of supporting optical zoom while ensuring clarity, lower the hardware threshold, and can be used by mobile phones, ordinary cameras, and professional cameras. It is simple to install, deploy, and operate, and reduces the user's usage threshold, attracting more video production users. At the same time, the spatial object can also assist in matte extraction and camera movement.
[0199] As Figure 10 shown, a device for determining the delay between a camera and a capture device is provided. The capture device is used to synchronously collect the capture device's attitude in the first coordinate system when the camera captures an image. The device may include an image attitude acquisition module 101, a first attitude sequence determination module 102, a second attitude sequence determination module 103, and a comparison module 104. Among them,
[0200] The image attitude acquisition module 101 is configured to obtain the curtain wall pictures of the curtain wall taken by the camera at different positions, as well as the first capture device attitude synchronously collected by the capture device, obtain the curtain wall video of the curtain wall taken by the camera during movement, and the capture device attitude sequence composed of the second capture device attitudes synchronously collected by the capture device during movement. The curtain wall is used to display markers, and the spatial coordinates of each marker in the second coordinate system have been determined in advance;
[0201] The first attitude sequence determination module 102 is configured to obtain a first attitude sequence according to each curtain wall picture and the curtain wall video. The first attitude sequence includes the first camera attitude of the camera in the second coordinate system when the camera captures each video frame in the curtain wall video;
[0202] The second attitude sequence determination module 103 is configured to obtain a second attitude sequence according to each curtain wall picture, each first capture device attitude, and the capture device attitude sequence. The second attitude sequence includes the second camera attitude of the camera in the second coordinate system when the capture device captures each second capture device attitude;
[0203] The comparison module 104 is configured to determine the delay between the camera and the capture device according to the first attitude sequence and the second attitude sequence.
[0204] The device of the embodiments of the present application can execute the method provided by the embodiments of the present application, and its implementation principle is similar. The actions performed by each module in the device of the embodiments of the present application correspond to the steps in the method of the embodiments of the present application. For the detailed function descriptions of each module of the device, reference can be specifically made to the descriptions in the corresponding methods shown above, and details are not repeated here.
[0205] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of a method for determining the delay between a camera and a capturer. Compared with the related art, it can achieve: on the basis that there is a rigid body relationship between the capturer and the camera and the frame rate of the capturer collecting its own attitude is the same as the frame rate of the camera collecting images, by obtaining curtain wall pictures taken by the camera at different positions and the first capturer attitude synchronously collected by the capturer, as well as the curtain wall video taken by the camera during movement and the sequence of capturer attitudes synchronously collected by the capturer during movement, where the first capturer attitude and the sequence of capturer attitudes are both the attitudes of the capturer itself in the first coordinate system, and for the curtain wall photographed by the curtain wall pictures and the curtain wall video, the spatial coordinates of each marker have been determined in the second coordinate system in advance. The present application can utilize the spatial coordinates and pixel coordinates of the markers in each curtain wall picture and curtain wall video, and use the PnP algorithm to obtain the first camera attitude of the camera in the second coordinate system when shooting each video frame of the curtain wall video. On the other hand, by using the curtain wall pictures, each first capturer attitude, and the sequence of capturer attitudes, the second camera attitude of the camera in the second coordinate system when the capturer collects a second capturer attitude can be obtained. By comparing the coincidence degree of the first attitude sequence and the second attitude sequence, the delay between the camera and the capturer can be determined, achieving the effect of accurately and quickly determining the delay.
[0206] In an alternative embodiment, an electronic device is provided, as Figure 11 shown. Figure 11 The electronic device 4000 shown includes a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.
[0207] The processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0208] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 11 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0209] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.
[0210] The memory 4003 is used to store the computer program for implementing the embodiments of the present application and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0211] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0212] The embodiments of the present application further provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0213] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than the illustrated or textually described order.
[0214] It should be understood that although the flowchart of the embodiments of the present application indicates each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless clearly stated herein, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.
[0215] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, other similar implementation means based on the technical idea of the present application also belong to the protection scope of the embodiments of the present application.
Claims
1. A method for determining the delay between a camera and a capture device, where the capture device is used to synchronously capture the pose of the capture device itself in a first coordinate system when the camera captures an image, characterized in that, The method includes: obtaining curtain wall pictures of the curtain wall taken by the camera at different positions, and the first capturer postures synchronously collected by the capturer, obtaining a curtain wall video of the curtain wall taken by the camera during movement, and a capturer posture sequence composed of the second capturer postures synchronously collected by the capturer during movement, where the curtain wall is used to display markers, and the spatial coordinates of each marker in the second coordinate system have been determined in advance; obtaining a first posture sequence according to each curtain wall picture and the curtain wall video, where the first posture sequence includes the first camera postures of the camera in the second coordinate system when the camera captures each video frame in the curtain wall video; obtaining a second posture sequence according to each curtain wall picture, each first capturer posture, and the capturer posture sequence, where the second posture sequence includes the second camera postures of the camera in the second coordinate system when the capturer captures each second capturer posture; determining the delay between the camera and the capturer according to the first posture sequence and the second posture sequence.
2. The method according to claim 1, characterized in that The determining the delay between the camera and the capturer according to the first posture sequence and the second posture sequence includes: taking one of the posture sequences in the first posture sequence and the second posture sequence as a reference sequence, and taking the other posture sequence as a target sequence; taking at least one camera posture in the reference sequence as a reference posture, and determining the camera posture with the highest similarity to the reference posture in the target sequence as the target posture of the reference posture; for each of the reference postures, determining the difference between the first time sequence and the second time sequence corresponding to the reference posture, where the first time sequence is the time sequence of the reference posture in the reference sequence, and the second time sequence is the time sequence of the target posture in the target sequence; determining the delay according to the differences corresponding to each of the reference postures.
3. The method according to claim 1, wherein The obtaining the first posture sequence according to each curtain wall picture and the curtain wall video includes: obtaining the internal parameters of the camera according to the spatial coordinates and pixel coordinates of the markers in each of the curtain wall pictures; obtaining the first posture sequence according to the curtain wall video and the internal parameters.
4. The method according to claim 1, wherein The obtaining the second posture sequence according to each curtain wall picture, each first capturer posture, and the capturer posture sequence includes: obtaining the offset of the camera relative to the capturer in the first coordinate system and the coordinate transformation relationship between the first coordinate system and the second coordinate system according to the internal parameters of the camera, each curtain wall picture, and the first capturer postures corresponding to each curtain wall picture; the internal parameters are determined according to the spatial coordinates and pixel coordinates of the markers in each of the curtain wall pictures; obtaining the second posture sequence according to the offset, the coordinate transformation relationship, and the capturer posture sequence.
5. The method according to claim 3, wherein The obtaining the first posture sequence according to the curtain wall video and the internal parameters includes: obtaining the first camera postures of the camera in the second coordinate system when the camera captures each video frame according to the internal parameters, the spatial coordinates, and the pixel coordinates of the markers in each video frame; Obtain the first pose sequence according to the time sequence of each video frame in the curtain wall video and the first camera pose corresponding to each video frame.
6. The method according to claim 4, wherein The obtaining of the offset of the camera relative to the capturer in the first coordinate system and the coordinate transformation relationship between the first coordinate system and the second coordinate system according to the internal parameters of the camera, each curtain wall picture, and the first capturer pose corresponding to each curtain wall picture includes: According to the internal parameters, the spatial coordinates and pixel coordinates of the markers in each curtain wall picture, obtain the third camera pose of the camera in the second coordinate system when shooting each curtain wall picture; Obtain the offset and coordinate transformation relationship according to the first capturer pose and the third camera pose corresponding to each curtain wall picture.
7. The method according to claim 6, characterized in that, The obtaining of the coordinate transformation relationship according to the first capturer pose and the third camera pose corresponding to each curtain wall picture includes: For each curtain wall picture, obtain the initial coordinate transformation relationship corresponding to the curtain wall picture according to the offset and the first capturer pose and the third camera pose corresponding to the curtain wall picture; Fit the initial coordinate transformation relationships corresponding to all curtain wall pictures to obtain the coordinate transformation relationship.
8. The method according to any one of claims 1-7, characterized in that, The marker is a two-dimensional code, and the two-dimensional code is used to record the spatial coordinates of the two-dimensional code in the second coordinate system.
9. A device for determining the delay between a camera and a capture device, the capture device being configured to synchronously capture the attitude of the capture device itself in a first coordinate system when the camera acquires an image, characterized in that, The device includes: An image pose acquisition module, configured to obtain curtain wall pictures of the curtain wall taken by the camera at different positions, and the first capturer pose synchronously acquired by the capturer, obtain a curtain wall video of the curtain wall taken by the camera during movement, and a capturer pose sequence composed of the second capturer poses synchronously acquired by the capturer during movement. The curtain wall is used to display markers, and the spatial coordinates of each marker in the second coordinate system have been determined in advance; A first pose sequence determination module, configured to obtain a first pose sequence according to each curtain wall picture and the curtain wall video. The first pose sequence includes the first camera pose of the camera in the second coordinate system when shooting each video frame in the curtain wall video; A second pose sequence determination module, configured to obtain a second pose sequence according to each curtain wall picture, each first capturer pose, and the capturer pose sequence. The second pose sequence includes the second camera pose of the camera in the second coordinate system when the capturer acquires each second capturer pose; A comparison module, configured to determine the delay between the camera and the capturer according to the first pose sequence and the second pose sequence.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-8.
Citation Information
Patent Citations
Low-cost motion capture method based on visual marker
CN111091587A
Camera and method for introducing light pulses in an image stream
CN112995455A