Real-time pose estimation method and device for studio multi-view infrared camera, equipment, medium and program product

CN122550684APending Publication Date: 2026-08-11BEIJING YOUKU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-03
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]在上述通过算法处理来获取摄像机位姿过程中,要想获取红外反射点准确的三维坐标,就需要有较多的红外摄像头看到该红外反射点,由此,为了保证影棚某些区域定位的鲁棒性,需要在现场人工调整这些红外摄像头,每次调整都会经历人工重新标定、验证覆盖有效性等验证流程,或者说,当发现影棚某个区域定位效果不好的时候,就需要调整更多红外摄像头对准这个区域, 但是目前在虚拟拍摄影棚中的红外摄像头并不能定量控制调整姿态,需要人工手动调整,因此,当人工调整红外摄像头姿态的时候,并不能实时知道调整了多少、符不符合最理想的调整结果,所以只能调整一次,重新验证一次,极大的浪费了时间,效率较低

Benefits of technology

[0017] According to various aspects of this disclosure, by matching the feature points and descriptors extracted from the current frame in real time with the feature points and descriptors extracted from the feature points of the previous frame during the process of adjusting the pose of the infrared camera, the relative pose of the current frame with respect to the previous frame is determined using the matched feature points, and then the current pose of the infrared camera is estimated using this relative pose. This method can achieve real-time and efficient determination of the pose of the infrared camera. In this way, the pose determined in real time can be used to verify whether the infrared camera has been adjusted to the ideal result, which helps to save the time of adjusting the pose of the infrared camera in the studio and improve the adjustment efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550684A_ABST
    Figure CN122550684A_ABST
Patent Text Reader

Abstract

This disclosure relates to a method, apparatus, device, medium, and program product for real-time attitude estimation of a multi-view infrared camera in a studio. The method includes: acquiring the current frame captured in real-time by the infrared camera during the attitude adjustment process; extracting feature points and corresponding descriptors from the current frame; determining matching feature points between the current frame and the previous frame based on the descriptors corresponding to the feature points in the current frame and the descriptors corresponding to the feature points in the previous frame, where the previous frame is a frame captured by the infrared camera before the current frame; determining the relative attitude of the infrared camera in the current frame relative to the previous frame based on the matching feature points between the current frame and the previous frame; and determining the current attitude of the infrared camera based on the relative attitude of the infrared camera in the current frame relative to the previous frame. Therefore, the attitude of the infrared camera can be estimated in real-time during the adjustment process, saving time spent adjusting the infrared camera in the studio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of virtual shooting, and in particular to a method, apparatus, device, medium and program product for real-time attitude estimation of a multi-view infrared camera in a studio. Background Technology

[0002] In virtual production, camera positioning is a crucial element in achieving the same effect as real-world shooting. Precise camera positioning ensures that the content captured by the camera at any position and angle is consistent with the content captured in the real world. In other words, a core aspect of virtual production is to obtain the accurate position and posture of the camera in real time within the studio performance area.

[0003] Currently, in virtual studio shooting scenarios, camera pose localization generally uses optical positioning, which often employs passive positioning devices to provide high-precision positioning results for the camera. This involves installing multiple infrared reflective points (actively emitting / passively reflecting) on ​​the camera, and then capturing the positions of these infrared reflective points using multiple multi-view infrared cameras installed at fixed locations in the studio (such as the studio ceiling). The camera pose is then obtained through algorithm processing. Specifically, the infrared cameras can obtain the two-dimensional coordinates of these infrared reflective points, and then the three-dimensional coordinates of the reflective points can be obtained through triangulation (i.e., triangulation method). Finally, the pose of the rigid body camera can be estimated using the three-dimensional coordinates of multiple reflective points. This pose obtained through optical positioning has high accuracy and low latency.

[0004] In the process of obtaining camera pose through algorithmic processing, to obtain accurate 3D coordinates of infrared reflection points, a large number of infrared cameras are needed to see the infrared reflection point. Therefore, to ensure the robustness of positioning in certain areas of the studio, these infrared cameras need to be manually adjusted on-site. Each adjustment involves manual recalibration and verification of coverage effectiveness. In other words, when the positioning effect in a certain area of ​​the studio is found to be poor, more infrared cameras need to be adjusted to aim at that area. However, currently, the infrared cameras in the virtual shooting studio cannot be quantitatively controlled to adjust their posture, requiring manual adjustment. Therefore, when manually adjusting the posture of the infrared cameras, it is not possible to know in real time how much has been adjusted or whether it meets the ideal adjustment result. So, it can only be adjusted once and then re-verified, which is a huge waste of time and has low efficiency. Summary of the Invention

[0005] In view of this, this disclosure proposes a method, apparatus, device, medium and program product for real-time attitude estimation of multi-view infrared cameras in a studio. Compared with the existing process of adjusting and verifying once, the solution of this disclosure can estimate the attitude of the infrared camera in real time during the adjustment process, saving the time of adjusting the infrared camera in the studio.

[0006] According to one aspect of this disclosure, a real-time attitude estimation method for a multi-view infrared camera is provided. The infrared camera is set at a fixed position in a virtual shooting studio. The method includes: acquiring the current frame captured in real time by the infrared camera during the attitude adjustment process; extracting feature points and corresponding descriptors in the current frame, wherein the descriptors represent feature vectors of a local region surrounding the feature points; determining matching feature points between the current frame and the previous frame based on the descriptors corresponding to the feature points in the current frame and the descriptors corresponding to the feature points in the previous frame, wherein the previous frame is a frame captured by the infrared camera before the current frame; determining the relative attitude of the infrared camera in the current frame relative to the previous frame based on the matching feature points between the current frame and the previous frame; and determining the current attitude of the infrared camera based on the relative attitude of the infrared camera in the current frame relative to the previous frame.

[0007] In one possible implementation, determining the current pose of the infrared camera based on its relative pose in the current frame relative to the previous frame includes: determining the orientation corresponding to the current frame based on its relative pose, wherein the orientation of the current frame represents the orientation of the infrared camera when acquiring the current frame; and determining whether a keyframe corresponding to the orientation is recorded in a preset two-dimensional grid, wherein the two-dimensional grid is used to record the acquired keyframes according to the orientation of the infrared camera, and the two-dimensional grid is obtained by measuring the three-dimensional shooting range of the infrared camera. The infrared camera's current pose is determined by the following methods: 1) Two-dimensional unfolding and rasterization; 2) When a keyframe corresponding to the orientation is recorded in the two-dimensional raster, the current pose of the infrared camera is determined based on the feature points in the current frame and the feature points in the keyframe; 3) When no keyframe corresponding to the orientation is recorded in the two-dimensional raster, the current frame is recorded as a keyframe in the two-dimensional raster according to the orientation corresponding to the current frame, and the current pose of the infrared camera is determined based on the relative pose of the infrared camera in the current frame relative to the previous frame. The position of the keyframe in the two-dimensional raster represents the orientation of the infrared camera when acquiring the keyframe.

[0008] In one possible implementation, determining the current pose of the infrared camera based on feature points in the current frame and feature points in the key frame includes: determining matching feature points between the current frame and the key frame based on descriptors corresponding to feature points in the current frame and descriptors corresponding to feature points in the key frame; determining the relative pose of the infrared camera in the current frame relative to the key frame based on the matching feature points between the current frame and the key frame; and determining the current pose of the infrared camera based on the relative pose of the infrared camera in the current frame relative to the key frame and the pose corresponding to the key frame.

[0009] In one possible implementation, the method further includes: fixing the pose of the first keyframe in the two-dimensional grid, optimizing the poses of each keyframe in the two-dimensional grid in a specified order starting from the first keyframe, where the first keyframe is the initial frame; adjusting the position of each keyframe in the two-dimensional grid according to the optimized poses of each keyframe; and after adjusting the position of each keyframe in the two-dimensional grid, if there are two or more keyframes at the same position, retaining only the keyframe with the earliest acquisition time at that position.

[0010] In one possible implementation, optimizing the pose of each keyframe in the two-dimensional grid in a specified order starting from the first keyframe includes: for the m-th keyframe in the specified order, determining the feature points that match between the m-th keyframe and at least one surrounding keyframe based on the descriptors of feature points in the m-th keyframe and the descriptors of feature points in at least one surrounding keyframe, where m is a positive integer and the m-th keyframe is a keyframe with an optimized pose; determining the relative pose of at least one surrounding keyframe relative to the m-th keyframe based on the feature points that match between the m-th keyframe and at least one surrounding keyframe; and determining the optimized pose of each of the at least one surrounding keyframe based on the relative poses of the at least one surrounding keyframe relative to the m-th keyframe and the optimized pose of the m-th keyframe.

[0011] In one possible implementation, the method further includes: if no feature points and their descriptors are extracted in the current frame, or if there are no matching feature points between the current frame and the previous frame, or the number of matching feature points is less than a specified number, adjusting the attitude of the infrared camera in a previously oriented direction; during the adjustment of the infrared camera's attitude in the previously oriented direction, acquiring a new current frame acquired in real time by the infrared camera, and extracting a global descriptor of the new current frame, the global descriptor representing the global feature vector of the new current frame; determining whether there is a target keyframe in the two-dimensional grid that matches the new current frame based on the global descriptor of the new current frame and the global descriptors of each keyframe in the two-dimensional grid; if there is a target keyframe in the two-dimensional grid that matches the new current frame, determining the current attitude of the infrared camera based on the feature points in the new current frame and the feature points in the target keyframe.

[0012] In one possible implementation, the virtual shooting studio further includes a camera for capturing video footage; if the positioning effect of the infrared camera in the virtual shooting studio on the camera in the designated area is lower than ideal, the attitude of the infrared camera in the virtual shooting studio is adjusted, and the attitude of the infrared camera during the adjustment process is estimated using the method described above. The adjusted infrared camera is then used to locate the camera in the designated area.

[0013] According to another aspect of this disclosure, a real-time attitude estimation device for a multi-view infrared camera is provided. The infrared camera is set at a fixed position in a virtual shooting studio. The device includes: an acquisition module, used to acquire the current frame captured in real time by the infrared camera during the attitude adjustment process; an extraction module, used to extract feature points and corresponding descriptors in the current frame, the descriptors representing feature vectors of a local area surrounding the feature points; a matching module, used to determine matching feature points between the current frame and the previous frame based on the descriptors corresponding to the feature points in the current frame and the descriptors corresponding to the feature points in the previous frame, the previous frame being a frame captured by the infrared camera before the current frame; a relative attitude determination module, used to determine the relative attitude of the infrared camera in the current frame relative to the previous frame based on the matching feature points between the current frame and the previous frame; and an attitude determination module, used to determine the current attitude of the infrared camera based on the relative attitude of the infrared camera in the current frame relative to the previous frame.

[0014] According to another aspect of this disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above-described method.

[0015] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the above-described method.

[0016] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0017] According to various aspects of this disclosure, by matching the feature points and descriptors extracted from the current frame in real time with the feature points and descriptors extracted from the feature points of the previous frame during the process of adjusting the pose of the infrared camera, the relative pose of the current frame with respect to the previous frame is determined using the matched feature points, and then the current pose of the infrared camera is estimated using this relative pose. This method can achieve real-time and efficient determination of the pose of the infrared camera. In this way, the pose determined in real time can be used to verify whether the infrared camera has been adjusted to the ideal result, which helps to save the time of adjusting the pose of the infrared camera in the studio and improve the adjustment efficiency.

[0018] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0019] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.

[0020] Figure 1 A schematic plan view of a virtual filming studio according to an embodiment of the present disclosure is shown.

[0021] Figure 2 A flowchart illustrating a real-time attitude estimation method for a multi-view infrared camera according to an embodiment of the present disclosure is shown.

[0022] Figure 3 A framework diagram of an infrared camera real-time attitude estimation system according to an embodiment of the present disclosure is shown.

[0023] Figure 4 A block diagram of a multi-view infrared camera real-time attitude estimation apparatus according to an embodiment of the present disclosure is shown.

[0024] Figure 5 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. Detailed Implementation

[0025] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0026] As used herein, the terms “comprising,” “including,” “having,” or variations thereof are open-ended and include one or more of the stated features, integrals, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, integrals, elements, steps, components, functions, or groups thereof.

[0027] When an element is referred to as “connected,” “coupled,” “responding,” or a variation thereof relative to another element, it may be directly connected, coupled, or responding to another element, or there may be an intermediate element present.

[0028] Although the terms first, second, third, etc., may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Therefore, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.

[0029] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0030] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0031] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions.

[0032] Figure 1 A schematic plan view of a virtual filming studio according to an embodiment of the present disclosure is shown, as follows: Figure 1As shown, the virtual shooting studio includes an LED screen 01, a shooting area 02 in front of the screen, multiple infrared cameras 03 fixedly arranged around the shooting area, and cameras 04 used for shooting within the shooting area. The shooting area may contain real sets and actors. The LED screen 01 can display a virtual background, and the cameras 04 can capture video footage of the virtual background displayed on the LED screen 01 and the sets and actors in front of the screen, thus achieving virtual shooting. Multiple infrared reflectors can be set on the cameras 04. During the shooting process, the multiple infrared cameras 03 can be used to optically position the cameras 04. The infrared cameras 03 can be multi-view infrared cameras (i.e., infrared cameras capable of adjusting their viewing angle or posture). The multiple infrared cameras 03 are typically fixedly installed at the top of the shooting area (i.e., the top of the virtual shooting studio). It should be understood that this embodiment does not limit the number or layout of the infrared cameras in the virtual shooting studio.

[0033] As mentioned above, in a virtual shooting studio, when the infrared camera's posture is manually adjusted, it is impossible to know in real time how much it has been adjusted or whether it meets the ideal adjustment result. Therefore, it can only be adjusted once and then recalibrated and verified, which is a huge waste of time and has low efficiency. Therefore, this disclosure proposes a real-time posture estimation method for multi-view infrared cameras, which can determine the posture of the infrared camera in real time and efficiently during the adjustment process. This allows the real-time determined posture to be used to verify whether the adjusted infrared camera has reached the ideal posture, which helps to save time in the studio for adjusting the infrared camera's posture.

[0034] As mentioned above, when it is found that the positioning effect in a certain area of ​​the studio is not good, it is necessary to adjust more infrared cameras to be aimed at this area. In other words, when the positioning effect of the infrared cameras in the virtual shooting studio on the cameras in the specified area is lower than the ideal effect, the attitude of the infrared cameras in the virtual shooting studio can be adjusted, and the real-time attitude estimation method proposed in the embodiments of this disclosure can be used to estimate the attitude of the infrared cameras during the adjustment process. The adjusted infrared cameras are used to locate the cameras in the specified area (such as the optical positioning mentioned above). The designated area may be the site area where the camera needs to be precisely located (or the site area where the camera is located during the shooting process, such as the center of the site). The positioning effect of the camera within the designated area can be characterized as the coverage of the designated area (such as the number of infrared cameras that can be observed in the designated area). The ideal effect can be understood as the coverage of the designated area that the user expects to achieve. The posture of at least some of the infrared cameras in the virtual shooting studio can be adjusted manually (e.g., manually bend the infrared cameras to face the designated area). However, as mentioned above, it is not possible to know how much has been adjusted in real time during the manual adjustment of the infrared cameras. Therefore, the method of this embodiment can be used to estimate the posture of the infrared cameras in real time during the adjustment process. Subsequently, the real-time estimated posture can be used to verify whether the adjusted camera meets the ideal situation. For example, the posture of each infrared camera in the site (including the posture of the adjusted infrared camera) can be combined to determine whether the coverage of the infrared cameras in the studio to the designated area has reached the ideal situation. For another example, the ideal posture of the infrared camera can be determined in advance, and it can be determined whether the current posture of the infrared camera has reached the ideal posture. It should be understood that the above verification process is one of the possible real-time methods provided by the embodiments of this disclosure. In fact, those skilled in the art can customize the subsequent verification process, and the embodiments of this disclosure do not limit this. The main focus of the embodiments of this disclosure is to protect the real-time attitude estimation of the infrared camera during the adjustment process.

[0035] The real-time attitude estimation method of this disclosure can be deployed on various terminal devices through software or hardware modifications. The terminal devices involved in this disclosure can refer to devices with wireless and / or wired connection functions. Wireless connection means that they can connect to other devices via wireless methods such as Wi-Fi and Bluetooth. The terminal devices involved in this disclosure can also communicate with other devices via wired connection functions. The terminal devices involved in this disclosure can be touchscreen, non-touchscreen, or screenless. Touchscreen devices can be controlled by clicking or swiping on the display screen using fingers or styluses. Non-touchscreen devices can connect to input devices such as mice, keyboards, and touch panels to control the terminal device. Screenless devices can be, for example, screenless Bluetooth speakers. For example, the terminal devices in this application can include, but are not limited to, user equipment (UE), mobile devices, user terminals, terminals, handheld devices, tablet computers, laptops, PDAs, and computing devices.

[0036] The real-time attitude estimation method of this disclosure can also be deployed on a server, which can be located in the cloud or locally, and can be a physical device or a virtual device, such as a virtual machine or container, with wireless communication capabilities. These wireless communication capabilities can be configured in the server's chip (system) or other components. This can refer to a device with wireless connectivity, meaning it can connect to other servers or terminal devices via wireless connections such as Wi-Fi or Bluetooth. The server involved in this disclosure can also have wired communication capabilities. For example, the server in this disclosure can be located in the cloud, communicating with terminal devices, receiving real-time frames captured by an infrared camera from the terminal device, and using the real-time attitude estimation method deployed on the server to determine the current attitude of the infrared camera based on the real-time captured frames, returning the result to the terminal device for subsequent real-time verification of whether the ideal attitude has been achieved.

[0037] Figure 2 A flowchart illustrating a real-time pose estimation method for a multi-view infrared camera according to an embodiment of the present disclosure is shown. As described above, the infrared camera is positioned at a fixed location in a virtual shooting studio, such as... Figure 2 As shown, the method includes steps S11 to S15.

[0038] In step S11, during the process of adjusting the attitude of the infrared camera, the current frame captured in real time by the infrared camera is obtained.

[0039] As described above, multiple infrared cameras can be fixedly set up in a virtual shooting studio. In actual situations, only some infrared cameras may need to be adjusted. Therefore, the frames captured in real time by the adjusted infrared camera can be obtained, and the method of this embodiment can be used to realize real-time pose estimation of the infrared camera during the adjustment process.

[0040] It should be understood that in practice, it may be that a single person adjusts the posture of the infrared cameras one by one, or multiple people adjust the posture of multiple infrared cameras simultaneously. Regardless of the former or the latter, the posture of each camera can be estimated in real time using the method of the embodiments of this disclosure. For example, when the posture of multiple infrared cameras is adjusted simultaneously, the method of the embodiments of this disclosure can be executed in parallel by multiple threads to achieve real-time posture estimation of multiple infrared cameras. The embodiments of this disclosure do not limit this.

[0041] In step S12, feature points and corresponding descriptors in the current frame are extracted. The descriptors represent the feature vectors of the local region surrounding the feature points.

[0042] In practical applications, feature extractors known in the art (such as SuperPoint (a deep learning-based image local feature extraction method) and Fast (Features from Accelerated Segment Test, a fast feature point extraction algorithm for corner detection) can be used to extract feature points (pixel locations in the image that are significant and unique, such as corners and key points in the frame) and corresponding descriptors in the current frame. This disclosure does not limit the scope of these applications. Feature points describe the location of features, and descriptors represent the content of features (i.e., the result of quantization encoding of the local region surrounding the feature point, converting the image information of the local region surrounding the feature point into a vector).

[0043] It should be understood that the feature extractor can extract one or more feature points and a descriptor for each feature point from the current frame, and this disclosure does not limit the scope of the embodiments.

[0044] In step S13, based on the descriptor corresponding to the feature point in the current frame and the descriptor corresponding to the feature point in the previous frame, the matching feature points between the current frame and the previous frame are determined. The previous frame is a frame captured by the infrared camera before the current frame.

[0045] It should be understood that the feature points and descriptors in the previous frame can also be extracted using the feature extractor described above. Assume the current frame and the previous frame are I... i , I i-1 The extracted feature points are p i , p i-1 The corresponding descriptor is f i , fi-1 By calculating the similarity between descriptors in two frames, we can obtain the matching feature points between the two frames. For example, we can target a specific feature point p in the current frame. i Calculate the feature point p i Description f i Compared with each feature point p in the previous frame i-1 descriptor f i-1 The similarity between the two frames is used to determine the feature points in the previous frame whose similarity exceeds a specified threshold, and these feature points are then used to match the current frame's feature points. It should be understood that the specific point matched between two frames can be understood as the pixel position of the same spatial point (such as a corner point on the same object) in the two frames.

[0046] In step S14, the relative pose of the infrared camera in the current frame relative to the previous frame is determined based on the feature points matched between the current frame and the previous frame.

[0047] It is understandable that the interval between the current frame and the previous frame captured by the infrared camera is small (e.g., one frame is captured every 50 milliseconds), and the movement of the infrared camera between two frames is also small. Therefore, it can be assumed that the translation of the infrared camera between two frames is 0, and the same spatial point represented by the matching feature points between two frames approximately satisfies pure rotational motion. In this case, the matching feature points between two frames can approximately satisfy... H represents the homography matrix. Therefore, the homography matrix of the current frame relative to the previous frame can be determined by using the feature points matched between the current frame and the previous frame. Then, by decomposing the homography matrix, the rotation matrix can be obtained, which can characterize the relative pose of the current frame relative to the previous frame.

[0048] The homography matrix described above describes the projection transformation of spatial points within the field of view of an infrared camera between two frames. In other words, for any spatial point in three-dimensional space, its pixel coordinates in the current frame (current camera view) and the previous frame (previous camera view) satisfy the following... Therefore, the solution can be obtained by using the matched feature points between two frames. We obtain the homography matrix. Given the camera intrinsic parameter K of the infrared camera, and knowing that the homography matrix is ​​caused by pure rotation, the rotation matrix can be expressed as... R represents the rotation matrix. Thus, the rotation matrix can be obtained by decomposing the homography matrix. For example, the rotation matrix can be obtained by performing singular value decomposition (SVD) on the homography matrix. This disclosure does not limit the scope of the embodiments.

[0049] In step S15, the current pose of the infrared camera is determined based on the relative pose of the infrared camera in the current frame relative to the previous frame.

[0050] In practical applications, the initial pose of the infrared camera before adjustment is known. For example, the pose of the first frame captured by the infrared camera (i.e., the initial frame captured when the pose of the infrared camera is adjusted) is a known initial pose. Then, starting from the second frame, the relative pose of the second frame with respect to the first frame can be calculated through the above steps S11 to S14. Then, the pose of the second frame can be obtained by using the pose of the first frame and the relative pose of the second frame with respect to the first frame. Similarly, the pose of the infrared camera in the current frame can be obtained by using the relative pose of the current frame with respect to the previous frame and the pose calculated in the previous frame. In other words, the relative pose of the current frame with respect to the previous frame, the relative pose of the previous frame with respect to the frame before that, and so on up to the relative pose of the second frame with respect to the initial frame. Then, combined with the pose of the initial frame, the current pose of the infrared camera is determined, which means that the pose of the infrared camera is estimated in real time.

[0051] In practical applications, after estimating the infrared camera's attitude in real time, the estimated attitude can be sent to the subsequent verification process. The verification process determines whether the infrared camera's attitude has reached the ideal attitude (or whether the desired adjustment result has been achieved). As described above, this disclosure protects the real-time attitude estimation of the infrared camera during the adjustment process. It does not limit the specific implementation of the verification process. Those skilled in the art can customize and design any verification process, and this disclosure does not limit this.

[0052] According to the method of this disclosure, during the process of adjusting the pose of the infrared camera, the feature points and descriptors extracted from the current frame acquired in real time by the infrared camera are matched with the feature points and descriptors extracted from the feature points of the previous frame. The relative pose of the current frame with respect to the previous frame is determined by the matched feature points, and then the current pose of the infrared camera is estimated by using the relative pose. This method can achieve real-time and efficient determination of the pose of the infrared camera. In this way, the pose determined in real time can be used to verify whether the infrared camera has been adjusted to the ideal result, which helps to save the time of adjusting the pose of the infrared camera in the studio and improve the adjustment efficiency.

[0053] It is understandable that step S14, which calculates the homography matrix based on the matched feature points between two frames to infer the relative pose, may contain errors. Therefore, in step S15, directly using the relative pose between two frames and the pose of the previous frame to determine the pose of the current frame will also contain errors, and these errors accumulate frame by frame. Thus, to reduce the accumulated error in pose estimation, in some embodiments, step S15 determines the current pose of the infrared camera based on its relative pose in the current frame relative to the previous frame, including:

[0054] Step S151: Determine the orientation of the current frame based on the relative pose of the current frame with respect to the previous frame. The orientation of the current frame represents the orientation of the infrared camera when acquiring the current frame.

[0055] Step S152: Based on the orientation of the current frame, determine whether the key frame corresponding to the orientation is recorded in the preset two-dimensional grid. The two-dimensional grid is used to record the key frames collected according to the orientation of the infrared camera. The two-dimensional grid is obtained by two-dimensionally unfolding and rasterizing the three-dimensional shooting range of the infrared camera.

[0056] Step S153: If the key frame corresponding to the orientation is recorded in the two-dimensional grid, determine the current pose of the infrared camera based on the feature points in the current frame and the feature points in the key frame.

[0057] Step S154: If no keyframe corresponding to the orientation is recorded in the two-dimensional grid, the current frame is recorded as a keyframe in the two-dimensional grid according to the orientation corresponding to the current frame. The current orientation of the infrared camera is determined according to the relative orientation of the infrared camera in the current frame relative to the previous frame. The position of the keyframe in the two-dimensional grid represents the orientation of the infrared camera when acquiring the keyframe.

[0058] In step S151, as described above, given the attitude of the initial frame, the relative attitude of the second frame with respect to the initial frame, and the relative attitude of the current frame with respect to the previous frame, the attitude corresponding to the current frame can be calculated. The attitude corresponding to the current frame can be represented as a rotation matrix. Optionally, the orientation corresponding to the current frame can be directly represented as the rotation matrix corresponding to the current frame, or the rotation matrix can be converted into Euler angles (i.e., yaw angle, pitch angle, roll angle) as the orientation corresponding to the current frame. Alternatively, the vector of the optical axis direction of the infrared camera can be extracted from the rotation matrix as the orientation corresponding to the current frame. This embodiment of the present disclosure does not limit this.

[0059] In step S152, it can be understood that the infrared camera captures images in real time during the adjustment process. Through the above steps S11 to S15, the attitude of the infrared camera in each frame can be calculated in real time. In order to reduce the cumulative error of attitude estimation, it is necessary to obtain the information of historically acquired frames (such as attitude, feature points, descriptors, etc.). However, if the information of each frame acquired by the infrared camera is stored, it will occupy a lot of memory space and affect the calculation efficiency. Therefore, this embodiment of the present disclosure uses a two-dimensional grid to record some historically acquired key frames and store the information of these key frames. This can greatly reduce the amount of frame information that needs to be stored, and can also use the information of key frames recorded in the two-dimensional grid to reduce the estimation error of the attitude corresponding to the current frame.

[0060] As described above, the infrared camera's position in space is fixed, but its orientation can change, meaning it can rotate in space. This is equivalent to the infrared camera moving back and forth within a 360° sphere. Therefore, the spherical range of motion of the infrared camera is its three-dimensional shooting range, or its orientation in space. This spherical range can be unfolded into a two-dimensional plane (similar to unfolding the surface of a globe into a planar diagram, i.e., two-dimensional unfolding, for example, unfolding along the direction corresponding to the center of the camera's image (i.e., the optical axis direction) into a two-dimensional plane). This unfolded two-dimensional plane is then divided into grids (i.e., rasterized), resulting in the aforementioned two-dimensional grid. Since this embodiment estimates a rotation matrix, the planar shooting range of the infrared camera can be described using this two-dimensional grid, thus improving the speed of subsequent keyframe lookup. The size of the grid cells can be customized according to actual conditions, and this embodiment does not impose any limitations on this.

[0061] It should be understood that the range of a two-dimensional grid can be interpreted as a planar range obtained by mapping the spherical range of the infrared camera's motion in two dimensions. Therefore, the orientation of the infrared camera in space corresponds to (i.e., mapped to, or falls into) a specific cell in the two-dimensional grid. Each cell represents an orientation within a certain range. If the orientation of the current frame falls within the orientation range covered by a cell, it can be considered that the orientation of the current frame falls within that cell. Once a keyframe is captured in each cell, no new keyframes are recorded. Therefore, each cell contains only one keyframe, which is the earliest frame to fall within the range of that cell. Whenever the infrared camera captures a current frame, it can... The system searches the 2D grid to see if a keyframe has been recorded for the orientation of the current frame (i.e., whether the cell corresponding to the orientation of the current frame is empty). If a keyframe has been recorded (i.e., the cell corresponding to the orientation is not empty), it can be assumed that the infrared camera has been to this orientation (or has been to the same position). At this point, step S153 can be executed to determine the orientation of the infrared camera using the keyframe and the feature points in the current frame. Compared with using the previous frame and the current frame, this method can offset some of the cumulative error from the keyframe to the previous frame, thereby improving the accuracy of the infrared camera's orientation estimation.

[0062] If the orientation corresponding to the current frame is not recorded as a keyframe in the two-dimensional grid (i.e., the grid corresponding to the orientation is empty), it can be assumed that the infrared camera has not generated this orientation (or has not been oriented towards this position). In this case, step S154 above can be executed, that is, according to the orientation corresponding to the current frame, the current frame is recorded as a keyframe and recorded (i.e. inserted) into the corresponding grid in the two-dimensional grid (the grid into which the keyframe is inserted is also the position of the keyframe in the two-dimensional grid). At the same time, the pose corresponding to the newly inserted keyframe (i.e., the pose calculated using the previous frame), the extracted feature points and their descriptors, frame identifiers (such as the timestamp of the acquisition time, frame number, etc.) and other information can be stored. It should be understood that in this case, the current pose of the infrared camera is the pose of the current frame calculated in step S151 above using the relative pose of the current frame with respect to the previous frame. That is, the current pose of the infrared camera is determined directly based on the relative pose of the infrared camera in the current frame with respect to the previous frame.

[0063] In some embodiments, step S153 above, determining the current pose of the infrared camera based on feature points in the current frame and feature points in the key frame, may include:

[0064] Step S1531: Determine the matching feature points between the current frame and the key frame based on the descriptors corresponding to the feature points in the current frame and the descriptors corresponding to the feature points in the key frame.

[0065] Step S1532: Determine the relative pose of the infrared camera in the current frame relative to the key frame based on the feature points matched between the current frame and the key frame.

[0066] Step S1533: Determine the current pose of the infrared camera based on the relative pose of the infrared camera in the current frame relative to the key frame, and the pose corresponding to the key frame.

[0067] In step S1531, it is assumed that the feature points extracted from the keyframe are p. k The corresponding descriptor is f k The feature points extracted in the current frame are p i The corresponding descriptor is f i Based on the similarity between the descriptors of the keyframe and the current frame, matching feature points can be found in both frames. In other words, matching feature points between two frames can be obtained by calculating the similarity between their descriptors. For example, for a specific feature point p in the current frame... i Calculate the feature point p i descriptor f i With each feature point p in the keyframe k descriptor f kThe similarity between the features is used to select feature points in the keyframes whose similarity exceeds a specified threshold as the feature points that match the current feature point.

[0068] In step S1532, optionally, the relative pose of the infrared camera in the current frame relative to the key frame can be determined by estimating the homography matrix to determine the relative pose based on the feature points matched between the current frame and the key frame, as described in step S14 above. That is, the homography matrix can be determined first based on the feature points matched between the current frame and the key frame, and then the homography matrix can be decomposed to obtain the relative pose of the current frame relative to the key frame. Considering that the method of determining the relative pose by estimating the homography matrix in step S14 above is based on the assumption that the translation between the two frames is 0, and that the accuracy requirement for determining the relative pose using the two frames above is not particularly high, and that the key frame and the current frame are not necessarily adjacent (may be separated by a long time), and also in order to achieve more accurate pose estimation, it is optional to use the method proposed in the paper "Direct Optimization of Frame-to-Frame Rotation" to calculate the rotation matrix (i.e., relative pose, i.e., relative rotation) to obtain a more accurate relative pose. Specifically, the constraint equation shown in formula (1) can be constructed using the feature points matched between the current frame and the key frame. The basic principle is that when there is translational motion, the normal of the polar plane formed by the camera center point of the infrared camera and the spatial point is located in the same plane, and pure rotational motion also satisfies this constraint. Therefore, whether it is pure rotation or rotation + translational motion, it satisfies the constraint that the minimum eigenvalue of matrix M is 0. Then, R can be solved according to this constraint.

[0069] (1)

[0070] in, This represents the rotation matrix of the current frame relative to the keyframe, where n represents the total number of feature points in the current frame. Represents the j-th feature point in the current frame. This represents the feature point in the keyframe that matches the j-th feature point in the current frame. represent transpose, Represents the smallest eigenvalue of matrix M. Representative solution The rotation matrix R that minimizes (e.g., is 0) is then solved to find the rotation matrix R that minimizes (e.g., is 0). The rotation matrix R at its minimum value is also the relative pose of the current frame with respect to the keyframe. It should be understood that those skilled in the art can use any known optimization algorithm in the art, such as the branch and bound method, to solve the above formula (1) to obtain the rotation matrix R, and this embodiment of the disclosure does not limit this. A more accurate relative pose can be obtained using this method.

[0071] In step S1533, it should be understood that, given the relative pose of the current frame with respect to the key frame and the pose corresponding to the key frame, the pose corresponding to the current frame can be obtained, which is also the current pose of the infrared camera.

[0072] In this embodiment of the disclosure, recording keyframes using a two-dimensional grid can save a lot of memory space and reduce the impact on computational efficiency. At the same time, using the current frame and keyframes to calculate the current pose of the infrared camera can reduce the impact of cumulative errors and improve the accuracy of the calculated pose.

[0073] As described above, the keyframes in the aforementioned two-dimensional grid are recorded in the two-dimensional grid based on the orientation determined by the relative pose of the current frame with respect to the previous frame. Therefore, the pose (i.e., orientation) corresponding to the keyframes recorded in the two-dimensional grid may also contain errors. To further improve the pose estimation accuracy, pose optimization can be performed on the keyframes in the two-dimensional grid. Thus, in some embodiments, the method may further include:

[0074] Step S21: By fixing the pose of the first keyframe in the two-dimensional grid, optimize the pose of each keyframe in the two-dimensional grid in a specified order starting from the first keyframe. The first keyframe is the initial frame.

[0075] Step S22: Adjust the position of each keyframe in the two-dimensional grid according to the optimized pose of each keyframe in the two-dimensional grid.

[0076] Step S23: After adjusting the position of each keyframe in the two-dimensional grid, if there are two or more keyframes at the same position, only the keyframe with the earliest acquisition time is retained at that position.

[0077] It is understandable that keyframes in any position (i.e., any cell) of a 2D grid may also have keyframes in adjacent positions. The grid division in a 2D grid can be relatively dense. Therefore, the images of two adjacent frames in a 2D grid actually overlap to some extent. For example, there is some overlap between a keyframe at a certain position and the keyframe below it. Then, there will also be matching feature points and matching descriptors between a keyframe at a certain position and the keyframes at its adjacent positions. That is, the relative pose between the keyframe and the keyframes at its adjacent positions can be calculated. Thus, the relative pose between adjacent keyframes can be used to optimize the keyframes in the entire 2D grid.

[0078] In step S21, fixing the pose of the first keyframe (i.e., the initial frame) in the 2D grid, that is, fixing the position of the first keyframe in the 2D grid, helps to avoid overall positional shifts during optimization, which is equivalent to defining the position of the origin. Since the position of the first keyframe is fixed, the keyframes around the first keyframe can be optimized starting from the first keyframe, and then the optimized keyframes around the first keyframe can be used to optimize the keyframes around it. In this way, starting from the first keyframe, each keyframe in the 2D grid can be precisely optimized with the pose of the first keyframe as the starting point, and the pose of the keyframe at any position is constrained by the adjacent keyframes above, below, left, and right, thereby achieving global optimization of the pose of each keyframe in the 2D grid. The specified order can be customized according to the actual situation, as long as it is the order of optimization from the first keyframe to the surrounding keyframes. For example, if there are keyframes adjacent to the bottom and right sides of the first keyframe, after optimizing the bottom and right keyframes using the first keyframe, the right keyframe of the first keyframe can be used as the second keyframe and the bottom keyframe as the third keyframe in the order from left to right and from top to bottom. Then, the surrounding adjacent keyframes are optimized based on the second keyframe and the surrounding adjacent keyframes are optimized based on the third keyframe. The order of the surrounding keyframes is then determined in the order from left to right and from top to bottom. This disclosure does not limit the implementation of this embodiment.

[0079] In step S21, starting from the first keyframe, the poses of each keyframe in the 2D grid are optimized in a specified order, including:

[0080] Step S211: For the m-th keyframe in the specified order, based on the descriptor of the feature points in the m-th keyframe and the descriptor of the feature points in at least one surrounding keyframe, determine the feature points that match between the m-th keyframe and at least one surrounding keyframe respectively, where m is a positive integer and the m-th keyframe is a keyframe with optimized pose.

[0081] Step S212: Based on the feature points matched between the m-th keyframe and at least one surrounding keyframe, determine the relative pose of at least one surrounding keyframe relative to the m-th keyframe.

[0082] Step S213: Determine the optimized pose of each of the at least one surrounding keyframe based on the relative pose of each of the at least one surrounding keyframe with respect to the m-th keyframe and the optimized pose of the m-th keyframe.

[0083] In step S211, the m-th keyframe is the keyframe after pose optimization (if it is the first keyframe, then it is itself an optimized keyframe). At least one surrounding keyframe may include the keyframe adjacent to the right and / or bottom of the m-th keyframe (this allows optimization from the top left to the bottom right in the 2D grid), or it may be the keyframe adjacent to the left and / or top of the m-th keyframe (this allows optimization from the bottom right to the top left in the 2D grid). This embodiment does not limit the scope of the discussion. The similarity between the descriptor of the feature point in the m-th keyframe and the descriptors of feature points in each surrounding keyframe can be calculated to obtain the feature points that match between the m-th keyframe and at least one surrounding keyframe. For example, for a feature point in the m-th keyframe, the similarity between the descriptor of that feature point and the descriptors of each feature point in the bottom adjacent keyframe can be calculated. Feature points in the bottom adjacent keyframes with a similarity exceeding a specified threshold are considered as the feature points that match the feature point in the m-th keyframe. It should be understood that for each keyframe, it is possible to match and associate it with its surrounding keyframes (such as the right keyframe and / or the bottom keyframe) to find matching feature points between the two frames.

[0084] In step S212, the implementation of step S1532 above can be referred to to determine the relative pose of at least one surrounding keyframe relative to the m-th keyframe based on the feature points matched between the m-th keyframe and at least one surrounding keyframe. That is, the relative pose of each surrounding keyframe relative to the m-th keyframe can be determined by estimating the homography matrix. More precisely, the relative pose of each surrounding keyframe relative to the m-th keyframe can be determined by constructing the constraint equation shown in formula (1) above. This disclosure embodiment does not limit this.

[0085] In step S213, given the relative pose of each surrounding keyframe with respect to the m-th keyframe and the optimized pose of the m-th keyframe, the optimized pose of each surrounding keyframe can be calculated.

[0086] It should be understood that the optimized pose of each keyframe in the two-dimensional grid can be calculated according to steps S211 to S213. Then, in step S22, the optimized orientation (i.e., optimized position) of each keyframe in the two-dimensional grid can be determined based on the optimized pose of each keyframe. As mentioned above, the pose can be represented as a rotation matrix. The optimized rotation matrix of the keyframe can be used as the optimized orientation of the keyframe, or the optimized rotation matrix can be converted into Euler angles (i.e., yaw angle, pitch angle, roll angle) as the optimized orientation, or the vector of the optical axis direction can be extracted from the optimized rotation matrix as the optimized orientation. This embodiment of the present disclosure does not limit this. Having obtained the optimized orientation of each keyframe in the two-dimensional grid, the position of each keyframe in the two-dimensional grid can be adjusted according to the optimized orientation. For example, according to the optimized orientation of a certain keyframe, that keyframe may need to be moved to another position, or its position may remain unchanged.

[0087] In step S23, after adjusting the positions of each keyframe in the two-dimensional grid, a keyframe may be moved to a blank position (i.e., a position where no keyframe is recorded) or a non-blank position (i.e., a position where a keyframe is recorded). This may result in multiple keyframes existing in the same grid. In this case, newly added keyframes can be deleted. That is, if there are two or more keyframes at the same position, only the keyframe with the earliest acquisition time is retained at that position, and the other keyframes are deleted. The keyframe with the earliest acquisition time may be a keyframe that originally existed at this position, or it may be one of the multiple keyframes that were moved to this position later. This embodiment of the present disclosure does not limit this.

[0088] It should be noted that there is no sequential execution order between the optimization process of key frame pose in the two-dimensional grid (steps S21 to S23) and the process of determining the infrared camera pose (steps S11 to S15, steps S151 to S154, etc.). Those skilled in the art can customize the triggering time of the optimization process of key frame pose in the two-dimensional grid. For example, the optimization process can be triggered every time a new key frame is inserted, or the optimization process can be triggered periodically at regular intervals (e.g., every second). This embodiment of the present disclosure does not limit this.

[0089] In this embodiment of the disclosure, by optimizing the pose of keyframes in the two-dimensional grid as a whole, it is beneficial to improve the accuracy of estimating the pose of the infrared camera using keyframes.

[0090] Considering that during manual adjustment of the infrared camera, scenarios such as the camera being obstructed, or the camera moving too fast leading to image blurring or underexposure may occur, the method of determining the pose by matching the current frame with the previous frame or with keyframes may fail in these situations. This is because the current frame may not contain valid feature points and their descriptors, or the number of matched feature points may be too small to calculate the relative pose (for example, the method of determining the pose by calculating the homography matrix may require at least four matched feature points). In this case, the current frame can be considered lost, or the pose of the infrared camera can be lost, meaning the current pose of the infrared camera cannot be determined. Therefore, to recover the pose of the infrared camera, a re-adjustment process can be performed. The basic idea of ​​the positioning process is that after the current frame is lost, the adjustment personnel can be notified to adjust the infrared camera back to its previous position. During this process, the newly acquired current frame from the infrared camera can be used to find a keyframe in the 2D grid that is similar to the new current frame. Feature point matching can then be performed using these similar keyframes to estimate the infrared camera's attitude, thereby achieving attitude repositioning. Specifically, a global descriptor can be calculated for each keyframe, that is, a string of vectors describing the entire frame image. When the current frame fails to match the previous frame or a keyframe, the global descriptor is matched for each keyframe, and the optimal matching global descriptor is used for keyframe repositioning. Therefore, specifically, the method may further include:

[0091] Step S31: If no feature points and their descriptors are extracted in the current frame, or if there are no matching feature points between the current frame and the previous frame, or if the number of matching feature points is less than a specified number, adjust the orientation of the infrared camera in the direction it has previously faced.

[0092] Step S32: During the process of adjusting the attitude of the infrared camera to the previously facing direction, the new current frame acquired in real time by the infrared camera is obtained, and the global descriptor of the new current frame is extracted. The global descriptor represents the global feature vector of the new current frame.

[0093] Step S33: Based on the global descriptor of the new current frame and the global descriptors of each keyframe in the two-dimensional grid, determine whether there is a target keyframe in the two-dimensional grid that matches the new current frame.

[0094] Step S34: If there is a target keyframe in the two-dimensional grid that matches the new current frame, determine the current pose of the infrared camera based on the feature points in the new current frame and the feature points in the target keyframe.

[0095] In step S31, the failure to extract feature points and their descriptors in the current frame may be due to the frame being too blurry. The absence of matching feature points between the current and previous frames, or the number of matching feature points being less than the specified number, may be due to the infrared camera moving too fast, resulting in a large difference between the two frames. It should be understood that when these situations occur, the orientation may be determined through steps S13 to S15, which may cause the current frame to fail to match the previous frame, and it may also be impossible to find the key frame from the two-dimensional grid, resulting in the current frame failing to match the key frame. In this case, it can be assumed that the orientation of the infrared camera cannot be determined at present. In this case, the adjustment personnel can adjust the orientation of the infrared camera in the direction it was previously facing, that is, adjust the orientation of the infrared camera back. This may allow the infrared camera to capture the previously captured image. During this process, the infrared camera is also acquiring frames in real time. Therefore, in step S32, the new current frame acquired in real time by the infrared camera can be obtained and the global descriptor of the new current frame can be extracted.

[0096] In step S32, it can be understood that during the process of adjusting the infrared camera's attitude in the previously facing direction, the infrared camera also acquires frames in real time. The new current frame is a frame acquired in real time during the process of adjusting the infrared camera back. For example, during the normal adjustment of the infrared camera's attitude in step S11 above, it is found that the i-th frame acquired by the infrared camera (the current frame in the above process) cannot match the feature points of the (i-1)-th frame. At this time, the attitude of the infrared camera can be adjusted back, which allows the calculation of the global descriptor to start from the (i+1)-th frame acquired by the infrared camera (that is, the new current frame). Then, it is matched with the keyframe in the two-dimensional grid through step S33. If the target keyframe that matches the (i+1)-th frame can be determined from the two-dimensional grid in step S33, then step S34 can be executed to determine the current attitude. If the target keyframe that matches the i-th frame cannot be determined in step S33, then the process returns to step S32 to calculate the global descriptor of the (i+2)-th frame acquired by the infrared camera (another new current frame) until the target keyframe can be matched.

[0097] In step S32, those skilled in the art can employ any known global descriptor extraction method. For example, NetVLAD (Neural Network-based VLAD, a differentiable global descriptor extraction method for image retrieval in computer vision) can be used to extract the global descriptor of the new current frame. This embodiment of the present disclosure does not limit this approach. Furthermore, the global descriptor of each keyframe in the two-dimensional grid can also be pre-extracted using the above method. For example, the global descriptor of each keyframe can be extracted and stored every time a keyframe is inserted into the two-dimensional grid. This embodiment of the present disclosure does not limit this approach.

[0098] In step S33, based on the global descriptor of the new current frame and the global descriptors of each keyframe in the two-dimensional grid, a target keyframe in the two-dimensional grid that matches the new current frame is determined. For example, this may include: calculating the similarity between the global descriptor of the new current frame and the global descriptor of each keyframe in the two-dimensional grid, determining whether there are keyframes in the two-dimensional grid whose similarity exceeds a preset threshold, and determining the keyframe with the highest similarity exceeding the preset threshold as the target keyframe that matches the new current frame.

[0099] In step S34, the implementation methods of steps S12 to S15 above can be referred to to determine the current pose of the infrared camera based on the feature points in the new current frame and the feature points in the target keyframe. That is, the feature points and corresponding descriptors in the new current frame can be extracted. Based on the descriptors corresponding to the feature points in the new current frame and the descriptors corresponding to the feature points in the target keyframe, the matching feature points between the new current frame and the target keyframe can be determined. Based on the matching feature points between the new current frame and the target keyframe, the relative pose of the new current frame with respect to the target keyframe can be determined. Based on the relative pose of the new current frame with respect to the target keyframe, the current pose of the infrared camera can be determined. The implementation method of step S1532 above can be referred to to determine the relative pose of the new current frame with respect to the target keyframe based on the matching feature points between the new current frame and the target keyframe. Knowing the relative pose of the new current frame with respect to the target keyframe and the pose corresponding to the target keyframe, the current pose of the infrared camera can be obtained.

[0100] It should be understood that after the infrared camera's orientation is repositioned, the operator can continue to adjust the infrared camera toward the ideal orientation (or toward the designated area), that is, return to the process of steps S11 to S15 above until the orientation adjustment of the infrared camera is completed. This disclosure does not limit this process.

[0101] In the embodiments of this disclosure, the above-described repositioning process can help recover the pose of the infrared camera when it is lost due to abnormal situations such as the infrared camera being blocked, moving too fast and causing image blurring, or insufficient exposure. This helps to improve the robustness of the entire pose estimation method.

[0102] The real-time attitude estimation method proposed in the above embodiments of this disclosure. Figure 3 This diagram illustrates a framework of an infrared camera real-time attitude estimation system according to an embodiment of the present disclosure, such as... Figure 3As shown, the system mainly includes a tracking thread and a construction thread. The adjacent frame pose tracking module in the tracking thread executes steps S11 to S15, where step S15 can be performed by estimating the homography matrix to determine the pose. The keyframe pose tracking module executes steps S151 to S153, and the keyframe selection module executes step S154 to check if a 2D grid is inserted in the current frame. If the adjacent frame pose tracking module succeeds, the keyframe pose tracking module can be executed. If the adjacent frame pose tracking module fails or is shut down... If the keyframe pose tracking module fails (i.e., it fails to determine the current pose using the previous frame or keyframe), the relocalization module can be executed. This module performs the relocalization process shown in steps S31 to S34 above, and during its execution, NetVLAD is used for image retrieval (i.e., keyframe matching based on global descriptors). The keyframe optimization module in the mapping thread performs the two-dimensional raster optimization process shown in steps S21 to S23 above. The keyframe library stores keyframes and their information (such as pose, feature points and descriptors, global descriptors, frame identifiers, etc.). Using this system, real-time and accurate pose estimation of the infrared camera during adjustment can be achieved, while also exhibiting high robustness and reliability.

[0103] Figure 4 A block diagram of a multi-view infrared camera real-time attitude estimation apparatus according to an embodiment of the present disclosure is shown, as follows: Figure 4 As shown, the device includes:

[0104] The acquisition module 301 is used to acquire the current frame captured in real time by the infrared camera during the process of adjusting the attitude of the infrared camera;

[0105] Extraction module 302 is used to extract feature points and corresponding descriptors in the current frame, wherein the descriptors represent feature vectors of local regions surrounding the feature points;

[0106] Matching module 303 is used to determine the matching feature points between the current frame and the previous frame based on the descriptor corresponding to the feature points in the current frame and the descriptor corresponding to the feature points in the previous frame, wherein the previous frame is a frame captured by the infrared camera before the current frame;

[0107] The relative pose determination module 304 is used to determine the relative pose of the infrared camera in the current frame relative to the previous frame based on the feature points matched between the current frame and the previous frame.

[0108] The attitude determination module 305 is used to determine the current attitude of the infrared camera based on the relative attitude of the infrared camera in the current frame relative to the previous frame.

[0109] In some embodiments, determining the current pose of the infrared camera based on the relative pose of the infrared camera in the current frame relative to the previous frame includes:

[0110] Based on the relative posture of the current frame with respect to the previous frame, the orientation corresponding to the current frame is determined, and the orientation corresponding to the current frame represents the orientation of the infrared camera when acquiring the current frame;

[0111] Based on the orientation corresponding to the current frame, it is determined whether a key frame corresponding to that orientation is recorded in a preset two-dimensional grid. The two-dimensional grid is used to record the key frames collected according to the orientation of the infrared camera. The two-dimensional grid is obtained by two-dimensionally unfolding and rasterizing the three-dimensional shooting range of the infrared camera.

[0112] If the keyframe corresponding to the orientation is recorded in the two-dimensional grid, the current pose of the infrared camera is determined based on the feature points in the current frame and the feature points in the keyframe.

[0113] If no keyframe corresponding to the orientation is recorded in the two-dimensional grid, the current frame is recorded as a keyframe in the two-dimensional grid according to the orientation corresponding to the current frame. The current orientation of the infrared camera is determined according to the relative orientation of the current frame with respect to the previous frame. The position of the keyframe in the two-dimensional grid represents the orientation of the infrared camera when acquiring the keyframe.

[0114] In some embodiments, determining the current pose of the infrared camera based on feature points in the current frame and feature points in the key frame includes:

[0115] Based on the descriptors corresponding to feature points in the current frame and the descriptors corresponding to feature points in the key frame, determine the matching feature points between the current frame and the key frame;

[0116] The relative pose of the infrared camera in the current frame with respect to the key frame is determined based on the feature points matched between the current frame and the key frame.

[0117] The current pose of the infrared camera is determined based on the relative pose of the infrared camera in the current frame with respect to the key frame, and the pose corresponding to the key frame.

[0118] In some embodiments, the apparatus further includes an optimization module, configured to:

[0119] By fixing the pose of the first keyframe in the two-dimensional grid, the poses of each keyframe in the two-dimensional grid are optimized in a specified order starting from the first keyframe, where the first keyframe is the initial frame;

[0120] Based on the optimized pose of each keyframe in the two-dimensional grid, adjust the position of each keyframe in the two-dimensional grid.

[0121] After adjusting the position of each keyframe in the two-dimensional grid, if there are two or more keyframes at the same position, only the keyframe with the earliest acquisition time is retained at that position.

[0122] In some embodiments, optimizing the pose of each keyframe in the 2D grid in a specified order, starting from the first keyframe, includes:

[0123] For the m-th keyframe in the specified order, based on the descriptor of the feature points in the m-th keyframe and the descriptor of the feature points in at least one surrounding keyframe, the feature points matched between the m-th keyframe and at least one surrounding keyframe are determined respectively, where m is a positive integer and the m-th keyframe is a keyframe with optimized pose.

[0124] Based on the feature points matched between the m-th keyframe and at least one surrounding keyframe, determine the relative pose of at least one surrounding keyframe relative to the m-th keyframe.

[0125] Based on the relative pose of at least one surrounding keyframe with respect to the m-th keyframe, and the optimized pose of the m-th keyframe, determine the optimized pose of each of the at least one surrounding keyframe.

[0126] In some embodiments, the apparatus further includes: a repositioning module, configured to:

[0127] If no feature points and their descriptors are extracted in the current frame, or if there are no matching feature points between the current frame and the previous frame, or if the number of matching feature points is less than a specified number, the orientation of the infrared camera is adjusted in the direction it was previously facing.

[0128] During the process of adjusting the attitude of the infrared camera to the direction it was previously facing, a new current frame acquired in real time by the infrared camera is obtained, and a global descriptor of the new current frame is extracted. The global descriptor represents the global feature vector of the new current frame.

[0129] Based on the global descriptor of the new current frame and the global descriptors of each keyframe in the two-dimensional grid, determine whether there is a target keyframe in the two-dimensional grid that matches the new current frame;

[0130] If there is a target keyframe in the two-dimensional grid that matches the new current frame, the current pose of the infrared camera is determined based on the feature points in the new current frame and the feature points in the target keyframe.

[0131] In some embodiments, the virtual shooting studio further includes a camera for capturing video footage; if the positioning effect of the infrared camera in the virtual shooting studio on the camera in the designated area is lower than ideal, the attitude of the infrared camera in the virtual shooting studio is adjusted, and the device is used to estimate the attitude of the infrared camera during the adjustment process. The adjusted infrared camera is used to locate the camera in the designated area.

[0132] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0133] This disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.

[0134] This disclosure also provides a non-volatile computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.

[0135] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0136] Figure 5 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 may be provided as a server or a terminal device. (Refer to...) Figure 5 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.

[0137] Electronic device 1900 may also include a power supply component 1926 configured to perform power management of electronic device 1900, a wired or wireless network interface 1950 configured to connect electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). Electronic device 1900 can operate on an operating system, such as Windows Server, stored in memory 1932. TM Mac OS X TM Unix TM Linux TM FreeBSD TM Or similar.

[0138] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of an electronic device 1900 to perform the above-described method.

[0139] Computer-readable storage media can be tangible devices capable of holding and storing programs / instructions used by instruction execution devices. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0140] The computer program (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage medium in the respective computing / processing device.

[0141] The computer program (or computer program instructions) used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions to implement various aspects of this disclosure.

[0142] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0143] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0144] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0145] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0146] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A real-time attitude estimation method for a multi-view infrared camera, wherein the infrared camera is set at a fixed position in a virtual shooting studio, characterized in that, The method includes: During the process of adjusting the attitude of the infrared camera, the current frame captured in real time by the infrared camera is obtained; Extract feature points and corresponding descriptors from the current frame, wherein the descriptors characterize the feature vectors of the local region surrounding the feature points; Based on the descriptor corresponding to the feature point in the current frame and the descriptor corresponding to the feature point in the previous frame, the matching feature points between the current frame and the previous frame are determined, where the previous frame is a frame captured by the infrared camera before the current frame. Based on the feature points matched between the current frame and the previous frame, the relative pose of the infrared camera in the current frame relative to the previous frame is determined; The current pose of the infrared camera is determined based on the relative pose of the infrared camera in the current frame with respect to the previous frame.

2. The method according to claim 1, characterized in that, Determining the current pose of the infrared camera based on its relative pose in the current frame relative to the previous frame includes: Based on the relative posture of the current frame with respect to the previous frame, the orientation corresponding to the current frame is determined, and the orientation corresponding to the current frame represents the orientation of the infrared camera when acquiring the current frame; Based on the orientation corresponding to the current frame, it is determined whether a key frame corresponding to that orientation is recorded in a preset two-dimensional grid. The two-dimensional grid is used to record the key frames collected according to the orientation of the infrared camera. The two-dimensional grid is obtained by two-dimensionally unfolding and rasterizing the three-dimensional shooting range of the infrared camera. If the keyframe corresponding to the orientation is recorded in the two-dimensional grid, the current pose of the infrared camera is determined based on the feature points in the current frame and the feature points in the keyframe. If no keyframe corresponding to the orientation is recorded in the two-dimensional grid, the current frame is recorded as a keyframe in the two-dimensional grid according to the orientation corresponding to the current frame. The current orientation of the infrared camera is determined according to the relative orientation of the current frame with respect to the previous frame. The position of the keyframe in the two-dimensional grid represents the orientation of the infrared camera when acquiring the keyframe.

3. The method according to claim 2, characterized in that, Determining the current pose of the infrared camera based on feature points in the current frame and feature points in the key frame includes: Based on the descriptors corresponding to feature points in the current frame and the descriptors corresponding to feature points in the key frame, determine the matching feature points between the current frame and the key frame; The relative pose of the infrared camera in the current frame with respect to the key frame is determined based on the feature points matched between the current frame and the key frame. The current pose of the infrared camera is determined based on the relative pose of the infrared camera in the current frame with respect to the key frame, and the pose corresponding to the key frame.

4. The method according to claim 2 or 3, characterized in that, The method further includes: By fixing the pose of the first keyframe in the two-dimensional grid, the poses of each keyframe in the two-dimensional grid are optimized in a specified order starting from the first keyframe, where the first keyframe is the initial frame; Based on the optimized pose of each keyframe in the two-dimensional grid, adjust the position of each keyframe in the two-dimensional grid. After adjusting the position of each keyframe in the two-dimensional grid, if there are two or more keyframes at the same position, only the keyframe with the earliest acquisition time is retained at that position.

5. The method according to claim 4, characterized in that, The step of optimizing the pose of each keyframe in the 2D grid in a specified order, starting from the first keyframe, includes: For the m-th keyframe in the specified order, based on the descriptor of the feature points in the m-th keyframe and the descriptor of the feature points in at least one surrounding keyframe, the feature points matched between the m-th keyframe and at least one surrounding keyframe are determined respectively, where m is a positive integer and the m-th keyframe is a keyframe with optimized pose. Based on the feature points matched between the m-th keyframe and at least one surrounding keyframe, determine the relative pose of at least one surrounding keyframe relative to the m-th keyframe. Based on the relative pose of at least one surrounding keyframe with respect to the m-th keyframe, and the optimized pose of the m-th keyframe, determine the optimized pose of each of the at least one surrounding keyframe.

6. The method according to claim 3, characterized in that, The method further includes: If no feature points and their descriptors are extracted in the current frame, or if there are no matching feature points between the current frame and the previous frame, or if the number of matching feature points is less than a specified number, the orientation of the infrared camera is adjusted in the direction it was previously facing. During the process of adjusting the attitude of the infrared camera to the direction it was previously facing, a new current frame acquired in real time by the infrared camera is obtained, and a global descriptor of the new current frame is extracted. The global descriptor represents the global feature vector of the new current frame. Based on the global descriptor of the new current frame and the global descriptors of each keyframe in the two-dimensional grid, determine whether there is a target keyframe in the two-dimensional grid that matches the new current frame; If there is a target keyframe in the two-dimensional grid that matches the new current frame, the current pose of the infrared camera is determined based on the feature points in the new current frame and the feature points in the target keyframe.

7. The method according to any one of claims 1 to 6, characterized in that, The virtual shooting studio also includes a camera for capturing video footage. If the positioning effect of the infrared camera in the virtual shooting studio on the camera in the designated area is lower than ideal, the attitude of the infrared camera in the virtual shooting studio is adjusted, and the attitude of the infrared camera during the adjustment process is estimated using the method described above. The adjusted infrared camera is then used to locate the camera in the designated area.

8. A real-time attitude estimation device for a multi-view infrared camera, wherein the infrared camera is positioned at a fixed location in a virtual shooting studio, characterized in that... The device includes: The acquisition module is used to acquire the current frame captured in real time by the infrared camera during the process of adjusting the attitude of the infrared camera; The extraction module is used to extract feature points and corresponding descriptors in the current frame, wherein the descriptors represent feature vectors of the local region surrounding the feature points; The matching module is used to determine the matching feature points between the current frame and the previous frame based on the descriptors corresponding to the feature points in the current frame and the descriptors corresponding to the feature points in the previous frame. The previous frame is a frame captured by the infrared camera before the current frame. The relative pose determination module is used to determine the relative pose of the infrared camera in the current frame relative to the previous frame based on the feature points matched between the current frame and the previous frame. An attitude determination module is used to determine the current attitude of the infrared camera based on the relative attitude of the infrared camera in the current frame relative to the previous frame.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.

10. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.