Image processing methods, related apparatus, equipment and storage media

CN115984445BActive Publication Date: 2026-08-14ZHEJIANG SENSETIME TECH DEV CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-15
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]然而,就目前的位姿跟踪而言,因为位姿跟踪算法的精度不高,且计算的噪声较多,使得跟踪得到的位姿并不平滑,这也使得在三维物体在被渲染到屏幕上时发生抖动,严重影响用户体验

Benefits of technology

[0033]上述方案,通过基于历史位姿调整初始位姿,得到调整位姿,并且设置调整位姿与历史位姿间的位姿变化量小于初始位姿与历史位姿间的位姿变化量,以此可以减少目标物相对于当前图像帧与历史图像帧之间的位姿变化,故减少目标物在当前图像帧和历史图像帧中的显示位置出现突变的可能性,即使得后续在显示该目标物时,可以减少抖动的现象,改善用户体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115984445B_ABST
    Figure CN115984445B_ABST
Patent Text Reader

Abstract

This application discloses an image processing method and related apparatus, devices, and storage medium. The method includes acquiring an initial pose of a target object relative to a current image frame and a historical pose of the target object relative to a first historical image frame; adjusting the initial pose based on the historical pose to obtain an adjusted pose, wherein the pose change between the adjusted pose and the historical pose is less than the pose change between the initial pose and the historical pose; obtaining a target pose of the target object relative to the current image frame based on the adjusted pose; and determining the display position of the target object in the current image frame based on the target pose of the target object relative to the current image frame. This method can reduce jitter and improve user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image processing method and related apparatus, devices and storage media. Background Technology

[0002] Currently, with the rapid development of technologies such as Virtual Reality (VR) and Augmented Reality (AR), rich visual experiences can be obtained by using pose information to render three-dimensional objects.

[0003] However, current pose tracking algorithms are not very accurate and have a lot of computational noise, resulting in uneven poses. This causes the 3D objects to jitter when rendered on the screen, which seriously affects the user experience.

[0004] Therefore, reducing jitter is of great significance for improving user experience. Summary of the Invention

[0005] This application provides an image processing method, related apparatus, device, and storage medium.

[0006] The first aspect of this application provides an image processing method, which includes obtaining an initial pose of a target object relative to a current image frame and a historical pose of the target object relative to a first historical image frame; adjusting the initial pose based on the historical pose to obtain an adjusted pose, wherein the pose change between the adjusted pose and the historical pose is less than the pose change between the initial pose and the historical pose; obtaining a target pose of the target object relative to the current image frame based on the adjusted pose; and determining the display position of the target object in the current image frame based on the target pose of the target object relative to the current image frame.

[0007] Therefore, by adjusting the initial pose based on the historical pose, the adjusted pose is obtained. Furthermore, the pose change between the adjusted pose and the historical pose is set to be less than the pose change between the initial pose and the historical pose. This reduces the pose change of the target object relative to the current image frame and the historical image frames, thereby reducing the possibility of abrupt changes in the display position of the target object in the current image frame and the historical image frames. This also reduces the jitter phenomenon when displaying the target object later, thus improving the user experience.

[0008] The above-mentioned adjustment of the initial pose based on the historical pose to obtain the adjusted pose includes: weighting the historical pose and the initial pose based on the first weight coefficient of the initial pose and the second weight coefficient of the historical pose to obtain the adjusted pose.

[0009] Therefore, by using the first and second weighting coefficients to weight the initial pose and the historical pose respectively, the adjusted pose can be obtained.

[0010] Wherein, the sum of the first weighting coefficient and the second weighting coefficient is a preset value; and / or, the first weighting coefficient is positively correlated with the pose change rate of the target object, and the second weighting coefficient is negatively correlated with the pose change rate. The pose change rate indicates how fast the pose of the target object changes relative to at least three reference image frames. The reference image frames include the current image frame or the second historical image frame.

[0011] Because the human eye is more sensitive to jitter when the target object moves slowly, by setting the first weighting coefficient to be positively correlated with the target object's pose change rate, α is smaller when the target object moves slowly, thereby reducing the amount of pose change between the adjusted pose and the historical pose, and thus reducing jitter. Furthermore, because the human eye is more sensitive to delay when the target object moves rapidly, by setting the first weighting coefficient to be positively correlated with the target object's pose change rate, the adjusted pose can be closer to the initial pose when the target object moves rapidly, thereby reducing delay.

[0012] The above-mentioned method of obtaining the target pose of the target object relative to the current image frame based on the adjusted pose includes: using the adjusted pose as the target pose of the target object relative to the current image frame; or, using the target compensation pose corresponding to the current image frame to compensate the adjusted pose to obtain the target pose of the target object relative to the current image frame.

[0013] Therefore, by directly using the adjusted pose as the target pose of the object relative to the current image frame, jitter can be reduced when displaying the object using the target pose. Furthermore, by compensating the adjusted pose using the target compensation pose corresponding to the current image frame to obtain the target pose of the object relative to the current image frame, both jitter and latency can be reduced.

[0014] Before compensating the adjusted pose using the target compensation pose corresponding to the current image frame to obtain the target pose of the target object relative to the current image frame, the method further includes: determining the target compensation pose based on the pose change rate of the target object, wherein the target compensation pose is positively correlated with the pose change rate, and the pose change rate indicates how fast the pose of the target object changes relative to at least two reference image frames, the reference image frames including the current image frame or a second historical image frame.

[0015] Therefore, by setting the target compensation pose to be positively correlated with the pose change rate, the larger the target compensation pose is when the target moves rapidly, the better the effect of reducing latency.

[0016] The aforementioned determination of the target compensation pose based on the pose change rate of the target object includes: taking the pose change between the initial pose and the adjusted pose as the initial compensation pose; and determining the target compensation pose based on the initial compensation pose and the pose change rate.

[0017] Therefore, by utilizing the initial compensated pose and the pose change rate, a correlation can be established between the target compensated pose and the rate of pose change of the target object.

[0018] The aforementioned determination of the target compensation pose based on the initial compensation pose and the pose change rate includes: determining the third weighting coefficient of the initial compensation pose and the fourth weighting coefficient of the target compensation pose corresponding to the first historical image frame based on the pose change rate, wherein the sum of the third weighting coefficient and the fourth weighting coefficient is a preset value, and the third weighting coefficient is positively correlated with the pose change rate of the target object; and performing weighted processing on the initial compensation pose and the target compensation pose corresponding to the first historical image frame based on the third weighting coefficient and the fourth weighting coefficient to obtain the target compensation pose.

[0019] Therefore, by setting the third weight coefficient to be positively correlated with the pose change rate of the target object, and the fourth weight coefficient to be negatively correlated with the pose change rate of the target object, the delay compensation is greater when the pose change is faster, so that the delay phenomenon can be reduced better when the target object moves quickly.

[0020] Wherein, when one of the reference image frames includes the current image frame, the pose of the target object relative to the reference image frame includes the initial pose of the target object relative to the current image frame; when one of the reference image frames includes a second historical image frame, the pose of the target object relative to the reference image frame includes the target pose of the target object relative to the second historical image frame.

[0021] Therefore, when one of the reference image frames includes the current image frame, the pose change rate can be calculated using the initial pose of the target object relative to the current frame image. Furthermore, when one of the reference image frames includes a second historical image frame, the pose change rate can be calculated using the target pose of the target object relative to the second historical frame image.

[0022] The aforementioned adjustment of pose includes rotation and translation components; the aforementioned compensation of the adjusted pose with the target compensation pose corresponding to the current image frame to obtain the target pose of the target object relative to the current image frame includes: compensating the translation component of the adjusted pose with the target compensation pose corresponding to the current image frame to obtain the compensated translation component; and obtaining the target pose of the target object relative to the current image frame based on the compensated translation component and the rotation component.

[0023] Since translation has a greater impact on latency than rotation, processing the translation component can focus more on reducing latency. Therefore, the compensated translation component can be obtained using the target compensated pose corresponding to the current image frame. Conversely, rotation has a greater impact on jitter than translation, so processing the rotation can focus more on eliminating jitter. Therefore, the target pose can be directly obtained using the rotation component of the adjusted pose. This also avoids jitter that might be caused by minor noise in the target compensated pose corresponding to the current image frame, further improving the display effect.

[0024] The aforementioned historical pose, initial pose, and adjusted pose all include rotation components, which are matrices composed of multiple elements. The process of adjusting the initial pose based on the historical pose to obtain the adjusted pose includes: adjusting the corresponding elements in the rotation components of the initial pose based on each element in the rotation components of the historical pose to obtain the corresponding elements in the rotation components of the adjusted pose.

[0025] Therefore, by adjusting the corresponding elements in the rotation component of the initial pose based on the elements in the rotation component of the historical pose, the corresponding elements in the rotation component of the adjusted pose can be obtained.

[0026] Wherein, the aforementioned historical pose is the target pose of the target object relative to the first historical image frame; and / or, obtaining the initial pose of the target object relative to the current image frame includes: processing the current image frame using a pose tracking algorithm to obtain the initial pose of the target object relative to the current image frame.

[0027] Therefore, by using a pose tracking algorithm to process the current image frame, the initial pose of the current image frame can be obtained, and the initial pose can then be optimized to reduce jitter.

[0028] Wherein, the target object mentioned above is a three-dimensional object; and / or the above-mentioned determination of the display position of the target object in the current image frame based on the target pose of the target object relative to the current image frame includes: determining the projection position of the target object in the current image frame based on the target pose of the target object relative to the current image frame, so as to serve as the display position; and / or, after determining the display position of the target object in the current image frame based on the target pose of the target object relative to the current image frame, the method further includes: displaying the target object at the display position in the current image frame.

[0029] Therefore, by determining the target's projection position in the current image frame based on its pose relative to the current image frame, the target can be displayed at that projection position in subsequent images. Furthermore, displaying the target at its designated display position in the previous image frame achieves the desired display of the target.

[0030] A second aspect of this application provides an image processing apparatus, comprising an acquisition module, a first determination module, a second determination module, and a third determination module. The acquisition module is used to acquire an initial pose of a target object relative to a current image frame and a historical pose of the target object relative to a first historical image frame. The first determination module is used to adjust the initial pose based on the historical pose to obtain an adjusted pose, wherein the pose change between the adjusted pose and the historical pose is less than the pose change between the initial pose and the historical pose. The second determination module is used to obtain a target pose of the target object relative to the current image frame based on the adjusted pose. The third determination module is used to determine the display position of the target object in the current image frame based on the target pose of the target object relative to the current image frame.

[0031] A third aspect of this application provides an electronic device including a processor and a memory coupled to each other, wherein the processor is configured to execute a computer program stored in the memory to perform the method described in the first aspect above.

[0032] A fourth aspect of this application provides a computer-readable storage medium having program instructions stored thereon, which, when executed by a processor, implement the method described in the first aspect above.

[0033] The above scheme obtains the adjusted pose by adjusting the initial pose based on the historical pose, and sets the pose change between the adjusted pose and the historical pose to be less than the pose change between the initial pose and the historical pose. This reduces the pose change of the target object relative to the current image frame and the historical image frames, thus reducing the possibility of abrupt changes in the display position of the target object in the current image frame and the historical image frames. This also reduces jitter when displaying the target object later, improving the user experience. Attached Figure Description

[0034] Figure 1 This is a schematic flowchart of an embodiment of the image processing method of this application;

[0035] Figure 2 This is a flowchart illustrating another embodiment of the image processing method of this application;

[0036] Figure 3 This is a flowchart illustrating another embodiment of the image processing method of this application;

[0037] Figure 4 This is a schematic diagram of the image processing flow of the image processing method of this application;

[0038] Figure 5 This is a schematic diagram of the image transmission process according to an embodiment of the image processing method of this application;

[0039] Figure 6 This is a schematic diagram of the framework of an embodiment of the image processing apparatus of this application;

[0040] Figure 7 This is a schematic diagram of the framework of an embodiment of the electronic device of this application;

[0041] Figure 8 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0042] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0043] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0044] In this paper, the terms "system" and "network" are often used interchangeably. The term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this paper means two or more.

[0045] The device used for the image processing method in this application can be an electronic device such as a computer, mobile phone, tablet computer, or smart glasses.

[0046] Please see Figure 1 , Figure 1 This is a schematic flowchart of an embodiment of the image processing method of this application. Specifically, it may include the following steps:

[0047] Step S11: Obtain the initial pose of the target object relative to the current image frame and the historical pose of the target object relative to the first historical image frame.

[0048] In this application, the pose of the target object relative to the image frame, or the pose of the target object relative to the current image frame, can be considered as the pose of the target object relative to the capturing device when the current image frame was captured, or simply the pose of the image frame. The target object is, for example, a three-dimensional object, such as a virtual three-dimensional object, or a three-dimensional object that is further rendered based on a real three-dimensional object.

[0049] In one implementation, a pose tracking algorithm can be used to process the current image frame to obtain the initial pose of the target object relative to the current image frame. The pose tracking algorithm can be a common tracking algorithm in the art, such as a 3D tracking algorithm based on edge detection. By using the pose tracking algorithm to process the current image frame, the initial pose of the current image frame can be obtained, and the initial pose can then be optimized to reduce jitter.

[0050] Step S12: Adjust the initial pose based on the historical pose to obtain the adjusted pose.

[0051] The historical pose can be based on pose information obtained from image frames preceding the current image frame, such as the historical pose obtained from the image frame before the current image frame. In this application, the pose change between the adjusted pose and the historical pose is less than the pose change between the initial pose and the historical pose.

[0052] In one implementation, the pose change between the initial pose and the historical pose can be determined first. Then, this pose change can be reduced to obtain the pose change between the adjusted pose and the historical pose. Finally, the adjusted pose is obtained based on the pose change between the adjusted pose and the historical pose, and the historical pose. In another implementation, relevant weight coefficients can be set for the initial pose and the historical pose respectively, so that the pose change between the adjusted pose and the historical pose is smaller than the pose change between the initial pose and the historical pose.

[0053] In one specific implementation, the pose change can be the change of the translation component or the rotation component in the pose, or the pose change obtained by combining the translation component and the rotation component.

[0054] Step S13: Based on the adjusted pose, obtain the target pose of the target object relative to the current image frame.

[0055] In one specific implementation, the target pose can be directly the adjusted pose.

[0056] In one specific implementation, the adjusted pose can be further optimized to obtain the target pose. Further optimization of the adjusted pose can be achieved by setting the target pose to be smaller than the initial pose, meaning the pose change between the target pose and the historical pose is less than the pose change between the initial pose and the historical pose.

[0057] In one implementation, the above-mentioned process of adjusting the initial pose based on the historical pose, obtaining the adjusted pose, and obtaining the target pose of the target object relative to the current image frame based on the adjusted pose can be processed using a software filter.

[0058] Step S14: Determine the display position of the target object in the current image frame based on the target pose of the target object relative to the current image frame.

[0059] After determining the target pose relative to the current image frame, the target's display position in the current image frame can be determined using methods common in the field, such as projection. The 3D object can then be displayed at that position.

[0060] In one specific implementation, the aforementioned historical pose is the target pose of the object relative to the first historical image frame. That is, the historical pose can be an adjusted pose, rather than a pose directly obtained using a pose tracking algorithm. This ensures that when determining the target pose of the current image frame, the target pose is also based on the target pose of the historical image frames, thereby further reducing the amount of pose variation between different image frames and improving the jitter phenomenon of 3D objects.

[0061] Therefore, by adjusting the initial pose based on the historical pose, the adjusted pose is obtained. Furthermore, the pose change between the adjusted pose and the historical pose is set to be less than the pose change between the initial pose and the historical pose. This reduces the pose change of the target object relative to the current image frame and the historical image frames, thereby reducing the possibility of abrupt changes in the display position of the target object in the current image frame and the historical image frames. This also reduces the jitter phenomenon when displaying the target object later, thus improving the user experience.

[0062] In one embodiment, the step of "adjusting the initial pose from the historical pose to obtain the adjusted pose" mentioned above specifically includes: weighting the historical pose and the initial pose based on the first weight coefficient of the initial pose and the second weight coefficient of the historical pose to obtain the adjusted pose.

[0063] In one specific implementation, the sum of the first weighting coefficient and the second weighting coefficient is a preset value, such as 1.

[0064] In one specific implementation, the adjusted pose can be obtained by the following formula (1).

[0065]

[0066] in, To adjust the pose, X i This is the initial pose. For the historical pose, α is the first weight coefficient and is less than 1, and (1-α) is the second weight coefficient.

[0067] Therefore, by using the first and second weighting coefficients to weight the initial pose and the historical pose respectively, the adjusted pose can be obtained.

[0068] In one implementation, the first weighting coefficient is positively correlated with the pose change rate of the target object, and the second weighting coefficient is negatively correlated with the pose change rate. The pose change rate represents how quickly the pose of the target object changes relative to at least three reference image frames. The pose of the reference image frames can be the initial pose or the target pose of the reference image frames. When the reference image frames include the current image frame, the pose of the reference image frames can be the initial pose; when the reference image frames are historical image frames, the pose of the reference image frames can be the target pose. The change in the pose of the target object relative to at least three reference image frames can determine two pose change quantities. For example, based on the first reference image frame P1 and the second reference image frame P2, one pose change quantity can be obtained, and based on the first reference image frame P2 and the second reference image frame P3, another pose change quantity can be obtained. By comparing the two pose change quantities, the rate of pose change can be determined, and thus the pose change rate can be determined. In one specific implementation, the pose change rate can be obtained by differentiating the pose of the target object with respect to the image frames.

[0069] In one specific implementation, by transforming formula (1), we can obtain the equivalent formula (2) as follows:

[0070]

[0071] Because the human eye is more sensitive to jitter when the target object moves slowly, setting a first weighting coefficient positively correlated with the target object's pose change rate ensures that α is smaller when the target object moves slowly. This reduces the amount of pose change between the adjusted pose and the historical pose, thus reducing jitter. Furthermore, because the human eye is more sensitive to delay when the target object moves rapidly, setting a first weighting coefficient positively correlated with the target object's pose change rate allows the adjusted pose to be closer to the initial pose when the target object moves rapidly, thereby reducing delay.

[0072] In one specific implementation, when one of the reference image frames includes the current image frame, the pose of the target object relative to the reference image frame includes the initial pose of the target object relative to the current frame image. In this case, the pose of the target object relative to other reference image frames can also be the initial pose, i.e., obtained through a pose tracking algorithm. Therefore, the pose change rate can be calculated using the initial pose of the target object relative to the current frame image.

[0073] In one specific embodiment, when one of the reference image frames is a second historical image frame, the pose of the target object relative to the reference image frame includes the target pose of the target object relative to the second historical image frame. The second historical image frame may include the first historical image frame. In another specific embodiment, when the reference image frames do not include the current image frame, all reference image frames are second historical image frames. In this case, the pose change rate can be calculated using the target pose of the target object relative to the second historical image frame. Therefore, the pose change rate can be calculated using the target pose of the target object relative to the second historical image frame.

[0074] In one embodiment, the step of "obtaining the target pose of the target object relative to the current image frame based on the adjusted pose" mentioned above specifically includes step S131 or step S132.

[0075] Step S131: Adjust the pose as the target pose of the object relative to the current image frame.

[0076] In this step, the adjusted pose can be directly used as the target pose of the object relative to the current image frame, which can reduce jitter when displaying the target pose in the future.

[0077] Step S132: Use the target compensation pose corresponding to the current image frame to compensate the adjustment pose, and obtain the target pose of the target object relative to the current image frame.

[0078] In this step, because the pose change between the adjusted pose and the historical pose is less than the pose change between the initial pose and the historical pose, a delay occurs. Therefore, the adjusted pose can be compensated using the target compensation pose corresponding to the current image frame, thereby obtaining the target pose of the object relative to the current image frame. The target compensation pose can be predetermined or obtained through a software filter. Understandably, the method of obtaining the target compensation pose is unrestricted.

[0079] Therefore, by using the target compensation pose corresponding to the current image frame to compensate for the adjusted pose, the target pose of the target object relative to the current image frame can be obtained, which can reduce jitter and delay at the same time.

[0080] In one implementation, before "compensating the adjusted pose using the target compensation pose corresponding to the current image frame to obtain the target pose of the object relative to the current image frame," the following can also be performed: determining the target compensation pose based on the pose change rate of the object. Furthermore, the target compensation pose is set to be positively correlated with the pose change rate. This ensures that when the object moves rapidly, a larger target compensation pose results in a better effect on reducing latency.

[0081] Please see Figure 2 , Figure 2 This is a schematic flowchart of another embodiment of the image processing method of this application. In this embodiment, the "determining the target compensation pose based on the pose change rate of the target object" mentioned in the above steps specifically includes steps S21 and S22.

[0082] Step S21: Use the pose change between the initial pose and the adjusted pose as the initial compensation pose.

[0083] Since the pose change between the initial pose and the adjusted pose can be considered as a real delay, the pose change between the initial pose and the adjusted pose can be used as the initial compensation pose, which facilitates the subsequent determination of the target compensation pose.

[0084] Step S22: Determine the target compensation pose based on the initial compensation pose and pose change rate.

[0085] In one specific implementation, the target compensation pose can be obtained by the following formula (3).

[0086]

[0087] in, To compensate for the target pose, X i This is the initial pose. To adjust the posture.

[0088] By utilizing the initial compensated pose and the pose change rate, a correlation can be established between the target compensated pose and the rate of pose change of the target object.

[0089] In one implementation, the step of "determining the target compensation pose based on the initial compensation pose and the pose change rate" mentioned above specifically includes steps S221 and S222.

[0090] Step S221: Based on the pose change rate, determine the third weight coefficient of the initial compensated pose and the fourth weight coefficient of the target compensated pose corresponding to the first historical image frame.

[0091] In this embodiment, a third weighting coefficient can be set for the initial compensation pose, and a fourth weighting coefficient can be set for the target compensation pose corresponding to the first historical image frame. Furthermore, the sum of the third and fourth weighting coefficients is a preset value, such as 1. Alternatively, the third weighting coefficient can be set to be positively correlated with the pose change rate of the target object, in which case the fourth weighting coefficient is negatively correlated with the pose change rate of the target object. Through the above settings, the target compensation pose corresponding to the current image frame can be correlated with the pose change rate.

[0092] Step S222: Based on the third and fourth weighting coefficients, the initial compensation pose and the target compensation pose corresponding to the first historical image frame are weighted to obtain the target compensation pose.

[0093] In one specific implementation, the target compensation pose can be obtained by the following formula (4).

[0094]

[0095] in, For the initial compensation pose, Let be the target compensation pose corresponding to the first historical image frame, β be the third weighting coefficient, and (1-β) be the fourth weighting coefficient. The target pose is compensated. In some implementations, the first weighting coefficient described above may be the same as the third weighting coefficient.

[0096] Therefore, since the third weighting coefficient is positively correlated with the pose change rate of the target object, and the fourth weighting coefficient is negatively correlated with the pose change rate of the target object, the delay compensation is greater when the pose change is faster, so that the delay phenomenon can be reduced better when the target object moves quickly.

[0097] In one embodiment, the weighting coefficient α (first weighting coefficient or third weighting coefficient) mentioned in the above embodiments can be obtained by the following formulas (5) to (7).

[0098]

[0099]

[0100]

[0101] Where τ is the amplitude of the signal change; T e f is the sampling time interval, i.e., the derivative of the frame rate; c For the cutoff frequency; The initial cutoff frequency is a predetermined value; γ is the pose change rate mentioned in the above embodiments; γ is the adjustment coefficient.

[0102] Through the above formulas (5) to (7), it can be determined that the weight coefficient α is positively correlated with the pose change rate. The faster (larger) the pose change rate, the smaller τ is and the larger α is; conversely, the slower (smaller) the pose change rate, the larger τ is and the smaller α is.

[0103] Please see Figure 3 , Figure 3This is a flowchart illustrating another embodiment of the image processing method of this application. In this embodiment, the pose adjustment mentioned in the above embodiments includes rotation and translation components. The step of "compensating the adjusted pose using the target compensation pose corresponding to the current image frame to obtain the target pose of the target object relative to the current image frame" mentioned above includes steps S31 and S32.

[0104] Step S31: Use the target compensation pose corresponding to the current image frame to compensate the translation component of the adjusted pose to obtain the compensated translation component.

[0105] In this embodiment, the target compensation pose corresponding to the current image frame can be the pose component that compensates for the translation component of the adjusted pose. Specifically, the change in translation component pose between the initial pose translation component and the adjusted pose translation component can be used as the initial translation component compensation pose. Then, the target compensation pose is determined based on the initial translation component compensation pose and the pose change rate. Here, the pose change rate can be the pose change rate of the translation component.

[0106] Step S32: Obtain the target pose of the target object relative to the current image frame based on the compensated translation and rotation components.

[0107] Once the compensated translation component is obtained, the compensated translation component and rotation component can be used as the translation component and rotation component of the target pose, thereby obtaining the target pose of the target object relative to the current image frame.

[0108] Since translation has a greater impact on latency than rotation, processing the translation component can focus more on reducing latency. Therefore, the compensated translation component can be obtained using the target compensated pose corresponding to the current image frame. Conversely, rotation has a greater impact on jitter than translation, so processing the rotation can focus more on eliminating jitter. Therefore, the target pose can be directly obtained using the rotation component of the adjusted pose. This also avoids jitter that might be caused by minor noise in the target compensated pose corresponding to the current image frame, further improving the display effect.

[0109] In one embodiment, the historical pose, initial pose, and adjusted pose mentioned in the above embodiments all include rotational components. The step of "adjusting the initial pose based on the historical pose to obtain the adjusted pose" specifically includes: adjusting the corresponding elements in the rotational component of the initial pose based on each element in the rotational component of the historical pose to obtain the corresponding elements in the rotational component of the adjusted pose.

[0110] The elements in the rotation components of the historical pose or the initial pose correspond to different expressions of the rotation components. When the rotation components are represented by a rotation matrix, the elements are the individual elements of the rotation matrix; when the rotation components are represented by quaternions, the elements are each number in the quaternion (a, bi, cj, and dk). Specific adjustment methods can be found in the relevant descriptions of the steps in the above embodiments, and will not be repeated here.

[0111] In one specific implementation, the rotation component is represented by a rotation matrix, and the adjustment process can be represented as follows:

[0112]

[0113] Where r11, r12, r13, r21, r22, r23, r31, r32 and r33 are the elements of the rotation matrix, F represents the adjustment of the elements, and F(R) is the corresponding element in the rotation component of the pose adjustment.

[0114] Therefore, by adjusting the corresponding elements in the rotation component of the initial pose based on the elements in the rotation component of the historical pose, the corresponding elements in the rotation component of the adjusted pose can be obtained.

[0115] In one embodiment, the step "determining the display position of the target object in the current image frame based on the target pose of the target object relative to the current image frame" mentioned in the above embodiment specifically includes: determining the projection position of the target object in the current image frame based on the target pose of the target object relative to the current image frame, and using this projection position as the display position. After determining the target pose of the target object relative to the current image frame, a projection method commonly used in the art can be used to project the target object onto the current image frame, thereby determining the projection position of the target object in the current image frame. For example, the target object can be projected using Open Graphics Library (OpenGL) to obtain the projection position, and the projection position is used as the display position. Therefore, by determining the projection position of the target object in the current image frame based on the target pose of the target object relative to the current image frame, the target object can be displayed at the projection position subsequently.

[0116] In one embodiment, after the step of "determining the display position of the target object in the current image frame based on the target pose of the target object relative to the current image frame," the following step can be performed: displaying the target object at the display position in the current image frame. Displaying the target object at the display position in the current image frame can be a rendering method commonly used in the art, such as rendering the target object using OpenGL, which will not be elaborated here. Therefore, by displaying the target object at the display position in the current image frame, the display of the target object is achieved.

[0117] Please see Figure 4 , Figure 4 This is a schematic diagram of the image processing flow of the image processing method of this application. In this embodiment, the relative pose between the three-dimensional object and the image frame is simply referred to as the pose of the image frame. In this embodiment, the original pose of the three-dimensional object 43 is P0, the initial pose of the first historical image frame 42 is P1, and the adjusted pose of the first historical image frame 42 is... The initial compensated pose of the first historical image frame 42 is The target compensation pose in the first historical image frame 42 is:

[0118] After the current image frame 41 is processed by the tracker 44 using a pose tracking algorithm, the initial pose P2 of the current image frame 41 is obtained. At this time, the initial pose P2 and the target pose of the first historical image frame 42 can be compared. (target pose) The historical pose is input to filter 1, which then adjusts the initial pose P2 based on the historical pose to obtain the adjusted pose. Then, the pose can be adjusted. Subtracting the initial pose P2 from the initial pose P2 yields the initial compensated pose. In the picture This indicates the subtraction of poses. Next, the initial compensated pose is... The target compensation pose corresponding to the first historical image frame 42 The input is fed into filter 2, which performs compensation on the initial pose based on the third and fourth weighting coefficients. Target compensation pose corresponding to the first historical image frame After weighted processing, the target compensation pose corresponding to the current image frame 41 is obtained. Then, the target can be used to compensate for the pose. and adjust posture Obtain the target pose of the current image frame 41 In the picture This represents the addition of poses. Finally, the target pose can be used. Rendering is performed on the current image frame 41 to obtain the rendered 3D object 43.

[0119] In this embodiment, the input pose of the tracker 44 is the initial pose obtained by the tracker 44 in the previous image frame. When rendering the 3D object 43, the target pose corresponding to the image frame is used for rendering. In this way, tracking and display can be separated, reducing the errors that may be caused by mixing the two and improving the display effect.

[0120] In one embodiment, when using a pose tracking algorithm to obtain the initial pose of the current image frame, steps S51 to S53 may be included.

[0121] Step S51: Generate an original reference image corresponding to the current image frame for the target object using the first processor, and extract pixel data of the first target region from the original reference image.

[0122] The current image frame may be obtained through an image acquisition device of the executing entity, such as a camera, or by another device acquiring the current image frame and then transmitting it to the executing entity. This application does not impose specific limitations on the method of acquiring the current frame.

[0123] The first processor can be a processor with image data processing capabilities, such as a graphics processing unit (GPU), a central processing unit (CPU), or a neural-network processing unit (NPU).

[0124] The first processor generates an original reference map corresponding to the current image frame for the target object. Specifically, this can be done by rendering the target object, for example, within an Open Graphics Library (OpenGL), thereby generating the original reference map corresponding to the current image frame. The target object can be, for example, a virtual 3D object. In one embodiment, the original reference map includes an original depth map. In another embodiment, the original reference map may include both an original depth map and an original mask map. In other embodiments, the original reference map can also be other types of images, such as feature maps extracted via a feature extraction network.

[0125] After obtaining the original reference image, pixel data of the first target region can be extracted from the original reference image, and the first target region includes the first projection region of the target object projected onto the current image frame. That is, the range of the first target region is not less than the range of the target object projected onto the current image frame. In this embodiment, the pixels in the original reference image are called first pixels, and the pixels in the current image frame are called second pixels. The first pixels and the second pixels have a corresponding relationship, for example, a one-to-one correspondence. The original reference image and the current image frame are the same size, and the first pixels and second pixels with the same pixel coordinates have a corresponding relationship. The pixel value of each first pixel in the original reference image represents the positional information between the second pixel corresponding to the first pixel in the current image frame and the target object. In a specific embodiment, the pixel value of each first pixel in the original reference image is, for example, depth information. In this case, the pixel value of each first pixel corresponding to the first projection region in the original depth image represents the distance between the corresponding projection point of the target object and the shooting device (e.g., a camera) of the current image frame. The depth value outside the first projection region can be 0. In one specific implementation, the pixel data also includes the pixel coordinates of the pixel and the position information of the pixel, such as the pixel being a vertex, the pixel being the top-left vertex, the pixel being the bottom-right vertex, etc.

[0126] Step S52: Transfer the pixel data of the first target area from the first processor to the second processor.

[0127] The second processor can be a GPU, CPU, or NPU, etc. In one specific embodiment, the first processor is a GPU and the second processor is a CPU. The data transfer method between different processors can be a common method in the computer field, which will not be elaborated here.

[0128] Step S53: The second processor generates a target reference map corresponding to the current image frame based on the pixel data of the first target region, and determines the initial pose of the current image frame based on the target reference map.

[0129] The first pixel of the original reference image corresponds to the third pixel of the target reference image. For example, the original reference image and the target reference image are the same size, and pixels with the same pixel coordinates correspond to each other.

[0130] After receiving the pixel data of the first target region, the second processor can use the pixel data of the first target region to generate a target reference map corresponding to the current image frame. For example, it can use the pixel coordinates and pixel values ​​of each pixel in the pixel data of the first target region to obtain the target reference map corresponding to the current image frame.

[0131] After receiving the pixel data of the first target area, the second processor indicates that it has acquired relevant data about the target object. Therefore, in this example, the target reference map can be used to determine the relative pose between the target object and the current image frame. The relative pose between the target object and the current image frame, that is, the pose of the target object relative to the current image frame, can be understood as the relative pose between the target object and the imaging device when capturing the current image frame.

[0132] The target reference map, such as a target depth map or a target mask map, can be used to obtain the relative pose between the target object and the current image frame. The target reference map and the original reference map can be the same size.

[0133] Therefore, by limiting the transmission of pixel data of the target area to the second processor, instead of transmitting the entire original reference image to the second processor, the amount of data that needs to be transmitted can be reduced, the data transmission speed can be increased, and the second processor can obtain the target reference image more quickly, thereby improving the overall image data processing speed of the second processor.

[0134] In one embodiment, the above step "generating a target reference map corresponding to the current image frame based on the pixel data of the first target region" includes steps S531 and S532.

[0135] Step S531: Determine the target location of the first target area in the original reference image.

[0136] In one implementation, when the first target region is a regular shape, the pixel coordinates of each vertex of the first target region can be determined, and then the pixel coordinates of the pixels within the area enclosed by the vertices can be determined, thereby determining the target position of the first target region in the original reference image. In a specific implementation, when the first target region is rectangular, the pixel coordinates of the left vertex of the first target region can also be determined, and then the pixel coordinates of the four vertices of the first target region can be determined based on the width and height of the first target region, thereby determining the target position of the first target region in the original reference image.

[0137] In one specific implementation, the target position of the first target region in the original reference image is determined. Specifically, the pixel coordinates of each first pixel point within the first target region can be determined to determine the target position of the first target region in the original reference image.

[0138] In one specific implementation, the target position of the first pixel point at the edge of the first target region in the original reference image can be determined, and then the pixel coordinates of the first pixel point in the area enclosed by all the first pixel points at the edge can be determined, thereby determining the target position of the first target region in the original reference image.

[0139] In one specific implementation, the first target region is a rectangle with width w and height h. The coordinates of the top-left vertex pixel of the first target region in the original reference image are (u, v). The width of the original reference image is c and the height is r. The pixel located at (x, y) is represented as p(x, y). The target position can be represented as R(u, v, w, h).

[0140] In one specific implementation, the target location of the first target region in the original reference image can be determined by the following formula (1).

[0141]

[0142] Where n = (x2-x1)×(y2-y1).

[0143] Each pixel within the first target region can be represented as:

[0144] p0=p(x1,y1), p1=p(x1,y1+1),...,p n =p(x2, y2)

[0145] The ranges of values ​​for x1, y1, x2, and y2 are as follows:

[0146] x1 = min(u, r)

[0147] y1 = min(v, c)

[0148] x2 = max(u + h, r)

[0149] y2 = max(v + w, c)

[0150] Where x1 = min(u, r) and y1 = min(v, c) indicate that the coordinates of the top-left vertex pixel of the first target region can only be the bottom-right vertex of the original reference image at most. x2 = max(u+h, r) and y2 = max(v+w, c) indicate that the coordinates of the bottom-right vertex pixel of the first target region can only be the bottom-right vertex of the original reference image at most. This ensures that the first target region is always the region on the original reference image. By determining the coordinates of each pixel in the first target region, the target position of the first target region in the original reference image can be determined.

[0151] In one implementation, step S531 specifically includes steps S5311 and S5312.

[0152] Step S5311: Based on the sequence number of each pixel value in the pixel data, determine the first position of the first pixel point corresponding to each pixel value.

[0153] In this embodiment, the pixel data includes several pixel values ​​with different serial numbers. Each serial number represents the pixel value of a first pixel in the first target area. The serial number is determined based on the first position of the corresponding first pixel in the original reference image. For example, the serial number can be the arrangement order of the first pixels in the first target area. Starting from the first pixel on the left in the first row of the first target area, its serial number is 1; the second pixel on the left has the serial number 2, and so on. After determining the serial number of the first pixel in the first row, the serial number is determined again starting from the first pixel on the left in the second row. For example, if the serial number of the rightmost pixel in the first row is 50, then the serial number of the first pixel on the left in the second row is 51. Alternatively, starting from the left, the serial number of the first pixel in the first row is 11; the serial number of the second pixel in the first row is 12; the serial number of the first pixel in the second row is 21; the serial number of the second pixel in the second row is 22, and so on.

[0154] Therefore, based on the coordinates of the top-left vertex of the first target region and the sequence number of each pixel value in the pixel data, the first position of the first pixel corresponding to the pixel value can be determined, i.e., the pixel coordinates of each first pixel. In other words, by determining the sorting method, the first position of the first pixel corresponding to the pixel value can be determined.

[0155] For example, the first target area is a rectangle with width w and height h. The coordinates of the top-left vertex pixel of the first target area in the original reference image are (u, v). The width of the original reference image is c and the height is r. x1 = u, y1 = v, x2 = u+, y2 = v+w. The coordinates of the left vertex pixel of the first target area are p0 = p(x1, y1). We know that the pixel value of p0 is index 1. Then the coordinates of the pixel with index 2 can be determined as p1 = p(x1, y1+1). When y1+1 = v+w, we can determine that the pixel value corresponding to this pixel value is the last pixel of the first row, and all the previous pixels are pixels of the first row. Then, we can continue to determine the first position of the pixel in the second row, for example, p(x1+1, y1), until y1+1=v+w. We can then determine that the pixel corresponding to this pixel value is the last pixel in the second row. By analogy, we can determine the first position of the pixel corresponding to each pixel value, that is, determine the coordinates of each pixel.

[0156] Step S5312: Obtain the target position based on the first position.

[0157] Understandably, after determining the first position of each first pixel, the target position of the first target region composed of the first pixels with determined first positions in the original reference image can be determined accordingly. For the specific method of determining the target position, please refer to step S131 above, which will not be repeated here.

[0158] Step S532: Based on the pixel data, determine the pixel value of the third pixel in the second target region of the target reference image located at the target position.

[0159] In this embodiment, the pixels in the target reference image and the original reference image have a corresponding relationship, and the pixels in the target reference image are referred to as third pixels. Therefore, when the target position of the first target region in the original reference image is determined, a second target region corresponding to the target position can also be determined in the target reference image accordingly. Specifically, the second target region can be determined by using the pixel coordinates of the first pixel of the first target region in the original reference image to identify the region composed of the corresponding third pixels in the target reference image. For example, if the pixel coordinates of a certain pixel in the first target region in the original reference image are (50, 50), then the pixel with pixel coordinates (50, 50) in the target reference image can be identified as the corresponding third pixel based on these pixel coordinates.

[0160] After determining the correspondence between the pixels in the first target region and the second target region, the pixel value of the third pixel in the second target region, where the target reference image is located, can be determined based on the pixel data of the first target region. For example, all corresponding third pixels in the second target region can be set to the same pixel value as the first pixel in the first target region.

[0161] Therefore, by determining the target position of the first target region in the original reference image and determining the pixel value of the third pixel in the second target region of the target reference image at the target position, the transmission of image data of the first projection region of the target object projected onto the current image frame between the first processor and the second processor is realized.

[0162] In one specific implementation, in response to the target reference map including a target depth map, the pixel value of a third pixel in the target depth map located outside the target location can be set to a preset depth value. The preset depth value is, for example, 0. It is understood that since the third pixel outside the target location is not associated with the target object, the pixel value of the third pixel in the target depth map located outside the target location can be directly set to the preset depth value, thereby obtaining a complete target depth map.

[0163] In one specific implementation, in response to the target reference image including a target mask image, the pixel value of the third pixel in the target mask image located at the target position can be set to a first pixel value, and the pixel value of the third pixel in the target mask image located outside the target position can be set to a second pixel value. The second pixel value indicates that the third pixel does not belong to the projection point of the target object, and the second pixel value is, for example, 0. For the target mask image, the third pixels outside the target position are not associated with the target object. Therefore, the pixel value of the third pixel in the target mask image located outside the target position can be directly set to the second pixel value, and the second pixel value can be used to indicate that the third pixel does not belong to the projection point of the target object. In this way, the third pixels whose pixel value is not the second pixel value can be determined to be the projection points of the target object.

[0164] In one embodiment, steps S61 and S62 may be performed before performing the above-described step "extracting pixel data of the first target region from the original reference image".

[0165] Step S61: Generate a bounding box surrounding the target object.

[0166] In one implementation, the bounding box encloses a larger area than the target object occupies; the bounding box is, for example, a regular shape such as a cube or cuboid. Specifically, for example, OpenGL generates a bounding box around the target object. In another implementation, the bounding box can be directly the outline of the target object.

[0167] Step S62: Project the bounding box onto the second projection region of the current image frame as the first target region.

[0168] Since the bounding box can enclose the target object, the second projection area projected onto the current image frame must also include the object projected onto the current image frame. Therefore, the second projection area can be used as the first target area.

[0169] Therefore, by generating a bounding box surrounding the target object, it can be determined that the second projection region contains the first projection region when the second projection region is used as the first target region. When the bounding box has a regular shape, the second projection region also has a regular shape. Therefore, the second projection region can be determined by determining the pixel coordinates of each vertex of the second projection region, thus enabling rapid determination of the second projection region.

[0170] In one embodiment, the aforementioned step "determining the pixel value of the third pixel point in the second target region of the target reference image located at the target position based on pixel data" specifically includes steps S71 to S73.

[0171] In this embodiment, the original reference map includes an original depth map, and the pixel data includes first pixel data obtained by extracting the pixel values ​​of the first target region from the original depth map. The pixel values ​​of each first pixel point corresponding to the first projection region in the original depth map represent the distance between the corresponding projection point of the target object and the capturing device in the current image frame. When the first target region is larger than the first projection region, the pixel values ​​of the first pixel points in the regions outside the first projection region of the first target region can be preset depth values. In this embodiment, the target reference map includes a target depth map.

[0172] Step S71: Obtain the pixel value of each first pixel point in the first target region from the first pixel data, and use it as the pixel value of the third pixel point corresponding to the first pixel point in the second target region of the target depth map located at the target position.

[0173] Each first pixel in the first target region and the third pixel in the second target region located at the target position can have a corresponding relationship, such as being in the same position. For example, the pixel coordinates of each first pixel in the first target region are the same as the pixel coordinates of the corresponding third pixel in the second target region. Therefore, by obtaining the pixel values ​​of each first pixel in the first target region from the first pixel data, the pixel values ​​of each first pixel in the first target region can be used as the pixel values ​​of the third pixels in the second target region located at the target position. For example, if the coordinates of a first pixel in the first target region are (50, 50) and its pixel value is A, then the target position can also be determined to be (50, 50), and thus the pixel value of the target position can also be determined to be A. In one specific embodiment, the pixel values ​​of the third pixels outside the second target region can also be set to a preset depth value. Thus, by using the pixel values ​​of each pixel in the first target region as the pixel values ​​of the corresponding pixels in the second target region located at the target position, the reconstruction of the target depth map can be achieved through the second processor, thereby realizing the transmission of the depth map between the first and second processors.

[0174] In one specific implementation, a target depth map of the same size as the original depth map can be created. Then, the pixel value of each third pixel in the target depth map is set to a preset depth value, such as 0. At this time, the pixel value of each first pixel in the first target region is set to... Where loc(x, y) represents the coordinates of the first pixel in the first target region, and the pixel value D of the corresponding third pixel in the second target region is... map (x, y) can be

[0175] In one implementation, the target reference image further includes a target mask image, and the original reference image does not include the original mask image. The aforementioned step "determining the pixel value of the third pixel point in the second target region of the target reference image located at the target position based on pixel data" further includes steps S72 and S73.

[0176] Step S72: Determine the second position of the first projection area in the target depth map.

[0177] The second position of the first projected region in the target depth map can be determined based on the pixel data of the third pixel of the first projected region. For example, the pixel data of the third pixel of the first projected region includes pixel coordinates, and thus the second position of the first projected region in the target depth map can be determined accordingly based on the pixel data of the third pixel of the first projected region, including pixel coordinates.

[0178] In one specific implementation, the aforementioned step "determining the second position of the first projection area in the target depth map" specifically includes steps S721 and S722.

[0179] Step S721: Find the third pixel point whose pixel value is not a preset depth value from the target depth map.

[0180] Step S722: Determine the second position based on the position of the found third pixel in the target depth map.

[0181] In the original depth map, areas outside the first projection region of the target object can be considered to have no depth value. Therefore, the depth value of the first pixel in areas outside the first projection region can be set to a preset depth value. Thus, a third pixel whose pixel value is not a preset depth value can be found in the target depth map. The position of this third pixel in the target depth map is then determined. Furthermore, based on the position of this third pixel in the target depth map, the second position of the first projection region in the entire target depth map is determined. Therefore, by finding a third pixel whose pixel value is not a preset depth value in the target depth map, the second position of the first projection region in the target depth map can be determined.

[0182] Step S73: Determine the region located at the second position in the target mask as the first projection region, and set the pixel value of the third pixel in the first projection region of the target mask as the first pixel value.

[0183] The pixels in the target mask and the target depth map have a corresponding relationship; for example, both the target mask and the target depth map are images of the same size. In this case, the region in the target mask located at the second position in the first projection region of the target depth map can be determined. Then, the pixel value of the third pixel in the first projection region of the target mask can be set to the first pixel value. The first pixel value indicates that the second pixel corresponding to the third pixel belongs to the projection point of the target object; for example, the first pixel value is 1. The pixel value of the third pixel in the region outside the first projection region of the target mask can be set to the second pixel value; for example, the second pixel value is 0. This is how the target mask is obtained.

[0184] In one specific embodiment, the depth value of the region corresponding to the first projection region in the second target region of the target depth map is not a preset depth value (the preset depth value is, for example, 0). In this case, the pixel value of the pixel in the first projection region of the target mask map can be set as the first pixel value according to the following formula (2), where the first pixel value is, for example, 1.

[0185]

[0186] Among them, D map (x, y) represents the depth value of each pixel in the target depth image, and (x, y) represents the coordinates of the pixel. M map (x, y) represents the pixel value of the pixel at coordinates (x, y) in the target mask image. Therefore, by determining whether the depth value of each pixel in the target depth image is a preset depth value, the second position can be determined, and the pixel value of the pixel in the first projection region of the target mask image can be set as the first pixel value.

[0187] Therefore, by setting the pixel value of the third pixel in the first projection area of ​​the target mask to the first pixel value, and setting the first pixel value to indicate that the second pixel corresponding to the third pixel belongs to the projection point of the target object, the reconstruction of the target mask is achieved through the second processor, thereby realizing the transmission of the mask between the first processor and the second processor.

[0188] In one embodiment, the target reference image further includes a target mask image, and the original reference image includes the original mask image. The aforementioned pixel data also includes second pixel data obtained by extracting pixel values ​​of the first target region from the original mask image. In this case, the aforementioned step of "determining the pixel value of the third pixel point in the second target region of the target reference image located at the target location based on the pixel data" may further include: obtaining the pixel values ​​of each first pixel point in the first target region from the second pixel data, so as to respectively serve as the pixel values ​​of the third pixel points corresponding to the first pixel points in the second target region of the target mask image located at the target location.

[0189] In this embodiment, the pixel value of the first pixel in the first projection region of the first target region in the original mask image can be a first pixel value, and the pixel value of the first pixel outside the first projection region can be set to a second pixel value, such as 0. At this time, by obtaining the pixel values ​​of each first pixel in the first target region from the second pixel data, and using them as the pixel values ​​of the third pixel corresponding to the first pixel in the second target region of the target mask image located at the target position, data transmission of the pixel at the target position is achieved. Furthermore, since the first target region includes the first projection region, data transmission of the pixel at the first projection region is also achieved, thereby realizing the reconstruction of the mask image.

[0190] The image data processing method of this application is illustrated below with reference to the accompanying drawings.

[0191] See Figure 5 , Figure 5 This is a schematic diagram of the image transmission process according to an embodiment of the image data processing method of this application. In this embodiment, the first processor is a GPU and the second processor is a CPU. The target object 501 has a corresponding bounding box 502. By using the rendering module 509, such as OpenGL, to render, an original depth map 503 and an original mask map 504 can be obtained. The box 505 in the original depth map 503 is the first target region obtained by projecting the bounding box 502, and the region 506 within the box 505 is the first projection region obtained by projecting the target object. By transferring the pixel data of the first target region from the GPU to the CPU, the CPU can obtain a target depth map 507 based on the pixel data of the first target region, and then obtain a target mask map 508 based on the target depth map 507, thereby realizing the data transmission of the depth map and the mask map between the GPU and the CPU.

[0192] In one embodiment, after the step "generating a target reference map corresponding to the current image frame based on the pixel data of the first target region by the second processor", at least one step of the following steps S81 and S82 can be performed.

[0193] Step S81: The second processor determines the relative pose between the target object and the current image frame based on the target reference map.

[0194] The relative pose between the target object and the current image frame can be understood as the relative pose between the target object and the imaging device when capturing the current image frame.

[0195] For a target reference map including a target depth map, the pixel coordinates of the third pixel in the first projection region of the target reference map can be used, combined with the projection matrix and depth value, to obtain the 3D coordinates of the 3D point corresponding to the second pixel in the first projection region. Then, the 3D coordinates of the 3D point are used to determine the relative pose between the target object and the current image frame. Specific pose calculation methods are those applicable to this field and will not be elaborated here.

[0196] If the target reference map also includes a target mask map, the target mask map can be used to obtain the outline of the target object, and then the pose calculation method commonly used in this field can be used to solve the relative pose between the target object and the current image frame.

[0197] Step S82: The second processor renders the current image frame based on the relative pose between the target object and the current image frame, so as to display the target object on the first projection area of ​​the current image frame.

[0198] After determining the relative pose between the target object and the imaging device, the current image frame is rendered based on the relative pose between the target object and the current image frame, and the target object is displayed on the first projection area of ​​the current image frame. The rendering method can be a rendering method commonly used in this field, and will not be described in detail here.

[0199] Therefore, by obtaining the target reference image, the relative pose between the target object and the current image frame can be determined. Furthermore, by determining the relative pose between the target object and the current image frame, the target object can be rendered using this relative pose, so that the target object is displayed on the first projection area of ​​the current image frame.

[0200] See Figure 6 , Figure 6 This is a schematic diagram of the framework of an embodiment of the image processing apparatus of this application. The image processing apparatus 60 includes an acquisition module 61, a first determination module 62, a second determination module 63, and a third determination module 64. The acquisition module 61 is used to acquire the initial pose of the target object relative to the current image frame and the historical pose of the target object relative to a first historical image frame; the first determination module 62 is used to adjust the initial pose based on the historical pose to obtain an adjusted pose, wherein the pose change between the adjusted pose and the historical pose is less than the pose change between the initial pose and the historical pose; the second determination module 63 is used to obtain the target pose of the target object relative to the current image frame based on the adjusted pose; the third determination module 64 is used to determine the display position of the target object in the current image frame based on the target pose of the target object relative to the current image frame.

[0201] The first determining module 62 mentioned above is used to adjust the initial pose based on the historical pose to obtain the adjusted pose, including: weighting the historical pose and the initial pose based on the first weight coefficient of the initial pose and the second weight coefficient of the historical pose to obtain the adjusted pose.

[0202] Wherein, the sum of the first weight coefficient and the second weight coefficient mentioned above is a preset value; the first weight coefficient is positively correlated with the pose change rate of the target object, and the second weight coefficient is negatively correlated with the pose change rate. The pose change rate indicates how fast the pose of the target object changes relative to at least three reference image frames, and the reference image frames include the current image frame or the second historical image frame.

[0203] The second determining module 63 is used to obtain the target pose of the target object relative to the current image frame based on the adjusted pose, including: using the adjusted pose as the target pose of the target object relative to the current image frame; or, using the target compensation pose corresponding to the current image frame to compensate the adjusted pose to obtain the target pose of the target object relative to the current image frame.

[0204] Before the second determining module 63 uses the target compensation pose corresponding to the current image frame to compensate the adjusted pose and obtain the target pose of the target object relative to the current image frame, the target compensation pose determining module of the image processing device 60 is used to determine the target compensation pose based on the pose change rate of the target object. The target compensation pose is positively correlated with the pose change rate, which indicates how fast the pose of the target object changes relative to at least two reference image frames. The reference image frames include the current image frame or the second historical image frame.

[0205] The aforementioned target compensation pose determination module is used to determine the target compensation pose based on the pose change rate of the target object, including: taking the pose change between the initial pose and the adjusted pose as the initial compensation pose; and determining the target compensation pose based on the initial compensation pose and the pose change rate.

[0206] The aforementioned target compensation pose determination module is used to determine the target compensation pose based on the initial compensation pose and the pose change rate, including: determining the third weight coefficient of the initial compensation pose and the fourth weight coefficient of the target compensation pose corresponding to the first historical image frame based on the pose change rate, wherein the sum of the third weight coefficient and the fourth weight coefficient is a preset value, and the third weight coefficient is positively correlated with the pose change rate of the target object; and performing weighted processing on the initial compensation pose and the target compensation pose corresponding to the first historical image frame based on the third weight coefficient and the fourth weight coefficient to obtain the target compensation pose.

[0207] Wherein, if one of the reference image frames includes the current image frame, the pose of the target object relative to the reference image frame includes the initial pose of the target object relative to the current image frame; if one of the reference image frames is a second historical image frame, the pose of the target object relative to the reference image frame includes the target pose of the target object relative to the second historical image frame.

[0208] The aforementioned whole pose includes rotation and translation components; the second determining module 63 is used to compensate the adjusted pose using the target compensation pose corresponding to the current image frame to obtain the target pose of the target object relative to the current image frame, including: compensating the translation component of the adjusted pose using the target compensation pose corresponding to the current image frame to obtain the compensated translation component; and obtaining the target pose of the target object relative to the current image frame based on the compensated translation component and the rotation component.

[0209] The aforementioned historical pose, initial pose, and adjusted pose all include rotation components, which are matrices composed of multiple elements. The aforementioned first determining module 62 is used to adjust the initial pose based on the historical pose to obtain the adjusted pose, including: adjusting the corresponding elements in the rotation component of the initial pose based on each element in the rotation component of the historical pose to obtain the corresponding elements in the rotation component of the adjusted pose.

[0210] The aforementioned historical pose is the target pose of the target object relative to the first historical image frame; the aforementioned acquisition module 61 is used to acquire the initial pose of the target object relative to the current image frame, including: processing the current image frame using a pose tracking algorithm to obtain the initial pose of the target object relative to the current image frame.

[0211] Wherein, the target object mentioned above is a three-dimensional object; and / or, the third determining module 64 is used to determine the display position of the target object in the current image frame based on the target pose of the target object relative to the current image frame, including: determining the projection position of the target object in the current image frame based on the target pose of the target object relative to the current image frame, as the display position; after the third determining module 64 determines the display position of the target object in the current image frame based on the target pose of the target object relative to the current image frame, the display module of the image processing device 60 is also used to display the target object at the display position in the current image frame.

[0212] Please see Figure 7 , Figure 7This is a schematic diagram of a framework of an embodiment of the electronic device of this application. The electronic device 70 includes a memory 701 and a processor 702 coupled to each other. The processor 702 is used to execute program instructions stored in the memory 701 to implement the steps of any of the above-described image processing method embodiments. In a specific implementation scenario, the electronic device 70 may include, but is not limited to, a microcomputer or a server. In addition, the electronic device 70 may also include mobile devices such as laptops and tablets, which are not limited here.

[0213] Specifically, processor 702 controls itself and memory 701 to implement the steps of any of the above-described image processing method embodiments. Processor 702 may also be referred to as a CPU (Central Processing Unit). Processor 702 may be an integrated circuit chip with signal processing capabilities. Processor 702 may also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor. Furthermore, processor 702 may be implemented using integrated circuit chips.

[0214] Please see Figure 8 , Figure 8 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. The computer-readable storage medium 80 stores program instructions 801 that can be executed by a processor. The program instructions 801 are used to implement the steps of any of the above-described image processing method embodiments.

[0215] The above scheme obtains the adjusted pose by adjusting the initial pose based on the historical pose, and sets the pose change between the adjusted pose and the historical pose to be less than the pose change between the initial pose and the historical pose. This reduces the pose change of the target object relative to the current image frame and the historical image frames, thus reducing the possibility of abrupt changes in the display position of the target object in the current image frame and the historical image frames. This also reduces jitter when displaying the target object later, improving the user experience.

[0216] This disclosure relates to the field of augmented reality (AR). It involves acquiring image information of target objects in a real-world environment and then using various visual algorithms to detect or identify the relevant features, states, and attributes of these objects, thereby achieving an AR effect that combines virtual and real elements to suit specific applications. For example, target objects may include human features such as faces, limbs, gestures, and movements; objects such as signs and markers; or venues such as sand tables, display areas, or displayed items. Visual algorithms may include visual localization, SLAM, 3D reconstruction, image registration, background segmentation, keypoint extraction and tracking of objects, and pose or depth detection. Specific applications can include interactive scenarios related to real-world scenes or objects, such as guided tours, navigation, explanations, reconstruction, and virtual effect overlay displays, as well as human-related special effects processing, such as makeup enhancement, body enhancement, special effects displays, and virtual model displays.

[0217] Convolutional neural networks (CNNs) can be used to detect or identify the relevant features, states, and attributes of target objects. The aforementioned CNNs are network models obtained through training using deep learning frameworks.

[0218] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0219] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0220] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0221] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0222] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0223] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. An image processing method, characterized in that, The method includes: Obtain the initial pose of the target object relative to the current image frame and the historical pose of the target object relative to the first historical image frame; The initial pose is adjusted based on the historical pose to obtain the adjusted pose, wherein the pose change between the adjusted pose and the historical pose is less than the pose change between the initial pose and the historical pose. Based on the adjusted pose, the target pose of the target object relative to the current image frame is obtained; Based on the target pose of the target object relative to the current image frame, determine the display position of the target object in the current image frame; The step of adjusting the initial pose based on the historical pose to obtain the adjusted pose includes: Based on the first weight coefficient of the initial pose and the second weight coefficient of the historical pose, the historical pose and the initial pose are weighted to obtain the adjusted pose. The first weight coefficient is positively correlated with the pose change rate of the target object, and the second weight coefficient is negatively correlated with the pose change rate. The pose change rate indicates how fast the pose of the target object changes relative to at least three reference image frames, including the current image frame or the second historical image frame.

2. The method according to claim 1, characterized in that, The sum of the first weighting coefficient and the second weighting coefficient is a preset value.

3. The method according to any one of claims 1 to 2, characterized in that, The step of obtaining the target pose of the target object relative to the current image frame based on the adjusted pose includes: The adjusted pose is taken as the target pose of the target object relative to the current image frame; Alternatively, the adjusted pose can be compensated using the target compensation pose corresponding to the current image frame to obtain the target pose of the target object relative to the current image frame.

4. The method according to claim 3, characterized in that, Before compensating the adjusted pose using the target compensation pose corresponding to the current image frame to obtain the target pose of the target object relative to the current image frame, the method further includes: The target compensation pose is determined based on the pose change rate of the target object, wherein the target compensation pose is positively correlated with the pose change rate, and the pose change rate indicates how fast the pose of the target object changes relative to at least two reference image frames, wherein the reference image frames are the current image frame or the second historical image frame.

5. The method according to claim 4, characterized in that, Determining the target compensated pose based on the pose change rate of the target object includes: The pose change between the initial pose and the adjusted pose is used as the initial compensation pose. The target compensated pose is determined based on the initial compensated pose and the pose change rate.

6. The method according to claim 5, characterized in that, Determining the target compensated pose based on the initial compensated pose and the pose change rate includes: Based on the pose change rate, a third weighting coefficient of the initial compensated pose and a fourth weighting coefficient of the target compensated pose corresponding to the first historical image frame are determined, wherein the sum of the third weighting coefficient and the fourth weighting coefficient is a preset value, and the third weighting coefficient is positively correlated with the pose change rate of the target object. Based on the third and fourth weighting coefficients, the initial compensation pose and the target compensation pose corresponding to the first historical image frame are weighted to obtain the target compensation pose.

7. The method according to claim 2 or 4, characterized in that, When the reference image frame includes the current image frame, the pose of the target object relative to the reference image frame includes the initial pose of the target object relative to the current image frame; when the reference image frame includes the second historical image frame, the pose of the target object relative to the reference image frame includes the target pose of the target object relative to the second historical image frame.

8. The method according to claim 3, characterized in that, The pose adjustment includes rotation and translation components; The step of compensating the adjusted pose using the target compensation pose corresponding to the current image frame to obtain the target pose of the target object relative to the current image frame includes: The translation component of the adjusted pose is compensated using the target compensation pose corresponding to the current image frame to obtain the compensated translation component; The target pose of the target object relative to the current image frame is obtained based on the compensated translation component and the rotation component.

9. The method according to any one of claims 1 to 8, characterized in that, The historical pose, initial pose, and adjusted pose all include a rotation component, which is a matrix composed of multiple elements. The step of adjusting the initial pose based on the historical pose to obtain the adjusted pose includes: Based on each element in the rotation component of the historical pose, the corresponding elements in the rotation component of the initial pose are adjusted to obtain the corresponding elements in the rotation component of the adjusted pose.

10. The method according to any one of claims 1 to 9, characterized in that, The historical pose is the target pose of the target object relative to the first historical image frame; And / or, obtaining the initial pose of the target object relative to the current image frame includes: The current image frame is processed using a pose tracking algorithm to obtain the initial pose of the target object relative to the current image frame.

11. The method according to any one of claims 1 to 10, characterized in that, The target object is a three-dimensional object; And / or, Determining the display position of the target object in the current image frame based on the target pose of the target object relative to the current image frame includes: determining the projection position of the target object in the current image frame based on the target pose of the target object relative to the current image frame, and using it as the display position; And / or, after determining the display position of the target object in the current image frame based on the target pose relative to the current image frame, the method further includes: displaying the target object at the display position in the current image frame.

12. An image processing apparatus, characterized in that, include: The acquisition module is used to acquire the initial pose of the target object relative to the current image frame and the historical pose of the target object relative to the first historical image frame; The first determining module is used to adjust the initial pose based on the historical pose to obtain the adjusted pose, wherein the pose change between the adjusted pose and the historical pose is less than the pose change between the initial pose and the historical pose. The second determining module is used to obtain the target pose of the target object relative to the current image frame based on the adjusted pose. The third determining module is used to determine the display position of the target object in the current image frame based on the target pose of the target object relative to the current image frame; The step of adjusting the initial pose based on the historical pose to obtain the adjusted pose includes: Based on the first weight coefficient of the initial pose and the second weight coefficient of the historical pose, the historical pose and the initial pose are weighted to obtain the adjusted pose. The first weight coefficient is positively correlated with the pose change rate of the target object, and the second weight coefficient is negatively correlated with the pose change rate. The pose change rate indicates how fast the pose of the target object changes relative to at least three reference image frames, including the current image frame or the second historical image frame.

13. An electronic device, characterized in that, This includes interconnected processors and memory, where, The processor is configured to execute the computer program stored in the memory to perform the method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, The device stores a computer program that can be executed by a processor, the computer program being used to implement the method as described in any one of claims 1 to 11.