An augmented reality method and its related device

By calculating the pose change of real objects and virtual objects, and adjusting the orientation of virtual objects to adapt to real objects, the user experience problem caused by the fixed orientation of virtual objects in augmented reality technology is solved, and the vividness of the image is improved.

CN113066125BActive Publication Date: 2025-07-29HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110221723.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-27
Publication Date
2025-07-29
Estimated Expiration
2041-02-27

AI Technical Summary

Technical Problem

Existing augmented reality technology cannot dynamically adjust the orientation of virtual objects to adapt to the movement changes of real objects, resulting in poor user experience.

Method used

By obtaining the position information of the real object and the virtual object in the camera coordinate system, the position and orientation of the virtual object are calculated, and the position and orientation of the virtual object are adjusted according to the change amount, so that it is associated with the orientation of the real object.

Benefits of technology

The dynamic correlation between the orientation of virtual objects and real objects is realized, improving the fidelity and user experience of the target image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113066125B_ABST
    Figure CN113066125B_ABST
Patent Text Reader

Abstract

The present application provides an augmented reality method and related device, which can dynamically associate the orientation of a virtual object with the orientation of a real object, making the combination of the real object and the virtual object presented in the target image more realistic and improving the user experience. The method of the present application includes: obtaining a target image captured by a camera and first position information of a first object in the target image; obtaining second position information of a second object in a three-dimensional coordinate system corresponding to the camera and third position information of a third object in the three-dimensional coordinate system, where the second object is a reference object for the first object; obtaining a pose change amount of the first object relative to the second object according to the first position information and the second position information; transforming the third position information according to the pose change amount to obtain fourth position information of the third object in the three-dimensional coordinate system; and rendering the third object in the target image according to the fourth position information to obtain a new target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to an augmented reality method and related devices. Background Art

[0002] Augmented reality (AR) technology can accurately "embed" virtual objects that do not exist in the real environment into the real environment, so that the virtual objects are integrated with the real environment, thereby presenting a new environment with a realistic sensory effect to the user to achieve the enhancement of reality.

[0003] Currently, AR technology can additionally render a virtual object on the image presenting the real environment, thereby obtaining a new image for the user to view and use. For example, if the target image presents a person in a room, based on the user's needs, a virtual wing can be added behind the person, so that the new target image presents more vivid and interesting content.

[0004] The orientation of the real object presented in the target image is usually not fixed, but will change with the shooting angle of the camera or the movement of the real object. For example, when the real object faces the camera directly, or faces the camera sideways, or faces the camera with its back, the orientation presented in the target image is different. However, the orientation of the virtual object rendered by the current AR technology is fixed, that is, the orientation of the virtual object cannot change with the change of the orientation of the real object, resulting in the target image being unable to present realistic content and poor user experience. Summary of the Invention

[0005] Embodiments of this application provide an augmented reality method and related devices, which can dynamically associate the orientation of the virtual object with the orientation of the real object, so that the target image presents realistic content and improves the user experience.

[0006] The first aspect of the embodiments of this application provides an augmented reality method, which includes:

[0007] When responding to a user's request to add a virtual object to the target image presenting the real environment, first obtain the target image captured by the camera. The real environment presented in the target image includes multiple real objects (that is, the target image is composed of images of multiple real objects), such as human bodies, trees, houses, and other objects. The multiple real objects include a first object. For example, the first object can be the real object that the user is concerned about. Further, the virtual object to be added can be regarded as a third object, and the reference object of the real object can be regarded as a second object. For example, if the user wants to add a virtual wing to the real human body in the target image, the three-dimensional standard human model can be used as the reference object of the real human body.

[0008] Next, obtain the first position information of the first object in the target image (which can be understood as a two-dimensional coordinate system constructed based on the target image, i.e., the image coordinate system), the second position information of the second object in the three-dimensional coordinate system corresponding to the camera (i.e., the camera coordinate system), and the third position information of the third object in the camera coordinate system. The second position information and the third position information are preset and associated information. Still as in the above example, in the camera coordinate system, the standard human model can be preset at a certain position and with a certain orientation, and the wings can also be preset at another position and with a certain orientation (the orientations preset for the model and the wings are usually the same or similar). Therefore, the second position information of the standard human model in the camera coordinate system and the third position information of the wings in the camera coordinate system can be obtained. The second position information includes the three-dimensional coordinates of each surface point of the standard human model in the camera coordinate system, and the third position information includes the three-dimensional coordinates of each surface point of the wings in the camera coordinate system. In addition, since the real human body occupies a certain position in the target image, the first position information of the real human body in the image coordinate system can also be obtained. The first position information includes the two-dimensional coordinates of each surface point of the real human body in the target image.

[0009] Then, since the second object is a reference object for the first object, the pose change amount of the first object relative to the second object can be determined according to the second position information of the second object and the first position information of the first object. This pose change amount is used to indicate the change in position and the change in orientation between the position of the second object and the position of the first object in the camera coordinate system. Still as in the above example, by calculating according to the second position information of the standard human model and the first position information of the real human body, the pose change amount of the real human body relative to the standard human model can be determined. This pose change amount is used to indicate the position change and the orientation change between the real human body and the standard human model in the camera coordinate system.

[0010] Since the pose change amount of the first object relative to the second object is used to indicate the position change and the orientation change from the second object to the first object in the camera coordinate system, the fourth position information of the third object in the camera coordinate system can be determined by using this pose change amount and the third position information of the third object, that is, the final position and the final orientation of the third object are determined. Still as in the above example, after obtaining the pose change amount of the real human body, the fourth position information of the wings can be calculated based on this pose change amount and the third position information of the wings, that is, the wings are translated from their original position to the final position and rotated from their original orientation to the final orientation. It can be seen that the rotation and translation operations of the wings and the rotation and translation operations of transforming the standard human model to the real human body are the same.

[0011] Finally, render the third object in the target image according to the fourth position information of the third object to obtain a new target image. In this way, the new target image can present both the third object and the first object, thus meeting the user's needs. Still taking the above example, after obtaining the fourth position information of the wings, the wings can be rendered in the target image according to the fourth position information of the wings. At this point, wings can be displayed near the real human body in the target image, and the orientation of the real human body and the orientation of the wings are dynamically associated, thus meeting the user's needs.

[0012] As can be seen from the above method: when it is necessary to add a third object to the target image, the second position information of the second object in the camera coordinate system, the third position information of the third object in the camera coordinate system, and the first position information of the first object in the target image in the image coordinate system can be obtained first. The second position information and the third position information are preset information. Since the second object is a reference object for the first object, the pose change amount of the first object relative to the second object can be determined according to the second position information and the first position information. This pose change amount is used to indicate the position change and orientation change from the second object to the first object in the camera coordinate system. Then, the third object can be made to undergo the same position change and orientation change, that is, the fourth position information of the third object in the camera coordinate system is determined according to this pose change amount and the third position information of the third object. In this way, the final orientation of the third object can be associated with the orientation of the first object (for example, their orientations are the same or their orientations are similar). Finally, render the third object in the target image according to the fourth position information. In the obtained new target image, the orientation of the third object can adapt to the orientation of the first object, and realistic content can be presented, thus improving the user experience.

[0013] In a possible implementation manner, determining the pose change amount of the first object relative to the second object according to the second position information of the second object and the first position information of the first object includes: first obtaining the depth information of the first object, and this depth information is used to indicate the distance from the first object to the camera. Then, according to the first position information of the first object in the image coordinate system and the depth information of the first object, determine the fifth position information of the first object in the camera coordinate system. Finally, calculate the change amount between the second position information of the second object in the camera coordinate system and the fifth position information of the first object in the camera coordinate system, so as to accurately obtain the pose change amount of the first object relative to the second object.

[0014] In a possible implementation, determining the pose change amount of the first object relative to the second object according to the second position information of the second object and the first position information of the first object includes: transforming the second position information of the second object in the camera coordinate system (equivalent to performing a rotation and translation operation on the second object) to obtain the fifth position information of the first object in the camera coordinate system. Then, projecting the fifth position information of the first object in the camera coordinate system onto the image coordinate system (i.e., the target image) to obtain the sixth position information of the first object in the image coordinate system. Finally, if the change amount between the sixth position information and the first position information meets a preset condition, the transformation matrix for transforming the second position information of the second object in the camera coordinate system is determined as the pose change amount of the first object relative to the second object.

[0015] In a possible implementation, rendering the third object in the target image according to the fourth position information of the third object in the camera coordinate system to obtain a new target image includes: first, performing pinhole imaging according to the fourth position information of the third object in the camera coordinate system to obtain an image of the third object. Then, obtaining the occlusion relationship between the third object and the first object. Finally, fusing the image of the third object and the image of the first object according to the occlusion relationship to obtain a new target image. In the foregoing implementation, when fusing the image of the third object into the target image, considering the occlusion relationship between the third object and the first object can enable the new target image to correctly present the relative position relationship between the third object and the first object, that is, make the content presented by the new target image more realistic and further improve the user experience.

[0016] In a possible implementation, obtaining the occlusion relationship between the third object and the first object includes: first, calculating the first distance between the first object and the origin of the camera coordinate system according to the fifth position information of the first object in the camera coordinate system. Then, calculating the second distance between the third object and the origin of the camera coordinate system according to the fourth position information of the third object in the camera coordinate system. Finally, comparing the first distance and the second distance can accurately obtain the occlusion relationship between the third object and the first object. For example, if the first distance is less than or equal to the second distance, the third object is occluded by the first object; if the first distance is greater than the second distance, the first object is occluded by the third object.

[0017] In a possible implementation, obtaining the occlusion relationship between the third object and the first object includes: first obtaining the correspondence between multiple surface points of the first object and multiple surface points of the second object. Then, according to this correspondence, determining the distribution of multiple surface points of the first object on the second object. Finally, according to this distribution, determining the occlusion relationship between the third object and the first object. For example, among all the surface points of the first object, if the number of surface points on the front of the second object is greater than or equal to the number of surface points on the back of the second object, the third object is occluded by the first object; if the number of surface points on the front of the second object is less than the number of surface points on the back of the second object, the first object is occluded by the third object.

[0018] In a possible implementation, the pose change amount of the first object relative to the second object includes the position of the first object and the orientation of the first object. Obtaining the occlusion relationship between the third object and the first object includes: first determining the front of the first object according to the orientation of the first object. Then, according to the angle between the orientation from the center point of the front of the first object to the origin of the camera coordinate system and the orientation of the first object, determining the occlusion relationship between the third object and the first object. For example, if this angle is less than or equal to 90°, the third object is occluded by the first object; if this angle is greater than 90°, the first object is occluded by the third object.

[0019] In a possible implementation, according to the occlusion relationship, the method further includes: inputting the target image into the first neural network to obtain an image of the first object.

[0020] The second aspect of the embodiments of the present application provides an augmented reality method, which includes: obtaining a target image and the third position information of the third object in the camera coordinate system, the target image including an image of the first object; inputting the target image into the second neural network to obtain the pose change amount of the first object relative to the second object, the second neural network being trained according to the second position information of the second object in the camera coordinate system, the second object being a reference object for the first object, and the second position information and the third position information being preset information; determining the fourth position information of the third object in the camera coordinate system according to the pose change amount and the third position information; rendering the third object in the target image according to the fourth position information to obtain a new target image.

[0021] As can be seen from the above method: when a third object needs to be added to the target image, the target image and the third position information of the third object in the camera coordinate system can be obtained first, and the target image includes the image of the first object. Then, the target image is input into the second neural network to obtain the pose change amount of the first object relative to the second object. The second neural network is trained according to the second position information of the second object in the camera coordinate system. The second object is the reference object of the first object, and the second position information and the third position information are preset information. Specifically, the pose change amount of the first object relative to the second object is used to indicate the position change and orientation change from the second object to the first object in the camera coordinate system. Therefore, the third object can be made to undergo the same position change and orientation change, that is, the fourth position information of the third object in the camera coordinate system is determined according to the pose change amount and the third position information of the third object. In this way, the final orientation of the third object can be associated with the orientation of the first object. Finally, the third object is rendered in the target image according to the fourth position information. In the obtained new target image, the orientation of the third object can be adapted to the orientation of the first object, presenting realistic content and thus improving the user experience.

[0022] In a possible implementation manner, rendering the third object in the target image according to the fourth position information to obtain a new target image includes: performing pinhole imaging according to the fourth position information to obtain an image of the third object; obtaining the occlusion relationship between the third object and the first object; and fusing the image of the third object and the image of the first object according to the occlusion relationship to obtain a new target image.

[0023] In a possible implementation manner, the method further includes: transforming the second position information according to the pose change amount to obtain the fifth position information of the first object in the three-dimensional coordinate system corresponding to the camera.

[0024] In a possible implementation manner, obtaining the occlusion relationship between the third object and the first object includes: calculating the first distance between the first object and the origin of the three-dimensional coordinate system according to the fifth position information; calculating the second distance between the third object and the origin of the three-dimensional coordinate system according to the fourth position information; and comparing the first distance and the second distance to obtain the occlusion relationship between the third object and the first object.

[0025] In a possible implementation manner, the pose change amount includes the orientation change of the first object relative to the second object. Obtaining the occlusion relationship between the third object and the first object includes: determining the front of the first object according to the orientation change of the first object relative to the second object; and obtaining the occlusion relationship between the third object and the first object according to the included angle between the orientation from the center point of the front of the first object to the origin of the three-dimensional coordinate system and the orientation of the first object.

[0026] In a possible implementation, according to the occlusion relationship, the method further includes: inputting the target image into a first neural network to obtain an image of a first object.

[0027] A third aspect of the embodiments of the present application provides a model training method, the method including: obtaining a to-be-trained image, the to-be-trained image including an image of a first object; inputting the to-be-trained image into a to-be-trained model to obtain a pose change amount of the first object relative to a second object; calculating, by a preset target loss function, a deviation between the pose change amount of the first object relative to the second object and a true pose change amount of the first object, the true pose change amount of the first object being determined according to second position information of the second object in a camera coordinate system, the second object being a reference object for the first object, and the second position information being preset information; updating parameters of the to-be-trained model according to the deviation until a model training condition is satisfied to obtain a second neural network.

[0028] It can be seen from the above method that: the second neural network trained by this method can accurately obtain the pose change amount of the object in the target image.

[0029] A fourth aspect of the embodiments of the present application provides a model training method, the method including: obtaining a to-be-trained image; obtaining an image of a first object through the to-be-trained model; calculating, by a preset target loss function, a deviation between the image of the first object and a true image of the first object; updating parameters of the to-be-trained model according to the deviation until a model training condition is satisfied to obtain a first neural network.

[0030] It can be seen from the above method that: the first neural network trained by this method can accurately obtain the image of the first object in the target image.

[0031] A fifth aspect of the embodiments of the present application provides an augmented reality device, the device including: a first acquisition module, configured to acquire a target image captured by a camera and first position information of a first object in the target image; a second acquisition module, configured to acquire second position information of a second object in a three-dimensional coordinate system corresponding to the camera and third position information of a third object in the three-dimensional coordinate system, the second object being a reference object for the first object, and the second position information and the third position information being preset information; a third acquisition module, configured to acquire a pose change amount of the first object relative to the second object according to the first position information and the second position information; a transformation module, configured to transform the third position information according to the pose change amount to obtain fourth position information of the third object in the three-dimensional coordinate system; and a rendering module, configured to render the third object in the target image according to the fourth position information to obtain a new target image.

[0032] As can be seen from the above device: when a third object needs to be added to the target image, the second position information of the second object in the camera coordinate system, the third position information of the third object in the camera coordinate system, and the first position information of the first object in the target image in the image coordinate system can be obtained first. The second position information and the third position information are preset information. Since the second object is a reference object for the first object, the pose change amount of the first object relative to the second object can be determined according to the second position information and the first position information. This pose change amount is used to indicate the position change and orientation change from the second object to the first object in the camera coordinate system. Then, the third object can be made to undergo the same position change and orientation change, that is, the fourth position information of the third object in the camera coordinate system is determined according to this pose change amount and the third position information of the third object. In this way, the final orientation of the third object can be associated with the orientation of the first object (for example, their orientations are the same or their orientations are similar). Finally, the third object is rendered in the target image according to the fourth position information. In the obtained new target image, the orientation of the third object can be adapted to the orientation of the first object, presenting realistic content and thus improving the user experience.

[0033] In a possible implementation manner, the third acquisition module is configured to: acquire the depth information of the first object; according to the first position information and the depth information, acquire the fifth position information of the first object in the three-dimensional coordinate system; calculate the change amount between the second position information and the fifth position information to obtain the pose change amount of the first object relative to the second object.

[0034] In a possible implementation manner, the third acquisition module is configured to: transform the second position information to obtain the fifth position information of the first object in the three-dimensional coordinate system; project the fifth position information onto the target image to obtain the sixth position information; if the change amount between the sixth position information and the first position information meets the preset condition, the pose change amount of the first object relative to the second object is the transformation matrix used to transform the second position information.

[0035] In a possible implementation manner, the rendering module is configured to: perform pinhole imaging according to the fourth position information to obtain the image of the third object; acquire the occlusion relationship between the third object and the first object; fuse the image of the third object and the image of the first object according to the occlusion relationship to obtain a new target image.

[0036] In a possible implementation manner, the rendering module is configured to: calculate the first distance between the first object and the origin of the three-dimensional coordinate system according to the fifth position information; calculate the second distance between the third object and the origin of the three-dimensional coordinate system according to the fourth position information; compare the first distance and the second distance to obtain the occlusion relationship between the third object and the first object.

[0037] In a possible implementation, a rendering module is configured to: obtain the correspondence between multiple surface points of a first object and multiple surface points of a second object; obtain the distribution of the multiple surface points of the first object on the second object according to the correspondence; and obtain the occlusion relationship between a third object and the first object according to the distribution.

[0038] In a possible implementation, the pose change amount includes the orientation change of the first object relative to the second object. The rendering module is configured to: determine the front side of the first object according to the orientation change of the first object relative to the second object; and obtain the occlusion relationship between the third object and the first object according to the included angle between the orientation of the center point of the front side of the first object to the origin of the three-dimensional coordinate system and the orientation of the first object.

[0039] In a possible implementation, the rendering module is further configured to input a target image into a first neural network to obtain an image of the first object.

[0040] A sixth aspect of the embodiments of the present application provides an augmented reality device, including: a first acquisition module, configured to acquire a target image captured by a camera and third position information of a third object in a three-dimensional coordinate system corresponding to the camera, where the target image includes an image of a first object; a second acquisition module, configured to input the target image into a second neural network to obtain a pose change amount of the first object relative to a second object, where the second neural network is trained according to second position information of the second object in the three-dimensional coordinate system, the second object is a reference object of the first object, and the second position information and the third position information are preset information; a transformation module, configured to transform the third position information according to the pose change amount to obtain fourth position information of the third object in the three-dimensional coordinate system; and a rendering module, configured to render the third object in the target image according to the fourth position information to obtain a new target image.

[0041] As can be seen from the above device: when it is necessary to add a third object to the target image, the target image and the third position information of the third object in the camera coordinate system can be obtained first, and the target image includes the image of the first object. Then, the target image is input into the second neural network to obtain the pose change amount of the first object relative to the second object. The second neural network is trained according to the second position information of the second object in the camera coordinate system. The second object is the reference object of the first object, and the second position information and the third position information are preset information. Specifically, the pose change amount of the first object relative to the second object is used to indicate the position change and orientation change from the second object to the first object in the camera coordinate system. Therefore, the third object can be made to undergo the same position change and orientation change, that is, the fourth position information of the third object in the camera coordinate system is determined according to the pose change amount and the third position information of the third object. In this way, the final orientation of the third object can be associated with the orientation of the first object. Finally, the third object is rendered in the target image according to the fourth position information. In the obtained new target image, the orientation of the third object can be adapted to the orientation of the first object, and realistic content can be presented, thereby improving the user experience.

[0042] In a possible implementation manner, the rendering module is configured to: perform pinhole imaging according to the fourth position information to obtain the image of the third object; obtain the occlusion relationship between the third object and the first object; and fuse the image of the third object and the image of the first object according to the occlusion relationship to obtain a new target image.

[0043] In a possible implementation manner, the transformation module is further configured to: transform the second position information according to the pose change amount to obtain the fifth position information of the first object in the three-dimensional coordinate system corresponding to the camera.

[0044] In a possible implementation manner, the rendering module is configured to: calculate the first distance between the first object and the origin of the three-dimensional coordinate system according to the fifth position information; calculate the second distance between the third object and the origin of the three-dimensional coordinate system according to the fourth position information; and compare the first distance and the second distance to obtain the occlusion relationship between the third object and the first object.

[0045] In a possible implementation manner, the pose change amount includes the orientation change of the first object relative to the second object. The rendering module is configured to: determine the front of the first object according to the orientation change of the first object relative to the second object; and obtain the occlusion relationship between the third object and the first object according to the angle between the orientation of the center point of the front of the first object to the origin of the three-dimensional coordinate system and the orientation of the first object.

[0046] In a possible implementation manner, the rendering module is further configured to input the target image into the first neural network to obtain the image of the first object.

[0047] The seventh aspect of the embodiments of the present application provides a model training device, which includes: an acquisition module, configured to acquire an image to be trained, where the image to be trained includes an image of a first object; a determination module, configured to input the image to be trained into a model to be trained, and obtain a pose change amount of the first object relative to a second object; a calculation module, configured to calculate, through a preset target loss function, a deviation between the pose change amount of the first object relative to the second object and the true pose change amount of the first object, where the true pose change amount of the first object is determined according to second position information of the second object in a camera coordinate system, the second object is a reference object for the first object, and the second position information is preset information; an update module, configured to update parameters of the model to be trained according to the deviation until a model training condition is satisfied, so as to obtain a second neural network.

[0048] It can be seen from the above device that: the second neural network trained by this device can accurately obtain the pose change amount of the object in the target image.

[0049] The eighth aspect of the embodiments of the present application provides a model training device, which includes: a first acquisition module, configured to acquire an image to be trained; a second acquisition module, configured to acquire an image of a first object through a model to be trained; a calculation module, configured to calculate, through a preset target loss function, a deviation between the image of the first object and the true image of the first object; an update module, configured to update parameters of the model to be trained according to the deviation until a model training condition is satisfied, so as to obtain a first neural network.

[0050] It can be seen from the above device that: the first neural network trained by this device can accurately obtain the image of the first object in the target image.

[0051] The ninth aspect of the embodiments of the present application provides an image processing method, which includes:

[0052] The terminal device responds to a first operation of the user and displays a target image, where a first object is presented in the target image, and the first object is an object in the real environment. For example, after the user operates the terminal device, the terminal device can be made to display a target image, and the image presents a person who is dancing.

[0053] The terminal device responds to a second operation of the user and presents a virtual object in the target image, where the virtual object is superimposed on the first object. For example, the user operates the terminal device again, so that the terminal device adds a virtual wing near the human body presented in the target image.

[0054] In response to the movement of the first object, the terminal device updates the pose of the first object in the target image and the pose of the virtual object in the target image to obtain a new target image; wherein, the pose includes position and orientation, and the orientation of the virtual object is associated with the orientation of the first object. For example, after the terminal device determines that a human body has moved (translated and / or rotated), it updates the pose of the human body and the pose of the virtual wings in the target image, so that the virtual wings remain near the human body and the orientation of the virtual wings is associated with the orientation of the human body.

[0055] It can be seen from the above method that when the terminal device responds to the user's operation, the terminal device can display the target image. Since the target image presents the first object that is moving, when the terminal device adds a virtual object indicated by the user to the target image, it will update the pose of the first object and the pose of the virtual object in the target image as the first object moves, so that the orientation of the virtual object is associated with the orientation of the first object, thereby enabling the new target image to present realistic content and improving the user experience.

[0056] In a possible implementation manner, in the new target image, the orientation of the virtual object is the same as the orientation of the first object, and the relative position between the virtual object and the first object remains unchanged after updating the poses of the first object and the virtual object.

[0057] In a possible implementation manner, the first object is a human body and the virtual object is virtual wings.

[0058] In a possible implementation manner, the first operation is an operation to start an application, and the second operation is an operation to add special effects.

[0059] The tenth aspect of the embodiments of the present application provides an image processing device, which includes: a display module, configured to display a target image in response to a first operation of the user, where a first object is presented in the target image, and the first object is an object in the real environment; a presentation module, configured to present a virtual object in the target image in response to a second operation of the user, where the virtual object is superimposed on the first object; an update module, configured to update the pose of the first object in the target image and the pose of the virtual object in the target image in response to the movement of the first object to obtain a new target image; wherein, the pose includes position and orientation, and the orientation of the virtual object is associated with the orientation of the first object.

[0060] As can be seen from the above device: when the image processing device responds to a user operation, the image processing device can display a target image. Since the target image presents a first object that is moving, when the image processing device adds a virtual object indicated by the user to the target image, the pose of the first object and the pose of the virtual object in the target image will be updated as the first object moves, so that the orientation of the virtual object is associated with the orientation of the first object, thereby enabling the new target image to present realistic content and improving the user experience.

[0061] In a possible implementation, in the new target image, the orientation of the virtual object is the same as the orientation of the first object, and the relative position between the virtual object and the first object remains unchanged after updating the poses of the first object and the virtual object.

[0062] In a possible implementation, the first object is a human body and the virtual object is a virtual wing.

[0063] In a possible implementation, the first operation is an operation to open an application, and the second operation is an operation to add special effects.

[0064] The eleventh aspect of the embodiments of the present application provides an augmented reality device, including a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the augmented reality device executes the method described in the first aspect, any possible implementation manner in the first aspect, the second aspect, or any possible implementation manner in the second aspect.

[0065] The twelfth aspect of the embodiments of the present application provides a model training device, including a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the model training device executes the method described in the third aspect or the fourth aspect.

[0066] The thirteenth aspect of the embodiments of the present application provides a circuit system, which includes a processing circuit configured to execute the method described in the first aspect, any possible implementation manner in the first aspect, the second aspect, any possible implementation manner in the second aspect, the third aspect, or the fourth aspect.

[0067] The fourteenth aspect of the embodiments of the present application provides a chip system, which includes a processor for calling a computer program or computer instruction stored in a memory, so that the processor executes the method described in the first aspect, any possible implementation manner in the first aspect, the second aspect, any possible implementation manner in the second aspect, the third aspect, or the fourth aspect.

[0068] In a possible implementation, the processor is coupled to the memory through an interface.

[0069] In a possible implementation, the chip system further includes a memory, and a computer program or computer instructions are stored in the memory.

[0070] The fifteenth aspect of the embodiments of the present application provides a computer storage medium, which stores a computer program. When the program is executed by a computer, the computer implements the methods described in the first aspect, any possible implementation of the first aspect, the second aspect, any possible implementation of the second aspect, the third aspect, or the fourth aspect.

[0071] The sixteenth aspect of the embodiments of the present application provides a computer program product, which stores instructions. When the instructions are executed by a computer, the computer implements the methods described in the first aspect, any possible implementation of the first aspect, the second aspect, any possible implementation of the second aspect, the third aspect, or the fourth aspect.

[0072] In the embodiments of the present application, when it is necessary to add a third object to the target image, the second position information of the second object in the camera coordinate system, the third position information of the third object in the camera coordinate system, and the first position information of the first object in the target image in the image coordinate system can be obtained first. The second position information and the third position information are preset information. Since the second object is a reference object for the first object, the pose change amount of the first object relative to the second object can be determined according to the second position information and the first position information. The pose change amount is used to indicate the position change and orientation change from the second object to the first object in the camera coordinate system. Then, the third object can be made to undergo the same position change and orientation change, that is, the fourth position information of the third object in the camera coordinate system is determined according to the pose change amount and the third position information of the third object. In this way, the final orientation of the third object can be associated with the orientation of the first object (for example, their orientations are the same or their orientations are similar). Finally, the third object is rendered in the target image according to the fourth position information. In the obtained new target image, the orientation of the third object can be adapted to the orientation of the first object, and realistic content can be presented, thereby improving the user experience. Description of the Drawings

[0073] Figure 1 It is a schematic structural diagram of an artificial intelligence main framework;

[0074] Figure 2a It is a schematic structural diagram of an image processing system provided by the embodiments of the present application;

[0075] Figure 2b It is another schematic structural diagram of an image processing system provided by the embodiments of the present application;

[0076] Figure 2c A schematic diagram of related devices for image processing provided by an embodiment of the present application;

[0077] Figure 3a A schematic diagram of the architecture of system 100 provided by an embodiment of the present application;

[0078] Figure 3b A schematic diagram of dense human pose estimation;

[0079] Figure 4 A flowchart of an augmented reality method provided by an embodiment of the present application;

[0080] Figure 5 A schematic diagram of a target image provided by an embodiment of the present application;

[0081] Figure 6 A schematic diagram of a standard human model and wings provided by an embodiment of the present application;

[0082] Figure 7 A schematic diagram of the amount of pose change provided by an embodiment of the present application;

[0083] Figure 8 A schematic diagram of the process of determining the occlusion relationship provided by an embodiment of the present application;

[0084] Figure 9 Another flowchart of an augmented reality method provided by an embodiment of the present application;

[0085] Figure 10 A schematic diagram of controlling a standard human model provided by an embodiment of the present application;

[0086] Figure 11 A schematic diagram of an application example of an augmented reality method provided by an embodiment of the present application;

[0087] Figure 12 A flowchart of a model training method provided by an embodiment of the present application;

[0088] Figure 13 A schematic diagram of the structure of an augmented reality device provided by an embodiment of the present application;

[0089] Figure 14 Another schematic diagram of the structure of an augmented reality device provided by an embodiment of the present application;

[0090] Figure 15 A schematic diagram of the structure of a model training device provided by an embodiment of the present application;

[0091] Figure 16 A schematic diagram of the structure of an execution device provided by an embodiment of the present application;

[0092] Figure 17 A structural schematic diagram of a training device provided by an embodiment of the present application;

[0093] Figure 18 A structural schematic diagram of a chip provided by an embodiment of the present application. Detailed implementation manners

[0094] An embodiment of the present application provides an augmented reality method and related devices, which can dynamically associate the orientation of a virtual object with the orientation of a real object, so that the target image presents realistic content and improves the user experience.

[0095] Terms such as "first" and "second" in the description, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0096] Artificial intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, and is a theory, method, technology and application system that perceives the environment, acquires knowledge and uses knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making. Therefore, embodiments of the present application can implement AR technology through AI technology, so as to provide users with more vivid and interesting video content or image content.

[0097] First, the overall working process of the artificial intelligence system will be described. Please refer to Figure 1 , Figure 1It is a schematic diagram of the structure of an artificial intelligence subject framework. The above artificial intelligence subject framework will be elaborated from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general processes of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (provision and processing technology implementation) to the industrial ecological process of the system.

[0098] (1) Infrastructure

[0099] The infrastructure provides computing power support for the artificial intelligence system, enables communication with the external world, and is supported through the basic platform. It communicates with the external through sensors; the computing power is provided by intelligent chips (hardware acceleration chips such as CPU, NPU, GPU, ASIC, FPGA, etc.); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the external to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for calculation.

[0100] (2) Data

[0101] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves the Internet of Things data of traditional devices, including the business data of existing systems and the sensed data such as force, displacement, liquid level, temperature, humidity, etc.

[0102] (3) Data Processing

[0103] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.

[0104] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on data.

[0105] Reasoning refers to the process of simulating the intelligent reasoning method of humans in a computer or intelligent system, and using formal information for machine thinking and problem-solving according to the reasoning control strategy. The typical function is search and matching.

[0106] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, prediction, etc.

[0107] (4) General Capabilities

[0108] After the data is processed by the above-mentioned data processing, some general capabilities can be further formed based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0109] (5) Intelligent products and industrial applications

[0110] Intelligent products and industrial applications refer to the products and applications of artificial intelligence systems in various fields. It is the encapsulation of the overall artificial intelligence solution, productizing intelligent information decision-making and realizing practical applications. Its application fields mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, smart cities, etc.

[0111] Next, several application scenarios of this application will be introduced.

[0112] Figure 2a FIG. is a schematic structural diagram of an image processing system provided for an embodiment of this application. The image processing system includes a user device and a data processing device. Among them, the user device includes intelligent terminals such as mobile phones, personal computers, or information processing centers. The user device is the initiating end of image processing and is the initiator of the image processing request. Usually, the user initiates a request through the user device.

[0113] The above-mentioned data processing device can be a device or server with data processing functions such as a cloud server, a network server, an application server, and a management server. The data processing device receives an image enhancement request from the intelligent terminal through an interaction interface, and then performs image processing in ways such as machine learning, deep learning, searching, reasoning, and decision-making through a memory for storing data and a processor for data processing. The memory in the data processing device can be a general term, including local storage and a database for storing historical data. The database can be on the data processing device or on other network servers.

[0114] In Figure 2a In the image processing system shown, the user device can receive the user's instruction. For example, the user device can obtain an image input / selected by the user, and then initiate a request to the data processing device, so that the data processing device executes an image semantic segmentation application for the image obtained by the user device, thereby obtaining a corresponding processing result for the image. Exemplarily, the user device can obtain an image to be processed input by the user, and then initiate an image processing request to the data processing device, so that the data processing device performs an image processing application on the image (for example, image target detection, obtaining the change amount of the pose of an object in the image, etc.), thereby obtaining a processed image.

[0115] In Figure 2aAmong them, the data processing device can execute the augmented reality method of the embodiments of the present application.

[0116] Figure 2b Another structural schematic diagram of the image processing system provided by the embodiments of the present application. In Figure 2b Among them, the user device directly serves as the data processing device. This user device can directly obtain the input from the user and directly process it by the hardware of the user device itself. The specific process is similar to Figure 2a and can refer to the above description, which will not be elaborated here.

[0117] In Figure 2b In the shown image processing system, the user device can receive the user's instructions. For example, the user device can obtain a to-be-processed image selected by the user in the user device, and then the user device itself executes an image processing application for this image (such as image target detection, obtaining the pose change amount of an object in the image, etc.), so as to obtain the corresponding processing result for this image.

[0118] In Figure 2b Among them, the user device itself can execute the augmented reality method of the embodiments of the present application.

[0119] Figure 2c A schematic diagram of the related device for image processing provided by the embodiments of the present application.

[0120] The above Figure 2a and Figure 2b The user device in can specifically be Figure 2c the local device 301 or the local device 302 in, Figure 2a The data processing device in can specifically be Figure 2c the execution device 210 in. Among them, the data storage system 250 can store the to-be-processed data of the execution device 210. The data storage system 250 can be integrated on the execution device 210 or can be set on the cloud or other network servers.

[0121] Figure 2a and Figure 2b The processor in can perform data training / machine learning / deep learning through a neural network model or other models (such as a model based on a support vector machine), and use the model finally trained or learned from the data to execute an image processing application for the image, so as to obtain the corresponding processing result.

[0122] Figure 3a A schematic diagram of the architecture of the system 100 provided by the embodiments of the present application. In Figure 3aAmong them, the execution device 110 configures an input / output (I / O) interface 112 for data interaction with external devices. Users can input data to the I / O interface 112 through the client device 140. The input data in the embodiments of this application may include: each task to be scheduled, callable resources, and other parameters.

[0123] During the preprocessing of the input data by the execution device 110, or during the related processing such as calculation by the calculation module 111 of the execution device 110 (such as implementing the functions of the neural network in this application), the execution device 110 can call data, code, etc. in the data storage system 150 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing into the data storage system 150.

[0124] Finally, the I / O interface 112 returns the processing result to the client device 140 and provides it to the user.

[0125] It should be noted that the training device 120 can generate corresponding target models / rules based on different training data for different targets or tasks. The corresponding target models / rules can be used to achieve the above-mentioned targets or complete the above-mentioned tasks, so as to provide the required results for users. Among them, the training data can be stored in the database 130 and comes from the training samples collected by the data collection device 160.

[0126] In Figure 3a the shown situation, the user can manually give the input data, and this manual giving can be operated through the interface provided by the I / O interface 112. In another situation, the client device 140 can automatically send the input data to the I / O interface 112. If the client device 140 is required to automatically send the input data and user authorization is required, the user can set the corresponding permissions in the client device 140. The user can view the results output by the execution device 110 in the client device 140, and the specific presentation forms can be display, sound, action and other specific ways. The client device 140 can also be used as a data collection end to collect the input data input to the I / O interface 112 and the output result of the output I / O interface 112 as new sample data and store them in the database 130. Of course, it is also possible not to collect through the client device 140, but to directly store the input data input to the I / O interface 112 and the output result of the output I / O interface 112 as new sample data into the database 130.

[0127] It should be noted that Figure 3a is only a schematic diagram of a system architecture provided by the embodiments of this application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, inFigure 3a In this case, the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 can also be placed in the execution device 110. As Figure 3a shown, a neural network can be trained according to the training device 120.

[0128] An embodiment of the present application also provides a chip, which includes a neural network processor NPU. The chip can be set in the execution device 110 as Figure 3a shown to complete the computing work of the computing module 111. The chip can also be set in the training device 120 as Figure 3a shown to complete the training work of the training device 120 and output the target model / rule.

[0129] The neural network processor NPU is mounted on the main central processing unit (CPU) (host CPU) as a coprocessor, and tasks are allocated by the main CPU. The core part of the NPU is the arithmetic circuit, and the controller controls the arithmetic circuit to extract data from the memory (weight memory or input memory) and perform operations.

[0130] In some implementations, the arithmetic circuit includes multiple processing units (PEs) inside. In some implementations, the arithmetic circuit is a two-dimensional systolic array. The arithmetic circuit can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit is a general matrix processor.

[0131] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are stored in the accumulator.

[0132] The vector calculation unit can further process the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. For example, the vector calculation unit can be used for network calculations in non-convolutional / non-FC layers of a neural network, such as pooling, batch normalization, local response normalization, etc.

[0133] In some implementations, the vector computing unit can store the processed output vectors in the unified cache. For example, the vector computing unit can apply a non-linear function to the output of the arithmetic circuit, such as a vector of accumulated values, to generate activation values. In some implementations, the vector computing unit generates normalized values, combined values, or both. In some implementations, the processed output vectors can be used as activation inputs to the arithmetic circuit, such as for use in subsequent layers of a neural network.

[0134] The unified memory is used to store input data and output data.

[0135] The weight data directly transfers the input data in the external memory to the input memory and / or the unified memory, stores the weight data in the external memory into the weight memory, and stores the data in the unified memory into the external memory through the direct memory access controller (DMAC).

[0136] The bus interface unit (BIU) is used to interact between the main CPU, DMAC, and the instruction fetch memory through the bus.

[0137] The instruction fetch buffer connected to the controller is used to store the instructions used by the controller;

[0138] The controller is used to call the instructions cached in the instruction fetch memory to control the working process of the arithmetic accelerator.

[0139] Generally, the unified memory, input memory, weight memory, and instruction fetch memory are all on-chip memories, and the external memory is the memory outside the NPU. The external memory can be a double data rate synchronous dynamic random access memory (DDRSDRAM), a high bandwidth memory (HBM), or other readable and writable memories.

[0140] Since the embodiments of the present application involve a large number of neural network applications, for the sake of easy understanding, the following first introduces the relevant terms and related concepts such as neural networks involved in the embodiments of the present application.

[0141] (1) Neural network

[0142] A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes xs and an intercept of 1 as inputs. The output of this operation unit can be:

[0143]

[0144] where s = 1, 2, …, n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many such single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neural units.

[0145] The operation of each layer in the neural network can be described by the mathematical expression y = a(Wx + b): From a physical perspective, the operation of each layer in the neural network can be understood as completing the transformation from the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of the matrix) through five operations on the input space. These five operations include: 1. Dimensionality increase / dimensionality reduction; 2. Magnification / minification; 3. Rotation; 4. Translation; 5. "Bending". Among them, the operations of 1, 2, and 3 are completed by Wx, the operation of 4 is completed by +b, and the operation of 5 is implemented by a(). The reason for using the word "space" here is that the object to be classified is not a single thing, but a class of things, and the space refers to the set of all individuals of this class of things. Among them, W is the weight vector, and each value in this vector represents the weight value of a neuron in this layer of the neural network. This vector W determines the space transformation from the input space to the output space described above, that is, the weight W of each layer controls how to transform the space. The purpose of training the neural network, that is, ultimately obtaining the weight matrices of all layers of the trained neural network (the weight matrix formed by vectors W of many layers). Therefore, the training process of the neural network is essentially to learn the way to control the space transformation, and more specifically, to learn the weight matrix.

[0146] Since we hope that the output of the neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the real target value, and then update the weight vector of each layer of the neural network according to the difference between the two. (Of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the neural network). For example, if the predicted value of the network is too high, we adjust the weight vector to make it predict lower, and keep adjusting until the neural network can predict the real target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function. They are important equations for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the neural network becomes a process of minimizing this loss as much as possible.

[0147] (2) Backpropagation algorithm

[0148] The neural network can use the backpropagation (BP) algorithm to correct the size of the parameters in the initial neural network model during the training process, so that the reconstruction error loss of the neural network model becomes smaller and smaller. Specifically, forward propagating the input signal until the output will generate an error loss, and updating the parameters in the initial neural network model by backpropagating the error loss information, so as to make the error loss converge. The backpropagation algorithm is a backpropagation movement dominated by the error loss, aiming to obtain the optimal parameters of the neural network model, such as the weight matrix.

[0149] (3) Dense human pose estimation

[0150] Through the dense human pose estimation algorithm, a dense correspondence relationship can be established between the pixel points of the image and the surface points of the three-dimensional standard human model. As Figure 3b shown ( Figure 3b is a schematic diagram of dense human pose estimation), this algorithm can establish a correspondence relationship between each pixel point of the real human body in the two-dimensional image and the surface points in the three-dimensional standard human model.

[0151] The method provided in this application will be described below from the training side and the application side of the neural network.

[0152] The model training method provided by the embodiments of the present application relates to image processing, and can be specifically applied to data processing methods such as data training, machine learning, and deep learning. It performs symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on training data (such as the to-be-trained images in the present application), and finally obtains trained neural networks (such as the first neural network and the second neural network in the present application). Moreover, the augmented reality method provided by the embodiments of the present application can use the above-mentioned trained neural networks, input input data (such as the target image in the present application) into the trained neural networks, and obtain output data (such as the image of the first object, the pose change amount of the first object relative to the second object, etc. in the present application). It should be noted that the model training method and the augmented reality method provided by the embodiments of the present application are inventions based on the same concept, and can also be understood as two parts of a system, or two stages of an overall process: such as the model training stage and the model application stage.

[0153] Figure 4 It is a schematic flowchart of an augmented reality method provided by an embodiment of the present application. As Figure 4 shown, the method includes:

[0154] 401. Obtain the target image captured by the camera, and the first position information of the first object in the target image.

[0155] 402. Obtain the second position information of the second object in the three-dimensional coordinate system corresponding to the camera, and the third position information of the third object in the three-dimensional coordinate system. The second object is a reference object for the first object, and the second position information and the third position information are pre-set information.

[0156] When the user needs to add a virtual object (or add a real object additionally) to the target image presenting the real environment, the target image captured by the camera can be obtained first. Among them, the real environment presented by the target image includes multiple real objects (that is, the target image is composed of images of multiple real objects). For example, as Figure 5 shown ( Figure 5 is a schematic diagram of a target image provided by an embodiment of the present application) real human bodies, trees, houses and other objects. Among the multiple real objects, the real object that the user is concerned about can be regarded as the first object. Further, the real object of the virtual object to be added by the user can also be regarded as the third object, and the reference object of the real object that the user is concerned about can be regarded as the second object. For example, if the user wants to add a pair of wings to a real human body in the target image, a three-dimensional standard human model can be used as the reference object for the real human body.

[0157] Specifically, the second object is usually a standard object model corresponding to the first object. The second object can be obtained through the principal component analysis (PCA) algorithm or can be set manually, etc. As a reference object for the first object, the second position information of the second object in the camera coordinate system is preset, that is, the three-dimensional coordinates of each surface point of the second object (all surface points of the entire object) in the camera coordinate system have been preset. In this way, the second object is preset with a pose, that is, the second object is preset at a certain position in the camera coordinate system and is preset with an orientation (pose).

[0158] Similarly, the third position information of the third object in the camera coordinate system (i.e., the three-dimensional coordinate system corresponding to the camera that captures the target object) is also preset, that is, the three-dimensional coordinates of each surface point of the third object (all surface points of the entire object) in the camera coordinate system have been preset. It can be seen that the third object is also preset with a pose. It should be noted that the second position information of the second object and the third position information of the third object can be related, that is, in the camera coordinate system, the position set for the second object is related to the position set for the third object, and the orientation set for the second object is also related to the orientation set for the third object (for example, their orientations are the same or similar). Still as the above example, as Figure 6 shown( Figure 6 is a schematic diagram of the standard human model and the wings provided by the embodiment of the present application), the standard human model is set at the origin of the camera coordinate system, and the wings are set on the back of the standard human body. The orientation of the standard human model( Figure 6 in which the orientation of the standard human model points to the positive half-axis of the z-axis) and the orientation of the wings are the same. The orientation of the standard human model refers to the direction pointed by the front of the standard human model, and the orientation of the wings refers to the direction pointed by the connecting end of the wings.

[0159] Since the first object also occupies a certain position in the target image, the first position information of the first object in the image coordinate system (i.e., the two-dimensional coordinate system corresponding to the target image) can be obtained, that is, the two-dimensional coordinates of each surface point of the first object (which can also be called pixel points) in the image coordinate system. It should be noted that in this embodiment, each surface point of the first object refers to each surface point of the first object captured by the camera (i.e., each surface point of the first object presented in the target image). For example, if the camera captures the front of the first object, then here it refers to each surface point of the front of the first object. If the camera captures the back of the first object, then here it refers to each surface point of the back of the first object. If the camera captures the side of the first object, then here it refers to each surface point of the side of the first object, etc.

[0160] It can be seen that after obtaining the target image to be processed, the second position information of the second object and the third position information of the third object can be directly obtained, and the first position information of the first object can be obtained from the target image.

[0161] It should be understood that the aforementioned camera coordinate system is a three-dimensional coordinate system constructed with the camera that captures the target image as the origin. The aforementioned image coordinate system is a two-dimensional coordinate system constructed with the upper left corner of the target image as the origin.

[0162] It should also be understood that in this embodiment, the orientation of the first object is the direction pointed by the front of the first object. The orientation of the first object is usually different from the orientation preset for the second object, and the position of the first object is usually different from the position preset for the second object.

[0163] 403. Determine the pose change amount of the first object relative to the second object according to the first position information and the second position information.

[0164] After obtaining the second position information of the second object, the third position information of the third object, and the first position information of the first object, the pose change amount of the first object relative to the second object can be determined according to the second position information and the first position information. Among them, the pose change amount of the first object relative to the second object (which can also be called a transformation matrix) is used to indicate the position change from the second object to the first object (including the distance change on the x-axis, the distance change on the y-axis, and the distance change on the z-axis) and the orientation change from the second object to the first object (including the rotation angle change on the x-axis, the rotation angle change on the y-axis, and the rotation angle change on the z-axis). Specifically, the pose change amount of the first object relative to the second object can be determined in various ways, which will be introduced separately below:

[0165] In a possible implementation, determining the pose change amount of the first object relative to the second object according to the second position information and the first position information includes: (1) Obtaining the depth information of the first object. Specifically, the depth information of the first object includes the depth values of each surface point in the first object, and the depth value of each surface point is used to indicate the distance from the surface point to the camera. (2) Determining the fifth position information of the first object in the camera coordinate system according to the first position information of the first object in the image coordinate system and the depth information of the first object. Specifically, the two-dimensional coordinates of each surface point of the first object are combined with the depth value of each surface point to calculate the three-dimensional coordinates of each surface point of the first object in the camera coordinate system. In this way, the pose of the first object in the camera coordinate system is determined. (3) Calculating the change amount between the second position information of the second object and the fifth position information of the first object to obtain the pose change amount of the first object relative to the second object. Specifically, through the dense human pose estimation algorithm, the correspondence relationship between multiple surface points of the first object and multiple surface points of the second object (that is, the corresponding points of each surface point of the first object on the second object) can be determined. Therefore, based on this correspondence relationship, the distance between the three-dimensional coordinates of each surface point in the first object and the three-dimensional coordinates of the corresponding surface point of the second object can be calculated, so as to obtain the pose change amount of the first object relative to the second object.

[0166] In another possible implementation, determining the pose change amount of the first object relative to the second object based on the second position information and the first position information includes: (1) Transforming the second position information of the second object in the camera coordinate system to obtain the fifth position information of the first object in the camera coordinate system. Specifically, through a dense human pose estimation algorithm, the corresponding points of each surface point of the first object on the second object can be determined. Since the three-dimensional coordinates of this part of the points on the second object have been preset, the three-dimensional coordinates of this part of the points on the second object can be arbitrarily transformed (i.e., arbitrarily rotated and translated), and the three-dimensional coordinates of this part of the points after rotation and translation can be regarded as the three-dimensional coordinates of each surface point of the first object in the camera coordinate system. (2) Projecting the fifth position information of the first object in the camera coordinate system onto the image coordinate system to obtain the sixth position information of the first object in the image coordinate system. Specifically, the three-dimensional coordinates of this part of the points after rotation and translation can be projected onto the target image, so as to obtain the new two-dimensional coordinates of each surface point of the first object in the target image. (3) If the change amount between the sixth position information and the first position information meets the preset condition, the transformation matrix for performing the transformation operation on the second position information is determined as the pose change amount of the first object relative to the second object. Specifically, if the distance between the original two-dimensional coordinates and the new two-dimensional coordinates of each surface point of the first object in the target image meets the preset condition, the transformation operation in step (1) (that is, a transformation matrix) is determined as the pose change amount of the first object relative to the second object. If the preset condition is not met, steps (1) to (3) are re-executed until the preset condition is met.

[0167] Further, the aforementioned preset condition can be less than or equal to a preset distance threshold, or can be the minimum value in multiple rounds of calculations, etc. For example, steps (1) and (2) can be repeatedly executed to obtain ten distances to be judged, and the smallest distance value among them is selected as the distance that meets the preset condition, and the transformation operation corresponding to this distance is determined as the pose change amount of the first object relative to the second object.

[0168] To further understand the foregoing process of determining the pose change amount, the following is combined with Figure 7 for further introduction. Figure 7 FIG. is a schematic diagram of the pose change amount provided by the embodiment of the present application. It should be noted that, Figure 7 On the basis of Figure 5 constructed, that is, Figure 7 The real human body in Figure 5 is the real human body in Figure 7 As shown in Figure 5In the target image shown, after determining that wings need to be added behind the real human body, the position information of the real human body in the target image can be obtained, that is, the two-dimensional coordinates of each surface point of the real human body in the target image, and the position information of the standard human body model in the camera coordinate system can be obtained, that is, the three-dimensional coordinates set for each surface point of the standard human body model.

[0169] Then, it is necessary to determine the position information of the real human body in the target image in the camera coordinate system. For example, according to the two-dimensional coordinates of each surface point of the real human body in the target image and the depth value of each surface point of the real human body, the three-dimensional coordinates of each surface point of the real human body in the camera coordinate system can be determined. Another example is to determine the corresponding points of each surface point of the real human body on the standard human body model, and perform rotation and translation operations on this part of the points until the three-dimensional coordinates of this part of the points after the operation meet the requirements (for example, the two-dimensional coordinates obtained by projecting the three-dimensional coordinates of this part of the points onto the target image are the smallest distance from the two-dimensional coordinates of each surface point of the real human body in the target image, etc.), then the three-dimensional coordinates of this part of the points can be finally determined as the three-dimensional coordinates of each surface point of the real human body in the camera coordinate system. In this way, the position and orientation of the real human body in the camera coordinate system can be determined.

[0170] Finally, according to the position information of the standard human body model in the camera coordinate system and the position information of the real human body in the camera coordinate system, the pose change amount of the real human body relative to the standard human body model can be determined. Specifically, after obtaining the three-dimensional coordinates of each surface point of the real human body in the camera coordinate system, for any surface point of the real human body, the distance between the three-dimensional coordinates of this surface point and the three-dimensional coordinates of the corresponding point of this point on the standard human body model (that is, the three-dimensional coordinates set for this corresponding point) can be calculated. After calculating the above for each surface point of the real human body, the pose change amount of the real human body relative to the standard human body model can be obtained. As Figure 7 shown, the orientation of the real human body does not point to the positive half-axis of the z-axis (that is, the orientation of the real human body slightly deviates from the positive half-axis of the z-axis, rather than directly facing the positive half-axis of the z-axis), while the orientation of the standard human body model points to the positive half-axis of the z-axis. That is, there is a certain gap between the orientation of the real human body and the orientation of the standard human body model, and there is also a certain gap between the position of the real human body and the position of the standard human body model. The pose change amount of the real human body relative to the standard human body model can be used to represent the change between the orientation of the real human body and the orientation of the standard human body model, as well as the change between the position of the real human body and the position of the standard human body model.

[0171] 404. Transform the third position information according to the pose change amount to obtain the fourth position information of the third object in this three-dimensional coordinate system.

[0172] After obtaining the pose change amount of the first object relative to the second object, since this pose change amount can be used to represent the orientation change between the first object and the second object and the position change between the first object and the second object, the third object can be made to undergo the same orientation change and position change. That is, the third position information is transformed according to this pose change amount to obtain the fourth position information of the third object in the camera coordinate system, thereby determining the final position and final orientation of the third object. Specifically, for any surface point of the third object, the three-dimensional coordinates of this surface point can be transformed (for example, by performing matrix multiplication calculations) using the pose change amount of the first object relative to the second object to obtain the new three-dimensional coordinates of this surface point. After all the surface points of the third object have been subjected to the aforementioned transformation, the third object is translated from its original set position to the final position, and the third object is rotated from its original set orientation to the final orientation.

[0173] As Figure 7 shown, after obtaining the pose change amount of the real human body relative to the standard human body model, this pose change amount represents the orientation change from the standard human body model to the real human body and the position change from the standard human body model to the real human body. Therefore, the wings can be made to undergo the same orientation change and position change, that is, the wings are rotated and translated according to this pose change amount so that the rotated and translated wings are related to the real human body, that is, the wings are located on the back of the real human body and their orientations are the same.

[0174] 405. Render the third object in the target image according to the fourth position information to obtain a new target image.

[0175] After obtaining the fourth position information of the third object, small hole imaging can be first performed according to the fourth position information to obtain an image of the third object. Then, the occlusion relationship between the third object and the first object is obtained, and in the target image, the image of the third object and the image of the first object are fused according to the occlusion relationship to obtain a new target image. Specifically, the occlusion relationship between the third object and the first object can be obtained in various ways, which will be introduced separately below:

[0176] In a possible implementation, obtaining the occlusion relationship between the third object and the first object includes: (1) Calculating a first distance between the first object and the origin of the three-dimensional coordinate system according to the fifth position information. Specifically, according to the three-dimensional coordinates of each surface point of the first object in the camera coordinate system, the distance from each surface point of the first object to the origin can be calculated, and the average value of the distances from these points to the origin is used as the first distance from the first object to the origin. (2) Calculating a second distance between the third object and the origin of the three-dimensional coordinate system according to the fourth position information. Specifically, according to the new three-dimensional coordinates of each surface point of the third object in the camera coordinate system, the distance from each surface point of the third object to the origin can be calculated, and the average value of the distances from these points to the origin is used as the second distance from the third object to the origin. (3) Comparing the first distance and the second distance to obtain the occlusion relationship between the third object and the first object. Specifically, if the first distance is less than or equal to the second distance, the third object is occluded by the first object; if the first distance is greater than the second distance, the first object is occluded by the third object.

[0177] In another possible implementation, obtaining the occlusion relationship between the third object and the first object includes: (1) Obtaining the correspondence between multiple surface points of the first object and multiple surface points of the second object. (2) According to this correspondence, obtaining the distribution of multiple surface points of the first object on the second object. For example, how many surface points of the first object are located on the front of the second object, and how many surface points of the first object are located on the back of the second object. (3) According to this distribution, obtaining the occlusion relationship between the third object and the first object. Specifically, among the surface points of the first object, if the number of surface points located on the front of the second object is greater than or equal to the number of surface points located on the back of the second object, the third object is occluded by the first object; if the number of surface points located on the front of the second object is less than the number of surface points located on the back of the second object, the first object is occluded by the third object.

[0178] In another possible implementation, obtaining the occlusion relationship between the third object and the first object includes: (1) Determining the front of the first object according to the orientation change of the first object relative to the second object. Specifically, since the orientation set for the second object is known, the front of the second object is also known. Then, according to the orientation change of the first object relative to the second object and the orientation of the second object, the orientation of the first object can be determined, that is, the front of the first object can be determined. (2) As Figure 8 shown ( Figure 8A schematic diagram of the process for determining the occlusion relationship provided by an embodiment of this application. According to the angle between the orientation of the center point of the front surface of the first object to the origin (camera) of the camera coordinate system and the orientation of the first object (i.e., the direction indicated by the front surface of the first object), the occlusion relationship between the third object and the first object is obtained. For example, if the angle is less than or equal to 90°, the third object is occluded by the first object; if the angle is greater than 90°, the first object is occluded by the third object.

[0179] After obtaining the occlusion relationship between the third object and the first object, the first neural network can be used to perform salient object detection on the target image to obtain the image of the first object and the images of the other objects except the first object. If the third object is occluded by the first object, when fusing the image of the first object, the image of the third object, and the images of the other objects, the image of the first object can be made to cover the image of the third object to obtain a new target image. If the first object is occluded by the third object, when fusing the image of the first object, the image of the third object, and the images of the other objects, the image of the third object can be made to cover the image of the first object to obtain a new target image. For example, in the new target image, since the real human body occludes the wings (i.e., what the camera captures is mostly the front of the real human body), at the junction of the wings and the real human body, the image of the wings will be covered by the image of the real human body, so that the new target image shows more realistic content and improves the user experience.

[0180] In an embodiment of this application, when it is necessary to add a third object to the target image, the second position information of the second object in the camera coordinate system, the third position information of the third object in the camera coordinate system, and the first position information of the first object in the image coordinate system in the target image can be obtained first. The second position information and the third position information are preset information. Since the second object is a reference object for the first object, the pose change amount of the first object relative to the second object can be determined according to the second position information and the first position information. This pose change amount is used to indicate the position change and orientation change from the second object to the first object in the camera coordinate system. Then, the third object can be made to undergo the same position change and orientation change, that is, the fourth position information of the third object in the camera coordinate system is determined according to this pose change amount and the third position information of the third object. In this way, the final orientation of the third object can be associated with the orientation of the first object (for example, their orientations are the same or their orientations are similar). Finally, the third object is rendered in the target image according to the fourth position information. In the obtained new target image, the orientation of the third object can be adapted to the orientation of the first object, presenting realistic content, thereby improving the user experience.

[0181] Figure 9 Another process schematic diagram of the augmented reality method provided by an embodiment of this application. AsFigure 9 As shown, the method includes:

[0182] 901. Obtain a target image captured by a camera and third position information of a third object in a three-dimensional coordinate system corresponding to the camera, where the target image includes an image of a first object.

[0183] 902. Input the target image into a second neural network to obtain a pose change amount of the first object relative to a second object. The second neural network is trained according to second position information of the second object in the three-dimensional coordinate system. The second object is a reference object for the first object, and the second position information and the third position information are preset information.

[0184] For the introduction of the target image, the first object, the second object, the third object, the second position information of the second object, and the third position information of the third object, reference can be made to Figure 4 the relevant description parts in steps 401 and 402 in the embodiments shown, which will not be elaborated here.

[0185] In the obtained target image, the target image can be input into the second neural network to obtain control parameters of the second object, including shape parameters and pose parameters. Among them, the shape parameters are used to control the shape (such as tall, short, fat, thin, etc.) of the second object, and the pose parameters are used to control the pose (such as orientation, action, etc.) of the second object.

[0186] By adjusting the shape parameters and pose parameters in the above formula, the shape of the second object can be adjusted to be the same as or similar to the shape of the first object in the target image. For the sake of easy understanding, the following combines Figure 10 to further introduce the foregoing control parameters. Figure 10 This is a schematic diagram of a control standard human model provided by an embodiment of the present application. By changing the original shape parameter β and the original pose parameter θ of the standard human model, the size and posture of the standard human model can be adjusted, so that the shape of the adjusted model (the shape parameter of the adjusted model is β1, and the pose parameter of the adjusted model is θ1) is the same as or similar to the shape of the real human body in the target image.

[0187] After obtaining the control parameters of the second object, these control parameters can be calculated to obtain a pose change amount of the first object relative to the second object.

[0188] 903. Transform the third position information according to the pose change amount to obtain fourth position information of the third object in the three-dimensional coordinate system.

[0189] 904. Render the third object in the target image according to the fourth position information to obtain a new target image.

[0190] After obtaining the fourth position information of the third object, the pinhole imaging can be performed according to the fourth position information first to obtain the image of the third object. Then, the second position information of the second object can be obtained, and the second position information can be transformed according to the pose change amount of the first object relative to the second object to obtain the fifth position information of the first object in the camera coordinate system. Then, the occlusion relationship between the third object and the first object is obtained. Finally, in the target image, the image of the third object and the image of the first object are fused according to the occlusion relationship to obtain a new target image.

[0191] Among them, the occlusion relationship between the third object and the first object can be obtained in various ways:

[0192] In a possible implementation, obtaining the occlusion relationship between the third object and the first object includes: calculating the first distance between the first object and the origin of the three-dimensional coordinate system according to the fifth position information; calculating the second distance between the third object and the origin of the three-dimensional coordinate system according to the fourth position information; comparing the first distance and the second distance to obtain the occlusion relationship between the third object and the first object.

[0193] In another possible implementation, obtaining the occlusion relationship between the third object and the first object includes: determining the front of the first object according to the orientation change of the first object relative to the second object; obtaining the occlusion relationship between the third object and the first object according to the included angle between the orientation from the center point of the front of the first object to the origin of the three-dimensional coordinate system and the orientation of the first object.

[0194] In addition, before performing image fusion, the target image can be subjected to salient object detection through a first neural network to obtain the image of the first object and the images of the other objects except the first object. Then, based on the occlusion relationship between the third object and the first object, the image of the first object, the images of the other objects, and the image of the third object are fused to obtain a new target image.

[0195] Regarding the introduction of steps 903 and 904, reference can be made to Figure 4 the relevant description parts in steps 404 and 405 in the illustrated embodiment, which will not be elaborated here.

[0196] In an embodiment of the present application, when it is necessary to add a third object to a target image, the target image and the third position information of the third object in the camera coordinate system may be obtained first. The target image includes an image of a first object. Then, the target image is input into a second neural network to obtain the pose change amount of the first object relative to a second object. The second neural network is trained according to the second position information of the second object in the camera coordinate system. The second object is a reference object for the first object, and the second position information and the third position information are preset information. Specifically, the pose change amount of the first object relative to the second object is used to indicate the position change and orientation change from the second object to the first object in the camera coordinate system. Therefore, the third object can be made to undergo the same position change and orientation change, that is, the fourth position information of the third object in the camera coordinate system is determined according to the pose change amount and the third position information of the third object. In this way, the final orientation of the third object can be associated with the orientation of the first object. Finally, the third object is rendered in the target image according to the fourth position information. In the obtained new target image, the orientation of the third object can be adapted to the orientation of the first object, and realistic content can be presented, thereby improving the user experience.

[0197] To further understand the augmented reality method provided in the embodiments of the present application, the following is Figure 11 further introduced in conjunction with Figure 11 FIG. Figure 11 is a schematic diagram of an application example of the augmented reality method provided in the embodiments of the present application. As

[0198] shown, this application example includes: Figure 11 S1. The terminal device responds to the user's first operation and displays a target image, and a first object is presented in the target image. The first object is an object in the real environment. For example, when the user opens an application (such as a certain live broadcast software, a certain shooting software, etc.) on the terminal device, the terminal device may display a

[0199] target image as shown. In the content presented in the target image, there is a real human body walking leftward.

[0200]

[0200] S3. In response to the movement of the first object, update the pose of the first object in the target image and update the pose of the virtual object in the target image to obtain a new target image; wherein, the pose includes position and orientation, and the orientation of the virtual object is associated with the orientation of the first object. Still as in the above example, when the terminal device determines that the real human body starts to walk to the right, that is, the orientation of the real human body changes, the terminal device can update the pose (including position and orientation) of the real human body in the target image, and synchronously update the pose of the virtual wings in the target image, so that the virtual wings are still displayed on the back of the real human body, and the orientation of the virtual wings and the orientation of the real human body both face right, thereby obtaining a new target image. In this way, the orientation of the virtual wings can change with the change of the orientation of the real human body, the orientations of the two can be dynamically associated, and their relative positions remain unchanged.

[0201] It should be noted that the virtual object in this application example can be the third object in the foregoing embodiment, and the calculation method of the pose of the virtual object can be obtained according to Figure 4 Steps 401 to 404 in, or obtained according to Figure 9 Steps 901 to 903 in the shown embodiment, and there is no limitation here.

[0202] The above is a detailed description of the augmented reality method provided by the embodiments of the present application. Next, the model training method provided by the embodiments of the present application will be introduced. Figure 12 is a schematic flowchart of a model training method provided by an embodiment of the present application. As shown in Figure 12 The method includes:

[0203] 1201. Obtain an image to be trained, where the image to be trained includes an image of a first object.

[0204] Before model training, a batch of images to be trained can be obtained. Among them, each image to be trained includes an image of a first object. The forms of the first object in different images to be trained can be different or the same. For example, in the image to be trained A, the front of the real human body faces the camera, and in the image to be trained B, the back of the real human body faces the camera, etc.

[0205] For any image to be trained, the true control parameters for adjusting the second object corresponding to the image to be trained are known. The second object is a reference object for the first object. Using this part of the true control parameters, the form of the second object can be adjusted so that the form of the adjusted second object is the same as or similar to the form of the first object in the image to be trained. In this way, based on this part of the true control parameters, the true pose change amount corresponding to the image to be trained can also be determined (that is, for the image to be trained, the true pose change amount of the first object relative to the second object) is known.

[0206] It should be noted that the second position information of the second object in the camera coordinate system is preset, which is equivalent to setting an original form for the second object, that is, setting the original control parameters for the second object. And the true control parameters corresponding to each training image are determined based on the original control parameters of the second object. Therefore, for any training image, the true control parameters corresponding to the training image are determined based on the second position information of the second object, that is, the true pose change amount corresponding to the training image is determined according to the second position information of the second object.

[0207] It should be understood that for the introduction of the first object, the second object, and the second position information of the second object, reference can be made to Figure 4 the relevant description parts of steps 401 and 402 in the embodiments shown, which will not be elaborated here.

[0208] 1202. Input the training image into the training model to obtain the pose change amount of the first object relative to the second object.

[0209] For any training image, inputting the training image into the training model can enable the training model to output the pose change amount corresponding to the training image, which can also be understood as the predicted pose change amount of the first object relative to the second object in the training image.

[0210] 1203. Calculate the deviation between the pose change amount of the first object relative to the second object and the true pose change amount of the first object through a preset target loss function.

[0211] For any training image, after obtaining the predicted pose change amount corresponding to the training image, the deviation between the pose change amount corresponding to the training image and the true pose change amount corresponding to the training image can be calculated through a preset target loss function. In this way, the deviation corresponding to each training image in this batch of training images can be obtained.

[0212] 1204. Update the parameters of the training model according to the deviation until the model training conditions are met to obtain the second neural network.

[0213] For any training image, if the deviation corresponding to the training image is within the qualified range, the training image is regarded as a qualified training image; if it is outside the qualified range, it is regarded as an unqualified training image. If there are only a small number of qualified training images in this batch of training images, adjust the parameters of the training model and retrain with another batch of training images (that is, re - execute steps 1201 to 1204) until there are a large number of qualified training image frames to obtain Figure 9 the second neural network in the embodiments shown.

[0214] In the embodiments of the present application, the second neural network trained by this training method can accurately obtain the pose change amount of the object in the target image.

[0215] The embodiments of the present application also relate to a model training method, which includes: obtaining an image to be trained; obtaining an image of a first object through the model to be trained; calculating the deviation between the image of the first object and the real image of the first object through a preset target loss function; updating the parameters of the model to be trained according to the deviation until the model training condition is satisfied to obtain a first neural network.

[0216] Before model training, obtain a batch of images to be trained, and determine in advance the real image of the first object in each image to be trained. After starting the training, an image to be trained can be input into the model to be trained. Then, obtain the image (predicted image) of the first object in the frame of the image to be trained through the model to be trained. Finally, calculate the deviation between the image of the first object in the image to be trained output by the model to be trained and the real image of the first object in the image to be trained through the target loss function. If the deviation is within the qualified range, the image to be trained is regarded as a qualified image to be trained; if it is outside the qualified range, it is regarded as an unqualified image to be trained. For each training image in this batch of images to be trained, the foregoing process needs to be performed for each one, which will not be elaborated here. If there are only a small number of qualified images to be trained in this batch of images to be trained, adjust the parameters of the model to be trained and retrain with another batch of images to be trained until there are a large number of qualified images to be trained to obtain the first neural network in the embodiment as Figure 4 or Figure 9 shown.

[0217] In the embodiments of the present application, the first neural network trained by this method can accurately obtain the image of the first object in the target image.

[0218] The above is a detailed description of the model training method provided by the embodiments of the present application. The augmented reality device provided by the embodiments of the present application will be introduced below. Figure 13 As a schematic structural diagram of the augmented reality device provided by the embodiments of the present application, as Figure 13 shown, the device includes:

[0219] A first acquisition module 1301, configured to acquire a target image captured by a camera and first position information of a first object in the target image;

[0220] A second acquisition module 1302, configured to acquire second position information of a second object in the three-dimensional coordinate system corresponding to the camera and third position information of a third object in the three-dimensional coordinate system, where the second object is a reference object of the first object, and the second position information and the third position information are preset information;

[0221] A third acquisition module 1303, configured to acquire a pose change amount of a first object relative to a second object according to first position information and second position information;

[0222] A transformation module 1304, configured to transform third position information according to the pose change amount to obtain fourth position information of a third object in a three-dimensional coordinate system;

[0223] A rendering module 1305, configured to render the third object in a target image according to the fourth position information to obtain a new target image.

[0224] In an embodiment of the present application, when it is necessary to add a third object to a target image, the second position information of the second object in the camera coordinate system, the third position information of the third object in the camera coordinate system, and the first position information of the first object in the image coordinate system in the target image may be acquired first. The second position information and the third position information are preset information. Since the second object is a reference object for the first object, the pose change amount of the first object relative to the second object may be determined according to the second position information and the first position information. The pose change amount is used to indicate the position change and orientation change from the second object to the first object in the camera coordinate system. Then, the third object may be made to undergo the same position change and orientation change, that is, the fourth position information of the third object in the camera coordinate system is determined according to the pose change amount and the third position information of the third object. In this way, the final orientation of the third object may be associated with the orientation of the first object (for example, the orientations of the two are the same or the orientations of the two are similar). Finally, the third object is rendered in the target image according to the fourth position information. In the obtained new target image, the orientation of the third object may be adapted to the orientation of the first object, and realistic content can be presented, thereby improving the user experience.

[0225] In a possible implementation manner, the third acquisition module 1303 is configured to: acquire depth information of the first object; acquire fifth position information of the first object in a three-dimensional coordinate system according to the first position information and the depth information; calculate a change amount between the second position information and the fifth position information to obtain a pose change amount of the first object relative to the second object.

[0226] In a possible implementation manner, the third acquisition module 1303 is configured to: transform the second position information to obtain fifth position information of the first object in a three-dimensional coordinate system; project the fifth position information onto the target image to obtain sixth position information; if a change amount between the sixth position information and the first position information meets a preset condition, the pose change amount of the first object relative to the second object is a transformation matrix used to transform the second position information.

[0227] In a possible implementation, the rendering module 1305 is configured to: perform pinhole imaging based on the fourth position information to obtain an image of the third object; obtain the occlusion relationship between the third object and the first object; and fuse the image of the third object and the image of the first object according to the occlusion relationship to obtain a new target image.

[0228] In a possible implementation, the rendering module 1305 is configured to: calculate a first distance between the first object and the origin of the three-dimensional coordinate system according to the fifth position information; calculate a second distance between the third object and the origin of the three-dimensional coordinate system according to the fourth position information; and compare the first distance and the second distance to obtain the occlusion relationship between the third object and the first object.

[0229] In a possible implementation, the rendering module 1305 is configured to: obtain the correspondence between multiple surface points of the first object and multiple surface points of the second object; obtain the distribution of multiple surface points of the first object on the second object according to the correspondence; and obtain the occlusion relationship between the third object and the first object according to the distribution.

[0230] In a possible implementation, the pose change amount includes the orientation change of the first object relative to the second object. The rendering module 1305 is configured to: determine the front of the first object according to the orientation change of the first object relative to the second object; and obtain the occlusion relationship between the third object and the first object according to the included angle between the orientation of the center point of the front of the first object to the origin of the three-dimensional coordinate system and the orientation of the first object.

[0231] In a possible implementation, the rendering module 1305 is further configured to input the target image into a first neural network to obtain an image of the first object.

[0232] Figure 14 Another structural schematic diagram of the augmented reality device provided by the embodiments of the present application is as Figure 14 shown. The device includes:

[0233] A first acquisition module 1401, configured to acquire a target image captured by a camera and third position information of a third object in a three-dimensional coordinate system corresponding to the camera, where the target image includes an image of a first object;

[0234] A second acquisition module 1402, configured to input the target image into a second neural network to obtain a pose change amount of the first object relative to the second object. The second neural network is trained according to second position information of the second object in the three-dimensional coordinate system. The second object is a reference object for the first object, and the second position information and the third position information are preset information;

[0235] The transformation module 1403 is configured to transform the third position information according to the pose change amount to obtain the fourth position information of the third object in the three-dimensional coordinate system;

[0236] The rendering module 1404 is configured to render the third object in the target image according to the fourth position information to obtain a new target image.

[0237] In the embodiment of the present application, when it is necessary to add a third object to the target image, the target image and the third position information of the third object in the camera coordinate system may be obtained first, and the target image includes an image of the first object. Then, the target image is input into the second neural network to obtain the pose change amount of the first object relative to the second object. The second neural network is trained according to the second position information of the second object in the camera coordinate system. The second object is a reference object for the first object, and the second position information and the third position information are preset information. Specifically, the pose change amount of the first object relative to the second object is used to indicate the position change and orientation change from the second object to the first object in the camera coordinate system. Therefore, the third object can be made to undergo the same position change and orientation change, that is, the fourth position information of the third object in the camera coordinate system is determined according to the pose change amount and the third position information of the third object. In this way, the final orientation of the third object can be associated with the orientation of the first object. Finally, the third object is rendered in the target image according to the fourth position information. In the obtained new target image, the orientation of the third object can be adapted to the orientation of the first object, and realistic content can be presented, thereby improving the user experience.

[0238] In a possible implementation manner, the rendering module 1404 is configured to: perform pinhole imaging according to the fourth position information to obtain an image of the third object; obtain the occlusion relationship between the third object and the first object; and fuse the image of the third object and the image of the first object according to the occlusion relationship to obtain a new target image.

[0239] In a possible implementation manner, the transformation module 1403 is further configured to: transform the second position information according to the pose change amount to obtain the fifth position information of the first object in the three-dimensional coordinate system corresponding to the camera.

[0240] In a possible implementation manner, the rendering module 1404 is configured to: calculate a first distance between the first object and the origin of the three-dimensional coordinate system according to the fifth position information; calculate a second distance between the third object and the origin of the three-dimensional coordinate system according to the fourth position information; and compare the first distance and the second distance to obtain the occlusion relationship between the third object and the first object.

[0241] In a possible implementation, the pose change amount includes the orientation change of the first object relative to the second object. The rendering module 1404 is configured to: determine the front of the first object according to the orientation change of the first object relative to the second object; obtain the occlusion relationship between the third object and the first object according to the included angle between the orientation of the center point of the front of the first object to the origin of the three-dimensional coordinate system and the orientation of the first object.

[0242] In a possible implementation, the rendering module 1404 is further configured to input the target image into the first neural network to obtain the image of the first object.

[0243] An embodiment of the present application further relates to another image processing device, which includes: a display module, configured to display a target image in response to a first operation of the user, where the first object is presented in the target image, and the first object is an object in the real environment; a presentation module, configured to present a virtual object in the target image in response to a second operation of the user, where the virtual object is superimposed on the first object; an update module, configured to update the pose of the first object in the target image and the pose of the virtual object in the target image in response to the movement of the first object to obtain a new target image; where the pose includes position and orientation, and the orientation of the virtual object is associated with the orientation of the first object.

[0244] It can be seen from the above device that: when the image processing device responds to the operation of the user, the image processing device can display the target image. Since the target image presents the moving first object, when the image processing device adds the virtual object indicated by the user to the target image, it will update the pose of the first object and the pose of the virtual object in the target image as the first object moves, so that the orientation of the virtual object is associated with the orientation of the first object, thereby enabling the new target image to present realistic content and improving the user experience.

[0245] In a possible implementation, in the new target image, the orientation of the virtual object is the same as the orientation of the first object, and the relative position between the virtual object and the first object remains unchanged after updating the poses of the first object and the virtual object.

[0246] In a possible implementation, the first object is a human body, and the virtual object is a virtual wing.

[0247] In a possible implementation, the first operation is an operation to open the application, and the second operation is an operation to add special effects.

[0248] The above is a detailed description of the augmented reality device provided by the embodiments of the present application. The model training device provided by the embodiments of the present application will be introduced below. Figure 15 This is a schematic structural diagram of the model training device provided by the embodiments of the present application, as Figure 15As shown in the figure, the device includes:

[0249] An acquisition module 1501, configured to acquire an image to be trained, where the image to be trained includes an image of a first object;

[0250] A determination module 1502, configured to input the image to be trained into a model to be trained, and obtain a pose change amount of the first object relative to a second object;

[0251] A calculation module 1503, configured to calculate a deviation between the pose change amount of the first object relative to the second object and the true pose change amount of the first object through a preset target loss function, where the true pose change amount of the first object is determined according to second position information of the second object in a camera coordinate system, the second object is a reference object of the first object, and the second position information is preset information;

[0252] An update module 1504, configured to update parameters of the model to be trained according to the deviation until a model training condition is satisfied, so as to obtain a second neural network.

[0253] In an embodiment of the present application, the second neural network trained by the device can accurately obtain the pose change amount of an object in a target image.

[0254] An embodiment of the present application further relates to a model training device, where the device includes: a first acquisition module, configured to acquire an image to be trained; a second acquisition module, configured to acquire an image of a first object through a model to be trained; a calculation module, configured to calculate a deviation between the image of the first object and the true image of the first object through a preset target loss function; and an update module, configured to update parameters of the model to be trained according to the deviation until a model training condition is satisfied, so as to obtain a first neural network.

[0255] In an embodiment of the present application, the first neural network trained by the device can accurately obtain the image of the first object in a target image.

[0256] It should be noted that, for the information interaction, execution process, and other contents between the above-mentioned device modules / units, since they are based on the same concept as the method embodiment of the present application, the technical effects brought by them are the same as those of the method embodiment of the present application. For specific contents, reference may be made to the description in the method embodiment shown above in the embodiment of the present application, and details are not described herein again.

[0257] An embodiment of the present application further relates to an execution device Figure 16 which is a schematic structural diagram of the execution device provided in an embodiment of the present application. As Figure 16 shown, the execution device 1600 may specifically be a mobile phone, a tablet computer, a notebook computer, a smart wearable device, a server, etc., which is not limited herein. Among them, the execution device 1600 may be deployed with Figure 13 orFigure 14 The augmented reality device described in the corresponding embodiment is used to implement Figure 4 or Figure 9 the function of image processing in the corresponding embodiment. Specifically, the execution device 1600 includes: a receiver 1601, a transmitter 1602, a processor 1603, and a memory 1604 (where the number of processors 1603 in the execution device 1600 can be one or more, Figure 16 and one processor is taken as an example herein), where the processor 1603 may include an application processor 16031 and a communication processor 16032. In some embodiments of the present application, the receiver 1601, the transmitter 1602, the processor 1603, and the memory 1604 may be connected through a bus or other means.

[0258] The memory 1604 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1603. A part of the memory 1604 may also include a non-volatile random access memory (NVRAM). The memory 1604 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, where the operation instructions may include various operation instructions for implementing various operations.

[0259] The processor 1603 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system, where the bus system may include a power bus, a control bus, a status signal bus, etc. in addition to a data bus. However, for the sake of clear illustration, all kinds of buses are referred to as a bus system in the figure.

[0260] The method disclosed in the embodiments of the present application can be applied to or implemented by the processor 1603. The processor 1603 can be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed by the integrated logic circuit in hardware or instructions in software form in the processor 1603. The above-mentioned processor 1603 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 1603 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1604, and the processor 1603 reads the information in the memory 1604 and combines its hardware to complete the steps of the above method.

[0261] The receiver 1601 can be used to receive input digital or character information, and generate signal inputs related to the relevant settings and function controls of the execution device. The transmitter 1602 can be used to output digital or character information through the first interface; the transmitter 1602 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1602 can also include a display device such as a display screen.

[0262] In an embodiment of the present application, in one case, the processor 1603 is used to execute Figure 4 or Figure 9 the augmented reality method executed by the terminal device in the corresponding embodiment.

[0263] The embodiments of the present application also relate to a training device, Figure 17 which is a schematic structural diagram of the training device provided by the embodiments of the present application. As Figure 17As shown, the training device 1700 is implemented by one or more servers. The training device 1700 can vary significantly due to configuration or performance differences, and may include one or more central processing units (CPUs) 1717 (e.g., one or more processors) and a memory 1732, and one or more storage media 1730 (e.g., one or more mass storage devices) for storing application programs 1742 or data 1744. Among them, the memory 1732 and the storage media 1730 can be transient storage or persistent storage. The programs stored in the storage media 1730 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the training device. Further, the central processing unit 1717 can be configured to communicate with the storage media 1730 and execute a series of instruction operations in the storage media 1730 on the training device 1700.

[0264] The training device 1700 may further include one or more power supplies 1726, one or more wired or wireless network interfaces 1750, one or more input / output interfaces 1758; or, one or more operating systems 1741, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0265] Specifically, the training device can execute Figure 12 the steps in the corresponding embodiments.

[0266] The embodiments of the present application also relate to a computer storage medium, in which a program for signal processing is stored. When it runs on a computer, it causes the computer to execute the steps performed by the aforementioned execution device, or causes the computer to execute the steps performed by the aforementioned training device.

[0267] The embodiments of the present application also relate to a computer program product, which stores instructions that, when executed by a computer, cause the computer to execute the steps performed by the aforementioned execution device, or cause the computer to execute the steps performed by the aforementioned training device.

[0268] The execution device, training device, or terminal device provided in the embodiments of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin, or a circuit. The processing unit may execute the computer execution instructions stored in the storage unit to cause the chip in the execution device to execute the data processing method described in the above embodiments, or to cause the chip in the training device to execute the data processing method described in the above embodiments. Optionally, the storage unit is a storage unit inside the chip, such as a register, a cache, etc. The storage unit may also be a storage unit outside the chip in the radio access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0269] Specifically, please refer to Figure 18 , Figure 18 which is a schematic structural diagram of the chip provided in the embodiments of the present application. The chip may be embodied as a neural network processor NPU 1800. The NPU 1800 is mounted on the main CPU (Host CPU) as a coprocessor, and tasks are assigned by the Host CPU. The core part of the NPU is the arithmetic circuit 1803, and the arithmetic circuit 1803 is controlled by the controller 1804 to extract matrix data from the memory and perform multiplication operations.

[0270] In some implementations, the arithmetic circuit 1803 includes multiple processing units (Process Engine, PE) inside. In some implementations, the arithmetic circuit 1803 is a two-dimensional systolic array. The arithmetic circuit 1803 may also be a one-dimensional systolic array or other electronic circuits that can perform mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1803 is a general matrix processor.

[0271] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 1802 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 1801 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are saved in the accumulator 1808.

[0272] The unified memory 1806 is used to store input data and output data. The weight data is directly transferred through the Direct Memory Access Controller (DMAC) 1805, and the DMAC transfers it to the weight memory 1802. The input data is also transferred to the unified memory 1806 through the DMAC.

[0273] The BIU is the Bus Interface Unit, that is, the bus interface unit 1810, which is used for the interaction between the AXI bus, the DMAC, and the Instruction Fetch Buffer (IFB) 1809.

[0274] The bus interface unit 1810 (Bus Interface Unit, abbreviated as BIU) is used for the instruction fetch buffer 1809 to obtain instructions from the external memory, and is also used for the storage unit access controller 1805 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0275] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 1806, or transfer the weight data to the weight memory 1802, or transfer the input data to the input memory 1801.

[0276] The vector calculation unit 1807 includes multiple arithmetic processing units, which, if necessary, further process the output of the arithmetic circuit 1803, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolution / full connection layer network calculations in neural networks, such as Batch Normalization, pixel-level summation, upsampling of the feature plane, etc.

[0277] In some implementations, the vector calculation unit 1807 can store the processed output vector in the unified memory 1806. For example, the vector calculation unit 1807 can apply a linear function; or, a non-linear function to the output of the arithmetic circuit 1803, such as linear interpolation of the feature plane extracted by the convolutional layer, and another example is the vector of the accumulated value, to generate the activation value. In some implementations, the vector calculation unit 1807 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit 1803, such as for use in subsequent layers in the neural network.

[0278] The instruction fetch buffer 1809 connected to the controller 1804 is used to store the instructions used by the controller 1804;

[0279] The unified memory 1806, the input memory 1801, the weight memory 1802, and the fetch memory 1809 are all On-Chip memories. The external memory is private to the NPU hardware architecture.

[0280] Among them, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above programs.

[0281] In addition, it should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.

[0282] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by means of dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, software program implementation is a better implementation method in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in various embodiments of this application.

[0283] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0284] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

Claims

1. An augmented reality method, characterized in that, Including: In response to a first operation of the user, a target image is displayed, and a first object is presented in the target image, where the first object is an object in the real environment; In response to a second operation of the user, a virtual object is presented in the target image, and the virtual object is superimposed on the first object; In response to the movement of the first object, update the pose of the first object in the target image and update the pose of the virtual object in the target image to obtain a new target image; where the pose includes position and orientation, the orientation of the virtual object is associated with the orientation of the first object, the pose of the virtual object in the target image is determined according to the pose of the first object in the target image and the second position information of the second object in the three-dimensional coordinate system, the second object is a reference object of the first object in the virtual environment, and the pose change amount of the first object relative to the second object is determined by the pose of the first object in the target image and the second position information of the second object in the three-dimensional coordinate system, and the position of the virtual object in the target image is updated according to the pose change amount to obtain the pose of the virtual object in the target image.

2. The method according to claim 1, characterized in that In the new target image, the orientation of the virtual object is the same as the orientation of the first object, and the relative position between the virtual object and the first object remains unchanged after updating the poses of the first object and the virtual object.

3. The method according to claim 1 or 2, characterized in that, The first object is a human body, and the virtual object is a virtual wing.

4. The method according to any one of claims 1 to 2, characterized in that The first operation is an operation to start the application, and the second operation is an operation to add special effects.

5. An augmented reality method, characterized in that, The method includes: Obtain a target image and first position information of a first object in the target image; Obtain second position information of a second object in the three-dimensional coordinate system and third position information of a third object in the three-dimensional coordinate system, where the second object is a reference object of the first object in the virtual environment, and the second position information and the third position information are preset information; According to the first position information and the second position information, obtain the pose change amount of the first object relative to the second object; Transform the third position information according to the pose change amount to obtain fourth position information of the third object in the three-dimensional coordinate system; Render the third object in the target image according to the fourth position information to obtain a new target image.

6. The method according to claim 5, characterized in that, The obtaining the pose change amount of the first object relative to the second object according to the first position information and the second position information includes: Obtain the depth information of the first object; According to the first position information and the depth information, obtain fifth position information of the first object in the three-dimensional coordinate system; According to the second position information and the fifth position information, obtain the pose change amount of the first object relative to the second object.

7. The method according to claim 5, wherein The obtaining the pose change amount of the first object relative to the second object according to the first position information and the second position information includes: Transform the second position information to obtain the fifth position information of the first object in the three-dimensional coordinate system; Project the fifth position information onto the target image to obtain the sixth position information; If the change amount between the sixth position information and the first position information meets a preset condition, the pose change amount of the first object relative to the second object is the transformation matrix used to transform the second position information.

8. The method according to claim 6 or 7, characterized in that, The rendering the third object in the target image according to the fourth position information to obtain a new target image includes: Perform pinhole imaging according to the fourth position information to obtain an image of the third object; Obtain the occlusion relationship between the third object and the first object; Fuse the image of the third object and the image of the first object according to the occlusion relationship to obtain a new target image.

9. The method according to claim 8, characterized in that, The obtaining the occlusion relationship between the third object and the first object includes: Calculate a first distance between the first object and the origin of the three-dimensional coordinate system according to the fifth position information; Calculate a second distance between the third object and the origin of the three-dimensional coordinate system according to the fourth position information; Compare the first distance and the second distance to obtain the occlusion relationship between the third object and the first object.

10. The method according to claim 8, characterized in that The obtaining the occlusion relationship between the third object and the first object includes: Obtain the corresponding relationship between multiple surface points of the first object and multiple surface points of the second object; According to the corresponding relationship, obtain the distribution of multiple surface points of the first object on the second object; According to the distribution, obtain the occlusion relationship between the third object and the first object.

11. The method according to claim 8, characterized in that The pose change amount includes the orientation change of the first object relative to the second object, and the obtaining the occlusion relationship between the third object and the first object includes: Determine the front of the first object according to the orientation change of the first object relative to the second object; Obtain the occlusion relationship between the third object and the first object according to the included angle between the orientation from the center point of the front of the first object to the origin of the three-dimensional coordinate system and the orientation of the first object.

12. The method according to claim 8, wherein The method further includes: Input the target image into a first neural network to obtain an image of the first object.

13. An augmented reality method, characterized in that, The method includes: Obtain a target image and the third position information of the third object in the three-dimensional coordinate system, where the target image includes an image of the first object; Input the target image into a second neural network to obtain the pose change amount of the first object relative to the second object, where the second neural network is trained according to the second position information of the second object in the three-dimensional coordinate system, the second object is a reference object of the first object in the virtual environment, and the second position information and the third position information are preset information; Transform the third position information according to the pose change amount to obtain the fourth position information of the third object in the three-dimensional coordinate system; Render the third object in the target image according to the fourth position information to obtain a new target image.

14. The method according to claim 13, wherein The rendering of the third object in the target image according to the fourth position information to obtain a new target image includes: Perform pinhole imaging according to the fourth position information to obtain an image of the third object; Obtain the occlusion relationship between the third object and the first object; Fuse the image of the third object and the image of the first object according to the occlusion relationship to obtain a new target image.

15. The method according to claim 14, wherein The method further includes: Transform the second position information according to the pose change amount to obtain a fifth position information of the first object in the three-dimensional coordinate system corresponding to the camera.

16. The method according to claim 15, wherein The obtaining of the occlusion relationship between the third object and the first object includes: Calculate a first distance between the first object and the origin of the three-dimensional coordinate system according to the fifth position information; Calculate a second distance between the third object and the origin of the three-dimensional coordinate system according to the fourth position information; Compare the first distance and the second distance to obtain the occlusion relationship between the third object and the first object.

17. The method according to claim 14, wherein The pose change amount includes the orientation change of the first object relative to the second object, and the obtaining of the occlusion relationship between the third object and the first object includes: Determine the front of the first object according to the orientation change of the first object relative to the second object; Obtain the occlusion relationship between the third object and the first object according to the angle between the orientation from the center point of the front of the first object to the origin of the three-dimensional coordinate system and the orientation of the first object.

18. The method according to any one of claims 14 to 17, characterized in that The method further includes: Input the target image into a first neural network to obtain an image of the first object.

19. An augmented reality device, characterized in that, The device includes: A display module, configured to display a target image in response to a first operation of a user, with a first object presented in the target image, and the first object is an object in the real environment; A presentation module, configured to present a virtual object in the target image in response to a second operation of the user, and the virtual object is superimposed on the first object; An update module, configured to update the pose of the first object in the target image and update the pose of the virtual object in the target image in response to the movement of the first object to obtain a new target image; wherein, the pose includes position and orientation, the orientation of the virtual object is associated with the orientation of the first object, the pose of the virtual object in the target image is determined according to the pose of the first object in the target image and the second position information of the second object in the three-dimensional coordinate system, the second object is a reference object of the first object in the virtual environment, and the pose of the first object in the target image and the second position information of the second object in the three-dimensional coordinate system are used to determine the pose change amount of the first object relative to the second object, and the position of the virtual object in the target image is updated according to the pose change amount in the target image to obtain the pose of the virtual object in the target image.

20. The device according to claim 19, characterized in that, In the new target image, the orientation of the virtual object is the same as that of the first object, and the relative position between the virtual object and the first object remains unchanged after updating the poses of the first object and the virtual object.

21. The device according to claim 19 or 20, characterized in that, The first object is a human body, and the virtual object is a virtual wing.

22. The device according to any one of claims 19 to 20, characterized in that, The first operation is an operation to start an application, and the second operation is an operation to add special effects.

23. An augmented reality device, characterized in that, The device includes: A first acquisition module, configured to acquire a target image and first position information of a first object in the target image; A second acquisition module, configured to acquire second position information of a second object in a three-dimensional coordinate system and third position information of a third object in the three-dimensional coordinate system, where the second object is a reference object of the first object in a virtual environment, and the second position information and the third position information are pre-set information; A third acquisition module, configured to acquire a pose change amount of the first object relative to the second object according to the first position information and the second position information; A transformation module, configured to transform the third position information according to the pose change amount to obtain fourth position information of the third object in the three-dimensional coordinate system; A rendering module, configured to render the third object in the target image according to the fourth position information to obtain a new target image.

24. The device according to claim 23, wherein The third acquisition module is configured to: Acquire depth information of the first object; Acquire fifth position information of the first object in the three-dimensional coordinate system according to the first position information and the depth information; Obtain a pose change amount of the first object relative to the second object according to the second position information and the fifth position information.

25. The device according to claim 23, characterized in that, The third acquisition module is configured to: Transform the second position information to obtain fifth position information of the first object in the three-dimensional coordinate system; Project the fifth position information onto the target image to obtain sixth position information; If the change amount between the sixth position information and the first position information meets a pre-set condition, the pose change amount of the first object relative to the second object is a transformation matrix used to transform the second position information.

26. The device according to claim 24 or 25, characterized in that The rendering module is configured to: Perform pinhole imaging according to the fourth position information to obtain an image of the third object; Obtain an occlusion relationship between the third object and the first object; Fuse the image of the third object and the image of the first object according to the occlusion relationship to obtain a new target image.

27. The device according to claim 26, characterized in that, The rendering module is configured to: Calculate a first distance between the first object and the origin of the three-dimensional coordinate system according to the fifth position information; Calculate a second distance between the third object and the origin of the three-dimensional coordinate system according to the fourth position information; Compare the first distance and the second distance to obtain an occlusion relationship between the third object and the first object.

28. The device according to claim 26, characterized in that, The rendering module is configured to: Obtain a correspondence relationship between a plurality of surface points of the first object and a plurality of surface points of the second object; Obtain the distribution of multiple surface points of the first object on the second object according to the corresponding relationship; Obtain the occlusion relationship between the third object and the first object according to the distribution; 29. The device according to claim 26, characterized in that, The pose change amount includes the orientation change of the first object relative to the second object. The rendering module is used for: Determine the front of the first object according to the orientation change of the first object relative to the second object; Obtain the occlusion relationship between the third object and the first object according to the angle between the orientation of the center point of the front of the first object to the origin of the three-dimensional coordinate system and the orientation of the first object; 30. The device according to claim 26, wherein, The rendering module is further used to input the target image into the first neural network to obtain the image of the first object.

31. An augmented reality device, characterized in that, The device includes: A first acquisition module, configured to acquire a target image and the third position information of the third object in the three-dimensional coordinate system, where the target image includes the image of the first object; A second acquisition module, configured to input the target image into the second neural network to obtain the pose change amount of the first object relative to the second object. The second neural network is trained according to the second position information of the second object in the three-dimensional coordinate system. The second object is a reference object of the first object in the virtual environment, and the second position information and the third position information are preset information; A transformation module, configured to transform the third position information according to the pose change amount to obtain the fourth position information of the third object in the three-dimensional coordinate system; A rendering module, configured to render the third object in the target image according to the fourth position information to obtain a new target image.

32. The device according to claim 31, wherein The rendering module is used for: Perform pinhole imaging according to the fourth position information to obtain the image of the third object; Obtain the occlusion relationship between the third object and the first object; Fuse the image of the third object and the image of the first object according to the occlusion relationship to obtain a new target image.

33. The device according to claim 32, characterized in that, The transformation module is further used for: Transform the second position information according to the pose change amount to obtain the fifth position information of the first object in the three-dimensional coordinate system corresponding to the camera.

34. The apparatus according to claim 33, wherein The rendering module is used for: Calculate the first distance between the first object and the origin of the three-dimensional coordinate system according to the fifth position information; Calculate the second distance between the third object and the origin of the three-dimensional coordinate system according to the fourth position information; Compare the first distance and the second distance to obtain the occlusion relationship between the third object and the first object.

35. The device according to claim 32, characterized in that, The pose change amount includes the orientation change of the first object relative to the second object. The rendering module is used for: Determine the front of the first object according to the orientation change of the first object relative to the second object; Obtain the occlusion relationship between the third object and the first object according to the angle between the orientation of the center point of the front of the first object to the origin of the three-dimensional coordinate system and the orientation of the first object.

36. The device according to any one of claims 32 to 35, characterized in that The rendering module is further configured to input the target image into a first neural network to obtain an image of the first object.

37. An augmented reality device, characterized in that, It includes a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the augmented reality device performs the method according to any one of claims 1 to 18.

38. A computer storage medium, characterized in that, The computer storage medium stores a computer program, and when the program is executed by a computer, the computer implements the method according to any one of claims 1 to 18.

39. A computer program product, characterized in that, The computer program product stores instructions, and when the instructions are executed by a computer, the computer implements the method according to any one of claims 1 to 18.

Citation Information

Patent Citations

  • Virtual dress-up system, method, device and medium

    CN110363867A

  • Augmented reality visualization method based on depth camera and application

    CN112258658A

  • Methods and systems for enabling creation of augmented reality content

    US20150040074A1

  • Skeleton-based effects and background replacement

    US20200021752A1