Scene rendering method and device, AR equipment and storage medium

By obtaining the first depth data from the user's observation perspective and the second depth data of the target scene, accurately judge the positional relationship of objects in the AR scene, solving the problem of inaccurate occlusion relationship between the virtual environment and the real environment object, and improving the realism and interactivity of the rendering.

CN120013780AActive Publication Date: 2025-05-16ZHUHAI MOJIE TECH CO LTD

Patent Information

Application Number
CN202510498729.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-16
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

In AR scene rendering, the front and back position relationship between objects in the virtual environment and objects in the real environment is inaccurate, resulting in inaccurate occlusion relationships, distortion of rendering, and affecting the user's visual experience.

Method used

By obtaining the first depth data from the user's observation perspective and the second depth data of the target scene, based on these data, accurately judge the front and back position relationship between the object in the target scene and the objects in the related scene, and render it, ensuring that the depth of the object in the target scene is less than the depth of the object in the relevant scene at the corresponding position.

Benefits of technology

It effectively avoids the problem of inaccurate occlusion relationship between objects and objects in related scenes in target scenes, and enhances the realism and interactivity of scene rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013780A_ABST
    Figure CN120013780A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of rendering, and provides a scene rendering method and device, AR equipment and a storage medium, and the method comprises the steps: obtaining first depth data corresponding to a related scene, the first depth data being depth data under a coordinate system corresponding to a user observation visual angle; acquiring second depth data corresponding to the target scene; and based on the first depth data and the second depth data, the target scene is rendered, and the depth of an object in the rendered target scene is smaller than the depth of an object in the related scene at the corresponding position. According to the embodiment of the invention, the reality effect of scene rendering is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of rendering technology, and in particular to a scene rendering method, apparatus, AR device and storage medium. Background Art

[0002] Rendering technology refers to the process of converting a three-dimensional scene or model into a two-dimensional image. It is widely used in computer graphics, game development, film and television production, virtual reality (VR), augmented reality (AR) and other fields. Taking AR scene rendering as an example, the virtual environment is superimposed and integrated with the real environment, and presented to the user to obtain a realistic visual experience. During the rendering process, if the front-to-back position relationship between the objects in the virtual environment and the objects in the real environment is incorrect, the occlusion relationship between the virtual and real objects will be inaccurate, the rendering will be distorted, and the user will have a visual sense of unreality, affecting the user's visual experience. Summary of the invention

[0003] The present application provides a scene rendering method, apparatus, AR device and storage medium, aiming to solve the rendering distortion problem and enhance the realism effect of scene rendering.

[0004] To achieve the above object, the present application provides a scene rendering method, the scene rendering method comprising: Acquire first depth data corresponding to the relevant scene, where the first depth data is depth data in a coordinate system corresponding to the user's observation perspective; Acquire second depth data corresponding to the target scene; The target scene is rendered based on the first depth data and the second depth data, and the depth of the rendered object in the target scene is less than the depth of the object in the related scene at the corresponding position.

[0005] In addition, to achieve the above-mentioned purpose, the present application also provides a scene rendering device, which includes a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program and implement the steps of the scene rendering method as described above when executing the computer program.

[0006] In addition, to achieve the above-mentioned purpose, the present application also provides an AR device, which includes the scene rendering device as described above.

[0007] In addition, to achieve the above-mentioned purpose, the present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned scene rendering method are implemented.

[0008] The present application discloses a scene rendering method, apparatus, AR device and storage medium, which obtain first depth data corresponding to a related scene, wherein the first depth data is depth data in a coordinate system corresponding to a user's observation perspective, and obtain second depth data corresponding to a target scene, and based on the first depth data and the second depth data, accurately judge the front-to-back position relationship between an object in a target scene and an object in a related scene, and render the target scene. The depth of the rendered object in the target scene is smaller than the depth of the object in the related scene at the corresponding position, thereby avoiding the problem of inaccurate occlusion relationship between the object in the target scene and the object in the related scene, thereby enhancing the realism effect of the target scene rendering. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0010] Figure 1 A schematic diagram of a flow chart of a scene rendering method provided in an embodiment of the present application; Figure 2 A schematic diagram of a sub-process of a scene rendering method provided in an embodiment of the present application; Figure 3 A schematic diagram of a flow chart of another scene rendering method provided in an embodiment of the present application; Figure 4 A schematic diagram of a sub-process of another scene rendering method provided in an embodiment of the present application; Figure 5 A schematic diagram of a sub-process of another scene rendering method provided in an embodiment of the present application; Figure 6 A schematic diagram of a process of rendering a virtual scene under a mobile perspective provided in an embodiment of the present application; Figure 7 A schematic block diagram of a scene rendering device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0011] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0012] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.

[0013] It should be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0014] It should also be understood that the term “and / or” used in the specification and appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0015] The embodiments of the present application provide a scene rendering method, apparatus, AR device and storage medium for enhancing the realism effect of scene rendering.

[0016] See also Figure 1 , Figure 1 1 is a flow chart of a scene rendering method provided in an embodiment of the present application. The method can be applied to a scene rendering device or an AR device, and the application scenario of the method is not limited in the present application.

[0017] like Figure 1 As shown, the scene rendering method specifically includes steps S101 to S103.

[0018] S101. Acquire first depth data corresponding to a relevant scene, where the first depth data is depth data in a coordinate system corresponding to a user's observation perspective.

[0019] Among them, the related scene is a scene that does not need to be rendered, and the related scene will be superimposed and merged with the scene to be rendered. For example, the related scene can be a real scene in the real physical environment that the user is currently in. It should be noted that the related scene can also be other scenes, which is not limited in the embodiments of the present application.

[0020] Generally speaking, the relevant scene contains one or more objects, for example, the real scene contains multiple real objects. The depth data of the objects in the relevant scene in the coordinate system corresponding to the user's observation perspective is obtained. For the convenience of distinguishing the description, the depth data in the coordinate system corresponding to the user's observation perspective is referred to as the first depth data below. Since the first depth data is the depth data in the coordinate system corresponding to the user's observation perspective, the first depth data can accurately correspond to the user's observation perspective, so that the subsequent judgment of the occlusion relationship between objects based on the first depth data is more accurate and reliable.

[0021] In some embodiments, Figure 2 As shown, step S101 may include sub-step S1011 and sub-step S1012.

[0022] S1011, obtaining scene depth data corresponding to the relevant scene, and obtaining a transformation matrix between a coordinate system corresponding to the device and a coordinate system corresponding to the user's observation angle, wherein the scene depth data is depth data in the coordinate system corresponding to the device; S1012: Obtain first depth data based on the scene depth data and the transformation matrix.

[0023] For example, a depth map of a relevant scene is captured by a depth camera, a depth camera or other device, and then scene depth data of objects in the relevant scene in the device's corresponding coordinate system is obtained based on the depth map. For another example, scene depth data of objects in the relevant scene in the device's corresponding coordinate system is obtained by collecting data using a sensor device such as a depth detector. The present application does not limit the method for obtaining scene depth data corresponding to the relevant scene.

[0024] It is understandable that since the scene depth data is obtained through devices such as depth cameras, depth cameras, depth detectors, etc., the scene depth data is the depth data in the corresponding coordinate system of the corresponding device. For example, if the scene depth data is obtained through the depth image taken by the depth camera, the scene depth data is the depth data in the depth camera coordinate system.

[0025] In actual applications, the coordinate system corresponding to the device is likely to be inconsistent with the coordinate system corresponding to the user's observation perspective. In order to make the depth data corresponding to the objects in the relevant scene accurately correspond to the user's observation perspective, the obtained scene depth data is converted into a coordinate system through the transformation matrix between the coordinate system corresponding to the device and the coordinate system corresponding to the user's observation perspective, so as to obtain the first depth data corresponding to the user's observation perspective. Compared with using the third depth data corresponding to devices such as depth cameras to determine the occlusion relationship between objects, the first depth data corresponding to the user's observation perspective is used to determine the occlusion relationship between objects, and the determination result is more in line with the user's visual observation.

[0026] Exemplarily, the transformation matrix between the coordinate system corresponding to the device and the coordinate system corresponding to the user's observation angle includes a rotation matrix, a projection matrix, and a translation vector. Among them, the rotation matrix describes the rotation relationship between the coordinate system corresponding to the device and the coordinate system corresponding to the user's observation angle; in computer graphics, the projection matrix is ​​often used to project points in three-dimensional space onto a two-dimensional plane. The projection matrix in this application is used to perform appropriate projection transformation on scene depth data; the translation vector describes the position offset of the coordinate system corresponding to the device relative to the coordinate system corresponding to the user's observation angle.

[0027] For example, the transformation matrix between the coordinate system corresponding to the device and the coordinate system corresponding to the user's observation angle is shown in formula (1): Tobscam=R·P+t(1) Among them, Tobscam is the transformation matrix between the coordinate system corresponding to the device and the coordinate system corresponding to the user's observation angle, R is the rotation matrix, P is the projection matrix, and t is the translation vector.

[0028] It should be noted that the transformation matrix between the coordinate system corresponding to the device and the coordinate system corresponding to the user's observation perspective can be currently established or pre-established and saved. Currently, you only need to call it to obtain the transformation matrix between the coordinate system corresponding to the device and the coordinate system corresponding to the user's observation perspective.

[0029] Based on the transformation matrix between the coordinate system corresponding to the device and the coordinate system corresponding to the user's observation angle of view, the scene depth data in the coordinate system corresponding to the device is converted into first depth data in the coordinate system corresponding to the user's observation angle of view.

[0030] Exemplarily, the first depth data in the coordinate system corresponding to the user's observation angle is calculated by formula (2): Dobsdepth=Tobscam·Dcamdepth(2) Among them, Tobscam is the transformation matrix between the coordinate system corresponding to the device and the coordinate system corresponding to the user's observation angle, Dcamdepth is the scene depth data in the coordinate system corresponding to the device, and Dobsdepth is the first depth data in the coordinate system corresponding to the user's observation angle.

[0031] For example, the scene depth data includes a depth map matrix in the depth camera coordinate system obtained based on the depth map. Based on the depth map matrix and the transformation matrix between the depth camera coordinate system and the coordinate system corresponding to the user's observation angle, each pixel in the depth map is remapped to the coordinate system corresponding to the user's observation angle to obtain a remapped depth map adapted to the user's observation angle. In this way, in subsequent rendering and other operations, the occlusion relationship of the object can be accurately processed according to the remapped depth map, thereby improving the realism and interactivity of the scene rendering.

[0032] Exemplarily, obtaining the first depth data based on the scene depth data and the transformation matrix includes: based on the transformation matrix and bilinear interpolation, converting the scene depth data into a coordinate system corresponding to the user's observation perspective to obtain the first depth data.

[0033] The transformation matrix describes the conversion relationship from the coordinate system corresponding to the device to the coordinate system corresponding to the user's observation angle. Bilinear interpolation is to calculate the depth value at the new coordinate position by taking the weighted average of the depth values ​​at adjacent coordinate positions. Each pixel in the depth map is remapped to the coordinate system corresponding to the user's observation angle. There may be pixels whose new coordinates in the coordinate system corresponding to the user's observation angle are not integers. By using bilinear interpolation to calculate the weighted average of the depth values ​​corresponding to the pixels at their adjacent coordinate positions, as the depth value corresponding to the pixel, a more accurate and reliable remapped depth map can be obtained.

[0034] It should be noted that in addition to the above-mentioned method of obtaining the first depth data through the transformation matrix, the first depth data can also be obtained through other methods, for example, obtaining the first depth data through point cloud registration. The method of obtaining the first depth data is not specifically limited in this application.

[0035] S102: Acquire second depth data corresponding to the target scene.

[0036] The target scene is a scene to be rendered, for example, the target scene may be a virtual scene. It should be noted that the target scene may also be other scenes, which is not limited in the embodiments of the present application.

[0037] The target scene also includes one or more objects, for example, the virtual scene includes multiple virtual objects. For the target scene, the depth data corresponding to the objects in the target scene is obtained, which is referred to as the second depth data below for the convenience of distinguishing descriptions.

[0038] For example, taking a virtual scene as an example, based on the perspective projection matrix, the vertex coordinates of the virtual object in the virtual scene are multiplied by the perspective projection matrix, and the vertices of the virtual object are appropriately projected and transformed to obtain the depth value of the vertices of the virtual object in the virtual space. This depth value represents the position of the vertices of the virtual object relative to the near clipping plane and the far clipping plane, thereby determining the depth data of the virtual object in the virtual scene, that is, obtaining the second depth data.

[0039] It should be noted that step S102 may be executed after step S101, or before step S101, or step S102 and step S101 may be executed simultaneously, and there is no limitation on the order in which step S102 and step S101 are executed.

[0040] S103: Rendering a target scene based on the first depth data and the second depth data, wherein the depth of an object in the rendered target scene is less than the depth of an object in a related scene at a corresponding position.

[0041] By comparing the first depth data and the second depth data, the front-to-back position relationship between the relevant scene object and the target scene object at the corresponding position is determined. If the depth of the target scene object is less than the depth of the relevant scene object at the corresponding position, that is, the target scene object is in front and the relevant scene object is behind, and the target scene object is not blocked, then the corresponding pixel of the target scene object is displayed. Conversely, if the depth of the target scene object is greater than or equal to the depth of the relevant scene object at the corresponding position, that is, the target scene object is behind and the relevant scene object is in front, and the target scene object is blocked, then the corresponding pixel of the target scene object is hidden.

[0042] By obtaining first depth data corresponding to the relevant scene that is adapted to the user's observation perspective, and comparing the first depth data with the second depth data corresponding to the target scene, the depth comparison result is more accurate and reliable, avoiding the problem of inaccurate occlusion relationship between the target scene object and the relevant scene object, thereby improving the realism and interactivity of the target scene rendering.

[0043] Exemplarily, taking the relevant scene as a real scene and the target scene as a virtual scene as an example, the first depth data includes a first depth value corresponding to a real object in the real scene, and the second depth data includes a second depth value corresponding to a virtual object in the virtual scene. Based on the first depth data and the second depth data, rendering the target scene includes: comparing the second depth value of the virtual object with the first depth value of the real object at the corresponding position, and if the second depth value is less than the first depth value, rendering the pixel point corresponding to the virtual object.

[0044] By comparing the second depth value of the virtual object in the virtual scene with the first depth value of the real object in the corresponding real scene, if the second depth value is less than the first depth value, that is, the virtual object is in front and the real object is behind, and the virtual object is not blocked, then the pixel points corresponding to the virtual object are rendered and displayed. On the contrary, if the second depth value is greater than or equal to the first depth value, that is, the virtual object is behind and the real object is in front, and the virtual object is blocked, then the pixel points corresponding to the virtual object are hidden.

[0045] By comparing the second depth value with the first depth value, the front-to-back position relationship between the virtual object and the real object can be accurately determined, thereby avoiding the problem of inaccurate occlusion relationship between the virtual object and the real object, thereby improving the realism of virtual scene rendering.

[0046] In some embodiments, the first depth data and the second depth data are obtained by parallel processing based on different graphics processing unit (GPU) channels.

[0047] Exemplarily, a dual-channel processing architecture is constructed in a GPU shader. For example, the dual-channel processing architecture is as follows: void main() {float realDepth = texture2D(realDepthTex, uv).r; floatvirtualDepth = computeVirtualDepth();gl_FragColor = (virtualDepth <realDepth)? virtualObj : realObj;} The dual-channel processing architecture means that the depth comparison task is assigned to two different channels for parallel processing, where one channel can be dedicated to processing the second depth data corresponding to the target scene, and the other channel is dedicated to processing the first depth data corresponding to the related scene. For example, one channel is dedicated to processing the depth data of virtual objects in a virtual scene, and the other channel is dedicated to processing the depth data of real objects in a real scene.

[0048] Exemplarily, the obtained first depth data and second depth data are cached in a cache area, for example, the first depth data and second depth data are cached in a GPU buffer area. By comparing the first depth data and second depth data in the GPU buffer area, the occlusion relationship between the target scene object and the related scene objects is determined.

[0049] Through dual-channel parallel processing, the efficiency of depth data comparison is greatly improved, and the occlusion relationship between the target scene object and related scene objects can be determined in real time, so as to perform rendering operations more efficiently and decide which pixels should be displayed and which should be hidden.

[0050] It should be noted that in addition to the above-mentioned method of parallel processing through GPU dual channels, the first depth data and the second depth data can also be processed in parallel through other methods, for example, by using CPU multi-threading to simultaneously process the first depth data and the second depth data, thereby improving data processing efficiency.

[0051] In some embodiments, Figure 3 As shown, step S104 may be included before step S101, and step S101 may include sub-step S1013.

[0052] S104, determining the displacement of the user within a preset time period; S1013: If the displacement is greater than a preset displacement threshold, first depth data corresponding to the relevant scene is obtained, where the first depth data is depth data in a coordinate system corresponding to the user's observation angle of view.

[0053] In actual applications, the user may be in a fixed perspective, that is, the user will not move, or in a mobile perspective, that is, the user will move. In a fixed perspective, the first depth data in the coordinate system corresponding to the user's observation perspective is unchanged, so the first depth data in the coordinate system corresponding to the user's observation perspective is obtained once, and then the rendering operation can be performed. In a mobile perspective, the first depth data in the coordinate system corresponding to the user's observation perspective changes, and it is necessary to repeatedly obtain the first depth data in the coordinate system corresponding to the user's observation perspective, dynamically update the first depth data, and perform the rendering operation of the target scene based on the dynamically updated first depth data.

[0054] The first depth data can be dynamically updated by presetting an update cycle and performing the operation of acquiring the first depth data regularly based on the update cycle. The operation of acquiring the first depth data each time can refer to the description in the previous embodiment, so it will not be repeated here.

[0055] By dynamically updating the first depth data, the problem of inaccurate occlusion relationship between the target scene object and related scene objects under a mobile perspective is avoided, thereby improving the realism and interactivity of the target scene rendering under a mobile perspective.

[0056] Considering that if the user only moves very slightly, the change in the first depth data is very small and will not change the occlusion relationship between the target scene object and the related scene objects, there is actually no need to re-acquire the first depth data and update the first depth data.

[0057] Therefore, a preset time length for updating the first depth data is preset, and the displacement of the user within the preset time length is determined at intervals of the preset time length. It should be noted that the preset time length can be flexibly set according to actual conditions, and is not specifically limited in this application.

[0058] In addition, a preset displacement threshold for determining whether to update the first depth data is preset, for example, the preset displacement threshold is set to 5 cm. It should be noted that the preset displacement threshold can be flexibly set according to actual conditions, and is not specifically limited in this application.

[0059] If the displacement of the user within the preset time length is greater than the preset displacement threshold, it means that the user has moved a certain distance. In order to ensure the accuracy of the occlusion relationship between the target scene object and the related scene object, at this time, the operation of obtaining the first depth data corresponding to the related scene is performed. For example, if the displacement of the user within the preset time length is greater than the preset displacement threshold, the depth map recalculation is triggered to obtain a remapped depth map adapted to the user's observation angle.

[0060] If the user's displacement within the preset time is less than or equal to the preset displacement threshold, it means that the user has moved a small distance and will not change the occlusion relationship between the target scene object and the related scene object. At this time, the operation of obtaining the first depth data corresponding to the related scene is not performed, thereby avoiding unnecessary depth map calculation process and reducing energy consumption.

[0061] Therefore, whether to re-acquire the first depth data is triggered based on the user's displacement, which not only avoids the problem of misalignment between the target scene object and the related scene objects under the moving perspective, but also reduces energy consumption.

[0062] In some embodiments, Figure 4 As shown, step S104 may include sub-step S1041 and sub-step S1042.

[0063] S1041, obtaining user posture information corresponding to a preset duration, where the user posture information includes at least one of inertial measurement unit (IMU) data and simultaneous localization and mapping (SLAM) data; S1042: Determine displacement based on user posture information.

[0064] For example, the IMU data such as the acceleration and angular velocity of the user's movement can be obtained through the IMU measurement of the AR glasses worn by the user. Based on the IMU data, the user's motion state and position change can be inferred, and the user's displacement within a preset time period can be obtained.

[0065] For example, images are captured by the camera of AR glasses worn by the user, and through feature extraction, matching and other algorithm processing, a map of the user's surrounding environment is constructed in real time, and SLAM data such as the user's own position (such as x, y, z three-dimensional coordinates) and posture (such as rotation angle) in the map are determined, and then the user's displacement within a preset time period is determined based on the SLAM data.

[0066] For example, IMU data and SLAM data can be obtained and integrated to give full play to the advantages of the two data. IMU data can provide high-frequency motion information, while SLAM data can provide more accurate position and posture information. By integrating the two data, the user's displacement within a preset time period can be calculated more accurately.

[0067] In some embodiments, Figure 5 As shown, step S1041 may include sub-step S10411 and sub-step S10412.

[0068] S10411, capturing images of relevant scenes with a camera device based on a preset refresh rate, wherein the preset duration is determined by the preset refresh rate; S10412, obtaining first SLAM data corresponding to the current frame image and second SLAM data corresponding to the previous frame image.

[0069] The camera device includes but is not limited to a camera, a still camera, etc. The preset refresh rate of the camera device is, for example, 100 Hz, but it can also be other values, which is not specifically limited in this application.

[0070] For example, the image of the relevant scene is captured by a camera with a 100Hz refresh rate of the AR glasses worn by the user. Based on the current frame image obtained by the shooting, the SLAM data such as the position and posture of the user corresponding to the timestamp of the current frame image can be obtained. For the convenience of distinguishing the description, it is referred to as the first SLAM data below. And based on the previous frame image, the SLAM data such as the position and posture of the user corresponding to the timestamp of the previous frame image can be obtained. For the convenience of distinguishing the description, it is referred to as the second SLAM data below.

[0071] The displacement of the user within the time period corresponding to the timestamps of two adjacent frames of images, that is, the displacement of the user within the preset time period, can be calculated by obtaining the first SLAM data and the second SLAM data. Alternatively, the displacement of the user within the preset time period can be more accurately calculated by fusing the first SLAM data, the second SLAM data and the IMU data.

[0072] The following example takes the virtual scene rendering under mobile perspective as an example. Figure 6 As shown in the figure, the rendering process of the virtual scene under the mobile perspective is as follows: stepA: start; Step B: Use a depth camera to obtain a depth map of the real scene where the user is. stepC: Get the transformation matrix between the coordinate system corresponding to the depth camera and the coordinate system corresponding to the user's observation angle; stepD: If the user's displacement exceeds the preset displacement threshold (such as 5cm), the transformation matrix is ​​triggered to remap each pixel in the depth map to the coordinate system corresponding to the user's observation angle, and a remapped depth map adapted to the user's observation angle is generated; stepE: Based on the perspective projection matrix, calculate the depth data of the virtual object in the virtual scene; stepF: remaps the depth data of real objects in the real scene and the depth data of virtual objects in the virtual scene corresponding to the depth map through the GPU buffer cache; stepG: Based on the depth data, determine the occlusion relationship between the real object in the real scene and the virtual object in the virtual scene, and determine whether the depth of the virtual object is less than the depth of the real object; if so, execute stepH; if not, execute stepI; stepH: Render the pixel corresponding to the virtual object; stepI: End.

[0073] By comparing the depth of virtual objects in the virtual scene with the depth of real objects in the real scene, the front-to-back position relationship between the virtual objects in the virtual scene and the real objects in the real scene can be accurately judged, avoiding the problem of inaccurate occlusion relationship between virtual objects and real objects, thereby improving the realism of virtual scene rendering.

[0074] See also Figure 7 , Figure 7 It is a schematic block diagram of a scene rendering device provided in an embodiment of the present application. The scene rendering device can be configured in an AR device to execute the aforementioned scene rendering method.

[0075] like Figure 7 As shown, the scene rendering device 200 may include a processor 210 and a memory 220, wherein the processor 210 and the memory 220 are connected via a bus, such as an I2C (Inter-integrated Circuit) bus.

[0076] Specifically, the processor 210 may be a micro-controller unit (MCU), a central processing unit (CPU), or a digital signal processor (DSP).

[0077] Specifically, the memory 220 may be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB disk, or a mobile hard disk, etc. The memory 220 stores various computer programs for the processor 210 to execute.

[0078] The processor 210 is used to run a computer program stored in the memory, and implements the following steps when executing the computer program: Acquire first depth data corresponding to the relevant scene, where the first depth data is depth data in a coordinate system corresponding to the user's observation perspective; Acquire second depth data corresponding to the target scene; The target scene is rendered based on the first depth data and the second depth data, and the depth of the rendered object in the target scene is less than the depth of the object in the related scene at the corresponding position.

[0079] In some embodiments, when the processor 210 implements the acquisition of the first depth data corresponding to the relevant scene, it is used to implement: Acquire scene depth data corresponding to the relevant scene, and acquire a transformation matrix between a coordinate system corresponding to a device and a coordinate system corresponding to the user's observation angle, wherein the scene depth data is depth data in the coordinate system corresponding to the device; The first depth data is obtained based on the scene depth data and the transformation matrix.

[0080] In some embodiments, when the processor 210 implements the obtaining of the first depth data based on the scene depth data and the transformation matrix, it is configured to implement: Based on the transformation matrix and the bilinear interpolation method, the scene depth data is converted into a coordinate system corresponding to the user observation perspective to obtain the first depth data.

[0081] In some embodiments, before acquiring the first depth data corresponding to the relevant scene, the processor 210 is used to implement: Determine the user's displacement within a preset time period; When the processor 210 implements the acquisition of the first depth data corresponding to the relevant scene, it is used to implement: If the displacement is greater than a preset displacement threshold, the first depth data corresponding to the relevant scene is acquired.

[0082] In some embodiments, when implementing the determination of the displacement of the user within a preset time period, the processor 210 is used to implement: Acquire user posture information corresponding to the preset duration, wherein the user posture information includes at least one of IMU inertial measurement unit data and SLAM synchronous positioning and mapping data; The displacement is determined based on the user posture information.

[0083] In some embodiments, when the processor 210 implements the acquisition of the user posture information corresponding to the preset duration, it is used to implement: A camera device based on a preset refresh rate captures an image of the relevant scene, wherein the preset duration is determined by the preset refresh rate; Acquire the first SLAM data corresponding to the current frame image and the second SLAM data corresponding to the previous frame image.

[0084] In some embodiments, the first depth data and the second depth data are obtained by parallel processing based on different GPU graphics processor channels.

[0085] In some embodiments, the related scene is a real scene, the target scene is a virtual scene, the first depth data includes a first depth value corresponding to a real object in the real scene, and the second depth data includes a second depth value corresponding to a virtual object in the virtual scene. When the processor 210 implements rendering the target scene based on the first depth data and the second depth data, it is used to implement: The second depth value of the virtual object is compared with the first depth value of the real object at a corresponding position, and if the second depth value is less than the first depth value, the pixel point corresponding to the virtual object is rendered.

[0086] The scene rendering device 200 can execute the scene rendering method provided in the embodiment of the present application, and therefore can achieve the beneficial effects that can be achieved by the scene rendering method provided in the embodiment of the present application. Please refer to the previous embodiment for details and will not be repeated here.

[0087] An embodiment of the present application also provides an AR device, the AR device includes a scene rendering device, the scene rendering device can be Figure 7 The scene rendering device 200 shown in FIG. Therefore, the AR device can achieve the beneficial effects that can be achieved by the scene rendering method provided in the embodiment of the present application, as detailed in the previous embodiment, which will not be repeated here.

[0088] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the scene rendering method as described above are implemented.

[0089] The computer-readable storage medium may be an internal storage unit of the scene rendering device or AR device described in the aforementioned embodiment, such as a hard disk or memory of the scene rendering device or AR device. The computer-readable storage medium may also be an external storage device of the scene rendering device or AR device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital card (Secure Digital Card, SD Card), a flash card (Flash Card), etc., equipped on the scene rendering device or AR device.

[0090] Since the computer program stored in the storage medium can execute any scene rendering method provided in the embodiments of the present application, the beneficial effects that can be achieved by any scene rendering method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0091] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.

[0092] The above description is only a specific implementation mode of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and these modifications or substitutions should be included in the protection scope of the present application.

Claims

1. A scene rendering method, characterized in that: The scene rendering method comprises: Determine the user's displacement within a preset time period; If the displacement is greater than a preset displacement threshold, first depth data corresponding to the relevant scene is acquired, where the first depth data is depth data in a coordinate system corresponding to the user's observation angle; Acquire second depth data corresponding to the target scene; The target scene is rendered based on the first depth data and the second depth data, and the depth of the rendered object in the target scene is less than the depth of the object in the related scene at the corresponding position.

2. The scene rendering method according to claim 1, characterized in that: The obtaining of first depth data corresponding to the relevant scene includes: Acquire scene depth data corresponding to the relevant scene, and acquire a transformation matrix between a coordinate system corresponding to a device and a coordinate system corresponding to the user's observation angle, wherein the scene depth data is depth data in the coordinate system corresponding to the device; The first depth data is obtained based on the scene depth data and the transformation matrix.

3. The scene rendering method according to claim 2, characterized in that: The obtaining the first depth data based on the scene depth data and the transformation matrix includes: Based on the transformation matrix and the bilinear interpolation method, the scene depth data is converted into a coordinate system corresponding to the user observation perspective to obtain the first depth data.

4. The scene rendering method according to claim 1, characterized in that: The determining the displacement of the user within a preset time period includes: Acquire user posture information corresponding to the preset duration, wherein the user posture information includes at least one of IMU inertial measurement unit data and SLAM synchronous positioning and mapping data; The displacement is determined based on the user posture information.

5. The scene rendering method according to claim 4, characterized in that: The obtaining of user posture information corresponding to the preset duration includes: A camera device based on a preset refresh rate captures an image of the relevant scene, wherein the preset duration is determined by the preset refresh rate; Acquire the first SLAM data corresponding to the current frame image and the second SLAM data corresponding to the previous frame image.

6. The scene rendering method according to claim 1, characterized in that: The first depth data and the second depth data are obtained by parallel processing based on different GPU graphics processor channels.

7. The scene rendering method according to claim 1, characterized in that: The related scene is a real scene, the target scene is a virtual scene, the first depth data includes a first depth value corresponding to a real object in the real scene, the second depth data includes a second depth value corresponding to a virtual object in the virtual scene, and rendering the target scene based on the first depth data and the second depth data includes: The second depth value of the virtual object is compared with the first depth value of the real object at a corresponding position, and if the second depth value is less than the first depth value, the pixel point corresponding to the virtual object is rendered.

8. A scene rendering device, characterized in that: The scene rendering device includes a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program and implement the steps of the scene rendering method according to any one of claims 1 to 7 when executing the computer program.

9. An AR device, characterized in that: The AR device includes the scene rendering apparatus as claimed in claim 8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the scene rendering method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Virtual reality presenting method based on depth map optimization

    CN107067456A

  • Photographing method, terminal device and cloud server

    CN114125310A

  • Generation method and system for augmented reality scene presenting real occlusion relationship of object

    CN114863066A

  • Depth error detection method and device, computer equipment and storage medium

    CN114993623A

  • Scene rendering method and apparatus, device, computer readable storage medium, and product

    WO2024198855A1

Cited By

  • Depth information generation method, head-mounted display device, storage medium and product

    CN120614444A