Method and apparatus for handling occlusion in split rendering

By determining the depth information of real-world objects and rendering occluded objects using occlusion materials, the accuracy of occlusion relationships in augmented reality systems is solved, improving the realism and immersion of the AR experience.

CN115244584BActive Publication Date: 2026-04-03QUALCOMM INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-05
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In augmented reality systems, when real-world content occludes augmented content, it is difficult to accurately capture and process the occlusion relationship, resulting in an unrealistic and unimmersive AR experience.

Method used

By determining the depth information of real-world objects, occlusion objects are rendered on client devices using occlusion materials, and the occlusion surfaces are represented by a mesh, gradually improving the accuracy of the occlusion content over time.

Benefits of technology

It improves the accuracy of occluded content in the split rendering system, enhancing the realism and immersion of the AR experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115244584B_ABST
    Figure CN115244584B_ABST
Patent Text Reader

Abstract

This disclosure relates to methods and apparatus for graphics processing. Aspects of this disclosure can identify a first content group and a second content group in a scene. Furthermore, aspects of this disclosure can determine whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group. Additionally, this disclosure can represent the first content group and the second content group based on the determination that at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group. In some aspects, the first content group may include at least some real content, and the second content group may include at least some augmented content. This disclosure can also use occlusion materials to render at least a portion of the surface of the first content group.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims the benefits of U.S. Provisional Application No. 63 / 005,164, filed April 3, 2020, entitled “METHODS AND APPARATUS FOR HANDLING OCCLUSIONS IN SPLIT RENDERING”, and U.S. Patent Application No. 17 / 061,179, filed October 1, 2020, entitled “METHODS AND APPARATUS FOR HANDLING OCCLUSIONS IN SPLIT RENDERING”, which have been assigned to the assignee of this application and are hereby expressly incorporated herein by reference. Technical Field

[0003] In summary, this disclosure relates to processing systems, and more specifically, to one or more techniques for graphics processing. Background Technology

[0004] Computing devices typically utilize graphics processing units (GPUs) to accelerate the rendering of graphics data for display. Such devices can include, for example, computer workstations, mobile phones such as so-called smartphones, embedded systems, personal computers, tablet computers, and video game consoles. A GPU executes a graphics processing pipeline, which includes one or more processing stages that work together to execute graphics processing commands and output frames. A central processing unit (CPU) controls the operation of the GPU by issuing one or more graphics processing commands to it. Modern CPUs are typically capable of executing multiple applications concurrently, each of which may require the GPU to utilize its resources during execution. Devices that provide content for visual presentation on a display typically include GPUs.

[0005] Typically, a device's GPU is configured to perform processes within a graphics processing pipeline. However, with the advent of wireless communication and smaller handheld devices, the need for improved graphics processing has been continuously increasing. Summary of the Invention

[0006] The following provides a brief overview of one or more aspects to offer a basic understanding of such aspects. This overview is not a comprehensive summary of all anticipated aspects, nor is it intended to identify key elements of all aspects or to depict the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed descriptions that follow.

[0007] In one aspect of this disclosure, methods, computer-readable media (e.g., non-transitory-readable computer-readable media), and apparatus are provided. The apparatus may be a server, a client device, a central processing unit (CPU), a graphics processing unit (GPU), or any device capable of performing graphics processing. The apparatus may identify a first content group and a second content group in a scene. The apparatus may also determine whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group. Furthermore, the apparatus may represent the first and second content groups based on the determination that at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group. In some aspects, when at least a portion of the first content group occludes at least a portion of the second content group, the apparatus may use an occlusion material to render at least a portion of one or more surfaces of the first content group. The apparatus may also generate at least one first bulletin board according to at least one first plane equation and at least one second bulletin board according to at least one second plane equation. The apparatus may also generate at least one first mesh and a first shading texture, and at least one second mesh and a second shading texture. Additionally, the apparatus may transmit information associated with the first content group and information associated with the second content group to a client device.

[0008] In another aspect, a graphics processing method includes: identifying a first content group and a second content group in a scene; determining whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group; and representing the first content group and the second content group based on the determination that at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group.

[0009] In a further example, an apparatus for graphics processing is provided, the apparatus comprising: a transceiver; a memory configured to store instructions; and one or more processors communicatively coupled to the transceiver and the memory. This aspect may include one or more processors configured to: identify a first content group and a second content group in a scene; determine whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group; and represent the first content group and the second content group based on the determination regarding whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group.

[0010] In another aspect, an apparatus for graphics processing is provided, the apparatus comprising: a unit for identifying a first content group and a second content group in a scene; a unit for determining whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group; and a unit for representing the first content group and the second content group based on the determination that at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group.

[0011] In another aspect, a non-transitory computer-readable medium is provided, comprising one or more processors executing code for graphics processing, the code causing the processor, when executed by the processor, to: identify a first content group and a second content group in a scene; determine whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group; and represent the first content group and the second content group based on the determination regarding whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group.

[0012] Details of one or more examples of this disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of this disclosure will be apparent from the specification, drawings, and claims. Attached Figure Description

[0013] Figure 1 This is a block diagram illustrating an example content generation system based on one or more techniques according to this disclosure.

[0014] Figure 2 Example images or scenes illustrating one or more techniques according to this disclosure are shown.

[0015] Figure 3 An example of a deep-aligned bulletin board based on one or more technologies according to this disclosure is shown.

[0016] Figure 4A Example coloring atlases of one or more techniques according to this disclosure are shown.

[0017] Figure 4B An example color atlas organization based on one or more techniques according to this disclosure is shown.

[0018] Figure 5 Example images or scenes illustrating one or more techniques according to this disclosure are shown.

[0019] Figure 6 An example flowchart of an example method of one or more techniques according to this disclosure is shown. Detailed Implementation

[0020] Accurately capturing occlusion can be challenging in augmented reality (AR) systems. This can be especially true when real-world content occludes augmented content. Furthermore, this can be particularly true for AR systems with latency issues. Accurate occlusion of both real-world and augmented content can help users achieve a more realistic and immersive AR experience. Some aspects of this disclosure address the aforementioned occlusion of real-world or augmented objects by determining depth information about these objects. For example, a real-world surface can be considered an occluder for augmented purposes and can be streamed to a client device and rendered on the client using occlusion material. Objects with occlusion material or empty display values ​​can be considered rendered objects but may not be visible on the client's display. Therefore, aspects of this disclosure can determine or render values ​​for real-world occluded objects and provide occlusion material for these objects. In some aspects, this occlusion material can represent the color of a physically flat surface (e.g., as in a bulletin board). In other aspects, occlusion material can be assigned to surfaces described by a grid (e.g., a grid corresponding to the real-world object considered as an occluder). By doing so, the mesh and / or information about occluded objects can become increasingly accurate over time. Therefore, aspects of this disclosure can improve the accuracy of occluded content in split rendering systems.

[0021] The various aspects of the systems, apparatus, computer program products, and methods are described more fully below with reference to the accompanying drawings. However, this disclosure may be embodied in many different forms and should not be construed as limited to any particular structure or function presented herein. Rather, these aspects are provided so that this disclosure will be comprehensive and complete, and will fully convey the scope of this disclosure to those skilled in the art. Based on the teachings herein, those skilled in the art will recognize that the scope of this disclosure is intended to cover any aspect of the systems, apparatus, computer program products, and methods disclosed herein, whether that aspect is implemented independently of or in combination with other aspects of this disclosure. For example, an apparatus may be implemented or a method may be practiced using any number of the aspects set forth herein. Furthermore, the scope of this disclosure is intended to cover such apparatuses or methods practiced using structures, functions, or structures and functions other than or different from the aspects of this disclosure set forth herein. Any aspect disclosed herein may be embodied by one or more elements of the claims.

[0022] While various aspects are described herein, numerous variations and substitutions of these aspects fall within the scope of this disclosure. Although some potential benefits and advantages of the aspects of this disclosure are mentioned, the scope of this disclosure is not intended to be limited to a particular benefit, use, or objective. Rather, the aspects of this disclosure are intended to be broadly applicable to different wireless technologies, system configurations, networks, and transport protocols, some of which are illustrated by way of example in the accompanying drawings and the description below. The detailed description and drawings are illustrative only and not limiting of this disclosure, and the scope of this disclosure is defined by the appended claims and their equivalents.

[0023] Several aspects are given with reference to various apparatuses and methods. These apparatuses and methods are described in the specific embodiments below and illustrated in the accompanying drawings by various blocks, components, circuits, processes, algorithms, etc. (collectively, "elements"). These elements can be implemented using electronic hardware, computer software, or any combination thereof. Whether such elements are implemented in hardware or software depends on the specific application and the design constraints imposed on the overall system.

[0024] For example, an element, or any part of an element, or any combination of elements, can be implemented as a "processing system," which includes one or more processors (which may also be called processing units). Examples of processors include: microprocessors, microcontrollers, graphics processing units (GPUs), general-purpose GPUs (GPGPUs), central processing units (CPUs), application processors, digital signal processors (DSPs), reduced instruction set computing (RISC) processors, system-on-a-chip (SoCs), baseband processors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various functions described herein. One or more processors in a processing system can execute software. Whether referred to as software, firmware, middleware, microcode, hardware description languages, or others, software can be broadly interpreted as instructions, instruction sets, code, code segments, program code, programs, subroutines, software components, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc. The term application can refer to software. As described herein, one or more technologies can refer to an application (i.e., software) configured to perform one or more functions. In such an example, the application may be stored on memory (e.g., on-chip memory of a processor, system memory, or any other memory). The hardware described herein (such as a processor) may be configured to execute the application. For example, an application may be described as including code that, when executed by the hardware, causes the hardware to perform one or more technologies described herein. As an example, the hardware may access code from memory and execute the code accessed from memory to perform one or more technologies described herein. In some examples, components are identified in this disclosure. In such examples, a component may be hardware, software, or a combination thereof. A component may be a single component or a subcomponent of a single component.

[0025] Accordingly, in one or more examples described herein, the described functionality may be implemented in hardware, software, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on or encoded as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media. Storage media may be any available medium accessible by a computer. By way of example, and not limitation, such computer-readable media may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), optical disk storage, magnetic disk storage, other magnetic storage devices, combinations of computer-readable media of the types described above, or any other medium that may be used to store computer-executable code accessible by a computer in the form of instructions or data structures.

[0026] In summary, this disclosure describes techniques for having a graphics processing pipeline in a single device or multiple devices, thereby improving the rendering of graphics content and / or reducing the load on processing units (i.e., any processing unit, such as a GPU, configured to perform one or more of the techniques described herein). For example, this disclosure describes techniques for graphics processing in any device that utilizes graphics processing. Other example benefits are described throughout this disclosure.

[0027] As used herein, instances of the term "content" can refer to "graphic content," "a product of 3D graphic design," its representation (i.e., "image"), and vice versa. This is true regardless of whether the term is used as an adjective, noun, or other part of speech. In some examples, as used herein, the term "graphic content" can refer to content produced by one or more processes in a graphics processing pipeline. In some examples, as used herein, the term "graphic content" can refer to content produced by a processing unit configured to perform graphics processing. In some examples, as used herein, the term "graphic content" can refer to content produced by a graphics processing unit.

[0028] In some examples, as used herein, the term "display content" can refer to content generated by a processing unit configured to perform display processing. Graphical content can be processed to become display content. For example, a graphics processing unit can output graphical content, such as frames, to a buffer (which may be referred to as a frame buffer). The display processing unit can read graphical content (such as one or more frames) from the buffer and perform one or more display processing techniques on it to generate display content. For example, the display processing unit can be configured to perform compositing on one or more rendering layers to generate frames. As another example, the display processing unit can be configured to composite, blend, or otherwise combine two or more layers into a single frame. The display processing unit can be configured to perform scaling (e.g., zooming in or out) on frames. In some examples, a frame can refer to a layer. In other examples, a frame can refer to two or more layers that have already been blended together to form a frame; that is, a frame comprises two or more layers, and frames comprising two or more layers can subsequently be blended.

[0029] Figure 1This is a block diagram illustrating an example system 100 configured to implement one or more technologies of this disclosure. System 100 includes device 104. Device 104 may include one or more components or circuitry for performing the various functions described herein. In some examples, one or more components of device 104 may be components of a System-on-a-Chip (SOC). Device 104 may include one or more components configured to perform one or more technologies of this disclosure. In the illustrated example, device 104 may include a processing unit 120, a content encoder / decoder 122, and system memory 124. In some aspects, device 104 may include multiple optional components, such as a communication interface 126, a transceiver 132, a receiver 128, a transmitter 130, a display processor 127, and one or more displays 131. Reference to display 131 may refer to one or more displays 131. For example, display 131 may include a single display or multiple displays. Display 131 may include a first display and a second display. The first display may be a left-eye display, and the second display may be a right-eye display. In some examples, the first and second displays may receive different frames for presentation thereon. In other examples, the first and second displays may receive the same frames used for rendering on them. In further examples, the results of graphics processing may not be displayed on the devices; for example, the first and second displays may not receive any frames used for rendering on them. Instead, the frames or graphics processing results may be transmitted to another device. In some respects, this may be referred to as split rendering.

[0030] Processing unit 120 may include internal memory 121. Processing unit 120 may be configured to perform graphics processing, such as in a graphics processing pipeline 107. Content encoder / decoder 122 may include internal memory 123. In some examples, device 104 may include a display processor (such as display processor 127) to perform one or more display processing techniques on one or more frames generated by processing unit 120 prior to rendering by one or more displays 131. Display processor 127 may be configured to perform display processing. For example, display processor 127 may be configured to perform one or more display processing techniques on one or more frames generated by processing unit 120. One or more displays 131 may be configured to display or otherwise present frames processed by display processor 127. In some examples, one or more displays 131 may include one or more of the following: liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, projection display device, augmented reality display device, virtual reality display device, head-mounted display, or any other type of display device.

[0031] Memory (such as system memory 124) external to processing unit 120 and content encoder / decoder 122 may be accessible to processing unit 120 and content encoder / decoder 122. For example, processing unit 120 and content encoder / decoder 122 may be configured to read from and / or write to external memory such as system memory 124. Processing unit 120 and content encoder / decoder 122 may be communicatively coupled to system memory 124 via a bus. In some examples, processing unit 120 and content encoder / decoder 122 may be communicatively coupled to each other via a bus or a different connection.

[0032] Content encoder / decoder 122 can be configured to receive graphic content from any source, such as system memory 124 and / or communication interface 126. System memory 124 can be configured to store the received encoded or decoded graphic content. Content encoder / decoder 122 can be configured to receive encoded or decoded graphic content from system memory 124 and / or communication interface 126, for example, in the form of encoded pixel data. Content encoder / decoder 122 can be configured to encode or decode any graphic content.

[0033] Internal memory 121 or system memory 124 may include one or more volatile or non-volatile memories or storage devices. In some examples, internal memory 121 or system memory 124 may include RAM, SRAM, DRAM, erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic data media or optical storage media, or any other type of memory.

[0034] According to some examples, internal memory 121 or system memory 124 may be a non-transitory storage medium. The term "non-transitory" may indicate that the storage medium is not embodied in a carrier wave or a propagating signal. However, the term "non-transitory" should not be construed as meaning that internal memory 121 or system memory 124 is non-removable or that its contents are static. As one example, system memory 124 may be removed from device 104 and moved to another device. As another example, system memory 124 may be non-removable from device 104.

[0035] Processing unit 120 may be a central processing unit (CPU), a graphics processing unit (GPU), a general-purpose GPU (GPGPU), or any other processing unit that can be configured to perform graphics processing. In some examples, processing unit 120 may be integrated into the motherboard of device 104. In some examples, processing unit 120 may reside on a graphics card mounted in a port on the motherboard of device 104, or may otherwise be incorporated into a peripheral device configured to interoperate with device 104. Processing unit 120 may include one or more processors, such as one or more microprocessors, GPUs, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), arithmetic logic units (ALUs), digital signal processors (DSPs), discrete logic, software, hardware, firmware, other equivalent integrated or discrete logic circuits, or any combination thereof. If the technology is partially implemented in software, processing unit 120 may store instructions for software in a suitable non-transitory computer-readable storage medium (e.g., internal memory 121), and may execute the instructions in hardware using one or more processors to perform the technology of this disclosure. Any of the foregoing, including hardware, software, and combinations of hardware and software, can be considered as one or more processors.

[0036] The content encoder / decoder 122 can be any processing unit configured to perform content decoding. In some examples, the content encoder / decoder 122 can be integrated into the motherboard of device 104. The content encoder / decoder 122 can include one or more processors, such as one or more microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), arithmetic logic units (ALUs), digital signal processors (DSPs), video processors, discrete logic, software, hardware, firmware, other equivalent integrated or discrete logic circuits, or any combination thereof. If the technology is partially implemented in software, the content encoder / decoder 122 can store instructions for software in a suitable non-transitory computer-readable storage medium (e.g., internal memory 121), and can execute the instructions in hardware using one or more processors to perform the technology of this disclosure. Any of the foregoing, including hardware, software, combinations of hardware and software, etc., can be considered as one or more processors.

[0037] In some aspects, system 100 may include an optional communication interface 126. Communication interface 126 may include a receiver 128 and a transmitter 130. Receiver 128 may be configured to perform any of the receiving functions described herein with respect to device 104. Additionally, receiver 128 may be configured to receive information from another device (e.g., eye or head position information, rendering commands, or location information). Transmitter 130 may be configured to perform any of the transmitting functions described herein with respect to device 104. For example, transmitter 130 may be configured to send information to another device that may include a request for content. Receiver 128 and transmitter 130 may be combined into transceiver 132. In such an example, transceiver 132 may be configured to perform any of the receiving and / or transmitting functions described herein with respect to device 104.

[0038] Refer again Figure 1 In some aspects, the graphics processing pipeline 107 may include a determining component 198 configured to identify a first content group and a second content group in a scene. The determining component 198 may also be configured to determine whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group. The determining component 198 may also be configured to represent the first content group and the second content group based on the determination that at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group. The determining component 198 may also be configured to render at least a portion of one or more surfaces of the first content group using an occlusion material when at least a portion of the first content group occludes at least a portion of the second content group. The determining component 198 may also be configured to generate at least one first bulletin board according to at least one first plane equation and at least one second bulletin board according to at least one second plane equation. The determining component 198 may also be configured to transmit information associated with the first content group and information associated with the second content group to a client device.

[0039] As described herein, a device such as device 104 can refer to any device, apparatus, or system configured to perform one or more of the techniques described herein. For example, a device can be a server, base station, user equipment, client device, station, access point, computer (e.g., personal computer, desktop computer, laptop computer, tablet computer, computer workstation, or mainframe computer), end product, apparatus, telephone, smartphone, server, video game platform or console, handheld device (e.g., portable video game device or personal digital assistant (PDA)), wearable computing device (e.g., smartwatch, augmented reality device, or virtual reality device), non-wearable device, display or display device, television, set-top box, intermediate network device, digital media player, video streaming device, content streaming device, in-vehicle computer, any mobile device, any device configured to generate graphical content, or any device configured to perform one or more of the techniques described herein. The processes described herein may be described as being performed by a specific component (e.g., GPU), but in further embodiments, they may be performed using other components (e.g., CPU) consistent with the disclosed embodiments.

[0040] A GPU can process various types of data or data packets within its pipeline. For example, in some aspects, a GPU can process two types of data or data packets, such as context register packets and draw call data. Context register packets can be a collection of global state information that manages how the graphics context will be processed, such as information about global registers, shaders, or constant data. For example, a context register packet may include information about color formats. In some aspects of context register packets, there may be bits indicating which workload belongs to the context register. Furthermore, there may be multiple functions or programs running simultaneously and / or in parallel. For example, a function or program may describe a specific operation, such as a color mode or color format. Therefore, context registers can define multiple states of the GPU.

[0041] Context states can be used to determine how individual processing units (e.g., vertex extractors (VFDs), vertex shaders (VSs), shader processors, or geometry processors) operate, and / or in which mode a processing unit operates. For this purpose, the GPU can use context registers and programming data. In some aspects, the GPU can generate workloads in the pipeline based on the context register definitions of modes or states, such as vertex or pixel workloads. Certain processing units (e.g., VFDs) can use these states to determine certain functions, such as how to assemble vertices. Because these modes or states can change, the GPU may need to modify the corresponding context. Furthermore, the workload corresponding to a mode or state can follow a constantly changing mode or state.

[0042] GPUs can render images in a variety of different ways. In some cases, GPUs can render images using either rendering or tiled rendering. In tiled rendering GPUs, an image can be divided or segmented into different sections or tiles. After the image is divided, each section or tile can be rendered individually. Tiled rendering GPUs can divide computer graphics images into a grid format, allowing each part of the grid (i.e., a tile) to be rendered individually. In some aspects, during the binning pass, an image can be divided into different bins or tiles. Furthermore, in the binning pass, different primitives can be colored in certain bins, for example, using draw calls. In some aspects, during the binning pass, a visibility stream can be constructed, where visible primitives or draw calls can be identified.

[0043] In some aspects of rendering, multiple processing stages or paths can exist. For example, rendering can be performed in two paths (e.g., a visibility path and a rendering path). During the visibility path, the GPU can input a rendering workload, record the positions of primitives or triangles, and then determine which primitives or triangles fall into which part of the frame. In some aspects of the visibility path, the GPU can also identify or mark the visibility of each primitive or triangle in the visibility stream. During the rendering path, the GPU can input a visibility stream and process one part of the frame at a time. In some aspects, the visibility stream can be analyzed to determine which primitives are visible or invisible. Thus, visible primitives can be processed. By doing so, the GPU can reduce the unnecessary workload of processing or rendering invisible primitives.

[0044] In some respects, rendering can be performed in multiple locations and / or on multiple devices, for example, to divide the rendering workload among different devices. For instance, rendering can be split between server and client devices, which can be referred to as "split rendering." In some cases, split rendering can be a method for bringing content to a user device or head-mounted display (HMD), where a portion of the graphics processing can be performed outside of that device or HMD (e.g., at the server).

[0045] Split rendering can be performed for various types of applications, such as virtual reality (VR) applications, augmented reality (AR) applications, cloud gaming, and / or extended reality (XR) applications. In VR applications, the content displayed on the user's device can correspond to artificial or animated content, such as content rendered on a server or user's device. In AR or XR content, a portion of the content displayed on the user's device can correspond to real-world content (e.g., real-world objects), and a portion of the content can be artificial or animated content. Furthermore, artificial or animated content and real-world content can be displayed in optical or video perspective devices, allowing users to view real-world objects as well as artificial or animated content simultaneously. In some respects, artificial or animated content can be referred to as augmented content, or vice versa.

[0046] In AR applications, objects may occlude other objects from a vantage point on the user's device. Different types of occlusion can exist within AR applications. For example, augmented content may occlude real-world content; for instance, a rendered object may partially occlude a real-world object. Furthermore, real-world content may occlude augmented content; for example, a real-world object may partially occlude a rendered object. This overlap between real-world and augmented content (which produces the aforementioned occlusion) is one reason why augmented and real-world content can blend so seamlessly within AR. This can also lead to difficulties in resolving occlusion between augmented and real-world content, causing the edges of augmented and real-world content to incorrectly overlap.

[0047] In some aspects, augmented content or enhancements can be rendered on top of real-world or perspective content. Therefore, an enhancement can occlude any object behind it from a vantage point on the user's device. For example, pixels using red (R), green (G), and blue (B) (RGB) values ​​can be rendered to occlude a real-world object. Thus, the enhancement can occlude a real-world object behind it. In video perspective systems, the same effect can be achieved by compositing an enhancement layer onto the foreground. Therefore, an enhancement can occlude real-world content, and vice versa.

[0048] As noted above, accurately capturing occlusion can be challenging when using AR systems. This is especially true for AR systems with latency issues. In some respects, accurately capturing real-world objects that are occluding augmented content can be particularly difficult. Accurate occlusion of both real-world and augmented content can help users achieve a more realistic and immersive AR experience.

[0049] Figure 2 Example images or scenes 200 based on one or more technologies according to this disclosure are shown. Scene 200 includes augmented content 210 and real-world content 220 (which includes edges 222). More specifically, Figure 2 This displays a real-world object 220 (e.g., a door) that is occluding augmented content 210 (e.g., a person). Figure 2 As shown, the augmented content 210 slightly overlaps with the edge 222 of the real-world object 220, even when the real-world object 220 is intended to completely occlude the augmented content 210. For example, when a door is intended to completely occlude a person, a portion of the person slightly overlaps with the edge of the door.

[0050] As pointed out above, Figure 2 This demonstrates how accurately real-world content occlusion can enhance content in AR systems, which may be difficult to achieve precisely. For example... Figure 2 As shown, AR systems may struggle to accurately reflect when real-world objects occlude augmented content, or vice versa. In fact, when two objects are real-world content and augmented content, some AR systems may have difficulty processing the edges of these objects quickly and accurately. Therefore, there is a need to accurately depict when real-world content occludes augmented content.

[0051] To accurately simulate real-world content that enhances occlusion, or vice versa, the geometric information of real-world occluded objects can be determined. In some cases, this can be achieved through a combination of computer vision techniques, 3D reconstruction, and / or meshing. In one scenario, meshing can be performed in real-time on the client device or HMD, and the reconstructed geometry can be transmitted to a server or cloud server. In another scenario, the client device or HMD can capture images sequentially and send them to a server, which then performs 3D reconstruction and / or meshing algorithms to extract 3D mesh information about objects visible in the real-world scene. For example, in some aspects, a real-world 3D mesh can be determined. Thus, the geometry of the occluded object can be available on the client device and sent to the server, or vice versa. As more observations of the same one or more real-world objects become available, this geometric representation can be adjusted and become more accurate over time. By doing so, the mesh and / or information about the occluded object can become increasingly accurate over time.

[0052] In some aspects, to obtain an accurate description of augmented and real-world content, the augmented content may need to precisely rest at the edge of the real-world content, for example, if the real-world content is occluding the augmented content. Therefore, the real-world content occluding the augmented content may need to correspond to the unrendered portion of the augmented content. For example... Figure 2 As shown, the superimposed enhanced content 210 may need to remain at the edge that obscures the real content 220.

[0053] In some cases, aspects of this disclosure can determine the precise edges of occluded real-world content. Furthermore, this edge of the real-world content can be determined at rendering time or is known. By doing so, augmented content can precisely rest at the edges of real-world objects, which can improve the accuracy of the AR experience.

[0054] In some aspects, the aforementioned edge determination can be resolved during the visibility path, for example, within the rendering engine. In a split AR system, the rendering engine resides on a server. The visibility path can consider the pose of all relevant meshes (e.g., both the real-world and augmented meshes from the current camera pose) and / or determine which triangles are visible from the current camera pose, thus taking into account the depth of the real-world geometry and the desired depth / placement of the augmented content. In various aspects of this disclosure, visible triangles on the mesh of a real-world object can be rendered as occlusion material or represented via occlusion messages (e.g., having (0,0,0) values), and visible triangles on the mesh of a computer graphics object can be rendered as non-zero pixels (i.e., in terms of color).

[0055] Within the context of this disclosure, the term "camera" can be used interchangeably to refer to two distinct entities: a computer graphics camera (e.g., a camera object within a graphics rendering engine) or a real-world observer (e.g., the eyes of a user wearing an AR device). Any variation of the term "camera" suitable for its current use can be determined based on what is "observed" via the camera; for example, if real-world content is being discussed, "camera" can refer to a real-world observer. It should also be noted that a correspondence can exist between the two types of cameras in an XR system, as the pose of the real-world observer can be used to modify the pose of the graphics rendering engine's camera. Furthermore, the computations discussed in this disclosure can combine real-world objects and animated objects within the realm of the computer graphics world, such that both can be represented within the graphics rendering engine. When this occurs, the two concepts of the term "camera" can be merged to correspond to a computer graphics camera.

[0056] Splitting XR, AR, and cloud gaming systems can introduce latency when delivering rendered content to the client's display. In some aspects, this latency may even be higher when rendering on a server compared to client-side rendering, but it can also enable more complex XR or AR applications. Furthermore, there can be a non-negligible latency between calculating camera pose and the time it takes for content to appear on the client's display. For example, there will almost always be a certain amount of latency in split-type XR, AR, and cloud gaming systems.

[0057] This disclosure can consider two modalities for splitting XR, AR, and cloud gaming architectures: pixel-stream (PS) architecture and vector-stream (VS) architecture. In a PS system, a variant of asynchronous time warp can be used on the client before display to compensate for content latency. In some aspects, in splitting XR systems, PS frames streamed to the client can be time-warped based on view transformations between the view used to render content and the latest view that is used to display the same content in time. For example, aspects of this disclosure can perform time-compensated reprojection, allowing the occurrence of delayed frames to be determined based on the estimated distance of the camera position and the direction of camera travel during the compensated time period.

[0058] The methods for occlusion handling in split / remote rendering AR systems described herein may include techniques specifically designed for systems where the scene rendered in screen space may be part of the interface between the server and the client. These systems may be referred to as pixel-stream systems. Additionally, this disclosure may include techniques specifically designed for systems where final screen-space rendering may occur on the client, while the server may provide relevant scene geometry and pose-dependent shading in the object space of a potentially visible texture. These systems may be referred to as vector-stream systems. Descriptions of both techniques may be distributed throughout this disclosure. One way to determine which system is being described is to consider the interface in the description. For example, the interface (signaling medium) for a pixel-stream system may be a bulletin board, while the interface (signaling medium) for a vector-stream system may be a mesh or geometry and a shading atlas for shading surface textures.

[0059] In some cases, asynchronous time warp (ATW) on a PS client can distort the shape of rendered content or CGI. For example, ATW might inherently be a planar view transformation, while most content in a CGI scene may not be planar in nature, and even fewer content surfaces may follow coplanar relationships where the plane in question is strictly aligned with the camera. In some aspects, if a portion of the enhancement is rendered on the server and then sent to the client for final rendering, the rendered enhancement may undergo time warp. In these aspects, it can be difficult to maintain accurate occlusion boundaries between any real-world objects and the enhancement after time warp has been applied only to the enhancement.

[0060] In a pixel-stream architecture, the server or game engine determines the pixel information or eye buffer information for the scene and sends this information to the client device. The client device can then render the scene based on this pixel information. In some cases, this pixel information or eye buffer information may correspond to a specific pose of the XR camera.

[0061] In some aspects, and in split AR systems, the reprojection operation of augmented content by ATW may also move the edges of the augmented content relative to occluded real-world objects. For example, in the above... Figure 2 In the image, the half-open door 220 appears to be partially blocked by a person 210. Furthermore, there may be a gap between the partially open door 220 and the person 210 that is being blocked.

[0062] In a vector streaming system, a server can stream the appearance of augmented or CGI surfaces, along with the geometry of these augmented objects. Later in the device processing, for example, before transmitting information to the display, the client can rasterize the geometry of the visible augmented objects and / or texture map the corresponding textures from the streaming atlas to the rasterized polygons. This variant of pre-display processing inherently lacks long pose-to-render latency and does not distort the shape of the augmented objects. Combined with system modifications that are the subject of this disclosure, this variant of pre-display processing is also able to maintain accurate boundaries between foreground and background content, and thus maintain an accurate depiction of the mixed reality scene.

[0063] In some aspects, shading atlas information transmitted within a vector stream system can be sent as a time-evolved texture and encoded as video. In other aspects, geometric information can represent a mesh or partial mesh of a CGI object and can be transmitted raw or encoded using mesh compression methods. A client device can receive both shading atlas information and geometric information corresponding to the headset pose T0 for rendering server frames. It can decode this information and use it to render a final eye-buffer representation of the scene using the most recent pose T1. Because pose T1 can be much more up-to-date than pose T0, and because the final rendering step at pose T1 is a correct geometric raster rather than an approximation, the described system can have fewer latency constraints compared to pixel-stream latency.

[0064] The client can use the most recent pose T1 to rasterize the shading atlas information and / or geometric information, that is, to convert the information into pixels that can be displayed on the device. Furthermore, in some aspects, the server can receive a pose streamed from the client device or HMD and then perform visibility and shading calculations based on the received pose T0. Visibility calculations affect which meshes or portions of meshes can be sent to the client, and shading calculations may include a shading atlas corresponding to the current frame, which is also sent to the client. Aspects of this disclosure may also assume that the client device has information about these meshes or geometry (e.g., as a result of an offline scene loading step), such that the geometry does not need to be streamed to the client in real time. For example, the server may only calculate and stream shading on object surfaces, which may change based on camera vantage points or other dynamic elements of the scene, and only this shading information can be streamed to the client in real time.

[0065] In some respects, both pixel-stream and vector-stream architectures used in AR applications can signal information to the client about real-world occlusion and how it affects background enhancement. This disclosure can signal this information and simplify client operations while maintaining occlusion edge fidelity throughout latency compensation.

[0066] Some aspects of this disclosure can resolve the objects in a pixel-streaming system by determining depth information about the aforementioned occluded real-world objects. For example, using only the occlusion material as the surface material type, a real-world surface can be treated as another mixed reality asset and streamed to the client device along with some other properties of the real-world asset so that the client can incorporate it as part of the pre-display processing, with the aim of keeping the occlusion boundary between the real-world and augmented objects close to their ideal appearance. An object with occlusion material can be treated as another rendered object by the rendering client, but due to its surface material type, it may not be visible at the client's display; instead, its silhouette rendered from the pre-display headset pose T1 is used to occlude other augmented objects / assets that can be rendered in the background. Therefore, aspects of this disclosure can determine or render values ​​for real-world occluded objects and attach occlusion materials to these objects. For example, the presence of real-world occluded objects can be signaled to the vector-streaming client by including a separate signaling message type for the occlusion material patch and listing a unique patch identifier for the patch to be rendered using the occlusion material.

[0067] By incorporating real-world content represented using occlusion materials, content compositing devices can precisely position augmented content at the edges of occluded real-world objects, regardless of the latency introduced between pose T0 and pose T1 in a split rendering system. In effect, to prevent overlap between augmented and occluded real-world objects, this disclosure can use occlusion materials to render foreground real-world content "on top of" the background animation, effectively masking portions of the animation that should be obscured by the real-world content. Therefore, this disclosure treats real-world content as being rendered, similar to an animation layer, but users of AR devices may not be able to view this rendered content because they are actually viewing real-world content through a "transparent" mask. The purpose of using occlusion materials at the client-side to represent the surface of real-world objects can be to act as a "mask" during the pre-display processing stage; a mask that forces perspective (real-world) content to be visible in the corresponding pixels. This concept can help maintain a clear boundary between real-world foreground content and animated background content.

[0068] In some aspects, pixel information or eye buffer information for the scene can be determined or rendered on the server and sent to the client in separate bulletin boards or layers. For example, enhancements can be determined or rendered in one bulletin board or layer and sent to the client, while real-world objects can be determined and sent in another bulletin board or layer. In other aspects, each bulletin board or layer generated on the server and streamed to the client can contain disjoint planar projections of scene assets (real-world objects or animations) that appear at approximately the same distance / “depth” from the camera (given the nearest camera / headset pose). In such a system, there can be many (e.g., a dozen) different bulletin boards or layers sent to the PS client for pre-display processing.

[0069] Figure 3 An example depth-aligned bulletin board 300 according to one or more technologies based on this disclosure is shown. The bulletin board 300 includes bulletin boards or layers 310, 320, 330, 340, and 350. As noted above, bulletin boards 310, 320, 330, 340, and 350 can represent different content, such as real-world content or augmented content. Furthermore, Figure 3 The use of bulletin boards 310, 320, 330, 340, and 350 in the pixel flow architecture is shown.

[0070] In some cases, bulletin boards or layers can be used to characterize the aforementioned occlusion material for real-world content, and another bulletin board or layer can be used to characterize augmented content, both types of layers being composited on the client in a back-to-foreground manner during pre-display processing. In doing so, the real-world occlusion material and the augmented content can correspond to separate bulletin boards or layers. These bulletin boards or layers can approximate the real-world and animated content in the AR scene, for example, through their planar representation.

[0071] In some respects, these bulletin boards or layers, rendered at various depth values, can help maintain the expected occlusion / occlusion relationships in the mixed reality scene as the AR camera moves at the client device. Thus, one or more bulletin boards or layers can correspond to real-world content, and one or more other bulletin boards or layers can correspond to augmented content, ensuring that the relative occlusion between the real world and the augmented content remains accurate, even as the client's pose moves from pose T0 (content rendering) to pose T1 (content presentation).

[0072] Furthermore, bulletin boards can correspond to planes at different depths, which helps preserve parallax / depth information across different parts of the scene when the AR camera is time-adjusted. This depth information on the bulletin board can also implicitly determine which content (e.g., real-world or augmented content) is occluding other content through content compositing, which can be done from back to front. Additionally, bulletin board or layer representations of real-world objects / surfaces can improve the accuracy of edges or divisions between real-world and augmented content, regardless of which is the occluder.

[0073] In some aspects, real-world objects in a mixed reality scene can be represented on the client as a mesh / geometry describing the object's position, shape, and size, along with occlusion materials for the corresponding surfaces. For example, a mesh representation of the real-world scene can be determined at the server, and the corresponding geometry with occlusion materials can be sent to the client for rendering of the scene. Furthermore, this mesh representation can be determined or known at the server and / or client. If it is also known at the client, then any information needed to represent the real-world content does not need to be sent by the server. Additionally, from a latency perspective, having real-world meshing capabilities on the client has its advantages. If the real-world scene is not entirely static (containing moving or deformable objects), then the latency involved in meshing the world and bringing the mesh information into a rasterization system utilizing occlusion materials can begin to take effect. Clients with on-device meshing capabilities can have an advantage because their real-world representation components will result in lower latency and thus correspond more closely to the actual perspective of the scene.

[0074] During object rendering, the server may require several pieces of information. For example, the server may need a mesh or model of a real-world occluded object, which can then be rendered as a series of bulletin boards or alternatives (e.g., in a pixel-stream system), or it can be sent to the client as a mesh accompanying an empty surface baked into a shading atlas (e.g., in a vector-stream system). In some aspects, bulletin boards containing the screen-space projection of real-world occluders can be characterized by corresponding planar parameters: orientation and distance from the camera (depth). Furthermore, bulletin boards or layers including augmented and real-world objects can be rendered with precise visibility calculations (e.g., within the view frustum), or they can be at least partially over-shaded, for example, to include portions of the object's screen-space projection that are slightly outside the view frustum or not fully visible from the current viewpoint (due to occlusion from other objects). On the client side, bulletin boards or layers can be rendered or composited from back to front and ordered by depth.

[0075] Depending on network and processing latency, environmental meshing can be done entirely in the cloud (e.g., uplink data is real-time video feeds and IMU data captured on the client), partially in the cloud and partially on the client (e.g., uplink data may include some keyframes, accompanying metadata, and other environmental descriptors extracted locally on the client), or entirely on the client (e.g., the client may send some advanced meshing features to the server, but typically maintains meshing properties locally). These options offer a variety of trade-offs between compute offloading (from the client to the edge server or the cloud) and business requirements (from high to medium).

[0076] In some cases of pixel-stream systems, assuming rendering enhancements on the bulletin board (where at least some overshading exceeds strict visibility), and if most objects are not too close to the display camera, then the methods used to depict real-world occlusions onto the client display can tolerate some latency and camera motion without causing severe distortion of occlusion boundaries. Furthermore, the relative motion accumulated on the client side between the head pose used for rendering and the head pose at display can affect the quality of the synthesized content. For example, a combination of AR camera movement and network latency can affect the accuracy of the rendered content.

[0077] In some cases of pixel-stream systems, each real and / or augmented object may utilize a separate bulletin board or layer. This can complicate server and client processing and increase downlink traffic. In some aspects of this disclosure, the server can identify or flag real-world objects that partially intersect the view frustum toward the augmented object and are located between the camera and the augmentation. The remaining real-world objects may not be explicitly signaled on the downlink, thus saving transmission bandwidth and corresponding client processing.

[0078] Furthermore, bulletin boards or alternatives describing real-world objects can trace the outlines of these objects from the viewpoint of the rendering pose T0. In some aspects, when ATW is applied at the client as a means of latency compensation, the outlines of real-world occluded objects signaled can be adjusted or “stretched” proportionally to the pose offset T1-T0 and may not correspond to the occluded object outlines from the latest viewpoint (which may be different or more advanced than the occluded object outlines signaled for bulletin board rendering purposes). This can result in an unnatural appearance at the occlusion edges between the real-world foreground and the enhanced background. Additionally, under certain conditions, the perceived occlusion boundaries may shift relative to the foreground objects, which can also appear unnatural.

[0079] Most of these drawbacks can be naturally addressed in methods proposed for vector flow systems. In such systems, the client performs the final rasterization of the blended scene objects (animated and real-world objects) onto the display screen, for example, using the latest client poses. Since rasterization can use the 3D mesh / geometry of the original assets as input, shape distortion and / or distortion / motion occlusion boundary issues can be avoided, provided that the geometry meshes used for real-world and augmented objects are accurately rendered at the client. In addition to the mesh / geometry of potential visible actors in the blended reality scene, the client can also receive the corresponding shading surfaces packaged in the shading atlas.

[0080] Figure 4A A color atlas 400 illustrating one or more techniques according to this disclosure is shown. For example... Figure 4A As shown, the shading atlas 400 demonstrates an efficient way to store textures in object space rather than in image space. Figure 4A It is also shown that different portions of the color atlas 400 are colored at different resolutions; for example, depending on the distance from the camera, the color texture needs to be depicted with more or less detail. Furthermore, the dark gray portions of the color atlas 400 (e.g., Figure 4A The rightmost shading (420) can represent an unassigned portion of the shading atlas 400. In some cases, the shading atlas 400 can be efficiently encoded based on high temporal coherence. For example, a block in the shading atlas 400 can represent the same physical surface in a mixed real-world environment. In some aspects, blocks in the shading atlas 400 can remain in the same location, as long as they are potentially visible and occupy similar areas in screen space.

[0081] Figure 4B A color atlas organization 450 according to one or more techniques based on this disclosure is shown. For example... Figure 4BAs shown, the shaded atlas 460 includes superblocks 462, where each superblock includes columns 470. Columns of a specific width can be assigned within the same superblock, where each column includes column width A and column width B. Each column may also include a block 480. In some aspects, as shown in the shaded atlas organization 450, column width A may be half the value of column width B. Figure 4B As shown, the color atlas organization 450, from large to small values, can be superblock 462 to column 470 to block 480. Figure 4B This demonstrates how block 480 can be efficiently packed into shader atlas 460. Furthermore, this hierarchy of memory management allows for extensive parallelism when determining the mapping from block 480 to the corresponding locations in shader atlas 460.

[0082] The memory organization of a shaded atlas can be as follows. The atlas consists of superblocks, which are further divided into columns (with equal widths), and each column contains stacked blocks. The blocks have a rectangular shape (squares are a special case), and they contain a fixed number of triangles (e.g., triangle 490). This is in... Figure 4B The following section illustrates this: This section depicts a column with width B, where several blocks are shown as comprising multiple triangles (e.g., triangle 490). This number could be 1, 2, or 3. In some implementations, it is possible to include a block constellation with more than 3 triangles. Triangle-to-block assignments are selected once during asset loading and scene preprocessing (e.g., performed offline), and remain fixed for the duration of the game after that point. In some aspects, the triangle-to-block mapping operation is separate from the online memory organization of the block-to-shading atlas, such as... Figure 4BAs depicted in [the text]. In one case, triangles can be selected to be assigned to a block such that they are neighbors in the object mesh, their aspect ratios are "compatible" with the selected block constellation, their sizes / areas are similar, and their surface normals are similar. Triangle-to-block assignment can be purely based on mesh properties, and therefore exists in object space, and can thus be completed before the game starts. In contrast, block-to-superblock assignment can change over time and is tightly coupled to the appearance of objects in screen space, and therefore depends on the camera position and orientation. Several factors predetermine whether a block will appear in the atlas, and if so, in which superblock. First, a block is initiated into the shading atlas once it becomes part of a potentially visible set. This roughly means that a block is first added once it can be revealed by a camera moving near its current position. The initial superblock assignment of an added block depends on the block's appearance in screen space. Blocks closer to the camera appear larger in screen space and therefore occupy a larger portion of the atlas's base facet. Furthermore, blocks appearing at a large angle of incidence relative to the camera axis may appear tilted and therefore can be represented in the atlas by blocks with a high aspect ratio (e.g., narrow rectangles). Depending on camera movement, the appearance of triangles in screen space may gradually change, and therefore the size and aspect ratio of the corresponding blocks may also change. In the case of a shading atlas organization where column widths are limited to powers of 2, there may be a finite number of block sizes and aspect ratios available in the atlas, and therefore switching blocks from one column-width superblock to another may not be very frequent or continuous in time.

[0083] To maintain the aforementioned temporal coherence within the atlas, a block, once added to the atlas, may be constrained to occupy the same location for as long as possible, subject to certain constraints. However, its position within the atlas can change once: (i) its screen space area and aspect ratio become increasingly unsuitable for the current superblock column width, thus necessitating a switch, or (ii) the block is no longer in the potentially visible set for an extended period. In both cases, the block in question can be moved or removed entirely within the atlas, particularly in case (ii). The change may result in a temporary increase in the coding bitrate, but the new state can persist for a period, leading to an overall decrease in the coding bitrate. To facilitate fast memory lookups, both the horizontal and vertical block sizes are kept powers of 2. For example, a block could be 16 pixels high and 64 pixels wide. This block will be assigned to a superblock containing a block that is entirely 64 pixels wide. When the headphone position changes, causing the block's aspect ratio to become closer to 1:2 and slightly larger in size, the block will be removed from its current position, and it can be assigned 32 pixels in height and 64 pixels in width. In this case, it may be reassigned to a superblock with a width of 64, but its position will need to be changed.

[0084] In the proposed method for characterizing information about real-world occlusions to a vector streaming client, the corresponding mesh / geometry is streamed along with the remaining geometry (e.g., the geometry of animated objects) sent on the downlink. In systems where the split rendering client supports occlusion material rendering, the resulting real-world polygons may need to be accompanied by corresponding occlusion material messages. In clients that do not support occlusion material rendering, the corresponding polygon surfaces need to be characterized as textures in a shading atlas. As mentioned above, these textures rendered to the shading atlas can be uniform black, (0,0,0) values, or empty textures, because in some cases, rendering (0,0,0) values ​​on an augmented reality display may cause the corresponding areas to be perceived as completely transparent or see-through. Furthermore, since the shading textures for these specific objects can be uniform in color, their resolution does not need to follow the traditional constraint that polygons appearing closer to the camera in a 3D scene can be represented by higher resolution textures (i.e., larger rectangles in the shading atlas). In other words, the polygon textures of real-world objects in a mixed reality scene can be represented by small rectangles shaded using uniform empty textures. In some respects, a small rectangle can be 2 pixels by 2 pixels in size (i.e., a square with 2 pixels in each dimension).

[0085] Furthermore, some vectorflow systems that do not support rendering occluded materials on the client side can support a special message format for describing polygons of a uniform color, which can be called a monochrome message. These messages can be sent instead of explicitly coloring uniform black squares in the shading atlas, and these messages can include a header indicating that a monochrome message is being sent, followed by a number of unique surface patches being mapped to the same color. The body of the message can describe the uniform color, which is (0,0,0) black in the case of occluded surfaces. This can then be followed by enumerating all parts of the texture (tiles / triangles) that might render the described color. Using a unique unit surface identifier facilitates this enumeration. For example, all tiles present in a mixed reality scene are assigned unique tile IDs, and these tile IDs can be included in the monochrome message. From a bitrate, memory, and / or latency perspective, this transmission scheme can be more efficient (compared to schemes where (0,0,0) color information is explicitly signaled as a single rectangle in the shading atlas). Furthermore, monochrome messages can be used to describe certain surfaces, which is an alternative to explicitly rendering them in the shading atlas. These monochrome messages can eliminate most of the redundancy and inefficiency of describing them using standard video compression by simply removing all surfaces of uniform color from the video.

[0086] In some cases, the Vector Stream client may be unaware of the existence of real-world occluded objects. For example, there may be no difference in client-side operation when dealing with AR applications with or without real-world occlusion. The Vector Stream approach, which is an integral part of the pre-display process in all split-rendering systems, offers improved performance compared to Asynchronous Time Warp (ATW) and similar methods that can exist in pixel-stream-based split-rendering systems for novel view compositing. Vector Stream-based novel view compositing avoids shape distortion and occlusion boundary shift / distortion based on planar reprojection by re-rendering the entire potentially visible geometry from the novel viewpoint. In some aspects, if uplink traffic is delayed or interrupted, or assuming meshing on the client, the client device can locally reconstruct the appearance and geometry of occluded real-world objects and then add the corresponding mesh and all-zero-value textures to the existing list of draw calls. In this way, Vector Stream clients with meshing capabilities can provide greater robustness to unstable communication links and offer a solution with lower overall latency for incorporating real-world objects into complex mixed reality experiences. In some respects, the locally available mesh of occluded objects (reconstructed using the client) may not be highly refined and may require bundled adjustments at the server level to make it more accurate and conform to actual real-world geometry. However, in other respects, in cases of long network latency or service outages, it may still be possible to choose a coarse version of the real-world occluders rather than having any such representation.

[0087] In some cases, if throughput is slow or latency is high, clients can choose to use a local mesh, for example, for real-world content, instead of a mesh from the server. By using a local mesh, the client can render this mesh content on the client side without waiting to receive the mesh from the server. This can provide a more immediate response to real-world occlusion. For example, a local client-side mesh can be useful for real-world content that the client detects but the server is not yet aware of.

[0088] The client device can also perform a coarse real-world mesh and send that information to the server. The server can then perform a bundle adjustment, comparing real-world observations of real objects made by the client at many different points in time and refining that information into a single, compact, refined mesh for the future representation of the real-world object. Therefore, real-world content refinement can occur on either the client or the server.

[0089] In some cases, client devices can utilize additional pathways to render occluded objects unknown to the server, for example, by using up-to-date real-world object meshes and poses. This can be detected on the client if the occluded objects are not present in the downlink data. Furthermore, the client can operate autonomously, for example, without input from the server, even in the presence of occluded objects and / or without reliable communication network services.

[0090] As noted above, aspects of this disclosure can handle occluding objects in a pixel-stream system by rendering occluding objects via layers or bulletin boards. For example, this disclosure can determine which real-world objects are likely occluding objects by comparing their geometry with that of augmented objects in mixed reality space. For example, the geometry of real-world objects (e.g., detected and meshed in a world coordinate system relative to any point in the real world) can be “transformed” into mixed-world coordinates indicated by the animated objects using the pose of a common subject (camera) between the two worlds. A variant of the standard visibility test can then be performed in the mixed-world representation to mark selected real-world objects as potential occluders. Real-world objects that overlap with the animated portion in pixel space and appear to be closer to the virtual camera than at least one augmented object can be marked as potential occluders and rendered onto a separate bulletin board using the latest known object mesh. These occluding objects can be rasterized onto the bulletin board plane at an average object distance from the AR camera (or using some other method to determine the planar distance and orientation relative to the camera). Object surfaces in the bulletin board can be rendered using all-zero pixels (e.g., RGB values ​​of (0,0,0)).

[0091] Furthermore, this disclosure can handle occlusion objects similar to other objects in a vector flow system. Real-world objects with known meshes can be viewed in the same way as animated objects on the server. Therefore, these real-world objects can participate in visibility calculations. If a real-world object is found to occlude at least one animated object, this disclosure can treat the real-world object as an active occluding object and add the corresponding information to the combined set of known objects. This disclosure can also send the potential visible geometry or texture of these objects as messages with monochrome values ​​or RGB values ​​of (0,0,0). Since the mesh information of these objects can be further refined on the server, this disclosure can send updates on the object geometry.

[0092] As noted herein, this disclosure may also include estimating the geometry of real-world occlusion objects locally on the client device (rather than on the server) when messages received from the server do not include information about these occluded objects. This may occur for various reasons, such as reducing system latency, temporary uplink interruptions, and / or limitations on server workload. In this disclosure, the vector streaming client can add the locally estimated mesh of the real-world object to a list of rasterized meshes on the client using all-zero values ​​for the corresponding texture. Furthermore, the pixel streaming client can rasterize the locally estimated mesh in a separate path before combining the result with pixel streaming textures received over the network. Therefore, in pixel streaming systems, this disclosure can utilize one or more solutions to address real-world content occlusion enhancement. And in vector streaming systems, this disclosure can utilize multiple solutions to address real-world content occlusion enhancement.

[0093] Figure 5 Example images or scenes 500 based on one or more technologies according to this disclosure are shown. Scene 500 includes augmented content 510 and real-world content 520 (which includes edges 522). More specifically, Figure 5 A real-world object 520 (e.g., a door) is displayed, which is occluding augmented content 510 (e.g., a person). Figure 5 As shown, the augmented content 510 rests at the edge 522 of the real-world object 520 because the real-world object 520 completely occludes the augmented content 510. For example, no part of the person 510 overlaps with the edge 522 of the door 520.

[0094] As pointed out above, Figure 5 It accurately demonstrates the effect of real-world content (e.g., content 520) occluding augmented content (e.g., content 510) in an AR system. Figure 5 As shown, aspects of this disclosure can accurately determine which real-world objects and animated objects can occlude each other when they fully or partially overlap in screen space. Furthermore, aspects of this disclosure can render mixed reality scene content in a split rendering architecture, ensuring that the occlusion boundaries between occluded / foreground objects and occluded / background objects remain sharp and accurate despite a series of network latency and camera movements, when the objects are both real-world and augmented content. Therefore, aspects of this disclosure can accurately depict when real-world content fully or partially occludes augmented content, or when augmented content occludes real-world content, and how and where the boundaries between them can be depicted in an AR display.

[0095] Figure 5An example of the process described above for handling occluded content in split rendering with both real-world and augmented content is shown. Figure 5 As shown, various aspects of this disclosure (e.g., the server and client devices described herein) can perform multiple different steps or processes to render a mixed reality scene in an immersive manner in which real-world content and augmented content coexist and interact with each other. For example, the server and client devices described herein can identify a first group of content within real-world objects in the scene (e.g., Figure 5 Door 520 in the middle) and the second content group in the animation object (e.g., Figure 5 (Person 510 in the document). The server and client devices described herein can also determine, based on their positions in the view frustum, from the current camera pose or nearby camera poses whether one might potentially occlude the other (partially or completely), and for potential occluders / occluded objects, their appearance and geometric properties are represented in such a way that the occlusion boundary between them can be represented accurately when the blended scene is re-rendered (e.g., according to a dense set of nearby camera poses and orientations). In some aspects, the first content group (e.g., door 520) may include at least some real content, and the second content group (e.g., person 510) may include at least some augmented content. The server and client devices described herein can also determine whether at least a portion of the first content group (e.g., door 520) occludes or potentially occludes at least a portion of the second content group (e.g., person 510).

[0096] In some aspects, at least one primary bulletin board may be used (e.g., Figure 3 The bulletin board 320 in the middle can be used to represent the first content group (e.g., door 520), and at least one second bulletin board (e.g., Figure 3 The second content group (e.g., person 510) can be represented by bulletin board 310. The server and client devices described herein can also generate at least one first bulletin board (e.g., bulletin board 320) based on at least one first plane equation, and at least one second bulletin board (e.g., bulletin board 310) based on at least one second plane equation. In some aspects, the determination of whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group can be based on at least one of the following: the current camera pose, one or more camera poses having a nearby location or orientation, or one or more predicted camera poses (including recent poses). Furthermore, at least one first bulletin board (e.g., bulletin board 320) can be positioned and oriented in a manner corresponding to the first plane, and at least one second bulletin board (e.g., bulletin board 310) can be positioned and oriented in a manner corresponding to the second plane.

[0097] In some aspects, at least one first mesh and at least one first shading texture may be used to represent a first content group (e.g., door 520), and at least one second mesh and at least one second shading texture may be used to represent a second content group (e.g., person 510). The server and client devices described herein may also generate at least one first mesh and one first shading texture, as well as at least one second mesh and one second shading texture.

[0098] The server and client devices described herein can also render a first content group (e.g., door 520) and a second content group (e.g., person 510) based on whether at least a portion of the first content group might occlude at least a portion of the second content group from the current or nearby camera position and orientation. In some aspects, pixels in the AR headset that are not rendered with any content can behave as if the rendered value were (0,0,0), i.e., these pixels remain completely transparent to real-world content (in the sense that a viewer can “see through” these pixels). Alternatively, when at least a portion of the first content group occludes at least a portion of the second content group, the client devices described herein can use occlusion material surface properties to render at least a portion of one or more surfaces of the first content group. In some aspects, at least one first mesh can be associated with first geometric information, and at least one second mesh can be associated with second geometric information.

[0099] The server and client devices described herein can also transmit information associated with a first content group (e.g., door 520) and information associated with a second content group (e.g., person 510) to the client device. In some aspects, the client device can perform indistinguishable processing of elements of the first content group (e.g., door 520) and elements of the second content group (e.g., person 510). In some aspects, when at least one eye buffer is composited for display, the client device can process information associated with the first content group and information associated with the second content group. Additionally, the representations of the first content group (e.g., door 520) and the second content group (e.g., person 510) can be prepared at least partially by the server. Furthermore, the representations of the first content group (e.g., door 520) and the second content group (e.g., person 510) can be rendered at least partially by the client device into, for example, an eye buffer for display. In some aspects, the first content group can be represented using one or more first grid elements and one or more first pixels, and the second content group can be represented using one or more second grid elements and one or more second pixels. In some respects, the first content group can be represented by one or more first occlusion messages.

[0100] In some aspects, a first content group (e.g., door 520) may include one or more pixels, and a second content group (e.g., person 510) may be represented by one or more pixels. Furthermore, a scene (e.g., scene 500) may be composited and / or generated on a display at a client device located at least one of the following: split augmented reality (AR) architecture, split extended reality (XR) architecture, or cloud gaming with a server. Additionally, the determination of whether at least a portion of the first content group (e.g., door 520) occludes or potentially occludes at least a portion of the second content group (e.g., person 510) may be performed at least partially by a GPU or CPU. In some aspects, rendering at least a portion of one or more surfaces of the first content group using occlusion materials when at least a portion of the first content group occludes at least a portion of the second content group may be performed at least partially on a server.

[0101] As noted above, the use of various aspects of the pixel-streaming system in this disclosure may include bulletin boards, alternates, or layers, and real-world content can be rendered as a completely empty / black silhouette mask onto the corresponding bulletin board using the current camera pose. The use of various aspects of the vector-streaming system in this disclosure may also include client devices that process information representing real-world content and animated content in exactly the same way. For example, it may be assumed that each real-world occlusion object is known at the server. Furthermore, it may be assumed that each real-world occlusion object is represented at least in part by a message sent to the client. In some aspects, the modality representing real-world content within a real scene embedded in the sent message may be the same as the modality representing animated content in a virtual / mixed reality scene. In practice, the client device may not make any distinction regarding which representation corresponds to an augmented or real-world object.

[0102] Furthermore, aspects of this disclosure may include a vector flow architecture, where the server may be unaware of real-world occlusion objects, and / or downlink traffic to the client may not contain information about occlusion objects, but the client device may be aware of real-world occlusion objects: their shape, size, and pose relative to the camera / observer. In these cases, the client device may have a separate rendering path, where the client rasterizes the mesh of the real-world occlusion objects, thereby rendering their textures using occlusion materials. For example, the server may need several frames to become aware of the occlusion objects, but the client may become aware of them within a single frame. Therefore, the client device can determine the geometry of the occlusion objects faster than the server, thus enabling faster characterization of the presence of occlusion objects and minimizing the occurrence of obviously incorrect occlusion.

[0103] Figure 6An example flowchart 600 of an example method according to one or more techniques of this disclosure is shown. This method can be executed by a device (such as a server, client, CPU, GPU, or device for graphics processing). At 602, the device can identify a first content group and a second content group in the scene, as combined in… Figure 2 , Figure 3 Figure 4 and Figure 5 The examples described herein. In some aspects, the first content group may include at least some real content, and the second content group may include at least some enhanced content, such as content combined with... Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document / reference] is as follows.

[0104] At 604, the device can determine whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group (e.g., from the current or nearby camera pose), as combined with Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document / reference] is as follows.

[0105] In some aspects, if cross-content occlusion is possible, the first content group may be represented using at least one first bulletin board. The second content group may be represented using at least one second bulletin board, such as in combination with... Figure 2 , Figure 3 Figure 4 and Figure 5 As described in the example. At 606, the device can generate at least one first bulletin board according to at least one first planar equation, and at least one second bulletin board according to at least one second planar equation, as combined with... Figure 2 , Figure 3 Figure 4 and Figure 5 The example described herein. In some aspects, the determination of whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group is based on at least one of the following: the current camera pose, one or more camera poses having a nearby location or orientation, or one or more predicted camera poses including the recent pose, such as in combination with Figure 2 , Figure 3 Figure 4 and Figure 5 As described in the example. Additionally, at least one first bulletin board may correspond to a first plane, and at least one second bulletin board may correspond to a second plane, as in conjunction with... Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document / reference] is as follows.

[0106] In some aspects, the first content group may be represented using at least one first mesh and a first shading texture, and the second content group may be represented using at least one second mesh and a second shading texture, as combined in... Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document / reference] is as follows.

[0107] At 610, the device can represent the first content group and the second content group based on a determination of whether at least a portion of the first content group obscures or potentially obscures at least a portion of the second content group, as combined with... Figure 2 , Figure 3 Figure 4 and Figure 5 As described in the example. At 612, when at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group, the device can use an occlusion material to render at least a portion of one or more surfaces of the first content group, as combined with... Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document / reference] is as follows.

[0108] In some respects, the first content group may be represented by a first color atlas, and the second content group may be based on a second color atlas, such as in combination with... Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document] further illustrates this. Additionally, at least one first mesh can be associated with first geometric information, and at least one second mesh can be associated with second geometric information, as combined with [other information]. Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document / reference] is as follows.

[0109] At point 614, the device can transmit information associated with the first content group and information associated with the second content group to the client device, such as in combination with... Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document] illustrates this. In some aspects, for example, when at least one eye buffer is synthesized for display, the client device can process information associated with a first content group and information associated with a second content group, such as in combination with [other components]. Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document] further illustrates this. Additionally, the representation of the first content group and the representation of the second content group can be at least partially colored on the server, as combined with [other elements]. Figure 2 , Figure 3 Figure 4 and Figure 5The example described in [the document] further illustrates this. Additionally, the first and second content groups can be rendered at least partially by the client device, as combined in [the document]. Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document / reference] is as follows.

[0110] In some respects, the first content group may include one or more pixels, and the second content group may include one or more pixels, as combined in... Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document] illustrates this. In some aspects, the first content group may be represented by one or more first occlusion messages, as combined with [other messages]. Figure 2 , Figure 3 Figure 4 and Figure 5 The examples described herein. In some aspects, a first content group may be represented using one or more first grid elements and one or more first pixels, and a second content group may be represented using one or more second grid elements and one or more second pixels, as combined in... Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document] further illustrates this. Additionally, the scene can be rendered on a display at a client device located in at least one of the following: a split augmented reality (AR) architecture, a split extended reality (XR) architecture, or cloud gaming with a server, such as [the specific application / system]. Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document] further illustrates this. Furthermore, determining whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group (e.g., from the current or nearby camera position and orientation) can be performed at least partially by the GPU or CPU, as in [the document]. Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document / reference] is as follows.

[0111] In some respects, determining whether at least a portion of the first content group obscures or potentially obscures at least a portion of the second content group can be performed at least partially on the server, such as in conjunction with Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document] further illustrates this. Additionally, the use of occlusion materials to render at least a portion of one or more surfaces of the first content group when at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group can be at least partially performed on the server, as in [the context of] [the document]. Figure 2 , Figure 3 Figure 4 and Figure 5 The example described in [the document / reference] is as follows.

[0112] In one configuration, a method or apparatus for graphics processing is provided. The apparatus may be a server, client device, CPU, GPU, or some other processor capable of performing graphics processing. In one aspect, the apparatus may be a processing unit 120 within device 104, or some other hardware within device 104 or another device. The apparatus may include units for identifying a first content group and a second content group in a scene. The apparatus may also include units for determining whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group. The apparatus may also include units for representing the first content group and the second content group based on the determination that at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group. The apparatus may also include units for rendering at least a portion of one or more surfaces of the first content group using an occlusion material when at least a portion of the first content group occludes at least a portion of the second content group. The apparatus may also include units for generating at least one first bulletin board according to at least one first plane equation and a first shading texture, and generating at least one second bulletin board and a second shading texture according to at least one second plane equation. The device may also include a unit for transmitting information associated with a first content group and information associated with a second content group to a client device.

[0113] The themes described herein can be implemented to achieve one or more benefits or advantages. For example, the described graphics processing techniques can be used by a server, client, GPU, CPU, or some other processor capable of performing graphics processing to implement the split rendering techniques described herein. This can also be implemented at a lower cost compared to other graphics processing techniques. Furthermore, the graphics processing techniques described herein can improve or accelerate data processing or execution. Further, the graphics processing techniques described herein can improve resource or data utilization and / or resource efficiency. Moreover, aspects of this disclosure can utilize split rendering processes that can improve the accuracy of handling occluded content in split rendering with both real-world and augmented content.

[0114] According to this disclosure, unless otherwise specified in the context, the term "or" can be interpreted as "and / or". Additionally, while phrases such as "one or more" or "at least one" may be used for some features disclosed herein but not for others, the absence of such language in a feature can be interpreted as implying such a meaning unless otherwise specified in the context.

[0115] In one or more examples, the functionality described herein may be implemented in hardware, software, firmware, or any combination thereof. For example, although the term “processing unit” has been used throughout this disclosure, such a processing unit may be implemented in hardware, software, firmware, or any combination thereof. If any functionality, processing unit, technique, or other module described herein is implemented in software, the functionality, processing unit, technique, or other module described herein may be stored on or transmitted via a computer-readable medium as one or more instructions or code. A computer-readable medium may include a computer data storage medium or a communication medium, the communication medium including any medium that facilitates the transfer of a computer program from one place to another. In this way, a computer-readable medium may generally correspond to: (1) a tangible computer-readable storage medium that is non-transitory; or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described herein. By way of example and not limitation, such a computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices. As used herein, disks and optical discs include compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically copy data magnetically, while optical discs use lasers to copy data optically. Combinations of the foregoing should also be included within the scope of computer-readable media. Computer program products may include computer-readable media.

[0116] The code can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), arithmetic logic units (ALUs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Furthermore, the techniques can be sufficiently implemented in one or more circuit or logic elements.

[0117] The technology disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed technology, but they do not necessarily need to be implemented through different hardware units. Specifically, as described above, the various units can be combined in any hardware unit, or provided by a batch of interoperable hardware units (including one or more processors as described above) combined with suitable software and / or firmware.

[0118] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. A method for graphics processing at a client device, comprising: Receive information from the server associated with a first content group in the scene and information associated with a second content group in the scene, wherein at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group; as well as The first content group and the second content group are rendered at the client device based on the information associated with the first content group and the information associated with the second content group.

2. The method according to claim 1, wherein, The first content group includes at least some real content, and the second content group includes at least some enhanced content.

3. The method according to claim 1, wherein, Rendering the first content group and the second content group includes: When at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group, an occlusion material is used to render at least a portion of one or more surfaces of the first content group.

4. The method according to claim 1, wherein, The first content group is represented using at least one first bulletin board, and the second content group is represented using at least one second bulletin board.

5. The method according to claim 4, further comprising: The at least one first bulletin board is generated according to at least one first plane equation, and the at least one second bulletin board is generated according to at least one second plane equation.

6. The method according to claim 1, wherein, Whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group is determined based on at least one of the following: the current camera pose, one or more camera poses having a nearby location or orientation, or one or more predicted camera poses including the recent pose.

7. The method according to claim 6, further comprising: The physical model and motion prediction of scene actors, including a portion of the first content group and a portion of the second content group, are used to predict the position and orientation of the portion of the first content group and the portion of the second content group at the time of display.

8. The method according to claim 1, wherein, The first content group is represented using at least one first mesh and a first shading texture, and the second content group is represented using at least one second mesh and a second shading texture.

9. The method according to claim 8, wherein, The at least one first grid is associated with first geometric information, and the at least one second grid is associated with second geometric information.

10. The method according to claim 8, wherein, The second content group is rendered using either occlusion material surface properties or pure black pixels with RGB values ​​of (0,0,0), depending on the client's capabilities.

11. The method according to claim 1, wherein, When at least one eye buffer is synthesized for display, the client device processes the information associated with the first content group and the information associated with the second content group.

12. The method according to claim 1, wherein, The rendering of the first content group and the second content group is performed at least partially on the server.

13. The method according to claim 1, wherein, The first content group is represented by one or more pixels, and the second content group is represented by one or more pixels.

14. The method according to claim 1, wherein, The first content group is represented by one or more first occlusion messages.

15. The method according to claim 7, wherein, The first content group is represented using one or more first grid elements and one or more first pixels, and the second content group is represented using one or more second grid elements and one or more second pixels.

16. The method according to claim 1, wherein, The scene is rendered on a display at the client device located in at least one of the following: a split augmented reality (AR) architecture with the server, a split extended reality (XR) architecture, or cloud gaming.

17. The method according to claim 1, wherein, Whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group is determined at least partially by the graphics processing unit (GPU) or the central processing unit (CPU).

18. The method according to claim 1, wherein, Whether at least a portion of the first content group obscures or potentially obscures at least a portion of the second content group is determined at least partially by the server.

19. The method according to claim 3, wherein, The rendering of at least a portion of one or more surfaces of the first content group using the occlusion material when at least a portion of the first content group occludes at least a portion of the second content group is performed at least partially on the server.

20. An apparatus for graphics processing at a client device, comprising: Memory; as well as At least one processor, coupled to the memory and configured to cause the device to: Receive from the server information associated with a first content group in the scene and information associated with a second content group in the scene, wherein at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group; and The first content group and the second content group are rendered based on the information associated with the first content group and the information associated with the second content group.

21. The apparatus according to claim 20, wherein, The first content group includes at least some real content, and the second content group includes at least some enhanced content.

22. The apparatus according to claim 20, wherein, In order to render the first content group and the second content group, the at least one processor is configured to: When at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group, an occlusion material is used to render at least a portion of one or more surfaces of the first content group.

23. The apparatus according to claim 20, wherein, The first content group is represented using at least one first bulletin board, and the second content group is represented using at least one second bulletin board.

24. The apparatus according to claim 23, wherein, The at least one processor is further configured to: The at least one first bulletin board is generated according to at least one first plane equation, and the at least one second bulletin board is generated according to at least one second plane equation.

25. The apparatus according to claim 20, wherein, Whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group is determined based on at least one of the following: the current camera pose, one or more camera poses having a nearby location or orientation, or one or more predicted camera poses including the recent pose.

26. The apparatus according to claim 25, wherein, The at least one processor is further configured to: utilize a physical model and motion prediction of scene actors including a portion of the first content group and a portion of the second content group to predict the position and orientation of the portion of the first content group and the portion of the second content group at the time of display.

27. The apparatus according to claim 20, wherein, The first content group is represented using at least one first mesh and a first shading texture, and the second content group is represented using at least one second mesh and a second shading texture.

28. The apparatus according to claim 27, wherein, The at least one first grid is associated with first geometric information, and the at least one second grid is associated with second geometric information.

29. The apparatus according to claim 27, wherein, The second content group is rendered using either occlusion material surface properties or pure black pixels with RGB values ​​of (0,0,0), depending on the client's capabilities.

30. The apparatus according to claim 20, wherein, When at least one eye buffer is synthesized for display, the client device processes the information associated with the first content group and the information associated with the second content group.

31. The apparatus according to claim 20, wherein, The rendering of the first content group and the second content group is performed at least partially on the server.

32. The apparatus according to claim 20, wherein, The first content group is represented by one or more pixels, and the second content group is represented by one or more pixels.

33. The apparatus according to claim 20, wherein, The first content group is represented by one or more first occlusion messages.

34. The apparatus according to claim 20, wherein, The first content group is represented using one or more first grid elements and one or more first pixels, and the second content group is represented using one or more second grid elements and one or more second pixels.

35. The apparatus according to claim 20, wherein, The scene is rendered on a display at the client device located in at least one of the following: a split augmented reality (AR) architecture with the server, a split extended reality (XR) architecture, or cloud gaming.

36. The apparatus according to claim 20, wherein, Whether at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group is determined at least in part by the graphics processing unit (GPU) or the central processing unit (CPU).

37. The apparatus according to claim 20, wherein, Whether at least a portion of the first content group obscures or potentially obscures at least a portion of the second content group is determined at least partially by the server.

38. The apparatus according to claim 22, wherein, The use of occlusion materials to render at least a portion of one or more surfaces of the first content group when at least a portion of the first content group occludes at least a portion of the second content group is performed at least partially on the server.

39. An apparatus for graphics processing at a client device, comprising: A unit for receiving from a server information associated with a first content group in a scene and information associated with a second content group in the scene, wherein at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group; as well as A unit for rendering the first content group and the second content group at the client device based on the information associated with the first content group and the information associated with the second content group.

40. A computer-readable medium storing computer-executable code for graphics processing at a client device, the code causing the client device to perform the following operations when executed by a processor: Receive from the server information associated with a first content group in the scene and information associated with a second content group in the scene, wherein at least a portion of the first content group occludes or potentially occludes at least a portion of the second content group; and The first content group and the second content group are rendered based on the information associated with the first content group and the information associated with the second content group.

Citation Information

Patent Citations

  • Image display method and device and electronic device

    CN109544698A