Media processing method, apparatus and storage medium

By modifying the description parameters of the acoustic and visual scenes to make them consistent, the problem of inconsistency between acoustic and visual scenes in immersive media was solved, improving the user experience and rendering effect.

CN116490922BActive Publication Date: 2026-05-29TENCENT AMERICA LLC

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2022-09-16
Publication Date
2026-05-29

Smart Images

  • Figure CN116490922B_ABST
    Figure CN116490922B_ABST
Patent Text Reader

Abstract

Media content data of an object is received. It is determined whether a first parameter indicated by a first description of the object in an acoustic scene is inconsistent with a second parameter indicated by a second description of the object in a visual scene. Based on the first parameter indicated by the first description of the object in the acoustic scene being inconsistent with the second parameter indicated by the second description of the object in the visual scene, one of the first description of the object in the acoustic scene and the second description of the object in the visual scene is modified based on the other of the first description and the second description that is not modified, wherein the modified one of the first description and the second description is consistent with the other of the first description and the second description that is not modified.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Patent Application No. 17 / 945,024, filed September 14, 2022, entitled "Consistence of Acoustic and Visual Scenes," and U.S. Provisional Application No. 63 / 248,942, filed September 27, 2021, entitled "Consistence of Acoustic and Visual Scenes." The disclosures of these prior applications are incorporated herein by reference in their entirety. Technical Field

[0003] This disclosure describes embodiments that generally involve media processing. Background Technology

[0004] The background description provided herein is intended to provide a general overview of the context of this disclosure. With respect to the work described in this background section and to various aspects of this specification that may not qualify as prior art at the time of filing, the work of the currently designated inventors is neither expressly nor implicitly acknowledged as prior art to this disclosure.

[0005] In immersive media applications such as virtual reality (VR) or augmented reality (AR), visual information from a virtual scene and acoustic information from an acoustic scene can be provided to create or mimic the physical world through digital simulation. Perception is created by surrounding the user with images, videos, sounds, or other stimuli that provide an engaging overall environment. Immersive media can create environments that allow users to interact with content within their surroundings. Summary of the Invention

[0006] Various aspects of this disclosure provide methods and apparatus for media processing. In some examples, the apparatus for media processing includes processing circuitry.

[0007] According to one aspect of this disclosure, a method for performing media processing at a media processing device is provided. In this method, media content data of an object can be received. The media content data may include a first description of the object in an acoustic scene generated by an audio engine and a second description of the object in a visual scene generated by a visual engine. It can be determined whether a first parameter indicated by the first description of the object in the acoustic scene is inconsistent with a second parameter indicated by the second description of the object in the visual scene. In response to the inconsistency between the first parameter indicated by the first description of the object in the acoustic scene and the second parameter indicated by the second description of the object in the visual scene, one of the first description of the object in the acoustic scene and the second description of the object in the visual scene may be modified based on the other of the first and second descriptions, which has not been modified. The modified first and second descriptions may be consistent with the other of the first and second descriptions, which has not been modified. The media content data of the object can be provided to a receiver that renders the media content data of the object for a media application.

[0008] In some embodiments, the first parameter and the second parameter may both be associated with one of the object's object size, object shape, object position, object orientation, and object texture.

[0009] One of the first description of an object in the acoustic scene and the second description of an object in the visual scene can be modified based on the other. Therefore, the first parameter indicated by the first description of the object in the acoustic scene can be consistent with the second parameter indicated by the second description of the object in the visual scene.

[0010] In this method, a third description of an object in a unified scene can be determined based on at least one of a first description of an object in an acoustic scene or a second description of an object in a visual scene. In response to a first parameter indicated by the first description of the object in the acoustic scene being different from a third parameter indicated by the third description of the object in the unified scene, the first parameter indicated by the first description of the object in the acoustic scene can be modified based on the third parameter indicated by the third description of the object in the unified scene. Similarly, in response to a second parameter indicated by the second description of the object in the visual scene being different from a third parameter indicated by the third description of the object in the unified scene, the second parameter indicated by the second description of the object in the visual scene can be modified based on the third parameter indicated by the third description of the object in the unified scene.

[0011] In the examples, the object size in the third description of the unified scene can be determined based on the object size in the first description of the object in the acoustic scene. The object size in the third description of the unified scene can be determined based on the object size in the second description of the object in the visual scene. The object size in the third description of the unified scene can be determined based on the intersection size of the objects in the first description of the acoustic scene and the objects in the second description of the visual scene. The object size in the third description of the unified scene can be determined based on the size difference between the object size in the first description of the acoustic scene and the object size in the second description of the visual scene.

[0012] In the examples, the object shape in the third description of an object in a unified scene can be determined based on the object shape in the first description of the object in the acoustic scene. The object shape in the third description of an object in a unified scene can be determined based on the object shape in the second description of the object in the visual scene. The object shape in the third description of an object in a unified scene can be determined based on the intersection shape of the objects in the first description of the acoustic scene and the objects in the second description of the visual scene. The object shape in the third description of an object in a unified scene can be determined based on the shape difference between the object shape in the first description of the acoustic scene and the object shape in the second description of the visual scene.

[0013] In the examples, the object position in the third description of an object in a unified scene can be determined based on the object position in the first description of the object in the acoustic scene. The object position in the third description of an object in a unified scene can be determined based on the object position in the second description of the object in the visual scene. The object position in the third description of an object in a unified scene can also be determined based on the positional difference between the object position in the first description of the object in the acoustic scene and the object position in the second description of the object in the visual scene.

[0014] In the examples, the object orientation in the third description of an object in a unified scene can be determined based on the object orientation in the first description of the object in the acoustic scene. The object orientation in the third description of an object in a unified scene can be determined based on the object orientation in the second description of the object in the visual scene. The object orientation in the third description of an object in a unified scene can be determined based on the directional difference between the object orientation in the first description of the object in the acoustic scene and the object orientation in the second description of the object in the visual scene.

[0015] In the examples, the object texture in the third description of an object in a unified scene can be determined based on the object texture in the first description of an object in the acoustic scene. The object texture in the third description of an object in a unified scene can be determined based on the object texture in the second description of an object in the visual scene. The object texture in the third description of an object in a unified scene can be determined based on the texture difference between the object texture in the first description of an object in the acoustic scene and the object texture in the second description of an object in the visual scene.

[0016] In some embodiments, the description of an object in an anchor scene of media content data can be determined based on one of a first description of an object in an acoustic scene and a second description of an object in a visual scene. In response to determining the description of an object in the anchor scene based on the first description of the object in the acoustic scene, the second description of the object in the visual scene can be modified based on the first description of the object in the acoustic scene. Similarly, in response to determining the description of an object in the anchor scene based on the second description of the object in the visual scene, the first description of the object in the acoustic scene can be modified based on the second description of the object in the visual scene. Further, signaling information can be generated to indicate which of the first description of the object in the acoustic scene and the second description of the object in the visual scene is selected to determine the description of the anchor scene.

[0017] In some embodiments, signaling information may be generated to instruct which of the first parameter in the first description of an object in the acoustic scene and the second parameter in the second description of an object in the visual scene should be selected to determine the third parameter in the third description of an object in the unified scene.

[0018] According to another aspect of this disclosure, an apparatus is provided. The apparatus includes processing circuitry. The processing circuitry can be configured to perform any of the methods for media processing.

[0019] This disclosure also provides a non-volatile computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform any of the methods for media processing. Attached Figure Description

[0020] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0021] Figure 1 A schematic diagram of an environment using 6 degrees of freedom (6DoF) is shown in some examples.

[0022] Figure 2 A schematic diagram of an MPEG-I immersive audio evaluation platform and process according to an embodiment of the present disclosure is shown.

[0023] Figure 3An exemplary acoustic scenario of a test scenario according to an embodiment of the present disclosure is shown.

[0024] Figure 4 An exemplary visual scene from a test scene derived from the Unity engine, according to an embodiment of this disclosure, is shown.

[0025] Figure 5 A block diagram of a media system according to an embodiment of the present disclosure is shown.

[0026] Figure 6 A flowchart outlining some embodiments of the process according to this disclosure is shown.

[0027] Figure 7 It is a schematic illustration of a computer system according to an embodiment. Detailed Implementation

[0028] The MPEG-I immersive media standard suite, including "immersive audio," "immersive video," and "system support," can support virtual reality (VR) or augmented reality (AR) presentation (100), in which a user (102) can navigate and interact with the environment using six degrees of freedom (6DoF) (including spatial navigation (x, y, z) and user head orientation (yaw, pitch, roll)), for example... Figure 1 As shown.

[0029] The goal of MPEG-I presentation is to give users (e.g., (102)) the feeling of actually existing in a virtual world. Audio in the world (or scene) can be perceived as being in the real world, where sound originates from associated visual images. That is, sound can be perceived at the correct location and / or the correct distance. A user's physical movement in the real world can be perceived as having a matching movement in the virtual world. Furthermore, and importantly, users can interact with the virtual scene and have sounds perceived as real match or otherwise simulate the user's experience in the real world.

[0030] This disclosure relates to immersive media. When rendering immersive media, acoustic and visual scenes may exhibit inconsistencies, which can degrade the user's media experience. This disclosure provides aspects of methods and apparatuses to improve the consistency between acoustic and visual scenes.

[0031] In the MPEG-I immersive audio standard, visual scenes can be rendered by a first engine such as the Unity engine, and acoustic scenes can be described by a second engine. The second engine can be an audio engine, such as the MPEG-I immersive audio encoder.

[0032] Figure 2 This is a block diagram of an exemplary MPEG-I Immersive Audio Evaluation Platform (AEP) (200). Figure 2 As shown, AEP (200) may include an evaluation component (202) and an offline encoding / processing component (204). The evaluation component (202) may include Unity (212) and Max / MSP (210). Unity is a cross-platform game engine that can be used to create 3D and 2D games, as well as interactive simulations and other experiences. Max / MSP, also known as Max or Max / MSP / Jitter, is a visual programming language for, for example, music and multimedia. Max is a platform that can accommodate and connect various tools for sound, graphics, music and interactivity using a flexible patching and programming environment.

[0033] Unity (212) can display video in a head-mounted display (HMD) (214). In response to a positioning beacon, the head tracker in the HMD can connect back to the Unity engine (212), which can then send user position and orientation information to the Max / MSP (210) via Unity extOSC messages. The Max / MSP (210) can support multiple max externals running in parallel. In AEP (200), each max external can be a candidate immersive audio decoding and rendering engine running in real time. For example, such as... Figure 2 As shown, Max / MSP(210) can include renderers A, B, and C, which can run in parallel. Still... Figure 2 In this context, the evaluation component (202) may include a compiler (216) configured to compile, for example, Python programs and communicate with control and video data from controllers (218) such as smartphones and game controllers. The compiler (216) may also communicate with Max / MSP (210) via Unity extOSC messages.

[0034] In the offline encoding / processing component (204), MPEG-I immersive audio streams can be processed. MPEG-H 3D audio (206) can be a codec for all audio signals (e.g., (208)). Therefore, the audio signals can be shared by all supporting renderers (e.g., renderers A, B, and C). In the AEP (200), the audio signals can be "pre-encoded," which could mean that the raw audio signals can be encoded using MPEG-H 3D audio (206) and then decoded, and then these signals are provided to the Max / MSP (210) in the evaluation component (202) as well as to the individual Max externals (or renderers). Still refer to Figure 2In the offline encoding / processing component (204), the data in the scene file (220) and the data in the directional file (222) can be processed by the MPEG-I processor (or compressor) (228). The processed data can be further transmitted to Max / Msp (210). Furthermore, the data in the HRTF file (224) can be transmitted to the renderer of Max / Msp (210), and the data in the model file / video file (226) can be transmitted to Unity (212) for processing.

[0035] For immersive media rendering, acoustic and visual scenes can be described by different engines. For example, an acoustic scene can be described using the MPEG-I immersive audio encoder input format, while a visual scene can be described by the Unity engine. Because the acoustic and visual scenes are processed by two different engines or modules, inconsistencies may occur between them.

[0036] Figure 3 and Figure 4 Examples of inconsistencies between acoustic and visual scenes can be shown. Figure 3 An overview of the acoustic scene (300) of the test scenario in the MPEG-I immersive audio standard is shown. (As...) Figure 3 As can be seen, gaps may exist between the cubic wall elements (304) (e.g., (302)). Acoustic diffraction and obstruction effects can be evaluated by a tester around the corners and edges of the wall elements (304).

[0037] Figure 4 Showing the target Figure 3 An overview of the same test scenario from the Unity engine visual scene (400). Figure 4 As can be seen, the walls can be rendered as stone shapes with almost no gaps between the wall elements (402). This discrepancy, where visible edges in the Unity-rendered visual scene (400) do not correspond to the edges of the acoustic geometry in the acoustic scene (300), can degrade the media experience for testers (or users). In other words, the audio renderer can acoustically render the diffraction caused by the cubic wall elements in the acoustic scene (300), while the walls can display very different visual geometry in the visual scene (400). Because the wall geometry in the visual scene and the wall geometry in the acoustic scene are inconsistent, testers may be confused as the rendered audio behaves inconsistently with the visual expectations derived from the visual render.

[0038] Inconsistencies between the rendered audio and visual experiences can be caused by inconsistencies between acoustic and visual scene descriptions. This disclosure provides methods for improving the consistency between acoustic and visual scene descriptions.

[0039] One or more parameters of an object (e.g., a wall element (304)) indicated by either the description of the object in the acoustic scene or the description of the object in the visual scene can be modified based on one or more parameters of the object indicated by the other description of the object in the acoustic scene or the description of the object in the visual scene. The object's parameters can be at least one of the following: object size, object shape, object position, object orientation, object texture (or object material), etc.

[0040] In some embodiments, the acoustic scene description (or description of the acoustic scene) may be modified (or altered) to be consistent with the visual scene description (or description of the visual scene), and audio rendering may be provided based on the modified (or altered) acoustic scene description.

[0041] For example, an acoustic scene description can be examined against a visual scene description, and if one or more differences are identified between the visual and acoustic scene descriptions, or any differences in some cases, the acoustic scene description can be corrected to be consistent with the visual scene description. For example, differences can be identified based on one or more distinct parameters of the object and one or more thresholds.

[0042] When the object size in the acoustic scene description differs from the object size in the visual scene description, the object size in the acoustic scene description can be changed (or modified) based on the object size in the visual scene description. For example, the object size in the acoustic scene description can be changed to be the same as (or equal to) the object size in the visual scene description. Therefore, the object size in the acoustic scene description can be consistent with or more consistent with the object size in the visual scene description.

[0043] When the shape of an object in the acoustic scene description differs from the shape of an object in the visual scene description, the shape of the object in the acoustic scene description can be changed based on the shape of the object in the visual scene description. For example, the shape of an object in the acoustic scene description can be changed to be the same as the shape of an object in the visual scene description.

[0044] When the object position in the acoustic scene description differs from the object position in the visual scene description, the object position in the acoustic scene description can be changed based on the object position in the visual scene description. For example, the object position in the acoustic scene description can be changed to be the same as the object position in the visual scene description.

[0045] When the object orientation of an object in the acoustic scene description differs from that of an object in the visual scene description, the object orientation of the object in the acoustic scene description can be changed based on the object orientation of the object in the visual scene description. For example, the object orientation of an object in the acoustic scene description can be changed to be the same as that of an object in the visual scene description.

[0046] When the object material of an object in the acoustic scene description differs from the object material of an object in the visual scene description, the object material of the object in the acoustic scene description can be changed based on the object material of the object in the visual scene description. For example, the object material of an object in the acoustic scene description can be changed to be the same as the object material of an object in the visual scene description.

[0047] In some embodiments, if one or more differences are identified between the visual scene description and the acoustic scene description, or in some cases any differences are identified, the visual scene description can be corrected to be consistent with the acoustic scene description, and visual rendering can be provided based on the corrected visual scene description.

[0048] For example, the visual scene description can be checked against the acoustic scene description, and the visual scene description can be corrected to be consistent with or better consistent with the acoustic scene description.

[0049] When the object size in the visual scene description differs from the object size in the acoustic scene description, the object size in the visual scene description can be changed based on the object size in the acoustic scene description. For example, the object size in the visual scene description can be changed to be the same as the object size in the acoustic scene description.

[0050] When the shape of an object in the visual scene description differs from the shape of an object in the acoustic scene description, the shape of the object in the visual scene description can be changed based on the shape of the object in the acoustic scene description. For example, the shape of an object in the visual scene description can be changed to be the same as the shape of an object in the acoustic scene description.

[0051] When the object position in the visual scene description differs from the object position in the acoustic scene description, the object position in the visual scene description can be changed based on the object position in the acoustic scene description. For example, the object position in the visual scene description can be changed to be the same as the object position in the acoustic scene description.

[0052] When the object orientation of an object in the visual scene description differs from that in the acoustic scene description, the object orientation of the object in the visual scene description can be changed based on the object orientation of the object in the acoustic scene description. For example, the object orientation of an object in the visual scene description can be changed to be the same as that of the object in the acoustic scene description.

[0053] When the object material of an object in the visual scene description differs from the object material of an object in the acoustic scene description, the object material of the object in the visual scene description can be changed based on the object material of the object in the acoustic scene description. For example, the object material of an object in the visual scene description can be changed to be the same as the object material of an object in the acoustic scene description.

[0054] In some embodiments, acoustic scene descriptions and visual scene descriptions may be merged or otherwise combined to generate a unified scene description. When the acoustic scene description differs from the unified scene description, the acoustic scene description may be modified based on the unified scene description or modified to be consistent with the unified scene description, and audio rendering may apply the modified acoustic scene description. When the visual scene description differs from the unified scene description, the visual scene description may be modified based on the unified scene description or modified to be consistent with the unified scene description, and visual rendering may apply the modified visual scene description.

[0055] In an embodiment, when the object size of an object in the acoustic scene description differs from the object size of an object in the visual scene description, the unified object size in the scene description can be one of the following or based on the following: (1) the object size of the object in the acoustic scene description, (2) the object size of the object in the visual scene description, (3) the size of the intersection of the object in the acoustic scene description and the object in the visual scene description, or (4) a size based on the difference between the object size in the acoustic scene description and the object size in the visual scene description. In some examples, different weights can be applied to the sizes of objects in the acoustic scene description and the visual scene description.

[0056] In an embodiment, when the object shape in the acoustic scene description differs from the object shape in the visual scene description, the unified object shape in the scene description can be one of the following or based on the following: (1) the object shape in the acoustic scene description, (2) the object shape in the visual scene description, (3) the shape of the intersection of the object in the acoustic scene description and the object in the visual scene description, or (4) a shape based on the difference between the object shape in the acoustic scene description and the object shape in the visual scene description. In some examples, different weights can be applied to the shapes of objects in the acoustic scene description and the visual scene description.

[0057] In an embodiment, when the object position in the acoustic scene description differs from the object position in the visual scene description, the unified object position in the scene description can be one of the following or based on the following: (1) the object position in the acoustic scene description, (2) the object position in the visual scene description, or (3) a position based on the difference between the object position in the acoustic scene description and the object position in the visual scene description. In some examples, different weights can be applied to the object positions in the acoustic scene description and the visual scene description.

[0058] In an embodiment, when the object orientation of an object in the acoustic scene description differs from the object orientation of an object in the visual scene description, the unified object orientation in the scene description can be one of the following or based on the following: (1) the object orientation in the acoustic scene description, (2) the object orientation in the visual scene description, or (3) a direction based on the difference between the object orientation in the acoustic scene description and the object orientation in the visual scene description. In some examples, different weights can be applied to the object orientations of objects in the acoustic scene description and the visual scene description.

[0059] In embodiments, when the object material (e.g., object texture, object composition, or object substance) of an object in the acoustic scene description differs from the object material of an object in the visual scene description, the object material in the unified scene description can be one of the following or based on the following: (1) the object material in the acoustic scene description, (2) the object material in the visual scene description, or (3) a material based on the difference between the object material in the acoustic scene description and the object material in the visual scene description. In some examples, different weights may be applied to the materials of objects in the acoustic scene description and the visual scene description.

[0060] In some embodiments, the anchor scene description may be determined (or selected) based on either a visual scene description or an acoustic scene description. For example, the anchor scene may include anchors, which are objects that AR software can recognize and apply to the integration of the real and virtual worlds. The acoustic scene description may be modified to align with the visual scene description, or vice versa. Visual rendering or audio rendering may further be based on the selected (or determined) anchor scene description.

[0061] In one embodiment, an indication may be sent to the receiver (or client) as part of the bitstream associated with visual or audio data. This indication may specify whether the anchor scene description is based on a visual or acoustic scene description. In another embodiment, such an indication may be sent as part of system-level metadata.

[0062] In some embodiments, selection information, such as a selection message, may be signaled to the receiver (or client). The selection message may indicate how the unified scene description was generated. For example, the selection message may indicate whether the unified scene description was determined from a visual scene and / or an audio scene (or an acoustic scene). Thus, based on the selection message, the unified scene can be determined as, for example, either a visual scene or an audio scene. In other words, in some examples, either a visual scene or an audio scene can be selected as the unified scene. Visual rendering or audio rendering may be based on the selected unified scene description. For example, signaling information (e.g., the selection message) may be sent as part of the bitstream or as system-level metadata.

[0063] In this embodiment, the object size of an object in the unified scene description can be signaled as originating from either the visual or audio scene. Therefore, based on the signaling information, the object size of an object in the unified scene description can be determined as either the object size in the visual scene description or the object size in the acoustic scene description.

[0064] In this embodiment, the object shape of an object in the scene description can be signaled as originating from either a visual or audio scene. Therefore, based on the signaling information, the object shape of an object in the unified scene description can be determined as either the object shape in the visual scene description or the object shape in the acoustic scene description.

[0065] In this embodiment, the object orientation of an object in the scene description can be signaled as originating from either the visual or audio scene. Therefore, based on the signaling information, the object orientation of an object in the unified scene description can be determined as either the object orientation in the visual scene description or the object orientation in the acoustic scene description.

[0066] In this embodiment, the location of an object in the scene description can be signaled as originating from either the visual or audio scene. Therefore, based on the signaling information, the location of an object in the unified scene description can be determined either based on the location of the object in the visual scene description or the location of the object in the acoustic scene description.

[0067] In this embodiment, the object material of an object in the scene description can be signaled as originating from either a visual or audio scene. Therefore, based on the signaling information, the object material of an object in the unified scene description can be determined as either the object material of the object in the visual scene description or the object material of the object in the acoustic scene description.

[0068] Figure 5 A block diagram of a media system (500) according to an embodiment of the present disclosure is shown. The media system (500) can be used in a variety of applications, such as immersive media applications, augmented reality (AR) applications, virtual reality applications, video game applications, sports game animation applications, teleconferencing and telepresence applications, streaming media applications, etc.

[0069] The media system (500) includes a media server device (510) and multiple media client devices (such as those connected via a network (not shown)). Figure 5 The media client device (560) is shown. In this example, the media server device (510) may include one or more devices with audio and video encoding / decoding capabilities. In this example, the media server device (510) includes a single computing device, such as a desktop computer, laptop computer, server computer, tablet computer, etc. In another example, the media server device (510) includes one or more data centers, one or more server clusters, etc. The media server device (510) can receive media content data. The media content data may include video and audio content. The media content data may include descriptions of objects in an acoustic scene generated by an audio engine and descriptions of objects in a visual scene generated by a visual engine. The media server device (510) can compress the video and audio content into one or more encoded bitstreams according to a suitable media encoding / decoding standard. The encoded bitstreams can be transmitted to the media client device (560) via a network.

[0070] The media client device (560) may include one or more devices having video and audio codec capabilities for media applications. In the example, the media client device (560) may include computing devices such as desktop computers, laptop computers, server computers, tablet computers, wearable computing devices, HMD devices, etc. The media client device (560) can decode the encoded bitstream according to a suitable media codec standard. The decoded video and audio content can be used for media playback.

[0071] Media server equipment (510) can be implemented using any suitable technology. Figure 5 In the example, the media server device (510) includes processing circuitry (530) and interface circuitry (511) coupled together.

[0072] The processing circuitry (530) may include any suitable processing circuitry, such as one or more central processing units (CPUs), one or more graphics processing units (GPUs), application-specific integrated circuits (ASICs), etc. The processing circuitry (530) may be configured to include various encoders, such as audio encoders, video encoders, etc. In this example, one or more CPUs and / or GPUs may execute software to act as an audio encoder or a video encoder. In another example, an application-specific integrated circuit may be used to implement the audio encoder or the video encoder.

[0073] In some examples, the processing circuitry (530) includes a scene processor (531). The scene processor (531) can determine whether one or more parameters of an object indicated by a description of an object in an acoustic scene are different from one or more parameters of an object indicated by a description of an object in a visual scene. In response to the difference between the parameters of the object indicated by the description of an object in the acoustic scene and the parameters of the object indicated by the description of an object in the visual scene, the scene processor (531) can modify at least one of the descriptions of the object in the acoustic scene or the descriptions of the object in the visual scene such that the parameters of the object indicated by the description of the object in the acoustic scene are consistent with or more consistent with the parameters of the object indicated by the description of the object in the visual scene.

[0074] Interface circuitry (511) can connect media server device (510) to a network. Interface circuitry (511) may include a receiving section for receiving signals from the network and a transmitting section for transmitting signals to the network. For example, interface circuitry (511) can transmit signals carrying encoded bitstreams to other devices, such as media client device (560), via the network. Interface circuitry (511) can receive signals from media client devices, such as media client device (560).

[0075] The network is suitably coupled to the media server device (510) and the media client device (560) via wired and / or wireless connections (such as Ethernet, fiber optic, WiFi, cellular, etc.). The network may include network server devices, storage devices, network devices, etc. The components of the network are suitably coupled together via wired and / or wireless connections.

[0076] The media client device (560) can be configured to decode the encoded bitstream. In the example, the media client device (560) can perform video decoding to reconstruct a sequence of video frames that can be displayed, and can perform audio decoding to generate an audio signal for playback.

[0077] Media client devices (560) can be implemented using any suitable technology. Figure 5 The example shows a media client device (560), but it is not limited to an HMD with headphones as a user device that can be used by a user (520).

[0078] exist Figure 5 In this context, the media client device (560) may include, for example: Figure 5 The interface circuit (561) and processing circuit (570) are coupled together as shown.

[0079] The interface circuit (561) can connect the media client device (560) to the network. The interface circuit (561) may include a receiving section for receiving signals from the network and a transmitting section for transmitting signals to the network. For example, the interface circuit (561) can receive signals carrying data from the network, such as signals carrying encoded bitstreams.

[0080] The processing circuit (570) may include suitable processing circuitry, such as a CPU, GPU, application-specific integrated circuit, etc. The processing circuit (570) may be configured to include various components, such as a scene processor (571), a renderer (572), a video decoder (not shown), an audio decoder (not shown), etc.

[0081] In some examples, an audio decoder can decode audio content in an encoded bitstream by selecting a decoding tool suitable for the audio content encoding scheme, and a video decoder can decode video content in an encoded bitstream by selecting a decoding tool suitable for the video content encoding scheme. A scene processor (571) is configured to modify either the description of a visual scene or the description of an acoustic scene in the decoded media content. Therefore, one or more parameters of an object indicated by the description of an object in the acoustic scene are consistent with one or more parameters of an object indicated by the description of an object in the visual scene.

[0082] Furthermore, the renderer (572) can generate a final digital product suitable for a media client device (560) based on the audio and video content decoded from the encoded bitstream. It should be noted that the processing circuitry (570) may include other suitable components (not shown), such as mixers, post-processing circuitry, etc., for further media processing.

[0083] Figure 6 A flowchart outlining a process (600) according to an embodiment of the present disclosure is shown. The process (600) may be executed by a media processing device, such as a scene processor (531) in a media server device (510), a scene processor (571) in a media client device (560), etc. In some embodiments, the process (600) is implemented as software instructions, so that the processing circuitry executes the process (600) when the software instructions are executed. The process begins at (S601) and proceeds to (S610).

[0084] At (S610), media content data of the object can be received. The media content data may include a first description of the object in the acoustic scene generated by the audio engine and a second description of the object in the visual scene generated by the visual engine.

[0085] At (S620), it can be determined whether the first parameter indicated by the first description of the object in the acoustic scene is inconsistent with the second parameter indicated by the second description of the object in the visual scene.

[0086] At (S630), in response to a discrepancy between a first parameter indicated by a first description of an object in the acoustic scene and a second parameter indicated by a second description of an object in the visual scene, one of the first description of the object in the acoustic scene and the second description of the object in the visual scene may be modified based on the other of the first and second descriptions that has not been modified. The modified first and second descriptions may be consistent with the other of the first and second descriptions that has not been modified.

[0087] At (S640), the media content data of the object can be provided to the receiver, which renders the media content data of the object for the media application.

[0088] In some embodiments, the first parameter and the second parameter may both be associated with one of the object's object size, object shape, object position, object orientation, and object texture.

[0089] One of the first description of an object in the acoustic scene and the second description of an object in the visual scene can be modified based on the other. Therefore, the first parameter indicated by the first description of the object in the acoustic scene can be consistent with the second parameter indicated by the second description of the object in the visual scene.

[0090] In process (600), a third description of an object in a unified scene can be determined based on at least one of a first description of an object in an acoustic scene or a second description of an object in a visual scene. In response to a first parameter indicated by the first description of the object in the acoustic scene being different from a third parameter indicated by the third description of the object in the unified scene, the first parameter indicated by the first description of the object in the acoustic scene can be modified based on the third parameter indicated by the third description of the object in the unified scene. In response to a second parameter indicated by the second description of the object in the visual scene being different from a third parameter indicated by the third description of the object in the unified scene, the second parameter indicated by the second description of the object in the visual scene can be modified based on the third parameter indicated by the third description of the object in the unified scene.

[0091] In the examples, the object size in the third description of the unified scene can be determined based on the object size in the first description of the object in the acoustic scene. The object size in the third description of the unified scene can be determined based on the object size in the second description of the object in the visual scene. The object size in the third description of the unified scene can be determined based on the intersection size of the objects in the first description of the acoustic scene and the objects in the second description of the visual scene. The object size in the third description of the unified scene can be determined based on the size difference between the object size in the first description of the acoustic scene and the object size in the second description of the visual scene.

[0092] In the examples, the object shape in the third description of an object in a unified scene can be determined based on the object shape in the first description of the object in the acoustic scene. The object shape in the third description of an object in a unified scene can be determined based on the object shape in the second description of the object in the visual scene. The object shape in the third description of an object in a unified scene can be determined based on the intersection shape of the objects in the first description of the acoustic scene and the objects in the second description of the visual scene. The object shape in the third description of an object in a unified scene can be determined based on the shape difference between the object shape in the first description of the acoustic scene and the object shape in the second description of the visual scene.

[0093] In the examples, the object position in the third description of an object in a unified scene can be determined based on the object position in the first description of the object in the acoustic scene. The object position in the third description of an object in a unified scene can be determined based on the object position in the second description of the object in the visual scene. The object position in the third description of an object in a unified scene can also be determined based on the positional difference between the object position in the first description of the object in the acoustic scene and the object position in the second description of the object in the visual scene.

[0094] In the examples, the object orientation in the third description of an object in a unified scene can be determined based on the object orientation in the first description of the object in the acoustic scene. The object orientation in the third description of an object in a unified scene can be determined based on the object orientation in the second description of the object in the visual scene. The object orientation in the third description of an object in a unified scene can be determined based on the directional difference between the object orientation in the first description of the object in the acoustic scene and the object orientation in the second description of the object in the visual scene.

[0095] In the examples, the object texture in the third description of an object in a unified scene can be determined based on the object texture in the first description of an object in the acoustic scene. The object texture in the third description of an object in a unified scene can be determined based on the object texture in the second description of an object in the visual scene. The object texture in the third description of an object in a unified scene can be determined based on the texture difference between the object texture in the first description of an object in the acoustic scene and the object texture in the second description of an object in the visual scene.

[0096] In some embodiments, the description of an object in an anchor scene of media content data can be determined based on one of a first description of an object in an acoustic scene and a second description of an object in a visual scene. In response to determining the description of an object in the anchor scene based on the first description of the object in the acoustic scene, the second description of the object in the visual scene can be modified based on the first description of the object in the acoustic scene. Similarly, in response to determining the description of an object in the anchor scene based on the second description of the object in the visual scene, the first description of the object in the acoustic scene can be modified based on the second description of the object in the visual scene. Further, signaling information can be generated to indicate which of the first description of the object in the acoustic scene and the second description of the object in the visual scene is selected to determine the description of the anchor scene.

[0097] In some embodiments, signaling information may be generated to instruct which of the first parameter in the first description of an object in the acoustic scene and the second parameter in the second description of an object in the visual scene should be selected to determine the third parameter in the third description of an object in the unified scene.

[0098] Then, the process proceeds to (S699) and terminates.

[0099] The process (600) can be modified as appropriate. One or more steps in the process (600) can be modified and / or omitted. One or more additional steps can be added. Any suitable implementation order can be used.

[0100] The above-described technology can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 7A computer system (700) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0101] Computer software can be coded using any suitable machine code or computer language, which can be created through assembly, compilation, linking or similar mechanisms. This code includes instructions that can be executed directly or through interpretation, microcode execution or other means by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.

[0102] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0103] Figure 7 The components shown for the computer system (700) are exemplary in nature and are not intended to impose any limitation on the scope or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement on any component or combination thereof illustrated in the exemplary embodiments of the computer system (700).

[0104] The computer system (700) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users via, for example, tactile input (such as keystrokes, swipes, or movement of a data glove), audio input (such as speech or tapping), visual input (such as gestures), or olfactory input (not depicted). The human-machine interface device may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (such as speech, music, or ambient sound), images (such as scanned images or photographic images obtained from a still image camera), and video (such as two-dimensional video or three-dimensional video, including stereoscopic video).

[0105] The input human-machine interface device may include one or more of the following (only one of each is depicted): keyboard (701), mouse (702), touchpad (703), touch screen (710), data glove (not shown), joystick (705), microphone (706), scanner (707), camera (708).

[0106] The computer system (700) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., touch screen (710), data gloves (not shown), or tactile feedback of a joystick (705), but may also include tactile feedback devices that are not used as input devices), audio output devices (such as speakers (709), headphones (not depicted)), visual output devices (such as screens (710), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capability, each with or without tactile feedback capability—some of which are capable of outputting two-dimensional or more than three-dimensional visual output in a manner such as stereoscopic output; virtual reality glasses (not depicted), holographic displays, and smoke canisters (not depicted)), and printers (not depicted).

[0107] The computer system (700) may also include human-accessible storage devices and their associated media, such as optical media (721) including media such as CD / DVD ROM / RW (720) with CD / DVD, thumb drives (722), removable hard disk drives or solid-state drives (723), conventional magnetic media such as magnetic tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.

[0108] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other volatile signals.

[0109] The computer system (700) may also include an interface (754) to one or more communication networks (755). The network may be, for example, wireless, wired, or optical. The network may further be local, wide area, metropolitan area, vehicular and industrial, real-time, latency-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require external network interface adapters (such as USB ports on the computer system (700)) attached to certain general-purpose data ports or peripheral buses (749); other networks are typically integrated into the core of the computer system (700) by attaching to system buses as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (700) can communicate with other entities. Such communication can be unidirectional (receive-only, e.g., broadcasting TV), unidirectional (transmit-only, e.g., CANbus to certain CANbus devices), or bidirectional, e.g., to other computer systems using local or wide area digital networks. Certain protocols and protocol stacks can be used on each of those networks and network interfaces as described above.

[0110] The aforementioned human-machine interface devices, human-accessible storage devices, and network interfaces can be attached to the core (740) of the computer system (700).

[0111] The core (740) may include one or more central processing units (CPUs) (741), graphics processing units (GPUs) (742), dedicated programmable processing units (743) in the form of field-programmable gate arrays (FPGAs), hardware accelerators (744) for certain tasks, graphics adapters (750), etc. These devices, along with read-only memory (ROM) (745), random access memory (746), and internal mass storage (747) such as internal non-user-accessible hard disk drives, SSDs, etc., may be connected via a system bus (748). In some computer systems, the system bus (748) may be accessed as one or more physical plugs to allow for expansion by adding CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus (749) to the core's system bus (748). In the example, a screen (710) may be connected to a graphics adapter (750). Peripheral bus architectures include PCI, USB, etc.

[0112] The CPU (741), GPU (742), FPGA (743), and accelerator (744) can execute certain instructions, and combinations of these instructions can constitute the aforementioned computer code. This computer code can be stored in ROM (745) or RAM (746). Transient data can also be stored in RAM (746), while permanent data can be stored, for example, in internal mass storage (747). Fast storage and retrieval of any of the memory devices can be enabled by using cache memory, which can be closely associated with one or more CPUs (741), GPUs (742), mass storage (747), ROM (745), RAM (746), etc.

[0113] Computer-readable media may have computer code thereon for performing various computer-implemented operations. The media and computer code may be those specifically designed and constructed for the purposes of this disclosure, or they may be of types well known and available to those skilled in the art of computer software.

[0114] By way of example and not limitation, a computer system (700) having an architecture, and in particular a core (740), can provide functionality as a result of one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, as well as some memory of the core (740) having non-volatile properties, such as internal mass storage (747) or ROM (745). Software implementing various embodiments of this disclosure can be stored in such a device and executed by the core (740). Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can cause the core (740) and, in particular, the processors therein (including CPUs, GPUs, FPGAs, etc.) to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (746) and modifying such data structures according to software-defined processes. Alternatively or as an alternative, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., an accelerator (744)) that may replace or operate with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may cover logic, and vice versa. Where appropriate, references to computer-readable media may cover circuitry storing software for execution (such as integrated circuits (ICs)), circuitry embodying logic for execution, or both. This disclosure covers any suitable combination of hardware and software.

[0115] While several exemplary embodiments have been described in this disclosure, there are changes, substitutions, and various alternative equivalents that fall within the scope of this disclosure. Therefore, it should be understood that those skilled in the art will be able to design numerous systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and are therefore within its spirit and scope.

Claims

1. A method for performing media processing at a media processing device, characterized in that, include: The media content data of the received object includes a first description of the object in an acoustic scene and a second description of the object in a visual scene. The acoustic scene is generated by an audio engine, and the visual scene is generated by a visual engine. Determine whether a first parameter indicated by a first description of the object in the acoustic scene is inconsistent with a second parameter indicated by a second description of the object in the visual scene; In response to a discrepancy between the first parameter indicated by a first description of the object in the acoustic scene and the second parameter indicated by a second description of the object in the visual scene, one of the first description of the object in the acoustic scene and the second description of the object in the visual scene is modified, the modification being based on the other of the first and second descriptions that has not been modified, wherein the modified first and second descriptions are consistent with the other of the first and second descriptions that has not been modified; both the first parameter and the second parameter are associated with one of the object's object size, object shape, object position, object orientation, and object texture; as well as The media content data of the object is provided to the receiver, and the receiver renders the media content data of the object for a media application. The method further includes: A third description of the object in a unified scene is determined based on at least one of a first description of the object in the acoustic scene or a second description of the object in the visual scene; In response to a first parameter indicated by a first description of the object in the acoustic scene being different from a third parameter indicated by a third description of the object in the unified scene, the first parameter indicated by the first description of the object in the acoustic scene is modified based on the third parameter indicated by the third description of the object in the unified scene; and In response to the second parameter indicated by the second description of the object in the visual scene being different from the third parameter indicated by the third description of the object in the unified scene, the second parameter indicated by the second description of the object in the visual scene is modified based on the third parameter indicated by the third description of the object in the unified scene.

2. The method according to claim 1, characterized in that, The modifications include: The first description of the object in the acoustic scene and the second description of the object in the visual scene are modified based on the other of the first description of the object in the acoustic scene and the second description of the object in the visual scene, such that the first parameter indicated by the first description of the object in the acoustic scene is consistent with the second parameter indicated by the second description of the object in the visual scene.

3. The method according to claim 1, characterized in that, The determination further includes at least one of the following: The object size in the third description of the object in the unified scene is determined based on the object size in the first description of the object in the acoustic scene; The object size in the third description of the object in the unified scene is determined based on the object size in the second description of the object in the visual scene; The object size in the third description of the object in the unified scene is determined based on the intersection size between the object in the first description of the object in the acoustic scene and the object in the second description of the object in the visual scene; and The object size in the third description of the object in the unified scene is determined based on the size difference between the object size in the first description of the object in the acoustic scene and the object size in the second description of the object in the visual scene.

4. The method according to claim 1, characterized in that, The determination further includes at least one of the following: The object shape in the third description of the object in the unified scene is determined based on the object shape in the first description of the object in the acoustic scene; The object shape in the third description of the object in the unified scene is determined based on the object shape in the second description of the object in the visual scene; The object shape in the third description of the object in the unified scene is determined based on the intersection shape between the object in the first description of the object in the acoustic scene and the object in the second description of the object in the visual scene; as well as The object shape in the third description of the object in the unified scene is determined based on the shape difference between the object shape in the first description of the object in the acoustic scene and the object shape in the second description of the object in the visual scene.

5. The method according to claim 1, characterized in that, The determination further includes at least one of the following: The object position in the third description of the object in the unified scene is determined based on the object position in the first description of the object in the acoustic scene; The object position in the third description of the object in the unified scene is determined based on the object position in the second description of the object in the visual scene; as well as The object position in the third description of the object in the unified scene is determined based on the positional difference between the object position in the first description of the object in the acoustic scene and the object position in the second description of the object in the visual scene.

6. The method according to claim 1, characterized in that, The determination further includes at least one of the following: The object orientation in the third description of the object in the unified scene is determined based on the object orientation in the first description of the object in the acoustic scene; The object orientation in the third description of the object in the unified scene is determined based on the object orientation in the second description of the object in the visual scene; as well as The object orientation in the third description of the object in the unified scene is determined based on the directional difference between the object orientation in the first description of the object in the acoustic scene and the object orientation in the second description of the object in the visual scene.

7. The method according to claim 1, characterized in that, The determination further includes at least one of the following: The object texture in the third description of the object in the unified scene is determined based on the object texture in the first description of the object in the acoustic scene; The object texture in the third description of the object in the unified scene is determined based on the object texture in the second description of the object in the visual scene; as well as The object texture in the third description of the object in the unified scene is determined based on the texture difference between the object texture in the first description of the object in the acoustic scene and the object texture in the second description of the object in the visual scene.

8. The method according to claim 1, characterized in that, Further includes: The description of the object in the anchor scene of the media content data is determined based on one of the first description of the object in the acoustic scene and the second description of the object in the visual scene; In response to determining the description of the object in the anchor scene based on the first description of the object in the acoustic scene, the second description of the object in the visual scene is modified based on the first description of the object in the acoustic scene; In response to determining the description of the object in the anchor scene based on the second description of the object in the visual scene, the first description of the object in the acoustic scene is modified based on the second description of the object in the visual scene; as well as Generate signaling information indicating which of the first description of the object in the acoustic scene and the second description of the object in the visual scene should be selected to determine the description of the anchor scene.

9. The method according to claim 1, characterized in that, Further includes: Generate signaling information indicating which of the first parameter in the first description of the object in the acoustic scene and the second parameter in the second description of the object in the visual scene should be selected to determine the third parameter in the third description of the object in the unified scene.

10. An apparatus for media processing, characterized in that, include: Processor and memory, The memory is used to store program instructions; The processor is used to invoke the program instructions stored in the memory to implement the method according to any one of claims 1-9.

11. A non-volatile computer-readable storage medium for storing instructions, characterized in that, When executed by a computer for video decoding, the instructions cause the computer to perform the method as described in any one of claims 1 to 9.