Audio processing method and apparatus

By real-time adjustment of audio stream metadata in the decoding device, the problem that users cannot dynamically adjust the attributes of audio stream objects is solved, and the flexibility and user experience of audio output are improved.

WO2025066533A9PCT designated stage expired Publication Date: 2025-05-22HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/109053
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-28
Filing Date
2024-07-31
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

The object-based audio stream fixes the object's channel and other attributes during the code stream production process, and users cannot dynamically adjust the attributes of the object in the audio stream according to their subjective feelings, resulting in poor audio stream output flexibility.

Method used

It provides an audio processing method, which obtains metadata in the audio stream through the decoding device and displays a control group on the interactive interface, allowing users to adjust the attributes of the object's channel and other attributes according to their needs, generates metadata change information, and applies these changes during the decoding process to realize real-time adjustment of the audio stream.

Benefits of technology

Users can dynamically adjust the attributes of objects in the audio stream based on subjective feelings, which improves the flexibility of audio output without modifying the code stream production process, which reduces the requirements for user professionalism and reduces the overhead of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024109053_22052025_PF_FP_ABST
    Figure CN2024109053_22052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of encoding and decoding, and discloses an audio processing method and apparatus. A decoding device acquires an audio stream comprising metadata for describing object information of an object, and when decoding the audio stream, adjusts the metadata of the audio stream in real time on the basis of the change of the metadata by a user, so as to control the audio output of the audio stream on the basis of the metadata that is adjusted in real time. Therefore, the user does not need to change the metadata of the object in the audio stream in a code stream manufacturing process, and can also adjust the sound channel and other attributes of any object in the audio stream in real time on the basis of requirements, thereby implementing audio output meeting the user requirements, and improving the flexibility of the audio output.
Need to check novelty before this filing date? Find Prior Art

Description

Audio processing method and device

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on September 28, 2023, with application number 202311290467.X and application name “Audio Processing Method and Device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of coding and decoding, and in particular to an audio processing method and device. Background Art

[0003] Object-based audio streaming decomposes the audio signal into multiple independent objects within the audio stream, such as a singer's voice at a concert or the engine sound of a car on a racetrack. Each object contains a portion of the audio signal, such as the location, direction, and distance of the sound source. These objects can be independently encoded, transmitted, and processed, enabling more flexible and efficient audio processing and transmission.

[0004] When outputting an object-based audio stream, objects are rendered to different channels, allowing multi-channel audio to be played through speakers on different channels. However, the object rendering method in the audio stream is fixed during the audio stream encoding process. Users cannot dynamically adjust attributes such as the channels of the objects in the audio stream based on their subjective preferences, resulting in limited flexibility in audio stream output.

[0005] Summary of the Invention

[0006] The present application provides an audio processing method and apparatus, which solves the problem that users cannot adjust the object properties of object-based audio streams and the audio stream output flexibility is poor.

[0007] In a first aspect, an audio processing method is provided. The audio processing method is applied to a codec system, such as when the audio processing method is executed by a decoding device in the codec system. The audio processing method includes: first, the decoding device obtains an audio stream, the audio stream includes metadata, and the metadata is used to describe object information in the audio stream. Then, the decoding device displays a control group on an interactive interface based on the metadata, and generates metadata change information in response to an operation on the control group. When decoding the audio stream, the decoding device changes the metadata of the audio stream according to the metadata change information to obtain a decoding result, which includes the changed metadata and the audio data carried by the audio stream. Finally, the decoding device outputs the audio corresponding to the audio data according to the changed metadata.

[0008] In this embodiment, the decoding device provides an interactive interface to the user, displaying a set of controls. Based on the user's actions on the controls, the device modifies the metadata of objects during the audio stream decoding process. This eliminates the need for users to modify metadata during the stream production process. Users can adjust attributes such as the channel of any object in the audio stream in real time, achieving audio output that meets their needs and increasing audio output flexibility.

[0009] As a possible implementation, the audio stream includes at least one object. The metadata includes at least one object parameter, where an object parameter is used to describe object information of an object. The object parameter includes at least one of an object name, an object volume, a mute state, and an object sound direction. The control group includes a first control group and a second control group.

[0010] Optionally, when displaying the interactive interface, the decoding device displays a first control group based on the name of at least one object parameter, then determines a target object parameter in response to a first operation on the first control group, and then displays a second control group based on the target object parameter. One control in the second control group is used to adjust the object volume, mute state, or object sound direction in response to the second operation, thereby generating metadata change information.

[0011] For example, the first control group includes at least one control, one control is displayed as the name of an object parameter, and each control is used to determine the target object parameter selected by the user in response to the user's first operation.

[0012] For another example, the second control group includes at least one control, each control being used to adjust the object volume, mute state, or object sound direction corresponding to the target object parameter in response to the user's second operation, and generate metadata change information.

[0013] In the above implementation, the decoding device provides the user with controls for each object in the audio stream and different types of object parameters of each object in the interactive interface, thereby improving the user's flexibility in adjusting the audio stream based on the object.

[0014] As a possible implementation manner, taking the target object parameter among the at least one object parameter as an example, the metadata change information includes first metadata change information, and the first metadata change information includes the changed target object parameter.

[0015] Optionally, when decoding an audio stream, the decoding device parses the audio stream into object information, sound bed information, and metadata information, and then replaces the target object parameters in the metadata information with the changed target object parameters to obtain the changed metadata encoding, and finally outputs the decoding result. Among them, the changed metadata information is used to represent the changed metadata, and the decoding result includes object information, sound bed information, and changed metadata information. In this way, when decoding the audio stream, the encoding device changes the metadata obtained by decoding the audio stream according to the metadata change information. The user does not need to modify the metadata of the objects in the audio stream during the code stream production process, which reduces the professional requirements for the user and improves the versatility of the method. At the same time, compared with encapsulating the metadata change information into the audio stream to indicate the parsing of the audio stream, the metadata of the objects in the audio stream is adjusted during the parsing process of the audio stream, and there is no need to encapsulate the metadata change information, which reduces the bandwidth and computing resources required for audio stream parsing.

[0016] Optionally, before changing the metadata information, the decoding device determines the target object parameters in the metadata information according to the index of the target object parameters, thereby achieving precise adjustment for the object.

[0017] As a possible implementation manner, the decoding device can enable the user to adjust different objects based on the group division of objects, thereby providing the user with more audio output modes.

[0018] Optionally, the object parameters also include the group to which the object belongs and the group members. The decoding device displays a third control group on the interactive interface based on the metadata, and in response to a third user operation on the third control group, determines a target group and generates second metadata modification information. Thus, when decoding the audio stream, metadata of the audio stream is modified based on the second metadata modification information. The second metadata modification information includes at least one modified object parameter, wherein the object volume of the at least one modified object parameter belonging to an object outside the target group is minimized.

[0019] In the above implementation, the decoding device batch adjusts the object parameters of multiple objects in the audio stream based on the user's selection of the group, providing the user with a simple way to switch between different audio output modes and reducing the operation complexity.

[0020] As one possible implementation, the object parameters also include mutually exclusive objects. After the decoding device determines the target object parameters in response to a first operation on the first control group, it sets the object volume of the object parameters of mutually exclusive objects of the target object parameters to a minimum value. In this way, when a user selects to output the audio of a target object among multiple objects, the object diversion of mutually exclusive objects that interfere with the target object is set to a minimum value. This eliminates the need for the user to manually select and adjust parameters such as the object volume of mutually exclusive objects, simplifying user operations.

[0021] As a possible implementation method, the decoding device updates the interactive interface in real time according to different frames of the audio stream, so that the user can adjust the object parameters more accurately and quickly according to the changes in the preset object parameters of the object in the audio stream, thereby improving the real-time performance of the object parameter adjustment.

[0022] As a possible implementation method, when the decoding device outputs audio based on the decoding result, it mixes the object information to the corresponding sound bed according to the constraints of the changed metadata to obtain the original multi-channel encoded data, and then renders the original multi-channel encoded data with sound effects to obtain the rendered multi-channel encoded data, and finally outputs the audio based on the rendered multi-channel encoded data.

[0023] As a possible implementation manner, the object sound direction is represented by three-dimensional coordinates.

[0024] In a second aspect, an audio processing device is provided, which includes a module for implementing the audio processing method provided by any one of the implementation methods in the first aspect. Exemplarily, the audio processing device includes a code stream module, a display module, an adjustment module, and a rendering module. The code stream module is used to obtain an audio stream, and the audio stream includes metadata, and the metadata is used to describe object information in the audio stream. The display module is used to display a control group on an interactive interface based on the metadata, and to generate metadata change information in response to an operation on the control group. The adjustment module is used to change the metadata according to the metadata change information when decoding the audio stream, and obtain a decoding result, and the decoding result includes the changed metadata and the audio data carried by the audio stream. The rendering module is used to output audio corresponding to the audio data according to the changed metadata.

[0025] As a possible implementation, the audio stream includes at least one object, the metadata includes at least one object parameter, an object parameter is used to describe the object information of an object, and the object parameter includes at least one of the object name, object volume, mute state and object sound direction.

[0026] Optionally, the display module is specifically used to: display a first control group based on the name of at least one object parameter; determine the target object parameter in response to a first operation on the first control group; display a second control group based on the target object parameter, and a control in the second control group is used to adjust the object volume, mute state or object sound direction in response to the second operation, and generate metadata change information.

[0027] As a possible implementation manner, the metadata change information includes first metadata change information, and the first metadata change information includes the changed target object parameters.

[0028] Optionally, the adjustment module is specifically used to: when decoding the audio stream, parse the audio stream into object information, sound bed information and metadata information; replace the target object parameters in the metadata information with the changed target object parameters to obtain the changed metadata information, and the changed metadata information is used to represent the changed metadata; output the decoding result, which includes the object information, sound bed information and the changed metadata information.

[0029] As a possible implementation, the first metadata change information further includes an index of the target object parameter. The adjustment module is further configured to determine the target object parameter in the metadata information according to the index of the target object parameter.

[0030] As a possible implementation, the object parameters also include the group to which the object belongs and the group members. The display module is further configured to: display a third control group on the interactive interface based on the metadata; and in response to a third operation on the third control group, determine a target group and generate second metadata change information, the second metadata change information including at least one modified object parameter, wherein the object volume of the at least one modified object parameter belonging to an object outside the target group is minimized.

[0031] As a possible implementation manner, the object parameters further include mutually exclusive objects, and the adjustment module is further configured to: set the object volume of the object parameter of the mutually exclusive object of the target object parameter to a minimum value.

[0032] As a possible implementation method, the display module is also used to: when the metadata of the second frame of the audio stream is different from the metadata of the first frame, update the control group based on the metadata of the second frame, the second frame is the next n frames of the first frame, and n is a positive integer.

[0033] As a possible implementation method, the rendering module is specifically used to: mix the object information to the corresponding sound bed according to the constraints of the changed metadata to obtain the original multi-channel encoded data; render the original multi-channel encoded data with sound effects to obtain the rendered multi-channel encoded data; and output audio according to the rendered multi-channel encoded data.

[0034] As a possible implementation manner, the object sound direction is represented by three-dimensional coordinates.

[0035] In a third aspect, an electronic device is provided, comprising a memory and a processor, wherein the memory is used to store a set of computer instructions, and when the processor executes the set of computer instructions, it is used to execute the operating steps of the audio processing method in the first aspect or any possible implementation of the first aspect.

[0036] In a fourth aspect, a coding and decoding system is provided, comprising an encoding device and a decoding device, wherein the encoding device is communicatively connected to the decoding device, the encoding device is used to generate an audio stream, the audio stream comprising metadata, the metadata being used to describe object information in the audio stream, and the decoding device is used to execute the operating steps of the audio processing method in the first aspect or any possible implementation of the first aspect.

[0037] In addition, the technical effects of the audio processing device described in the second aspect, the electronic device described in the third aspect, and the encoding and decoding system described in the fourth aspect can refer to the technical effects of the obstacle detection method described in the first aspect, and will not be repeated here.

[0038] In a fifth aspect, a readable storage medium is provided, which includes a computer program or instructions that, when executed on a computer, causes the computer to execute the audio processing method described in any possible implementation of the first aspect.

[0039] In a sixth aspect, a computer program product is provided, which includes a computer program or instructions, and when the computer program or instructions are executed on a computer, causes the computer to execute the audio processing method described in any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] FIG1 is a schematic diagram of the principle of Dolby Atmos;

[0041] FIG2 is a schematic diagram of the principle of immersive sound;

[0042] FIG3 is a schematic structural diagram of an audio transmission system provided by the present application;

[0043] FIG4 is a schematic diagram of an audio encoding and decoding system provided by the present application;

[0044] FIG5 is a flow chart of an audio processing method provided by the present application;

[0045] FIG6 is a schematic diagram of an interactive interface provided by this application;

[0046] FIG7 is a schematic diagram of a group interaction interface provided by the present application;

[0047] FIG8a is a schematic diagram of another group interaction interface provided by the present application;

[0048] FIG8b is a schematic diagram of another group interaction interface provided by the present application;

[0049] FIG9 is a schematic diagram of an object provided by the present application that does not support sound direction adjustment;

[0050] FIG10 is a schematic diagram of an object-supported sound direction adjustment provided by the present application;

[0051] FIG11 is a schematic diagram of another embodiment provided by the present application that does not support sound direction adjustment;

[0052] FIG12 is a schematic diagram of another object-supported sound direction adjustment provided by the present application;

[0053] FIG13 is a schematic structural diagram of an audio processing device provided by the present application;

[0054] FIG14 is a schematic structural diagram of another audio processing device provided by the present application;

[0055] FIG15 is a schematic structural diagram of an electronic device provided in this application. DETAILED DESCRIPTION

[0056] The audio processing method and apparatus provided in the embodiments of the present application can be applied to multi-channel spatial audio playback scenarios. A brief introduction to the practical technologies of the present application is given below.

[0057] (1) Spatial Audio

[0058] Spatial audio can refer to the audio specifications of three-dimensional sound. Compared to traditional multi-channel audio, spatial audio has more channels than stereo and can also accommodate object-based sound information. Object-based sound information accurately records the position (static) and displacement (dynamic) of the source point relative to the listener in coordinate form. On playback devices, combined with local or cloud-based codecs, the amplitude-frequency response (timbre) and phase-frequency response (spatial perception) of an object in a channel or sound bed can be accurately calculated and played back.

[0059] In audio production and playback, a channel generally refers to independent audio signal paths captured or played back at different spatial locations. A sound bed is a production-side term for the corresponding channel. For example, in a 7.1.4-channel audio system using object mixing, any audio track of an object is assigned to a 7.1.4 sound bed channel based on the object's three-dimensional coordinates, thereby controlling the direction of the sound image in three-dimensional space.

[0060] An object is a recording of audio and video information in the form of three-dimensional coordinates. Each object contains a portion of the audio signal, such as a child crying, the commentary in a football game, or the cheers of the audience in a basketball game.

[0061] Spatial audio also includes object metadata. Metadata describes the object information in the audio stream, including its volume, position, and angle. This metadata can be adjusted during the stream production process to achieve better audiovisual effects.

[0062] Spatial audio technology also includes a rendering part, in which the renderer renders the objects in the audio. During the rendering process, the objects are rendered to speakers of different channels according to the metadata of the objects, restoring the intention of the stream producer or the on-site listening experience of the sound recording, allowing users to have an immersive audio experience.

[0063] (2) Codec

[0064] A codec is a general term for encoders and decoders. A codec is a device or program that transforms a signal or data stream. This transformation includes encoding a signal or data stream (usually for transmission, storage, or encryption) or extracting a coded stream, as well as restoring the coded stream to a form suitable for observation or manipulation.

[0065] For example, an audio encoder compresses a digital audio signal sampled one by one into a smaller data stream for easier transmission or storage. Alternatively, an audio decoder converts the compressed data stream back into the original digital audio signal for easier playback and processing.

[0066] Existing object-based audio stream processing includes standards such as Dolby Atmos and DTS:X.

[0067] Dolby Atmos is a widely used object-based spatial audio technology that supports playback systems with up to 128 independent audio tracks and 64 speakers. Applied to a sound bed, Dolby Atmos introduces the concept of sound objects, each of which has its own metadata, recording the object's X, Y, and Z coordinates, as well as its volume.

[0068] As shown in Figure 1, Dolby Atmos includes steps such as encoding, decoding / rendering, rendering / post-processing, and speaker output.

[0069] However, after the Dolby Atmos bitstream is produced, the metadata information remains fixed. Users cannot dynamically adjust the properties of the sound object based on their subjective feelings. There are also strict requirements on the number and arrangement of hardware such as speakers. It is necessary to match the metadata requirements during bitstream production to better restore the multi-channel playback effect.

[0070] Like Dolby Atmos, immersive audio is object-based, but it doesn't strictly calibrate the number of speakers. Playback rendering can be freely adapted to the number and placement of speakers. As shown in Figure 2, users can modify the metadata configuration of the source device's bitstream using a user-configured device such as a remote control. The adjusted metadata is then sent from the source device to the sink device, which decodes and renders it before outputting it through the speakers.

[0071] While immersive audio allows users to modify object properties, this requires re-encoding the bitstream. This re-encoding operation is typically performed on the production side, which is complex and requires a high level of expertise. Furthermore, the re-encoding process takes a long time from modifying an object's properties to outputting them to the speakers, hindering real-time adjustments to audio object properties.

[0072] The present application provides an audio processing method, specifically an audio processing method that "supports real-time modification of metadata of objects". First, a decoding device obtains an audio stream, and the audio stream includes metadata, and the metadata is used to describe the object information in the audio stream. Then, the decoding device displays a control group on an interactive interface based on the metadata, and generates metadata change information in response to the operation of the control group. When decoding the audio stream, the decoding device changes the metadata of the audio stream according to the metadata change information to obtain a decoding result, and the decoding result includes the changed metadata. Finally, the decoding device outputs the decoding result in audio according to the changed metadata. In this way, the user can change the metadata of the object in real time during the decoding process of the audio stream based on the interactive interface, without having to modify the metadata of the object in the audio stream in the code stream production process. It can also adjust the properties such as the channel of any object in the audio stream in real time as needed to achieve audio output that meets user needs, thereby improving the flexibility of audio output.

[0073] The implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0074] FIG3 is a schematic structural diagram of an audio transmission system provided by this application.

[0075] Figure 3 is a schematic diagram of an audio transmission system provided in an embodiment of the present application. The audio processing process includes audio acquisition, audio encoding, audio transmission, audio decoding and playback. The audio transmission system includes multiple terminal devices (such as terminal devices 311 to terminal devices 315 shown in Figure 3) and a network, wherein the network can realize the function of audio transmission. The network may include one or more network devices, which may be routers or switches, etc.

[0076] The terminal device shown in FIG3 may be, but is not limited to, a user equipment (UE), a mobile station (MS), a mobile terminal (MT), etc. The terminal device may be a mobile phone (such as the terminal device 314 shown in FIG3 ), a tablet computer, a computer with wireless transceiver function (such as the terminal device 315 shown in FIG3 ), a virtual reality (VR) terminal device (such as the terminal device 313 shown in FIG3 ), an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in a smart city, a wireless terminal in a smart home, etc.

[0077] As shown in Figure 3, the terminal devices are different in different audio processing processes. For example, in the audio acquisition process, the terminal device 311 can be a speaker with audio recording function, a camera (such as a video camera, a camera, etc.), a mobile phone, a tablet computer or a smart wearable device with audio acquisition function, etc. For another example, in the audio encoding process, the terminal device 312 can be a server or a data center, which can include one or more physical devices with encoding functions, such as a server, a mobile phone, a tablet computer or other encoding devices. For another example, in the audio decoding and playback process, the terminal device 313 can be a television, and the user can modify the metadata of the objects in the audio stream by selecting the control group of the interactive interface through the remote control; the terminal device 314 can be a mobile phone, and the user can modify the metadata of the objects in the audio stream by controlling the control group of the interactive interface through touch operation or air operation; the terminal device 315 can be a personal computer, and the user can modify the metadata of the objects in the audio stream by controlling the control group of the interactive interface through input devices such as a mouse or keyboard.

[0078] As you can understand, audio is a general term that includes a sequence of multiple consecutive frames, where one frame corresponds to a fixed length of audio data, for example, one frame corresponds to a 26 millisecond audio data segment.

[0079] Figure 3 is only a schematic diagram, and the audio transmission system may also include other devices, which are not shown in Figure 3. The embodiments of the present application do not limit the number and type of terminal devices included in the system.

[0080] Based on the audio transmission system shown in Figure 3, Figure 4 is a schematic diagram of an audio codec system provided in an embodiment of the present application. The audio codec system 400 includes an encoding device 410 and a decoding device 420. The encoding device 410 establishes a communication connection with the decoding device 420 through a communication channel 430.

[0081] The encoding device 410 described above can implement the audio encoding function. As shown in FIG3 , the encoding device 410 can be a terminal device 312 . The encoding device 410 can also be a data center with audio encoding capabilities, for example, the data center includes multiple servers.

[0082] The encoding device 410 may include a data source 411 , a pre-processing module 412 , an encoder 413 , and a communication interface 414 .

[0083] Data source 411 may include or may be any type of electronic device for collecting audio, and / or any type of source audio generating device, such as a computer sound card for generating computer animation scenes, or any type of device for acquiring and / or providing source audio, or computer-generated source audio. Data source 411 may be any type of memory or storage for storing the source audio. The source audio may include multiple audio streams collected by multiple audio collection devices (e.g., microphones, speakers, etc.).

[0084] The pre-processing module 412 is used to receive source audio and pre-process the source audio to obtain multi-channel audio (such as spatial audio). For example, the pre-processing performed by the pre-processing module 412 may include audio format conversion, audio splicing, etc.

[0085] The encoder 413 is used to receive audio and encode the audio to obtain encoded data. For example, the encoder 413 encodes the audio to obtain a code stream (also called an audio stream). In some optional situations, the encoded code stream can also be called a bit stream.

[0086] The communication interface 414 in the encoding device 410 can be used to: receive encoded data and send the encoded data (or a version of the encoded data after any other processing) to another device such as the decoding device 420 or any other device through the communication channel 430 for storage, playback or direct reconstruction, etc.

[0087] Optionally, the encoding device 410 includes a bitstream buffer, which is used to store bitstreams corresponding to one or more encoding units.

[0088] The above-mentioned decoding device 420 can realize the function of audio decoding. As shown in FIG3 , the decoding device 420 can be any one of the terminal devices 313 to 315 shown in FIG3 .

[0089] The decoding device 420 may include a playback device 421 , a post-processing module 422 , a decoder 423 and a communication interface 424 .

[0090] The communication interface 424 in the decoding device 420 is used to receive the encoded data (or a version of the encoded data after any other processing has been performed on the encoded data) from the encoding device 410 or any other encoding device such as a storage device.

[0091] Communication interface 414 and communication interface 424 can be used to send or receive encoded data through a direct communication link between the encoding device 410 and the decoding device 420, such as a direct wired or wireless connection, or through any type of network, such as a wired network, a wireless network or any combination thereof, any type of private network and public network or any combination thereof.

[0092] The communication interface 424 corresponds to the communication interface 414 and can be used, for example, to receive transmission data and process the transmission data using any type of corresponding transmission decoding or processing and / or decapsulation to obtain encoded data.

[0093] Both the communication interface 424 and the communication interface 414 can be configured as a unidirectional communication interface as indicated by the arrow pointing from the encoding device 410 to the corresponding communication channel 430 of the decoding device 420 in Figure 4, or a bidirectional communication interface, and can be used to send and receive messages, etc. to establish a connection, confirm and exchange any other information related to the communication link, or data transmission such as encoded compressed data transmission, etc.

[0094] The decoder 423 is used to receive encoded data and decode the encoded data to obtain decoded data (audio, etc.).

[0095] The post-processing module 422 is used to post-process the decoded data to obtain post-processed data (such as audio to be played). The post-processing performed by the post-processing module 422 may include sound bed rendering, sound effect rendering, etc., or any other processing used to generate data for playback by the playback device 421, etc.

[0096] The playback device 421 is used to receive the post-processed data for playback to a user or viewer. The playback device 421 can be or include any type of display screen and speaker, for example, a display screen with an integrated speaker. For example, the display screen can include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS) display, a digital light processor (DLP), or any other type of display screen. For another example, the speaker can include a multi-channel speaker of an electrodynamic, electrostatic, electromagnetic, or piezoelectric type.

[0097] As an optional implementation, the encoding device 410 and the decoding device 420 may transmit the encoded data via a data forwarding device, such as a router or a switch.

[0098] The implementation method of this application is described in detail below with reference to the accompanying drawings.

[0099] Here, the audio processing method of an embodiment of the present application is executed by the decoding device 420 shown in Figure 4 as an example for explanation. Figure 5 is a flow chart of an audio processing method provided by an embodiment of the present application, and the audio processing method includes the following steps 510 to 550.

[0100] Step 510: The decoding device 420 obtains the audio stream.

[0101] In one possible example, an audio stream includes data of at least one object. The audio stream also includes metadata, which is used to describe object information of multiple objects in the audio stream. The objects here can refer to, but are not limited to, sound objects. For example, the audio signal in the audio stream is decomposed into multiple independent objects, such as the voice of a singer in a concert, the engine sound of a car on a racetrack, etc. Each object contains part of the audio signal information, such as the location, direction, and distance of the sound source. These objects can be independently encoded, transmitted, and processed, thereby achieving more flexible and efficient audio processing and transmission.

[0102] As a possible implementation, the metadata includes at least one object parameter, where an object parameter is used to describe object information of an object, and the object parameter includes at least one field of object name, object volume, mute state, and object sound direction.

[0103] The object volume indicates the volume of the object. The mute state indicates whether the object is mute; a mute object has the minimum volume. The object sound direction indicates the object's sound position. For example, the object sound direction uses a Cartesian coordinate system to represent the object's sound position.

[0104] This application does not limit the format of the audio stream. For example, the audio stream can be in avs3-p3 format, etc.

[0105] Step 520: The decoding device 420 displays the control group on the interactive interface based on the metadata.

[0106] In one possible example, the control group includes a first control group and a second control group, the first control group is used to determine the target object parameters in response to the user's first operation, and the second control group is used to adjust the value of at least one field in the target object parameters in response to the user's second operation to generate metadata change information.

[0107] As a possible implementation, the decoding device 420 displays a first control group based on the name of at least one object parameter, determines a target object parameter in response to a user's first operation on the first control group, and then displays a second control group based on the target object parameter.

[0108] Optionally, the target object parameter is carried in metadata of the audio stream, and the metadata may be obtained by parsing the audio stream by the decoder 423 .

[0109] For example, as shown in Figure 6, the first area of ​​the interactive interface displays the object to be selected, and the decoding device 420 uses the object name of the object to which at least one object parameter belongs as a display identifier, displays the first control group name and the controls corresponding to each object in the first area, such as three controls whose object names are Object 1, Object 2 and Object 3 respectively.

[0110] Continuing with the example of the control for object 1, the control for object 1 includes a selected state and an unselected state. The selected state indicates that the object parameter corresponding to the control for object 1 is in the selected state, while the unselected state indicates that the object parameter corresponding to the control for object 1 is in the unselected state. When a user triggers the unselected control for object 1 by performing an operation such as selecting the control on a touch screen or using a remote control, the decoding device 420 switches the object parameter corresponding to the control for object 1 to the selected state in response to the user operation and determines that the target object parameter is the object parameter of object 1.

[0111] In a football game broadcast scenario, object 1 can be the commentary, object 2 can be the atmosphere of the stadium, and object 3 can be the broadcast of the stadium. The commentary can include Chinese commentary, English commentary, etc.

[0112] The second area of ​​the interactive interface displays the name of the second control group and the second control group for adjusting the target object parameters. The second control group includes controls for adjusting any field value of the target parameter object. The number and type of controls contained in the second control group are determined by the fields contained in the target object parameters. For example, when the target object parameter is object 1, the second control group includes three controls for indicating the mute state, object volume, and object sound direction of the object parameter. The control corresponding to the mute state can be displayed as a mute indicator, including a selected state and an unselected state. The selected state is used to indicate that the object volume of object 1 is set to the minimum value. The control corresponding to the object volume includes a sliding bar, which adjusts the object volume in response to user operation. When the sliding bar slides in a first direction (for example, toward 16), the object volume increases, and when it slides in a second direction (for example, toward 0), the object drainage decreases. The second direction is the opposite direction of the first direction.

[0113] Optionally, the interactive interface can also display the second control group in more areas. For example, the third area of ​​the interactive interface displays the controls corresponding to the object sound direction in the second control group. The controls corresponding to the object sound direction include multiple sliders, and the number of sliders is related to the position dimension of the sound direction. For example, if the object sound direction is represented in three dimensions in a Cartesian coordinate system, the controls corresponding to the object sound direction may include three sliders, and the three sliders respectively adjust the object sound direction to move left and right, front and back, or up and down in response to user operations. For example, the minimum value of the three sliders is -10, and the maximum value is 10. The slider for adjusting the object sound direction to move left and right is marked as left and right, respectively displayed on the left and right sides of the slider. The slider for adjusting the object sound direction to move back and forth is marked as back and front, respectively displayed on the left and right sides of the slider. The slider for adjusting the object sound direction to move up and down is marked as bottom and top, respectively displayed on the left and right sides of the slider.

[0114] As a possible implementation manner, the decoding device 420 may update the control group displayed on the interactive interface based on changes in metadata of different frames of the audio stream.

[0115] Optionally, when metadata of a second frame of the audio stream is different from metadata of a first frame, the decoding device 420 updates the control group based on metadata of the second frame, where the second frame is n frames after the first frame, and n is a positive integer.

[0116] Step 530: The decoding device 420 generates metadata change information in response to the operation on the control group.

[0117] As a possible example, the decoding device 420 generates metadata change information in response to the user's operations on the first control group and the second control group. The metadata change information includes target object parameters after one or more fields are changed.

[0118] For example, in response to the user adjusting the object volume of object 1 based on the interactive interface, the decoding device 420 increases the object volume value from 6 to 8. The object volume of the changed target object parameter in the metadata change information generated by the decoding device 420 is set to 8, and the values ​​of other fields remain unchanged. For another example, in response to the user adjusting the object sound direction of object 1 based on the interactive interface, the decoding device 420 generates metadata change information with an updated value for the coordinate value of the object sound direction of the changed target object parameter. The direction value (such as a left-right movement value, a front-back movement value, or an up-and-down movement value) in the control corresponding to the object sound direction is mapped to the coordinate system of the object sound direction in the metadata, thereby obtaining the coordinate value of the object sound direction of the changed target object parameter based on the direction value.

[0119] Step 540: When decoding the audio stream, the decoding device 420 changes the metadata according to the metadata change information to obtain a decoding result.

[0120] As a possible example, when using the decoder 423 to decode the audio stream into object information, sound bed information, and metadata information, the decoding device 420 replaces the target object parameters in the metadata information with the modified target object parameters to obtain a decoding result.

[0121] The decoding result includes the changed metadata information and the audio data carried by the audio stream, and the changed metadata information is used to represent the changed metadata.

[0122] Optionally, in addition to the metadata information, the audio stream also carries audio data, including object information and sound bed information. The object information, sound bed information, and metadata information may be in pulse code modulation (PCM) format, and thus may also be referred to as object encoding, sound bed encoding, and metadata encoding.

[0123] As a possible implementation manner, the decoding device 420 determines the target object parameter in the metadata information according to the index of the target object parameter, thereby changing the target object parameter.

[0124] Step 550: The decoding device 420 outputs the decoding result as audio according to the modified metadata.

[0125] As a possible example, the decoding device 420 mixes the object information to the corresponding sound bed according to the constraints of the changed metadata to obtain original multi-channel encoded data, then renders the original multi-channel encoded data with sound effects to obtain rendered multi-channel encoded data, and outputs audio based on the rendered multi-channel encoded data.

[0126] Optionally, the sound rendering may be implemented using algorithms such as dynamic range compression (DRC) and audio quality detection (AQ).

[0127] As a possible implementation, the decoding device 420 outputs the rendered multi-channel encoded data, which can be transmitted to an external speaker via a network or the like for playback, or can be played via the playback device 421 of the decoding device 420 .

[0128] In a possible embodiment of the present application, the metadata types supported by the decoding device 420 and the audio output method can refer to the international standard content of the audio definition model (ADM), which will not be repeated here.

[0129] Based on the above steps 510 to 550, the decoding device 420 provides the user with an interactive interface to display a control group, and changes the metadata of the object in the decoding process of the audio stream according to the user's operation on the control group, so that the user can visually indicate the modification of the metadata based on the control group of the interactive interface. In this way, whether it is an online playback scenario or a local playback scenario, the user does not need to re-encode the audio stream in the bitstream production process to modify the metadata of the object, and can also adjust the channel and other properties of any object in the audio stream in real time according to needs to achieve audio output that meets user needs, thereby improving the flexibility of audio output. At the same time, the user does not need to re-encode the audio stream, which reduces the overhead of computing resources and the complexity of operations.

[0130] The above describes the overall process of the audio processing method in combination with Figure 5. The audio processing method realizes the individual adjustment of the metadata of the object, but in scenarios such as sports event broadcasts and live concerts, users often need to adjust the metadata of multiple related objects at a time. Therefore, the audio processing method of the present application can also divide at least one object in the audio stream into groups, and a group contains object parameters of multiple related objects, thereby providing users with batch object adjustment functions of different groups (or modes).

[0131] The following describes in detail the implementation of the group function with reference to FIG7 .

[0132] When the metadata change information of the object parameters of a single object is referred to as the first metadata change information, the decoding device 420 displays a third control group on the interactive interface based on the group in the metadata, determines the target group and generates second metadata change information in response to a third operation on the third control group.

[0133] For example, as shown in Figure 7, the fourth area of ​​the interactive interface displays the third control group. The decoding device 420 uses the group name of multiple groups composed of at least one object as a display identifier, and displays the controls corresponding to each group in the third control group in the fourth area, such as three controls with object names Group 1, Group 2 and Group 3.

[0134] Continuing with the example of the controls of Group 1, the controls of Group 1 include a selected state and an unselected state. The selected state is used to indicate that the object parameter corresponding to the controls of Group 1 is in the selected state, and the unselected state is used to indicate that the object parameter corresponding to the controls of Group 1 is in the unselected state. When a user triggers the unselected controls of Group 1 by touching the screen or selecting with a remote control, the decoding device 420 switches the object parameter corresponding to the controls of Group 1 to the selected state in response to the user operation, determines that the target group is the object included in Group 1, and displays the first control group and the second control group corresponding to the objects included in Group 1 on the interactive interface. The circular controls in front of the groups in Figure 7 indicate the selected or unselected state of the groups. For example, a filled circle indicates that Group 1 is in the selected state, and an unfilled circle indicates that Group 2 and Group 3 are in the unselected state.

[0135] In a football game broadcast scenario, Group 1 may be a standard group, Group 2 may be a home and away atmosphere group, and Group 3 may be a live atmosphere group. For example, when a user selects the home and away atmosphere group via a touchscreen or remote control, the decoding device 420 displays controls corresponding to the home and away atmosphere group's commentary, stadium atmosphere, stadium broadcast, Team A fans, Team B fans, and other objects in the first area of ​​the interactive interface.

[0136] As shown in FIG8a , the home and away game range group includes objects such as Chinese commentary, English commentary, stadium atmosphere, stadium broadcast, fans of team A, and fans of team B.

[0137] As shown in FIG8b , the on-site atmosphere group includes objects such as stadium atmosphere and stadium broadcast.

[0138] Given that different groups typically represent different modes, when a user selects a target group for object metadata adjustment, it is necessary to mute objects not included in the target group. To simplify user operation, decoding device 420 generates second metadata modification information after determining the target group in response to an operation on the third control group. The second metadata modification information includes at least one modified object parameter, wherein the object volume of the object parameter not included in the target group is minimized.

[0139] As a possible implementation, when the commentary displayed in the first area of ​​the interactive interface includes Chinese and English commentary, it is considered that if the user selects the Chinese and English commentary separately for adjustment, the Chinese and English commentary sounds will be played simultaneously, which will interfere with the user's listening experience. The object parameters of the present application may also include a mutually exclusive object field for representing mutually exclusive objects of the object parameters. When the decoding device 420 determines the target object parameter in response to the user's first operation on the first control group, it sets the value of the object volume field of the object parameter of the mutually exclusive object of the target object parameter to the minimum value.

[0140] Optionally, when displaying mutually exclusive objects of a certain object, the decoding device 420 displays a group of mutually exclusive objects in a selection box.

[0141] For example, the mutually exclusive object of the Chinese commentary includes the English commentary. When the decoding device 420 determines that the target object parameter is the Chinese commentary in response to the user's first operation on the first control group, it queries the object parameter of the English commentary in the metadata according to the object index of the mutually exclusive object English commentary, and sets the value of the object volume field of the object parameter of the English commentary to the minimum value.

[0142] For another example, the mutually exclusive objects of Team A fans include Team B fans. When the decoding device 420 determines that the target object parameter is Team A fans in response to the user's first operation on the first control group, it queries the object parameters of Team B fans in the metadata based on the object index of the mutually exclusive object Team B fans, and sets the value of the object volume field of the object parameters of Team B fans to the minimum value.

[0143] As a possible implementation method, for any object parameter contained in the metadata, the values ​​of all fields of the object parameter may support changes, or the values ​​of some fields of the object parameter may not support changes. In this case, the decoding device 420 sets the controls corresponding to the fields that do not support changes to an inoperable state in the interactive interface.

[0144] For example, as shown in Figure 9, the English commentary object does not support object volume adjustment, and the decoding device 420 sets the control corresponding to the object volume in the interactive interface to an inoperable state. As shown in Figure 10, the English commentary object supports object volume adjustment, and the decoding device 420 sets the control corresponding to the object volume in the interactive interface to an operable state.

[0145] For example, as shown in Figure 11, the English commentary object does not support adjustment of the object sound direction, and the decoding device 420 sets the control corresponding to the object sound direction in the interactive interface to an inoperable state. For example, as shown in Figure 12, the Chinese commentary object supports adjustment of the object sound direction, and the decoding device 420 sets the control corresponding to the object sound direction in the interactive interface to an operable state.

[0146] In the above-mentioned Figures 9 to 12 , low grayscale is used to represent inoperable controls in the interactive interface, and high grayscale is used to represent operable controls in the interactive interface.

[0147] It is understandable that in order to implement the functions in the above embodiments, the encoding device and the decoding device include hardware structures and / or software modules corresponding to the execution of each function. It should be readily apparent to those skilled in the art that, in combination with the units and method steps of each example described in the embodiments disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in hardware or in a computer software-driven hardware manner depends on the specific application scenario and design constraints of the technical solution.

[0148] Figure 13 is a schematic diagram of the structure of an audio processing device provided in an embodiment of the present application. The audio processing device can be used to implement the functions of the decoding device in the above-mentioned method embodiment, thereby also achieving the beneficial effects of the above-mentioned audio processing method embodiment. In the embodiment of the present application, the audio processing device can be the decoding device 420 shown in Figure 4, or it can be a module (such as a chip) applied to the decoding device.

[0149] As shown in Figure 13, the audio processing device 1300 includes a bitstream module 1301, a display module 1302, an adjustment module 1303, and a rendering module 1304. The bitstream module 1301 is used to obtain an audio stream, which includes metadata describing object information within the audio stream. The display module 1302 is used to display a control group on an interactive interface based on the metadata and to generate metadata change information in response to operations on the control group. The adjustment module 1303 is used to modify the metadata according to the metadata change information when decoding the audio stream, thereby obtaining a decoding result including the modified metadata. The rendering module 1304 is used to output audio corresponding to the audio data based on the modified metadata.

[0150] As a possible implementation, the audio stream includes at least one object, the metadata includes at least one object parameter, an object parameter is used to describe the object information of an object, and the object parameter includes at least one of the object name, object volume, mute state and object sound direction.

[0151] The audio processing device 1300 can be used to implement the corresponding operating steps of each method in the aforementioned audio processing method embodiment, thereby also achieving the corresponding beneficial effects, which will not be described in detail here. For more information about the bit stream module 1301, display module 1302, adjustment module 1303, and rendering module 1304, please refer to the content of the above audio processing method embodiment and will not be described in detail here.

[0152] When the encoding device or decoding device implements any of the methods shown in the aforementioned figures through software, the encoding device or decoding device and its various units may also be software modules. The aforementioned methods are implemented by calling the software module through a processor. The processor may be a central processing unit (CPU) or other application-specific integrated circuit (ASIC), or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0153] For a more detailed description of the above-mentioned encoding device or decoding device, please refer to the relevant description in the embodiment shown in the aforementioned drawings, which will not be repeated here. It can be understood that the encoding device or decoding device shown in the aforementioned drawings is only an example provided by this embodiment, and the encoding device or decoding device may include more or fewer units, which is not limited by this application. For example, as shown in Figure 14, the audio processing device 1400 includes an interactive module 1401, a player module 1402, a metadata management module 1403, an audio decoder module 1404, a metadata adjustment module 1405, an object rendering module 1406, a sound effect rendering module 1407 and a playback module 1408. Each module is obtained by splitting or combining the modules of the audio processing device 1300. For the functions of each module, please refer to the description of functional steps 1-9 in Figure 14, which will not be repeated here.

[0154] When the audio processing device 1300 is implemented via hardware, the hardware may be implemented via a processor or a chip. The chip includes an interface circuit and a control circuit. The interface circuit is used to receive data from devices other than the processor and transmit it to the control circuit, or to transmit data from the control circuit to devices other than the processor.

[0155] The control circuit and the interface circuit are used to implement any possible implementation method of the above embodiments through logic circuits or execution code instructions. The beneficial effects can be found in the description of any aspect of the above embodiments, which will not be repeated here.

[0156] In addition, the audio processing device 1300 can also be implemented by an electronic device, where the electronic device may refer to the decoding device in the aforementioned embodiment, or, when the electronic device is a chip or chip system applied to a server, the audio processing device 1300 can also be implemented by the chip or chip system.

[0157] As shown in FIG15 , FIG15 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 1500 includes: a processor 1510, a bus 1520, a memory 1530, a memory unit 1550 (also referred to as a main memory unit), and a communication interface 1540. The processor 1510, the memory 1530, the memory unit 1550, and the communication interface 1540 are connected via the bus 1520.

[0158] It should be understood that in this embodiment, the processor 1510 may be a CPU, but may also be other general-purpose processors, DSPs, ASICs, FPGAs or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0159] The communication interface 1540 is used to implement communication between the electronic device 1500 and an external device or component. In this embodiment, the communication interface 1540 is used to exchange data with other electronic devices.

[0160] Bus 1520 may include a path for transmitting information between the aforementioned components (e.g., processor 1510, memory unit 1550, and storage 1530). In addition to a data bus, bus 1520 may also include a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as bus 1520 in the figure. Bus 1520 may be a Peripheral Component Interconnect Express (PCIe) bus, an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc.

[0161] As an example, electronic device 1500 may include multiple processors. The processor may be a multi-core (multi-CPU) processor. The processor herein may refer to one or more devices, circuits, and / or computing units for processing data (e.g., computer program instructions). Processor 1510 may invoke computer program instructions stored in memory 1530 to execute steps 510 to 550 as shown in FIG5 .

[0162] It is worth noting that Figure 15 only takes the electronic device 1500 including 1 processor 1510 and 1 memory 1530 as an example. Here, the processor 1510 and the memory 1530 are respectively used to indicate a type of device or equipment. In a specific embodiment, the number of each type of device or equipment can be determined according to business requirements.

[0163] The memory unit 1550 may correspond to the storage medium used to store information such as pre-upgrade data in the above-mentioned method embodiment. The memory unit 1550 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0164] The memory 1530 is used to store bit streams or coding units to be encoded, etc., and can be a solid-state drive or a mechanical hard drive.

[0165] It should be understood that the electronic device 1500 described above may be a DPU. The electronic device 1500 according to this embodiment may correspond to the encoding device or decoding device in this embodiment, and may correspond to executing the corresponding subjects according to FIG. 5 . The above and other operations and / or functions of the various modules in the encoding device or decoding device are respectively for implementing the corresponding processes in the aforementioned method embodiments, and for the sake of brevity, they are not further described here.

[0166] The method steps in this embodiment can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC. In addition, the ASIC can be located in a computing device. Of course, the processor and storage medium can also exist as discrete components in a network device or a terminal device.

[0167] The present application also provides a chip system, which includes a processor for implementing the functions of the encoding device or decoding device in the above method. In one possible design, the chip system also includes a memory for storing program instructions and / or data. The chip system can be composed of a chip alone or can include a chip and other discrete devices.

[0168] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the process or function described in the embodiments of the present application is performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device or other programmable device. The computer program or instruction can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instruction can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a tape; it can also be an optical medium, such as a digital video disc (DVD); it can also be a semiconductor medium, such as a solid state drive (SSD).

[0169] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. An audio processing method, characterized in that: include: Acquire an audio stream, wherein the audio stream includes metadata, and the metadata is used to describe object information in the audio stream; Displaying a control group on an interactive interface based on the metadata; generating metadata change information in response to an operation on the control group; When decoding the audio stream, the metadata is changed according to the metadata change information to obtain a decoding result, wherein the decoding result includes the changed metadata and the audio data carried by the audio stream; The audio corresponding to the audio data is output according to the modified metadata in the decoding result.

2. The method according to claim 1, characterized in that The audio stream includes at least one object, the metadata includes at least one object parameter, an object parameter is used to describe object information of an object, the object parameter includes at least one of an object name, an object volume, a mute state, and an object sound direction, the control group includes a first control group and a second control group, and displaying the control group on the interactive interface based on the metadata includes: displaying the first group of controls based on the name of the at least one object parameter; In response to a first operation on the first control group, determining a target object parameter; A second control group is displayed based on the target object parameter, and a control in the second control group is used to adjust the object volume, mute state or object sound direction in response to a second operation to generate the metadata change information.

3. The method according to claim 2, characterized in that The metadata change information includes first metadata change information, the first metadata change information includes changed target object parameters, and changing the metadata according to the metadata change information to obtain a decoding result includes: Parsing the audio stream into object information, sound bed information and metadata information; Replacing the target object parameter in the metadata information with the changed target object parameter to obtain the changed metadata information, wherein the changed metadata information is used to represent the changed metadata; The decoding result is output, where the decoding result includes the object information, the sound bed information, and the modified metadata information.

4. The method according to claim 3, characterized in that The first metadata change information also includes an index of the target object parameter, and the method further includes: The target object parameter is determined in the metadata information according to an index of the target object parameter.

5. The method according to any one of claims 2 to 4, characterized in that: The object parameters also include the group to which the object belongs and the group members, and the method further includes: Displaying a third control group on the interactive interface based on the metadata; In response to a third operation on the third control group, a target group is determined and second metadata change information is generated, wherein the second metadata change information includes at least one of the changed object parameters, wherein the object volume of the object parameters outside the target group in the at least one changed object parameter is a minimum value.

6. The method according to any one of claims 2 to 5, characterized in that: The object parameters also include mutually exclusive objects, and after determining the target object parameters in response to the first operation on the first control group, the method further includes: Sets the object volume of the object parameter of the mutually exclusive object of the target object parameter to the minimum value.

7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: When metadata of a second frame of the audio stream is different from metadata of a first frame, the control group is updated based on metadata of the second frame, where the second frame is n frames after the first frame, and n is a positive integer.

8. The method according to claim 3, characterized in that The step of outputting the audio corresponding to the audio information according to the modified metadata in the decoding result includes: Mixing the object information to the corresponding sound bed according to the constraints of the modified metadata to obtain original multi-channel encoded data; Performing sound effect rendering on the original multi-channel encoded data to obtain rendered multi-channel encoded data; Output audio according to the rendered multi-channel encoded data.

9. The method according to any one of claims 2 to 8, characterized in that: The object sound direction is represented by three-dimensional coordinates.

10. An audio processing device, characterized in that: The device comprises a module for executing the operation steps of any one of the methods of claims 1-9.

11. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory is used to store a group of computer instructions; when the processor executes the group of computer instructions, the operation steps of any one of the methods described in claims 1-9 are performed.

12. A coding and decoding system, characterized in that: The invention comprises an encoding device and a decoding device, wherein the encoding device is communicatively connected to the decoding device, the encoding device is used to generate an audio stream, the audio stream comprises metadata, and the metadata is used to describe object information in the audio stream, and the decoding device is used to implement the method described in any one of claims 1 to 9.

13. A readable storage medium, characterized in that: The readable storage medium includes a computer program or instructions. When the computer program or instructions are executed on a computer, the computer is enabled to execute the operation steps of any one of the methods described in claims 1 to 9.