Audio processing method and terminal

The audio processing method dynamically adjusts audio mixing based on user location within optimized listening areas, enhancing the VR music experience by maintaining high-quality audio effects during movement.

JP7741334B2Active Publication Date: 2025-09-17HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024544814
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-01-28
Filing Date
2022-12-20
Publication Date
2025-09-17
Estimated Expiration
2042-12-20

AI Technical Summary

Technical Problem

In VR music scenes, users' listening experience is compromised when they move freely due to the assumption that their position remains unchanged during audio mixing, limiting the effectiveness of six degrees of freedom (6DoF) experiences.

Method used

An audio processing method that decodes audio bitstreams to obtain metadata and mixing parameters for optimized listening areas, allowing dynamic audio mixing based on the user's location within these areas, enhancing the listening experience by providing audio optimization suitable for movement.

Benefits of technology

Improves the listening experience in VR music scenes by ensuring that users moving freely receive optimized audio mixing, maintaining high-quality musical effects across different positions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007741334000002
    Figure 0007741334000002
  • Figure 0007741334000003
    Figure 0007741334000003
  • Figure 0007741334000004
    Figure 0007741334000004
Patent Text Reader

Abstract

The embodiment of the present application discloses an audio processing method and a terminal for improving the listening effect obtained when a user moves freely. The embodiment of the present application provides an audio processing method, including: decoding an audio bitstream to obtain audio optimization metadata, basic audio metadata, and M decoded audio data, where the audio optimization metadata includes first metadata of a first optimized listening area and first decoded audio mixing parameters corresponding to the first optimized listening area, where M is a positive integer; rendering the M decoded audio data based on a current position of the user and the basic audio metadata to obtain M rendered audio data; performing first audio mixing on the M rendered audio data based on the first decoded audio mixing parameters to obtain M first audio mixing data when the current position of the user is within the first optimized listening area; and mixing the M first audio mixing data to obtain mixed audio data corresponding to the first optimized listening area.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to Chinese Patent Application No. 202210109139.4, entitled "Audio Processing Method and Terminal," filed with the State Intellectual Property Office of China on January 28, 2022, which is incorporated herein by reference in its entirety.

[0002] The present application relates to the field of audio technology, and in particular to an audio processing method and terminal. [Background technology]

[0003] Audio mixing is an essential process in music production, and the quality of audio mixing determines the success or failure of a musical work. The audio output after audio mixing allows listeners to hear subtle and layered musical effects that cannot be heard in live recordings, making the music more expressive.

[0004] Virtual reality (VR) technology is gradually being applied to the music field, leading to the emergence of VR music scenes. Currently, in the process of creating VR music scenes, when mixing music signals, creators typically assume that the user is located in the sweet area and their position remains unchanged. Therefore, this type of VR music scene can realize the effect of the user's head rotation (e.g., three degrees of freedom (3DoF)). The user can only enjoy a good music experience when they are within the sweet area. If the user's position changes, the user's listening effect will be reduced, further affecting the user's music experience. Summary of the Invention [Means for solving the problem]

[0005] The embodiments of the present application provide an audio processing method and a terminal for improving the listening experience obtained when a user moves freely.

[0006] In order to solve the aforementioned technical problems, the embodiments of the present application provide the following technical solutions.

[0007] According to a first aspect, an embodiment of the present application provides: decoding the audio bitstream to obtain audio-optimization metadata, basic audio metadata, and M decoded audio data, wherein the audio-optimization metadata includes first metadata of a first optimized listening area and first decoded audio mixing parameters corresponding to the first optimized listening area, where M is a positive integer; To obtain M rendered audio data, M Return of rendering the decoded audio data; performing first audio mixing on the M pieces of rendered audio data based on the first decoded audio mixing parameters to obtain M pieces of first audio mixing data when the current location is within a first optimized listening area; mixing the M first audio mixing data to obtain mixed audio data corresponding to a first optimized listening area; An audio processing method is provided, including:

[0008] In the aforementioned solution, in this embodiment of the present application, metadata of a first optimized listening area and a first decoded audio mixing parameter corresponding to the first optimized listening area can be obtained, and M pieces of rendered audio data are obtained based on the user's current location and basic audio metadata. Return ofThe decoded audio data is rendered. Then, when it is determined that the user's current location is within the first optimized listening area, first audio mixing is performed on the M rendered audio data based on the first decoded audio mixing parameters to obtain M first audio mixing data. Finally, the M first audio mixing data are mixed to obtain mixed audio data corresponding to the first optimized listening area. Therefore, in this embodiment of the present application, when the user's current location is located within the first optimized listening area, both audio mixing and data mixing are performed by using audio data corresponding to the first optimized listening area, so that audio optimization metadata suitable for the user to move freely in the first optimized listening area can be provided, and the listening experience obtained when the user moves freely can be improved.

[0009] In one possible implementation, the audio optimization metadata further includes second decoded audio mixing parameters corresponding to the first optimized listening area.

[0010] The method further includes performing second audio mixing on the mixed audio data based on the second decoded audio mixing parameters to obtain second audio mixing data corresponding to the first optimized listening area.

[0011] In the above solution, after obtaining the second decoded audio mixing parameters, the decoding terminal may further perform second audio mixing on the mixed audio data corresponding to the first optimized listening area based on the second decoded audio mixing parameters corresponding to the first optimized listening area to obtain second audio mixing data corresponding to the first optimized listening area. The second audio mixing data may be obtained through the second audio mixing. When the second audio mixing data is played, the user's listening experience may be improved.

[0012] In one possible implementation, the second decoded audio mixing parameters include at least one of an identifier of the second audio mixing data, an equalization parameter, a compressor parameter, and a reverberator parameter.

[0013] In the above solution, the second decoded audio mixing parameters may include an identifier of the second audio mixing data. The second decoded audio mixing parameters may further include equalization parameters. For example, the equalization parameters may include an equalization parameter identifier, a gain value for each frequency band, and a Q value. The Q value is an equalization filter parameter that represents a quality coefficient of the equalization filter and may be used to describe the bandwidth of the equalization filter. The second decoded audio mixing parameters may further include compressor parameters. For example, the compressor parameters may include a compressor identifier, a threshold, a compression ratio, a start time, a release time, and a gain compensation value. The second decoded audio mixing parameters may further include reverberator parameters. For example, the reverberator parameters may include a reverberation type, a reverberation time, a delay time, and a direct-to-reverberant ratio.

[0014] In one possible implementation, the audio optimization metadata further includes N-1 difference parameters of N-1 second decoded audio mixing parameters corresponding to N-1 optimized listening areas other than the first optimized listening area within the N optimized listening areas with respect to the second decoded audio mixing parameters corresponding to the first optimized listening area, where N is a positive integer.

[0015] In the above solution, the difference parameters are parameters of differences between the N-1 second decoded audio mixing parameters of the N-1 optimized listening areas other than the first optimized listening area among the N optimized listening areas and the second decoded audio mixing parameters corresponding to the first optimized listening area. The difference parameters are not the N-1 second decoded audio mixing parameters of the N-1 optimized listening areas. The audio-optimized metadata carries the difference parameters, so that the data amount of the audio-optimized metadata can be reduced, and data transmission efficiency and decoding efficiency can be improved.

[0016] In one possible implementation, the first decoded audio mixing parameters include at least one of an identifier of the rendered audio data, an equalization parameter, a compressor parameter, and a reverberator parameter.

[0017] In the above solution, the first decoded audio mixing parameters may include an identifier of the rendered audio data, e.g., an identifier of M pieces of rendered audio data. The first decoded audio mixing parameters may further include equalization parameters. For example, the equalization parameters may include an equalization parameter identifier, a gain value for each frequency band, and a Q factor. The first decoded audio mixing parameters may further include compressor parameters. For example, the compressor parameters may include a compressor identifier, a threshold, a compression ratio, a start time, a release time, and a gain compensation value. The first decoded audio mixing parameters may further include reverberator parameters. For example, the reverberator parameters may include a reverberation type, a reverberation time, a delay time, and a direct-to-reverberant ratio.

[0018] In one possible embodiment, the method comprises: decoding the video image bitstream to obtain decoded video image data and video image metadata, the video image metadata including video metadata and image metadata; rendering the decoded video image data based on the video image metadata to obtain rendered video image data; establishing a virtual scene based on the rendered video image data; identifying a first optimized listening area within the virtual scene based on the rendered video image data and the audio optimization metadata; Further includes:

[0019] In the above solution, the decoding terminal may render the decoded video image data based on the video image metadata to obtain rendered video image data, and the decoding terminal may establish a virtual scene by using the rendered video image data. Finally, the decoding terminal may identify a first optimized listening area in the virtual scene based on the rendered video image data and the audio optimization metadata, so that the decoding terminal displays the first optimized listening area in the virtual scene and guides the user to experience music in the optimized listening area, thereby improving the user's listening experience.

[0020] In one possible implementation, the first metadata includes at least one of a reference coordinate system of the first optimized listening area, a center position coordinate of the first optimized listening area, and a shape of the first optimized listening area.

[0021] In the above solution, the metadata of the first optimized listening area may include a reference coordinate system, or the metadata of the first optimized listening area may not include a reference coordinate system. For example, the first optimized listening area uses a default coordinate system. The metadata of the first optimized listening area may include descriptive information for describing the first optimized listening area, such as the center position coordinates of the first optimized listening area and information for describing the shape of the first optimized listening area. In this embodiment of the present application, there may be multiple shapes of the first optimized listening area. For example, the shape may be a sphere, a cube, a pillar, or any other shape.

[0022] In one possible implementation, the audio optimization metadata includes N-1 difference parameters of N-1 first decoded audio mixing parameters corresponding to N-1 optimized listening areas other than the first optimized listening area within the N optimized listening areas, relative to the first decoded audio mixing parameters corresponding to the first optimized listening area, where N is a positive integer.

[0023] In the above solution, the difference parameters are parameters of differences between the N-1 first decoded audio mixing parameters of the N-1 optimized listening areas other than the first optimized listening area among the N optimized listening areas and the first decoded audio mixing parameters corresponding to the first optimized listening area. The difference parameters are not the N-1 first decoded audio mixing parameters of the N-1 optimized listening areas. The audio-optimized metadata carries the difference parameters, so that the data amount of the audio-optimized metadata can be reduced, and data transmission efficiency and decoding efficiency can be improved.

[0024] According to a second aspect, an embodiment of the present application provides: receiving audio-optimization metadata, basic audio metadata, and M first audio data, wherein the audio-optimization metadata includes first metadata of a first optimized listening area and first audio mixing parameters corresponding to the first optimized listening area, where M is a positive integer; performing compression encoding on the audio-optimized metadata, the basic audio metadata, and the M first audio data to obtain an audio bitstream; transmitting an audio bitstream; The present invention further provides an audio processing method, including:

[0025] In the above solution, audio optimization metadata, basic audio metadata, and M first audio data are first received, and the audio optimization metadata includes metadata of a first optimized listening area and first audio mixing parameters of the first optimized listening area. Therefore, audio optimization metadata suitable for a user to move freely in the first optimized listening area can be provided, and the listening experience obtained when the user moves freely can be improved.

[0026] In one possible implementation, the audio optimization metadata further comprises a second audio mixing parameter change identifier.

[0027] The second audio mixing parameter change identifier indicates whether the second audio mixing parameters corresponding to the first audio data of the current frame have changed compared to the second audio mixing parameters corresponding to the first audio data of the previous frame.

[0028] In the above solution, the transmitting terminal may set a second audio mixing parameter change identifier in the audio optimization metadata. The second audio mixing parameter change identifier indicates whether the second audio mixing parameter corresponding to the first optimized listening area has changed. Therefore, the decoding terminal determines whether the second audio mixing parameter corresponding to the first optimized listening area has changed based on the second audio mixing parameter change identifier. For example, if the second audio mixing parameter corresponding to the first audio data of the current frame has changed compared to the second audio mixing parameter corresponding to the first audio data of the previous frame, the second audio mixing parameter change identifier is TRUE, and the transmitting terminal may further transmit modification information of the second audio mixing parameter corresponding to the first audio data. The decoding terminal receives the modification information of the second audio mixing parameter corresponding to the first audio data and obtains the modified second audio mixing parameter corresponding to the first audio data of the current frame based on the modification information.

[0029] In one possible implementation, the audio optimization metadata further includes second audio mixing parameters corresponding to the first optimized listening area.

[0030] In the above solution, when the production terminal performs audio mixing twice, the audio optimization metadata acquired by the production terminal may include first metadata of a first optimized listening area, first audio mixing parameters corresponding to the first optimized listening area, and second audio mixing parameters corresponding to the first optimized listening area. After the audio optimization metadata is acquired by the decoding terminal, the decoding terminal also needs to perform audio mixing twice, and the user's listening experience can be improved by performing audio mixing twice.

[0031] In one possible implementation, the audio optimization metadata further includes N-1 difference parameters of the second audio mixing parameters corresponding to a first optimized listening area among the N optimized listening areas and the N-1 second audio mixing parameters corresponding to N-1 optimized listening areas other than the first optimized listening area among the N optimized listening areas, relative to the second audio mixing parameters corresponding to the first optimized listening area, where N is a positive integer.

[0032] In the above solution, the difference parameters are parameters of differences between the N-1 second audio mixing parameters of the N-1 optimized listening areas other than the first optimized listening area among the N optimized listening areas and the second audio mixing parameters corresponding to the first optimized listening area. The difference parameters are not the N-1 second audio mixing parameters of the N-1 optimized listening areas. The audio-optimized metadata carries the difference parameters, so that the data amount of the audio-optimized metadata can be reduced, and data transmission efficiency and decoding efficiency can be improved.

[0033] In one possible implementation, the second audio mixing parameters include at least one of an identifier of the first audio data, an equalization parameter, a compressor parameter, and a reverberator parameter.

[0034] In the previous solution, the second audio mixing parameter is M piecesThe second audio mixing parameters may include an identifier of the first audio data, for example. The second audio mixing parameters may further include equalization parameters. For example, the equalization parameters may include an equalization parameter identifier, a gain value for each frequency band, and a Q value. The second audio mixing parameters may further include compressor parameters. For example, the compressor parameters may include a compressor identifier, a threshold, a compression ratio, a start time, a release time, and a gain compensation value. The second audio mixing parameters may further include reverberator parameters. For example, the reverberator parameters may include a reverberation type, a reverberation time, a delay time, and a direct sound to reverberation sound ratio.

[0035] In one possible implementation, the audio optimization metadata further includes N-1 difference parameters of N-1 first audio mixing parameters corresponding to N-1 optimized listening areas other than the first optimized listening area within the N optimized listening areas, relative to the first audio mixing parameter corresponding to the first optimized listening area.

[0036] In the above solution, the difference parameters are parameters of differences between the N-1 first audio mixing parameters of the N-1 optimized listening areas other than the first optimized listening area among the N optimized listening areas and the first audio mixing parameters corresponding to the first optimized listening area. The difference parameters are not the N-1 first audio mixing parameters of the N-1 optimized listening areas. The audio-optimized metadata carries the difference parameters, so that the data amount of the audio-optimized metadata can be reduced, and data transmission efficiency and decoding efficiency can be improved.

[0037] In one possible implementation, the first audio mixing parameters include at least one of an identifier of the first audio data, an equalization parameter, a compressor parameter, and a reverberator parameter.

[0038] In the above solution, the first audio mixing parameters may include an identifier of the first audio data, for example, identifiers of M pieces of the first audio data. The first audio mixing parameters may further include equalization parameters. For example, the equalization parameters may include an equalization parameter identifier, a gain value for each frequency band, and a Q value. The first audio mixing parameters may further include compressor parameters. For example, the compressor parameters may include a compressor identifier, a threshold, a compression ratio, a start time, a release time, and a gain compensation value. The first audio mixing parameters may further include reverberator parameters. For example, the reverberator parameters may include a reverberation type, a reverberation time, a delay time, and a direct-to-reverberant sound ratio.

[0039] In one possible implementation, the first metadata for the first optimized listening area includes at least one of a reference coordinate system for the first optimized listening area, a center position coordinate for the first optimized listening area, and a shape of the first optimized listening area.

[0040] In the above solution, the metadata of the first optimized listening area may include a reference coordinate system, or the metadata of the first optimized listening area may not include a reference coordinate system. For example, the first optimized listening area uses a default coordinate system. The metadata of the first optimized listening area may include descriptive information for describing the first optimized listening area, such as the center position coordinates of the first optimized listening area and information for describing the shape of the first optimized listening area. In this embodiment of the present application, there may be multiple shapes of the first optimized listening area. For example, the shape may be a sphere, a cube, a pillar, or any other shape.

[0041] In one possible implementation, the audio optimization metadata further includes a center position coordinate of a first optimized listening area among the N optimized listening areas, and position offsets of the center position coordinates of N-1 optimized listening areas other than the first optimized listening area among the N optimized listening areas relative to the center position coordinate of the first optimized listening area, where N is a positive integer.

[0042] In the above solution, the position offset is an offset between the center position coordinate of N-1 optimized listening areas other than the first optimized listening area among the N optimized listening areas and the center position coordinate of the first optimized listening area that is not other than the center position coordinate of the N-1 optimized listening areas. The audio-optimized metadata carries the position offset, so that the data amount of the audio-optimized metadata can be reduced, and data transmission efficiency and decoding efficiency can be improved.

[0043] In one possible implementation, the audio optimization metadata further includes an optimized listening area change identifier and / or a first audio mixing parameter change identifier.

[0044] The optimized listening area change identifier indicates whether the first optimized listening area has changed.

[0045] The first audio mixing parameter change identifier indicates whether the first audio mixing parameter corresponding to the first audio data of the current frame has changed compared to the first audio mixing parameter corresponding to the first audio data of the previous frame.

[0046] In the above solution, the transmitting terminal may set a first audio mixing parameter change identifier in the audio optimization metadata. The first audio mixing parameter change identifier indicates whether the first audio mixing parameter corresponding to the first audio data of the current frame has changed compared to the first audio mixing parameter corresponding to the first audio data of the previous frame, and the decoding terminal may then determine whether the first audio mixing parameter corresponding to the first optimized listening area has changed based on the first audio mixing parameter change identifier. The transmitting terminal may also set an optimized listening area change identifier in the audio optimization metadata. The optimized listening area change identifier indicates whether the optimized listening area determined by the producing terminal has changed, and the decoding terminal may then determine whether the optimized listening area has changed based on the optimized listening area change identifier.

[0047] According to a third aspect, an embodiment of the present application provides a method for manufacturing a semiconductor device, comprising: obtaining basic audio metadata and metadata for N optimized listening areas, where N is a positive integer and the N optimized listening areas include a first optimized listening area; rendering M pieces of target audio data based on the first optimized listening area and the basic audio metadata to obtain M pieces of rendered audio data corresponding to the first optimized listening area, where M is a positive integer; performing first audio mixing on the M pieces of rendered audio data to obtain M pieces of first audio mixing data and first audio mixing parameters corresponding to a first optimized listening area; generating audio-optimized metadata based on first metadata of the first optimized listening area and first audio mixing parameters, the audio-optimized metadata including the first metadata and the first audio mixing parameters; The present invention further provides an audio processing method, including:

[0048] In the above solution, the audio optimization metadata in this embodiment of the present application includes first metadata and first audio mixing parameters of the first optimized listening area. Therefore, audio optimization metadata suitable for the user to move freely in the first optimized listening area can be provided, and the listening experience obtained when the user moves freely can be improved.

[0049] In one possible embodiment, the method comprises: mixing the M first audio mixing data to obtain mixed audio data corresponding to a first optimized listening area; performing second audio mixing on the mixed audio data to obtain second audio mixing data corresponding to the first optimized listening area and second audio mixing parameters corresponding to the first optimized listening area; further comprising The step of generating audio-optimizing metadata based on first metadata for the first optimized listening area and first audio mixing parameters includes: generating audio-optimized metadata based on the first metadata for the first optimized listening area, the first audio mixing parameters, and the second audio mixing parameters; Includes.

[0050] In the above solution, when the production terminal performs audio mixing twice, the audio optimization metadata acquired by the production terminal may include first metadata of a first optimized listening area, first audio mixing parameters corresponding to the first optimized listening area, and second audio mixing parameters corresponding to the first optimized listening area. After the audio optimization metadata is acquired by the decoding terminal, the decoding terminal also needs to perform audio mixing twice, and the user's listening experience can be further improved by performing audio mixing twice.

[0051] In one possible implementation, the second audio mixing parameters include an identifier of the second audio mixing data, an equalization parameter, a compressor parameter, and a reverberator parameter.

[0052] In the above solution, the second audio mixing parameters may include an identifier of the second audio mixing data. The second audio mixing parameters may further include equalization parameters. For example, the equalization parameters may include an equalization parameter identifier, a gain value for each frequency band, and a Q value. The second audio mixing parameters may further include compressor parameters. For example, the compressor parameters may include a compressor identifier, a threshold, a compression ratio, a start time, a release time, and a gain compensation value. The second audio mixing parameters may further include reverberator parameters. For example, the reverberator parameters may include a reverberation type, a reverberation time, a delay time, and a direct-to-reverberant sound ratio.

[0053] In one possible implementation, the first audio mixing parameters include an identifier of the first audio mixing data, an equalization parameter, a compressor parameter, and a reverberator parameter.

[0054] In the above solution, the first audio mixing parameters may include an identifier of the first audio mixing data, for example, identifiers of M pieces of first audio mixing data. The first audio mixing parameters may further include equalization parameters. For example, the equalization parameters may include an equalization parameter identifier, a gain value for each frequency band, and a Q value. The first audio mixing parameters may further include compressor parameters. For example, the compressor parameters may include a compressor identifier, a threshold, a compression ratio, a start time, a release time, and a gain compensation value. The first audio mixing parameters may further include reverberator parameters. For example, the reverberator parameters may include a reverberation type, a reverberation time, a delay time, and a direct-to-reverberant sound ratio.

[0055] In one possible implementation, the step of obtaining metadata for the N optimized listening areas comprises: obtaining video image metadata and video image data, the video image metadata including video metadata and image metadata, and the video image data including video data and image data; Rendering video image data based on video image metadata to obtain video scene information; obtaining metadata of the N optimized listening areas based on the video scene information; Includes.

[0056] In the above solution, the production terminal configures N optimized listening areas based on the generated video scene information, so that metadata for the N optimized listening areas can be generated. The video scene information is used to generate metadata for the N optimized listening areas. Therefore, an optimized listening area that better matches the video scene can be selected.

[0057] In one possible implementation, the step of rendering the M pieces of processed audio data based on the first optimized listening area and the basic audio metadata comprises: adjusting the basic audio metadata based on mixed audio data corresponding to a first optimized listening area to obtain adjusted basic audio metadata, where the mixed audio data is obtained by mixing M first audio mixing data; Rendering the M pieces of audio data to be processed based on the first optimized listening area and the adjusted basic audio metadata; Includes.

[0058] In the above solution, the production terminal adjusts the basic audio metadata based on the mixed audio data corresponding to the first optimized listening area to obtain adjusted basic audio metadata. For example, parameters such as the frequency response of one or more audio signals in the audio data or the position and gain of the audio signals in the basic audio metadata are adjusted, thereby adjusting parameters such as the position and gain of the audio data. By adjusting the basic audio metadata, the user's listening experience can be further improved.

[0059] In one possible implementation, the first metadata for the first optimized listening area includes at least one of a reference coordinate system for the first optimized listening area, a center position coordinate for the first optimized listening area, and a shape of the first optimized listening area.

[0060] In the above solution, the metadata of the first optimized listening area may include a reference coordinate system, or the metadata of the first optimized listening area may not include a reference coordinate system. For example, the first optimized listening area uses a default coordinate system. The metadata of the first optimized listening area may include descriptive information for describing the first optimized listening area, such as the center position coordinates of the first optimized listening area and information for describing the shape of the first optimized listening area. In this embodiment of the present application, there may be multiple shapes of the first optimized listening area. For example, the shape may be a sphere, a cube, a pillar, or any other shape.

[0061] According to a fourth aspect, an embodiment of the present application provides a method for manufacturing a semiconductor device comprising: a decoding module configured to decode the audio bitstream to obtain audio-optimization metadata, basic audio metadata, and M decoded audio data, wherein the audio-optimization metadata includes first metadata of a first optimized listening area and first decoded audio mixing parameters corresponding to the first optimized listening area, where M is a positive integer; To obtain M rendered audio data, M Return of a rendering module configured to render the encoded audio data; an audio mixing module configured to perform first audio mixing on the M pieces of rendered audio data based on the first decoded audio mixing parameters to obtain M pieces of first audio mixing data when the current location is within the first optimized listening area; a mixing module configured to mix the M first audio mixing data to obtain mixed audio data corresponding to a first optimized listening area; The present invention further provides a decoding terminal including:

[0062] In a fourth aspect of the present application, the module included in the decoding terminal may further perform the steps described in the first aspect and possible implementations. For details, please refer to the description of the first aspect and possible implementations.

[0063] According to a fifth aspect, an embodiment of the present application comprises: a receiving module configured to receive audio-optimization metadata, basic audio metadata, and M pieces of first audio data, wherein the audio-optimization metadata includes first metadata of a first optimized listening area and first audio mixing parameters corresponding to the first optimized listening area, where M is a positive integer; an encoding module configured to perform compression encoding on the audio-optimized metadata, the basic audio metadata, and the M pieces of first audio data to obtain an audio bitstream; a transmitting module configured to transmit an audio bitstream; The present invention further provides a transmitting terminal including:

[0064] In a fifth aspect of the present application, the module included in the transmitting terminal may further perform the steps described in the second aspect and possible implementations of the second aspect. For details, please refer to the description of the second aspect and possible implementations.

[0065] According to a sixth aspect, an embodiment of the present application comprises: an acquisition module configured to acquire basic audio metadata and metadata for N optimized listening areas, where N is a positive integer and the N optimized listening areas include a first optimized listening area; a rendering module configured to render M pieces of target audio data based on the first optimized listening area and the basic audio metadata to obtain M pieces of rendered audio data corresponding to the first optimized listening area, where M is a positive integer; an audio mixing module configured to perform first audio mixing on the M pieces of rendered audio data to obtain M pieces of first audio mixing data and first audio mixing parameters corresponding to a first optimized listening area; a generation module configured to generate audio-optimized metadata based on first metadata of a first optimized listening area and first audio mixing parameters, the audio-optimized metadata including the first metadata and the first audio mixing parameters; and The present invention further provides a production terminal including:

[0066] In a sixth aspect of the present application, the module included in the production terminal may further perform the steps described in the third aspect and possible implementations. For details, please refer to the description of the third aspect and possible implementations.

[0067] According to a seventh aspect, an embodiment of the present application provides a computer-readable storage medium storing instructions that, when executed on a computer, enable the computer to perform the methods according to the first to third aspects.

[0068] According to an eighth aspect, an embodiment of the present application provides a computer program product comprising instructions which, when run on a computer, enable the computer to perform a method according to the first to third aspects.

[0069] According to a ninth aspect, an embodiment of the present application provides a communication device. The communication device may include an entity such as a terminal device or a chip, and the communication device includes a processor and a memory. The memory is configured to store instructions, and the processor is configured to execute the instructions in the memory such that the communication device performs a method according to any of the first to third aspects.

[0070] According to a tenth aspect, the present application provides a chip system. The chip system includes a processor configured to support a decoding terminal, a transmitting terminal, and a producing terminal in performing the functions of the aforementioned aspects, for example, in transmitting or processing data and / or information in the aforementioned methods. In one possible design, the chip system further includes a memory configured to store program instructions and data necessary for the decoding terminal, the transmitting terminal, and the producing terminal. The chip system may include a chip, or may include a chip and other discrete components.

[0071] According to an eleventh aspect, the present application provides a device including: a receiver configured to receive a bitstream obtained by using a method according to any implementation of the second aspect; and a memory configured to store the bitstream received by the receiver.

[0072] According to a twelfth aspect, the present application provides a computer-readable storage medium for storing a bitstream obtained by using a method according to any implementation of the second aspect.

[0073] According to the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:

[0074] In one embodiment of the present application, the audio bitstream is decoded to obtain audio optimization metadata, basic audio metadata, and M decoded audio data, where the audio optimization metadata includes first metadata of a first optimized listening area and first decoded audio mixing parameters corresponding to the first optimized listening area, and M is a positive integer. To obtain M rendered audio data, M pieces of audio optimization metadata are decoded based on the user's current location and the basic audio metadata. Return of When the user's current location is within the first optimized listening area, first audio mixing is performed on the M rendered audio data based on the first decoded audio mixing parameters to obtain M first audio mixing data. The M first audio mixing data are mixed to obtain mixed audio data corresponding to the first optimized listening area. In this embodiment of the present application, metadata of the first optimized listening area and first decoded audio mixing parameters corresponding to the first optimized listening area may be obtained, and the M first audio mixing data are mixed based on the user's current location and basic audio metadata to obtain M rendered audio data. Return ofThe decoded audio data is rendered. Then, when it is determined that the user's current location is within the first optimized listening area, first audio mixing is performed on the M rendered audio data based on the first decoded audio mixing parameters to obtain M first audio mixing data. Finally, the M first audio mixing data are mixed to obtain mixed audio data corresponding to the first optimized listening area. Therefore, in this embodiment of the present application, when the user's current location is located within the first optimized listening area, both audio mixing and data mixing are performed by using audio data corresponding to the first optimized listening area, so that audio optimization metadata suitable for the user to move freely in the first optimized listening area can be provided, and the listening experience obtained when the user moves freely can be improved.

[0075] In another embodiment of the present application, audio-optimization metadata, basic audio metadata, and M pieces of first audio data are received, where the audio-optimization metadata includes first metadata for a first optimized listening area and first audio mixing parameters corresponding to the first optimized listening area, where M is a positive integer. Compression encoding is performed on the audio-optimization metadata, basic audio metadata, and M pieces of first audio data to obtain an audio bitstream. The audio bitstream is transmitted. In this embodiment of the present application, the audio-optimization metadata, basic audio metadata, and M pieces of first audio data are first received, where the audio-optimization metadata includes metadata for the first optimized listening area and first audio mixing parameters for the first optimized listening area. Therefore, audio-optimization metadata suitable for a user to freely move around the first optimized listening area can be provided, and the listening experience obtained when the user freely moves around can be improved.

[0076] In yet another embodiment of the present application, basic audio metadata and metadata for N optimized listening areas are obtained, where N is a positive integer, and the N optimized listening areas include a first optimized listening area. M pieces of audio data to be processed are rendered based on the first optimized listening area and the basic audio metadata to obtain M pieces of rendered audio data corresponding to the first optimized listening area, where M is a positive integer. First audio mixing is performed on the M pieces of rendered audio data to obtain M pieces of first audio mixing data and first audio mixing parameters corresponding to the first optimized listening area. Audio-optimization metadata is generated based on the first metadata and first audio mixing parameters for the first optimized listening area, and the audio-optimization metadata includes the first metadata and first audio mixing parameters. In this embodiment of the present application, the audio-optimization metadata includes the first metadata and first audio mixing parameters for the first optimized listening area. Therefore, audio optimization metadata suitable for the user to move freely in the first optimized listening area can be provided, and the listening effect obtained when the user moves freely can be improved. [Brief explanation of the drawings]

[0077] [Figure 1] 1 is a schematic diagram of the configuration structure of an audio processing system according to an embodiment of the present application; [Figure 2A] 3 is a schematic flowchart of the interaction between a producing terminal, a sending terminal and a decoding terminal according to an embodiment of the present application; [Figure 2B] 3 is a schematic flowchart of the interaction between a producing terminal, a sending terminal and a decoding terminal according to an embodiment of the present application; [Figure 3]1 is a schematic flowchart of streaming data processing in a virtual reality streaming service system according to an embodiment of the present application; [Figure 4] 1 is an end-to-end flowchart of a 6DoFVR music scene according to an embodiment of the present application. [Figure 5] FIG. 1 is a schematic diagram of a VR concert scene supporting 6DoF according to an embodiment of the present application. [Figure 6] 10 is an end-to-end flowchart of another 6DoFVR music scene according to an embodiment of the present application. [Figure 7] FIG. 2 is a schematic diagram of the configuration structure of a decoding terminal according to an embodiment of the present application; [Figure 8] FIG. 2 is a schematic diagram of the configuration structure of a transmitting terminal according to an embodiment of the present application; [Figure 9] 1 is a schematic diagram of the configuration structure of a production terminal according to an embodiment of the present application; [Figure 10] FIG. 10 is a schematic diagram of the configuration structure of another decoding terminal according to an embodiment of the present application; [Figure 11] FIG. 10 is a schematic diagram of the configuration structure of another transmitting terminal according to an embodiment of the present application; [Figure 12] FIG. 10 is a schematic diagram of the configuration structure of another production terminal according to an embodiment of the present application; DETAILED DESCRIPTION OF THE INVENTION

[0078] The embodiments of the present application provide an audio processing method and a terminal for improving the listening experience obtained when a user moves freely.

[0079] The following describes embodiments of the present application with reference to the accompanying drawings.

[0080] In the specification, claims, and accompanying drawings of this application, terms such as "first," "second," etc. are used to distinguish between similar objects and do not necessarily dictate a particular order or sequence. It should be understood that terms so used are interchangeable under appropriate circumstances and are merely a distinguishing mechanism used when objects having the same attributes are described in the embodiments of this application. Additionally, the terms "include" and "have," as well as any variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, system, product, or device that includes a set of units is not necessarily limited to those units and may include other units not expressly listed or inherent in such process, method, system, product, or device.

[0081] The production process of a musical work includes the following steps: writing lyrics, arranging, recording, audio mixing, master tape, etc. Audio mixing is an essential step in music production, and the quality of audio mixing determines the success or failure of a musical work.

[0082] Audio mixing is the integration of multiple sound sources, such as violins, drums, humans, or other recorded sounds, into a single, stereo, or multi-channel track. The audio mixing process requires adjusting the frequency, dynamics, timbre, placement, and sound field of each source signal separately to optimize the signal for each audio track. Finally, the frequency and dynamics of the mixed signal are adjusted to optimize the auditory effect of the mixed signal. Mixers include equalizers, compressors, and reverberators. Audio mixing allows listeners to hear subtle and layered musical effects that cannot be heard in live recordings, making music more expressive.

[0083] With the development of virtual reality (VR), augmented reality (AR), and mixed reality (MR), virtual reality technology is gradually being applied to the music industry. Various VR music scenes have emerged, including VR music video scenes, VR concert live scenes, and various VR music programs. Compared to traditional music, these VR music scenes combine 3D spatial music effects with VR visual experiences, creating a more lively and immersive experience, significantly improving the user's music experience. In most current VR music scenes, the music effects in 3DoF scenes only support the user's head rotation effect, and do not support six degrees of freedom (6DoF) scenes.

[0084] VR hardware devices are becoming increasingly mature. Users have higher requirements for music experiences. Therefore, VR music scenes supporting 6DoF will become a trend in the music industry in the future. In traditional music production, the user is usually in a desk position during audio mixing on the production side, and the user's position is assumed to be constant. Audio mixing is completed before the music signal is transmitted, and the transmitted music signal is the mixed signal obtained through audio mixing. On the user side (in other words, the audio decoder side), the audio renderer only needs to adapt the user's playback device, so the user can experience the complete music effect. In a VR music scene supporting 6DoF, the user can move freely within the scene. A violin sound source is used as an example. The volume, timbre, and reverberation of the violin sound heard by a user 3 meters (m) away from the violin are significantly different from those heard 0.5 meters away. Because the user's position is constantly changing, the user's actual position cannot be determined during audio mixing. Therefore, traditional music production methods cannot guarantee that users can hear the full musical effect when they move freely, which may significantly affect the user experience in VR music scenes that support 6DoF.

[0085] According to an embodiment of the present application, the listening experience obtained when a user moves freely can be improved. For example, the listening experience of a user can be improved when the user moves freely within a virtual reality scene or an augmented reality scene. The following describes the embodiment of the present application in detail. As shown in FIG. 1 , one embodiment of the present application provides an audio processing system 100 including a production terminal 101, a transmission terminal 102, and a decoding terminal 103. The production terminal 101 can select one or more optimized listening areas from a virtual scene. The optimized listening areas may also be called "sweet spots," and the optimized listening areas are pre-selected listening areas from the virtual scene. The production terminal 101 can configure audio optimization metadata for each optimized listening area. Specifically, the production terminal 101 can generate a set of audio mixing parameters corresponding to each optimized listening area to ensure that a listener can hear the musical effect of a music signal obtained by audio mixing in the optimized listening area. In this way, the user's music experience in a 6DoF music scene is improved.

[0086] The production terminal 101 may communicate with the transmitting terminal 102, which may communicate with the decoding terminal 103. The transmitting terminal 102 may receive audio optimization metadata for each optimized listening area from the production terminal 101, and the transmitting terminal 102 may perform compression encoding on the audio optimization metadata to obtain an audio bitstream. The transmitting terminal 102 may transmit the audio bitstream to the decoding terminal 103. The decoding terminal 103 may obtain the audio optimization metadata for each optimized listening area, and the decoding terminal 103 may select an optimized listening area that matches the user's current location based on the user's current location (e.g., the matching optimized listening area is referred to as the first optimized listening area). Audio mixing is then performed by using first decoded audio mixing parameters corresponding to the first optimized listening area, so that the user hears the music effect of the music signal obtained by audio mixing, thereby improving the user's music experience in the 6DoF music scene.

[0087] In this embodiment of the present application, the production terminal may include 6DoF audio VR music software, 3D audio engine, etc. The production terminal may be used in VR terminal devices, chips, and wireless network devices.

[0088] The transmitting terminal may be used in a terminal device requiring audio communication, or a wireless device and a core network device requiring transcoding. For example, the transmitting terminal may be an audio encoder of a terminal device, a wireless device, or a core network device. For example, the audio encoder may include a media gateway in a radio access network or a core network, a transcoding device, a media resource server, a mobile terminal, a fixed network terminal, etc. Alternatively, the audio encoder may be an audio encoder applied to streaming media services in virtual reality technology.

[0089] Similarly, the decoding terminal may be used in a terminal device that requires audio communication, or in a wireless device and a core network device that requires transcoding. For example, the decoding terminal may be an audio decoder of a terminal device, a wireless device, or a core network device.

[0090] An audio processing method provided in an embodiment of the present application will be described first. The audio processing method is implemented based on the audio processing system of Fig. 1. Fig. 2A and Fig. 2B are schematic flowcharts of interactions between a production terminal, a transmitting terminal, and a decoding terminal according to an embodiment of the present application. The production terminal may communicate with the transmitting terminal, and the transmitting terminal may communicate with the decoding terminal. The production terminal performs the following steps 201 to 204, the transmitting terminal performs the following steps 205 to 207, and the decoding terminal performs the following steps 208 to 211.

[0091] 201: A production terminal obtains basic audio metadata and metadata of N optimized listening areas, where N is a positive integer, and the N optimized listening areas include a first optimized listening area.

[0092] The basic audio metadata is the basic metadata required when creating a VR music scene, and the components and contents of the basic audio metadata are not limited. For example, as shown in Table 1, the basic audio metadata includes at least one of sound source metadata, physical model metadata, acoustic metadata, moving object metadata, interaction metadata, and resource metadata.

[0093] [Table 1]

[0094] Specifically, sound source metadata is used to describe attributes of a sound source. For example, sound source metadata may include target audio metadata, multi-channel audio metadata, and scene audio metadata. The target audio metadata and multi-channel audio metadata include information such as the reference coordinate system, position, gain, volume, shape, directivity, attenuation mode, and playback control of the sound source. The scene audio metadata includes the position and reference coordinate system of the scene microphone, the gain of the scene audio, the effective area, the playback support degree of freedom type (0 / 3 / 6DoF), attenuation mode, and playback control.

[0095] The physical model metadata includes a sphere model, a cylinder model, a cube model, a triangle mesh model, etc. The sphere model, the cylinder model, and the cube model are used to describe the shape of objects in a virtual room, etc. The triangle mesh model can be used to describe rooms of arbitrary shapes and irregular objects in a scene.

[0096] Acoustic metadata includes acoustic material metadata and acoustic environment metadata. Acoustic material metadata is used to describe the acoustic properties of the surface materials of objects and rooms in a scene, while acoustic environment metadata is used to describe the reverberation information of rooms in a VR scene.

[0097] Moving object metadata is used to describe the movement information of sound sources, objects, etc. in the scene. Interaction metadata is used to describe the interaction behavior between the user and the VR scene.

[0098] Resource metadata is used to describe resource information required in a VR scene.

[0099] The metadata used in most VR music scenes can be specifically covered by the metadata in Table 1 above.

[0100] Also, in this embodiment of the present application, in addition to obtaining basic audio metadata, the production terminal may further obtain N optimized listening areas from the virtual scene. The value of N is not limited. For example, N may be equal to 1, or N may be greater than 1. The production terminal obtains metadata of the N optimized listening areas. The metadata of the optimized listening areas includes configuration parameters of the optimized listening areas. For example, the configuration parameters may be parameters such as the size, shape, or center position of the listening area. The configuration parameters included in the metadata of the optimized listening areas are not limited.

[0101] For example, N optimized listening areas may cover different locations of a user, and the N optimized listening areas may include a first optimized listening area, which may refer to an optimized listening area that coincides with the user's current location.

[0102] In some embodiments of the present application, the first optimized listening area may be any optimized listening area among the N optimized listening areas. The first metadata of the first optimized listening area includes at least one of a reference coordinate system of the first optimized listening area, a center position coordinate of the first optimized listening area, and a shape of the first optimized listening area.

[0103] In particular, the metadata for the first optimized listening area may include a reference coordinate system, or the metadata for the first optimized listening area may not include a reference coordinate system, e.g., the first optimized listening area uses a default coordinate system.

[0104] The metadata of the first optimized listening area may include descriptive information for describing the first optimized listening area, such as the center position coordinates of the first optimized listening area, and information for describing the shape of the first optimized listening area. In this embodiment of the present application, there may be multiple shapes of the first optimized listening area. For example, the shape may be a sphere, a cube, a pillar, or any other shape.

[0105] In some embodiments of the present application, the production terminal obtaining metadata of the N optimized listening areas in step 201 includes the following steps:

[0106] A1: The production terminal obtains video image metadata and video image data, where the video image metadata includes video metadata and image metadata, and the video image data includes video data and image data.

[0107] The production terminal may further acquire video image metadata and video image data in the virtual scene. The video image metadata may also be referred to as video and image metadata, and the video image data may also be referred to as video and image data. The video image data includes video and image data content, and the video image metadata is information used to describe attributes of the video and image content.

[0108] A2: The production terminal renders the video image data based on the video image metadata to obtain the video scene information.

[0109] The production terminal performs video scene rendering on the video image data by using the video image metadata to obtain video scene information, for example, the video scene may be a virtual scene.

[0110] A3: The production terminal obtains metadata of N optimized listening areas based on the video scene information.

[0111] The production terminal configures N optimized listening areas based on the generated video scene information, so that metadata for the N optimized listening areas can be generated. The video scene information is used to generate metadata for the N optimized listening areas. Thus, an optimized listening area that better matches the video scene can be selected.

[0112] 202: The production terminal renders M pieces of audio data to be processed based on the first optimized listening area and the basic audio metadata to obtain M pieces of rendered audio data corresponding to the first optimized listening area, where M is a positive integer.

[0113] The producing terminal obtains M pieces of first audio data to be processed. The M pieces of first audio data to be processed are audio data that need to be sent to the decoding terminal. The value of M is not limited. For example, M may be equal to 1, or M may be greater than 1.

[0114] After obtaining the M pieces of audio data to be processed, the production terminal renders each optimized listening area to obtain rendered audio data corresponding to each optimized listening area. For example, the production terminal renders the M pieces of audio data to be processed based on the first optimized listening area in the N optimized listening areas and the basic audio metadata to obtain M pieces of rendered audio data corresponding to the first optimized listening area in the N optimized listening areas.

[0115] It should be noted that the second audio data obtained by rendering may be a single-channel signal or a binaural rendering signal. N optimized listening areas have a total of N*M second audio data obtained by rendering, where * indicates a multiplication operation symbol.

[0116] It should be noted that in addition to the first optimized listening area, the N optimized listening areas may further include a second optimized listening area. The method provided in this embodiment of the present application may further include the following steps:

[0117] The production terminal renders M pieces of audio data to be processed based on the second optimized listening area and the basic audio metadata to obtain M pieces of rendered audio data corresponding to the second optimized listening area, where M is a positive integer.

[0118] The rendering performed by the production terminal based on the second optimized listening area is similar to the rendering performed based on the first optimized listening area in step 201, and the details will not be described again here. Similarly, the subsequent steps 203 and 204 are also processes performed for the first optimized listening area, and the same processes as those in steps 203 and 204 can also be performed for the second optimized listening area, and the details will not be described again here.

[0119] 203: The production terminal performs first audio mixing on the M pieces of rendered audio data to obtain M pieces of first audio mixing data and first audio mixing parameters corresponding to the first optimized listening area.

[0120] After obtaining the M pieces of rendered audio data corresponding to the first optimized listening area, the production terminal may further perform first audio mixing on the M pieces of rendered audio data corresponding to the first optimized listening area for the first optimized listening area to obtain the M pieces of first audio mixing data and first audio mixing parameters corresponding to the first optimized listening area. The first audio mixing parameters are used to record the audio mixing parameters used during the first audio mixing, and the audio mixing parameters may also be referred to as "audio mixing metadata." The aforementioned audio mixing step may be completed by the VR music scene production terminal or by the audio mixing terminal. This is not limited here.

[0121] It should be noted that the M pieces of first audio mixing data in step 203 are audio data obtained by performing first audio mixing by the producing terminal, and the M pieces of first audio mixing data and the M pieces of first audio mixing data obtained by subsequently performing first audio mixing by the decoding terminal are different audio data.

[0122] In some embodiments of the present application, the first audio mixing parameters include at least one of an identifier of the first audio mixing data, an equalization parameter, a compressor parameter, and a reverberator parameter.

[0123] The first audio mixing parameters may include an identifier of the first audio mixing data, e.g., identifiers of M pieces of first audio mixing data. The first audio mixing parameters may further include equalization parameters. For example, the equalization parameters may include an equalization parameter identifier, a gain value for each frequency band, and a Q value. The Q value is a parameter of an equalization filter, represents a quality coefficient of the equalization filter, and may be used to describe the bandwidth of the equalization filter. The first audio mixing parameters may further include compressor parameters. For example, the compressor parameters may include a compressor identifier, a threshold, a compression ratio, a start time, a release time, and a gain compensation value. The first audio mixing parameters may further include reverberator parameters. For example, the reverberator parameters may include a reverberation type, a reverberation time, a delay time, and a direct-to-reverberant sound ratio.

[0124] 204: The production terminal generates audio optimization metadata based on the first metadata and the first audio mixing parameters of the first optimized listening area, where the audio optimization metadata includes the first metadata and the first audio mixing parameters.

[0125] After obtaining the first audio mixing parameters corresponding to the first optimized listening area, the production terminal may generate audio optimization metadata for the first optimized listening area. The audio optimization metadata differs from the aforementioned basic audio metadata in that the audio optimization metadata includes first metadata of the first optimized listening area and first audio mixing parameters corresponding to the first optimized listening area. The audio optimization metadata is used to optimize the music signal to be heard by a user whose current location is within the first optimized listening area to improve the musical effect of the music signal.

[0126] In some embodiments of the present application, the audio processing method that may be performed by the production terminal further includes the following steps:

[0127] B1: The production terminal mixes M first audio mixing data to obtain mixed audio data corresponding to a first optimized listening area.

[0128] B2: Perform second audio mixing on the mixed audio data to obtain second audio mixing data corresponding to the first optimized listening area and second audio mixing parameters corresponding to the first optimized listening area.

[0129] Specifically, after the production terminal performs the first audio mixing in step 203, the production terminal may further perform steps B1 and B2 to further improve the audio mixing effect of the audio data. For the first optimized listening area, the production terminal may mix M first audio mixing data to obtain mixed audio data corresponding to the first optimized listening area. Then, the production terminal may perform second audio mixing on the mixed audio data to obtain second audio mixing data corresponding to the first optimized listening area and second audio mixing parameters corresponding to the first optimized listening area. The aforementioned audio mixing steps may be completed by the VR music scene production terminal or by the audio mixing terminal. This is not limited here.

[0130] In an implementation scenario that performs steps B1 and B2, in step 204, the production terminal generating audio optimization metadata based on the first metadata and first audio mixing parameters of the first optimized listening area includes:

[0131] The production terminal generates audio optimization metadata based on the first metadata, the first audio mixing parameters, and the second audio mixing parameters.

[0132] When the production terminal performs audio mixing twice, the audio optimization metadata acquired by the production terminal may include first metadata of a first optimized listening area, first audio mixing parameters corresponding to the first optimized listening area, and second audio mixing parameters corresponding to the first optimized listening area. After the audio optimization metadata is acquired by the decoding terminal, the decoding terminal also needs to perform audio mixing twice, and the user's listening experience can be further improved by performing audio mixing twice.

[0133] In some embodiments of the present application, the second audio mixing parameters include an identifier of the second audio mixing data, an equalization parameter, a compressor parameter, and a reverberator parameter.

[0134] The second audio mixing parameters may include an identifier of the second audio mixing data. The second audio mixing parameters may further include an equalization parameter. For example, the equalization parameter may include an equalization parameter identifier: each The second audio mixing parameters may include gain values ​​of frequency bands and Q values. The second audio mixing parameters may further include compressor parameters. For example, the compressor parameters may include a compressor identifier, a threshold, a compression ratio, a start time, a release time, and a gain compensation value. The second audio mixing parameters may further include reverberator parameters. For example, the reverberator parameters may include a reverberation type, a reverberation time, a delay time, and a direct-to-reverberant sound ratio.

[0135] In some embodiments of the present application, the audio processing method that may be performed by the production terminal further includes the following steps:

[0136] C1: The producing terminal sends M pieces of first audio data, basic audio metadata, and audio optimization metadata to the sending terminal.

[0137] The producing terminal may transmit the M pieces of first audio data, the basic audio metadata, and the audio-optimization metadata together to the transmitting terminal, or the producing terminal may transmit the M pieces of first audio data, the basic audio metadata, and the audio-optimization metadata separately to the transmitting terminal. The specific transmission manner is not limited in this specification. The transmitting terminal may further transmit the M pieces of first audio data, the basic audio metadata, and the audio-optimization metadata to the decoding terminal, and the decoding terminal receives the M pieces of first audio data, the basic audio metadata, and the audio-optimization metadata.

[0138] In some embodiments of the present application, the production terminal may further adjust the basic audio metadata. Specifically, in the aforementioned step 202, the M pieces of audio data to be processed are rendered based on the first optimized listening area and the basic audio metadata, including the following steps:

[0139] D1: The production terminal adjusts the basic audio metadata based on mixed audio data corresponding to the first optimized listening area to obtain adjusted basic audio metadata, and the mixed audio data is obtained by mixing M first audio mixing data.

[0140] D2: The production terminal renders the M pieces of audio data to be processed based on the first optimized listening area and the adjusted basic audio metadata.

[0141] Specifically, in D1, for a first optimized listening area, the production terminal may mix M first audio mixing data to obtain mixed audio data corresponding to the first optimized listening area. The production terminal adjusts the basic audio metadata based on the mixed audio data corresponding to the first optimized listening area to obtain adjusted basic audio metadata. For example, parameters such as the frequency response of one or more audio signals in the audio data or the position and gain of an audio signal in the basic audio metadata are adjusted, thereby adjusting parameters such as the position and gain of the audio data, and adjusting the musical effect of the music signal finally heard by the user. In D2, the production terminal renders M audio data to be processed by using the first optimized listening area and the adjusted basic audio metadata. The user's listening experience can be further improved by adjusting the basic audio metadata.

[0142] The producing terminal can obtain the audio-optimized metadata by performing the above steps 201 to 204, and then the producing terminal sends the audio-optimized metadata to the sending terminal, which performs the following steps 205 to 207.

[0143] 205: The transmitting terminal receives audio optimization metadata, basic audio metadata, and M pieces of first audio data, where the audio optimization metadata includes first metadata of a first optimized listening area and first audio mixing parameters corresponding to the first optimized listening area, and M is a positive integer.

[0144] The process of generating audio-optimized metadata is described in detail in steps 201 to 204, which are performed by the production terminal. After the production terminal generates the audio-optimized metadata, the production terminal may further transmit the audio-optimized metadata to the sending terminal, and the sending terminal receives the audio-optimized metadata from the production terminal. The production terminal may also further transmit basic audio metadata and M pieces of first audio data to the sending terminal, and the sending terminal receives the basic audio metadata and M pieces of first audio data from the production terminal.

[0145] 206: The sending terminal performs compression encoding on the audio optimization metadata, the basic audio metadata, and the M pieces of first audio data to obtain an audio bitstream.

[0146] After receiving the audio-optimized metadata, the basic audio metadata, and the M pieces of first audio data, the transmitting terminal may perform compression encoding on the audio-optimized metadata, the basic audio metadata, and the M pieces of first audio data by using a preset encoding algorithm to obtain an audio bitstream. The encoding algorithm used is not limited in this embodiment of the present application.

[0147] 207: The sending terminal sends the audio bitstream.

[0148] The sending terminal transmits the audio bitstream by using a transmission channel between the sending terminal and the decoding terminal.

[0149] In some embodiments of the present application, the audio processing method that may be performed by the transmitting terminal further includes the following steps.

[0150] E1: The sending terminal receives video image metadata and video image data from the producing terminal, where the video image metadata includes video metadata and image metadata, and the video image data includes video data and image data.

[0151] E2: The sending terminal performs compression encoding on the video image metadata and the video image data to obtain a video image bitstream.

[0152] E3: The sending terminal sends the video image bitstream to the decoding terminal.

[0153] The producing terminal may further transmit the video image metadata and the video image data to the transmitting terminal. After receiving the video image metadata and the video image data, the transmitting terminal may generate a video image bitstream. The video image bitstream carries the video image metadata and the video image data. Thus, after receiving the video image bitstream from the transmitting terminal, the decoding terminal may obtain the video image metadata and the video image data.

[0154] In some embodiments of the present application, the audio optimization metadata further includes a second audio mixing parameter change identifier.

[0155] The second audio mixing parameter change identifier indicates whether the second audio mixing parameters corresponding to the first audio data of the current frame have changed compared to the second audio mixing parameters corresponding to the first audio data of the previous frame.

[0156] The transmitting terminal may set a second audio mixing parameter change identifier in the audio optimization metadata. The second audio mixing parameter change identifier indicates whether the second audio mixing parameter corresponding to the first optimized listening area has changed. Thus, the decoding terminal determines whether the second audio mixing parameter corresponding to the first optimized listening area has changed based on the second audio mixing parameter change identifier. For example, if the second audio mixing parameter corresponding to the first audio data of the current frame has changed compared to the second audio mixing parameter corresponding to the first audio data of the previous frame, the second audio mixing parameter change identifier is TRUE, and the transmitting terminal may further transmit modification information of the second audio mixing parameter corresponding to the first audio data. The decoding terminal receives the modification information of the second audio mixing parameter corresponding to the first audio data and obtains the modified second audio mixing parameter corresponding to the first audio data of the current frame based on the modification information.

[0157] In some embodiments of the present application, the audio optimization metadata further includes second audio mixing parameters corresponding to the first optimized listening area.

[0158] When the production terminal performs audio mixing twice, the audio optimization metadata acquired by the production terminal may include first metadata of a first optimized listening area, first audio mixing parameters corresponding to the first optimized listening area, and second audio mixing parameters corresponding to the first optimized listening area. After the audio optimization metadata is acquired by the decoding terminal, the decoding terminal also needs to perform audio mixing twice, and the user's listening experience can be improved by performing audio mixing twice.

[0159] In some embodiments of the present application, the audio optimization metadata is 1st Optimized Listening Area and differential parameters of N-1 second audio mixing parameters of N-1 optimized listening areas other than the first optimized listening area within the N optimized listening areas relative to the second audio mixing parameters corresponding to the first optimized listening area.

[0160] The difference parameters are parameters of differences between the N-1 second audio mixing parameters of the N-1 optimized listening areas other than the first optimized listening area among the N optimized listening areas and the second audio mixing parameters corresponding to the first optimized listening area. The difference parameters are not the N-1 second audio mixing parameters of the N-1 optimized listening areas. For example, the second audio mixing parameters corresponding to the first optimized listening area include Parameter 1, Parameter 2, and Parameter 3. If the second audio mixing parameters corresponding to each of the N-1 second audio mixing parameters corresponding to the N-1 optimized listening areas include Parameter 1, Parameter 2, and Parameter 4, the difference parameter between the N-1 second audio mixing parameters corresponding to the N-1 optimized listening areas and the second audio mixing parameters corresponding to the first optimized listening area includes Parameter 4. The audio-optimized metadata carries differential parameters, so that the data amount of the audio-optimized metadata can be reduced, and data transmission efficiency and decoding efficiency can be improved.

[0161] In some embodiments of the present application, the second audio mixing parameters include an identifier of the first audio data, an equalization parameter, a compressor parameter, and a reverberator parameter.

[0162] The second audio mixing parameters may include an identifier of the first audio data, e.g. M pieces The second audio mixing parameters may include an identifier of the first audio data. The second audio mixing parameters may further include equalization parameters. For example, the equalization parameters may include an equalization parameter identifier, a gain value for each frequency band, and a Q value. The second audio mixing parameters may further include compressor parameters. For example, the compressor parameters may include a compressor identifier, a threshold, a compression ratio, a start time, a release time, and a gain compensation value. The second audio mixing parameters may further include reverberator parameters. For example, the reverberator parameters may include a reverberation type, a reverberation time, a delay time, and a direct sound to reverberation sound ratio.

[0163] In some embodiments of the present application, the audio optimization metadata further includes N-1 difference parameters of the N-1 first audio mixing parameters corresponding to N-1 optimized listening areas other than the first optimized listening area within the N optimized listening areas, relative to the first audio mixing parameter corresponding to the first optimized listening area.

[0164] The difference parameters are parameters of differences between the N-1 first mixing parameters of the N-1 optimized listening areas other than the first optimized listening area among the N optimized listening areas and the first mixing parameters corresponding to the first optimized listening area. The difference parameters are not the N-1 first mixing parameters of the N-1 optimized listening areas. For example, the first audio mixing parameters corresponding to the first optimized listening area include Parameter 1, Parameter 2, and Parameter 3. If the first audio mixing parameters corresponding to each of the N-1 first audio mixing parameters corresponding to the N-1 optimized listening areas include Parameter 1, Parameter 2, and Parameter 4, the difference parameter between the N-1 first audio mixing parameters corresponding to the N-1 optimized listening areas and the first audio mixing parameters corresponding to the first optimized listening area includes Parameter 4. The audio-optimized metadata carries differential parameters, so that the data amount of the audio-optimized metadata can be reduced, and data transmission efficiency and decoding efficiency can be improved.

[0165] In some embodiments of the present application, the first audio mixing parameters include an identifier of the first audio data, an equalization parameter, a compressor parameter, and a reverberator parameter.

[0166] The first audio mixing parameters may include an identifier of the first audio data, for example, identifiers of M pieces of the first audio data. The first audio mixing parameters may further include equalization parameters. For example, the equalization parameters may include an equalization parameter identifier, a gain value for each frequency band, and a Q value. The first audio mixing parameters may further include compressor parameters. For example, the compressor parameters may include a compressor identifier, a threshold, a compression ratio, a start time, a release time, and a gain compensation value. The first audio mixing parameters may further include reverberator parameters. For example, the reverberator parameters may include a reverberation type, a reverberation time, a delay time, and a direct-to-reverberant sound ratio.

[0167] In some embodiments of the present application, the metadata of the first optimized listening area includes at least one of a reference coordinate system of the first optimized listening area, a center position coordinate of the first optimized listening area, and a shape of the first optimized listening area.

[0168] In particular, the metadata for the first optimized listening area may include a reference coordinate system, or the metadata for the first optimized listening area may not include a reference coordinate system, e.g., the first optimized listening area uses a default coordinate system.

[0169] The metadata of the first optimized listening area may include descriptive information for describing the first optimized listening area, such as the center position coordinates of the first optimized listening area, and information for describing the shape of the first optimized listening area. In this embodiment of the present application, there may be multiple shapes of the first optimized listening area. For example, the shape may be a sphere, a cube, a pillar, or any other shape.

[0170] In some embodiments of the present application, the audio optimization metadata further includes a center position coordinate of a first optimized listening area among the N optimized listening areas, and position offsets of center position coordinates of N-1 optimized listening areas other than the first optimized listening area among the N optimized listening areas relative to the center position coordinate of the first optimized listening area.

[0171] The position offset is an offset between the center position coordinate of N-1 optimized listening areas other than the first optimized listening area among the N optimized listening areas and the center position coordinate of the first optimized listening area that is not other than the center position coordinate of the N-1 optimized listening areas. The audio-optimized metadata carries the position offset, so that the data amount of the audio-optimized metadata can be reduced, and data transmission efficiency and decoding efficiency can be improved.

[0172] In some embodiments of the present application, the audio optimization metadata further includes an optimized listening area change identifier and / or a first audio mixing parameter change identifier.

[0173] The optimized listening area change identifier indicates whether the first optimized listening area has changed.

[0174] The first audio mixing parameter change identifier indicates whether the first audio mixing parameter corresponding to the first audio data of the current frame has changed compared to the first audio mixing parameter corresponding to the first audio data of the previous frame.

[0175] The transmitting terminal may set a first audio mixing parameter change identifier in the audio optimization metadata. The first audio mixing parameter change identifier indicates whether a first audio mixing parameter corresponding to the first audio data of the current frame has changed compared to a first audio mixing parameter corresponding to the first audio data of the previous frame, and the decoding terminal then determines whether a first audio mixing parameter corresponding to the first optimized listening area has changed based on the first audio mixing parameter change identifier. The transmitting terminal may also set an optimized listening area change identifier in the audio optimization metadata. The optimized listening area change identifier indicates whether an optimized listening area determined by the producing terminal has changed, and the decoding terminal then determines whether the optimized listening area has changed based on the optimized listening area change identifier. For example, the optimized listening area metadata change identifier and the first audio mixing parameter change identifier are added to the encoded 6DoF audio optimization metadata to improve transmission efficiency of the 6DoF audio optimization metadata. When the VR music scene is initialized, the initial audio optimization metadata is transmitted. When the VR scene changes and the optimized listening area position and shape information changes, the optimized listening area change identifier is set to true and the optimized listening area change information is transmitted. When the first audio mixing parameter of the current frame changes, the first audio mixing parameter change identifier is set to true and the first audio mixing parameter change information is transmitted.

[0176] The transmitting terminal performs the above steps 205 to 207, and the decoding terminal performs the subsequent steps 208 to 211. It will be understood that the audio processing process performed by the decoding terminal is similar to the audio processing process performed by the producing terminal. The following describes the audio processing process performed by the decoding terminal.

[0177] 208: The decoding terminal decodes the audio bitstream to obtain audio optimization metadata, basic audio metadata, and M pieces of decoded audio data, where the audio optimization metadata includes first metadata of a first optimized listening area and first decoded audio mixing parameters corresponding to the first optimized listening area, and M is a positive integer.

[0178] The process of generating an audio bitstream is detailed in steps 205 to 207, which are performed by the transmitting terminal. The transmitting terminal transmits the audio bitstream to the decoding terminal, which then receives M Return of The present invention is directed to an audio terminal that receives an audio bitstream from a sending terminal to obtain decoded audio data, audio-optimized metadata, and basic audio metadata. The M decoded audio data correspond to the M audio data to be processed at the producing terminal. For descriptions of the M decoded audio data, audio-optimized metadata, and basic audio metadata, please refer to the above-mentioned embodiment, and the details will not be described again here.

[0179] 209: The decoding terminal performs M operations based on the user's current location and the basic audio metadata to obtain M pieces of rendered audio data. Return of Renders the encoded audio data.

[0180] M decoding terminals Return of After obtaining the decoded audio data, the audio optimization metadata, and the basic audio metadata, the decoding terminal performs M decoding operations based on the user's current location and the basic audio metadata to obtain M rendered audio data. Return of Renders the encoded audio data.

[0181] It should be noted that the M pieces of rendered audio data in step 209 are audio data obtained by a decoding terminal by performing rendering. The M pieces of rendered audio data and the M pieces of rendered audio data obtained by a production terminal by performing rendering are different audio data.

[0182] 210: When the user's current location is within the first optimized listening area, the decoding terminal performs first audio mixing on the M rendered audio data based on the first decoding audio mixing parameters to obtain M first audio mixing data.

[0183] The decoding terminal obtains an optimized listening area that matches the user's current location from the N optimized listening areas based on the user's current location, and the optimized listening area that matches the current location is referred to as the first optimized listening area. In step 208, the decoding terminal obtains audio optimization metadata, where the audio optimization metadata includes first decoded audio mixing parameters corresponding to the first optimized listening area. Thus, the first decoded audio mixing parameters corresponding to the first optimized listening area can be obtained from the audio optimization metadata. The decoding terminal performs first audio mixing on the M rendered audio data based on the first decoded audio mixing parameters corresponding to the first optimized listening area to obtain M first audio mixing data corresponding to the first optimized listening area. The first decoded audio mixing parameters correspond to first audio mixing parameters of the production terminal, and the first audio mixing parameters are used to record the audio mixing parameters used when the production terminal performs the first audio mixing. The aforementioned audio mixing step can be completed by a VR music scene production terminal or an audio mixing terminal, which is not limited here.

[0184] 211: The decoding terminal mixes the M first audio mixing data to obtain mixed audio data corresponding to a first optimized listening area.

[0185] After obtaining the M pieces of first audio mixing data corresponding to the first optimized listening area, the decoding terminal mixes the M pieces of first audio mixing data corresponding to the first optimized listening area to obtain mixed audio data corresponding to the first optimized listening area. Since the first optimized listening area is an optimized listening area that includes the user's current location, the decoding terminal mixes the M pieces of first audio mixing data corresponding to the first optimized listening area to obtain mixed audio data corresponding to the first optimized listening area. The first optimized listening area can be adapted to the user's actual location. Therefore, audio optimization metadata suitable for the user to move freely in the first optimized listening area can be provided, and the listening experience obtained when the user moves freely can be improved.

[0186] It should be noted that the mixed audio data may be directly used for playback, and the user's listening effect can be improved when the mixed audio data is played back.

[0187] In some embodiments of the present application, the audio optimization metadata further includes second decoded audio mixing parameters corresponding to the first optimized listening area.

[0188] The second decoded audio mixing parameters correspond to the second audio mixing parameters on the production terminal side, and the second audio mixing parameters are used to record the audio mixing parameters used during the second audio mixing.

[0189] The audio processing method that may be performed by the decoding terminal further includes the following steps:

[0190] F1: The decoding terminal performs second audio mixing on the mixed audio data based on the second decoding audio mixing parameters to obtain second audio mixing data corresponding to the first optimized listening area.

[0191] After obtaining the second decoded audio mixing parameters, the decoding terminal may further perform second audio mixing on the mixed audio data based on the second decoded audio mixing parameters to obtain second audio mixing data corresponding to the first optimized listening area. The second audio mixing data may be obtained through the second audio mixing. When the second audio mixing data is played, the user's listening experience may be improved. The audio mixing step may be completed by a VR music scene creation terminal or an audio mixing terminal. This is not limited here.

[0192] When the production terminal performs audio mixing twice, the audio optimization metadata acquired by the production terminal may include: first metadata of a first optimized listening area, first audio mixing parameters corresponding to the first optimized listening area, and second audio mixing parameters corresponding to the first optimized listening area. After the audio optimization metadata is acquired by the decoding terminal, the decoding terminal also needs to perform audio mixing twice, and the user's listening experience can be improved by performing audio mixing twice.

[0193] In some embodiments of the present application, the second decoded audio mixing parameters include an identifier of the second audio mixing data, an equalization parameter, a compressor parameter, and a reverberator parameter.

[0194] The second decoded audio mixing parameters may include an identifier of the second audio mixing data. The second decoded audio mixing parameters may further include equalization parameters. For example, the equalization parameters may include an equalization parameter identifier, a gain value for each frequency band, and a Q factor. The second decoded audio mixing parameters may further include compressor parameters. For example, the compressor parameters may include a compressor identifier, a threshold, a compression ratio, a start time, a release time, and a gain compensation value. The second decoded audio mixing parameters may further include reverberator parameters. For example, the reverberator parameters may include a reverberation type, a reverberation time, a delay time, and a direct-to-reverberant ratio.

[0195] In some embodiments of the present application, the audio optimization metadata further includes N-1 difference parameters of N-1 second decoded audio mixing parameters corresponding to N-1 optimized listening areas other than the first optimized listening area within the N optimized listening areas, relative to the second decoded audio mixing parameters corresponding to the first optimized listening area, where N is a positive integer.

[0196] The difference parameters are parameters of differences between the N-1 second decoded audio mixing parameters for the N-1 optimized listening areas other than the first optimized listening area among the N optimized listening areas and the second decoded audio mixing parameters corresponding to the first optimized listening area. The difference parameters are not the N-1 second decoded audio mixing parameters for the N-1 optimized listening areas. For example, the second decoded audio mixing parameters corresponding to the first optimized listening area include Parameter 1, Parameter 2, and Parameter 3. The second decoded audio mixing parameters corresponding to each of the N-1 second decoded audio mixing parameters corresponding to the N-1 optimized listening areas include Parameter 1, Parameter 2, and Parameter 4. The difference parameter between the N-1 second decoded audio mixing parameters corresponding to the N-1 optimized listening areas and the second decoded audio mixing parameter corresponding to the first optimized listening area includes Parameter 4. The audio-optimized metadata carries differential parameters, so that the data amount of the audio-optimized metadata can be reduced, and data transmission efficiency and decoding efficiency can be improved.

[0197] In some embodiments of the present application, the first decoded audio mixing parameters include at least one of an identifier of the rendered audio data, an equalization parameter, a compressor parameter, and a reverberator parameter.

[0198] The first decoded audio mixing parameters may include an identifier of the rendered audio data, e.g., an identifier of the M pieces of rendered audio data. The first decoded audio mixing parameters may further include equalization parameters. For example, the equalization parameters may include an equalization parameter identifier, a gain value for each frequency band, and a Q factor. The first decoded audio mixing parameters may further include compressor parameters. For example, the compressor parameters may include a compressor identifier, a threshold, a compression ratio, a start time, a release time, and a gain compensation value. The first decoded audio mixing parameters may further include reverberator parameters. For example, the reverberator parameters may include a reverberation type, a reverberation time, a delay time, and a direct-to-reverberant ratio.

[0199] In some embodiments of the present application, the audio optimization metadata further includes N-1 difference parameters of the N-1 first decoded audio mixing parameters corresponding to N-1 optimized listening areas other than the first optimized listening area within the N optimized listening areas, relative to the first decoded audio mixing parameters corresponding to the first optimized listening area, where N is a positive integer.

[0200] The difference parameters are parameters of differences between the N-1 first decoded mixing parameters of the N-1 optimized listening areas other than the first optimized listening area among the N optimized listening areas and the first decoded mixing parameters corresponding to the first optimized listening area. The difference parameters are not the N-1 first decoded mixing parameters of the N-1 optimized listening areas. For example, the first decoded audio mixing parameters corresponding to the first optimized listening area include Parameter 1, Parameter 2, and Parameter 3. The first decoded audio mixing parameters corresponding to each of the N-1 first decoded audio mixing parameters corresponding to the N-1 optimized listening areas include Parameter 1, Parameter 2, and Parameter 4. The difference parameter between the N-1 first decoded audio mixing parameters corresponding to the N-1 optimized listening areas and the first decoded audio mixing parameter corresponding to the first optimized listening area includes Parameter 4. The audio-optimized metadata carries differential parameters, so that the data amount of the audio-optimized metadata can be reduced, and data transmission efficiency and decoding efficiency can be improved.

[0201] In some embodiments of the present application, the audio processing method that may be performed by the decoding terminal further includes the following steps:

[0202] G1: The decoding terminal decodes the video image bitstream to obtain decoded video image data and video image metadata, where the video image metadata includes video metadata and image metadata.

[0203] G2: The decoding terminal renders the decoded video image data based on the video image metadata to obtain rendered video image data.

[0204] G3: The decoding terminal establishes a virtual scene based on the rendered video image data.

[0205] G 4 The decoding terminal identifies a first optimized listening area within the virtual scene based on the rendered video image data and the audio optimization metadata.

[0206] The transmitting terminal may generate a video image bitstream based on the video image metadata and the video image data, where the video image bitstream carries the video image metadata and the video image data. Thus, after receiving the video image bitstream from the transmitting terminal, the decoding terminal may obtain the video image metadata and the decoded video image data. The decoding terminal may render the decoded video image data based on the video image metadata to obtain rendered video image data, and the decoding terminal may establish a virtual scene by using the rendered video image data. Finally, the decoding terminal may identify a first optimized listening area within the virtual scene based on the rendered video image data and the audio optimization metadata, whereby the decoding terminal side displays the first optimized listening area within the virtual scene and guides the user to experience music in the optimized listening area, thereby improving the user's listening experience.

[0207] For example, the decoding terminal may identify a first optimized listening area within the virtual scene based on the rendered video image data and the audio optimization metadata. In a processing manner similar to that for the first optimized listening area, the decoding terminal may further identify N optimized listening areas within the virtual scene. The decoding terminal may generate an audio experience route based on the N optimized listening areas identified within the virtual scene to guide the user to better experience 6DoF music.

[0208] From the examples described in the foregoing embodiments, it can be seen that the decoding terminal may receive audio optimization metadata. In this embodiment of the present application, the audio optimization metadata includes metadata of a first optimized listening area and first audio mixing parameters corresponding to the first optimized listening area. The first optimized listening area is determined based on the user's current location. Therefore, to perform audio mixing, first decoding audio mixing parameters corresponding to the first optimized listening area may be obtained. Therefore, audio optimization metadata suitable for the user to move freely in the first optimized listening area may be provided, and the listening experience obtained when the user moves freely may be improved.

[0209] To better understand and implement the aforementioned solutions in the embodiments of the present application, a specific description is provided below by using corresponding application scenarios as examples.

[0210] In this embodiment of the present application, the optimized listening area can be represented as a sweet spot. In the audio processing method in this embodiment of the present application, the production terminal selects one or more sweet spots from the VR scene. The music effect of the rendered music signal heard by the user in each sweet spot is optimized, thereby improving the user's music experience in the 6DoF music scene.

[0211] Specifically, a 6DoF virtual music scene production method is performed on the production terminal side. It is assumed that processes such as VR video scene production, audio collection, and 6DoF basic audio metadata production have been completed in the 6DoF music scene. In this embodiment of the present application, several sweet spots are selected in the VR scene, and the sweet spots should match the listening area of ​​the user's interest as closely as possible. For each sweet spot, a rendering signal of each audio signal at the center position of the sweet spot is generated based on the audio signal and the 6DoF basic audio metadata. Then, audio mixing is performed on each rendering signal to adjust frequency, dynamics, sound quality, placement, sound field, etc. Audio mixing parameters for each audio mixing step corresponding to each audio signal are reserved. Finally, audio mixing is performed on the mixed audio signal, and audio mixing parameters for each audio mixing step corresponding to the mixed signal are reserved.

[0212] The 6DoF audio optimization metadata generated by the production terminal includes sweet spot metadata and audio mixing parameters corresponding to each sweet spot. The sweet spot metadata includes information such as the center position coordinates of each sweet spot and the shape of the sweet spot. Each sweet spot metadata corresponds to a group of audio mixing metadata including audio mixing parameters of an audio mixing parameter signal corresponding to a single audio signal. The production terminal transmits the 6DoF audio optimization metadata to the transmitting terminal, which generates an audio bitstream based on the 6DoF audio optimization metadata and transmits the audio bitstream to the decoding terminal.

[0213] Optionally, a sweet spot position change identifier and an audio mixing parameter change identifier are added to the 6DoF audio optimization metadata. When the VR music scene is initialized, the initial audio optimization metadata is sent. When the VR scene changes and the sweet spot position and shape information changes, the sweet spot metadata change identifier is true and the sweet spot change information is sent. When the audio mixing metadata of the current frame changes, the audio mixing parameter change identifier is true and the audio mixing metadata change information is sent.

[0214] The decoding terminal may include a video renderer and an audio renderer, and the decoding terminal may perform a 6DoF music scene audio rendering method. Specifically, the video renderer identifies a sweet spot based on the decoded sweet spot metadata and guides the user to properly experience 6DoF music. When the user's current position is within the sweet spot, the audio renderer renders an audio signal based on the user's position information, the 6DoF basic audio metadata, and the 6DoF audio optimization metadata to provide the user with a better music experience. The audio optimization metadata has a specific application range. The application range is determined based on the shape of the current sweet spot, and the shape is predetermined based on the music production scene. When the user's current position is outside the sweet spot, the audio renderer renders an audio signal based on the user's position information and the 6DoF basic audio metadata.

[0215] This embodiment of the present application is applicable to scene creation, audio metadata transmission, and user-side audio and video rendering in VR, AR, or MR applications. The terminal in this embodiment of the present application is particularly applicable to VR music software including 6DoF audio, 3D audio engines, etc. For example, the terminal may include a VR terminal device, a chip, a wireless network device, etc.

[0216] 3 is a schematic flowchart of streaming data processing in a virtual reality streaming service system according to an embodiment of the present application. This embodiment of the present application is applicable to a 6DoF audio rendering module (audio binaural rendering) in applications such as AR or VR, and specifically applies to the audio data preprocessing module, audio data encoding module, audio data decoding module, and audio rendering module of FIG. 3. The end-to-end audio signal processing process is as follows: After the audio signal passes through the VR scene collection or creation module, the audio signal undergoes an audio preprocessing operation. The preprocessing operation includes removing low-frequency components below 50 Hz in the signal and extracting 6DoF audio metadata (including 6DoF basic audio metadata, 6DoF audio optimization metadata, etc.). The preprocessed audio signal then undergoes audio encoding and encapsulation, and the processed signal is delivered to the decoder side. The decoder side first performs file / segment decapsulation, then audio decoding. The decoded audio signal is then rendered and mapped to the listener's speaker or headset device. The headset device can be a standalone headset or a headset on a glasses device.

[0217] 4 is an end-to-end flowchart of a 6DoFVR music scene according to an embodiment of the present application, which mainly includes a production terminal, a sending terminal, and a decoding terminal. The following provides separate explanations by using examples from the perspectives of different terminals.

[0218] The processes performed by the production terminal include VR video scene and metadata production, audio data and 6DoF basic audio metadata production, sweet spot selection, audio mixing and audio mixing parameter extraction, audio optimization metadata production, etc.

[0219] The production terminal includes a VR video and image data module, a VR video and image metadata module, a VR audio data module, a 6DoF basic audio metadata module, a video renderer rendering module, a sweet spot acquisition module, an audio renderer pre-rendering module, an audio mixing module, and an audio optimization metadata module.

[0220] The VR video and image data module is configured to obtain video and image data to be transmitted.

[0221] The VR video and image metadata module is configured to obtain video and image metadata produced in the VR scene.

[0222] The VR audio data module is configured to acquire audio data for transmission. Each audio data is object-based audio. -based The audio data may be multi-channel audio data, channel-based audio data, or scene-based audio data.

[0223] The 6DoF basic audio metadata module is configured to obtain 6DoF basic audio metadata produced in the VR scene. For example, metadata types that may be included in the 6DoF basic audio metadata may be one or more types of metadata in Table 1.

[0224] The video renderer rendering module is configured to perform rendering based on the VR video and image data and the VR video and image metadata to generate a first VR video scene.

[0225] The sweet spot acquisition module is configured to acquire sweet spot information in a VR scene. The sweet spot is selected based on the rendered VR video scene. There are 1 or N sweet spots. The sweet spot information includes a reference coordinate system, a center position coordinate, a shape, and other information. Optionally, the sweet spot information includes a center position coordinate and a shape. The shape of the sweet spot may be a sphere, a cube, a pillar, or any other shape.

[0226] The audio renderer pre-rendering module is configured to separately perform first rendering on the M pieces of audio data (referred to as first audio data) for each sweet spot based on the sweet spot center position coordinates, the VR audio data, and the 6DoF basic audio metadata to obtain M pieces of rendered audio data (referred to as second audio data). The audio signal obtained by the first rendering may be a single-channel signal or a binaural rendering signal. The N sweet spots have a total of N*M audio signals obtained by the first rendering.

[0227] The audio mixing module is configured to perform first audio mixing on each audio signal obtained by the first rendering and extract parameters for each audio mixing step of each audio signal in the audio mixing process, where the parameters are referred to as first audio mixing parameters. The audio data obtained by the audio mixing is referred to as third audio data. All audio signals in the third audio data are mixed to obtain a fourth audio signal. A second audio mixing is performed on the fourth audio signal, and parameters for each audio mixing step are reserved. The parameters are referred to as second audio mixing parameters. The audio mixing steps can be completed by the VR music scene production terminal or the audio mixing terminal.

[0228] The audio-optimized metadata module is configured to obtain the sweet-spot information, the first audio mixing parameters, and the second audio mixing parameters, and generate audio-optimized metadata according to a specific data structure.

[0229] The processes performed by the transmitting terminal include compressing, encoding and transmitting video scenes and metadata, compressing, encoding and transmitting audio data, compressing, encoding and transmitting 6DoF basic audio metadata, and compressing, encoding and transmitting audio optimization metadata.

[0230] The transmitting terminal includes a video and image metadata compression and transmission module, a video compression and transmission module, an image compression and transmission module, an audio optimized metadata compression and transmission module, an audio compression and transmission module, and a 6DoF basic audio metadata compression and transmission module.

[0231] The video and image metadata compression and transmission module is configured to perform compression encoding on the video and image metadata and transmit the generated bitstream.

[0232] The video compression and transmission module is configured to perform compression encoding on the video data in the VR scene and transmit the bitstream.

[0233] The image compression and transmission module is configured to perform compression encoding on the image data in the VR scene and transmit the bitstream.

[0234] The audio-optimized metadata compression and transmission module is configured to perform compression encoding on the audio-optimized metadata provided in this embodiment of the present application and transmit a bitstream.

[0235] The audio compression and transmission module is configured to perform compression encoding on the audio data in the VR scene and transmit the bitstream.

[0236] The 6DoF basic audio metadata compression and transmission module is configured to perform compression encoding on the 6DoF basic audio metadata and transmit a bitstream.

[0237] The processes performed by the decoding terminal (in other words, the user side) include obtaining the user's 6DoF position information, 6DoF video rendering, 6DoF audio rendering, etc. In this embodiment of the present application, the decoded audio optimization metadata is used for 6DoF video rendering and 6DoF audio rendering.

[0238] The decoding terminal includes an audio and video decoder, a video renderer, and an audio renderer.

[0239] The audio and video decoder is configured to decode the bitstream to obtain decoded VR video and image data, video and image metadata, audio data, 6DoF basic audio metadata, and audio optimization metadata.

[0240] The video renderer is configured to render the VR video scene based on the decoded video and image data, the decoded video and image metadata, and the user's position information.

[0241] Optionally, the video renderer identifies sweet spots based on sweet spot information in the decoded audio optimization metadata, and identifies recommended 6DoF music experience routes to guide the user to better experience the 6DoF music. The experience routes may be connecting lines between sweet spots, etc. This is not limited in this embodiment of the present application.

[0242] Similar to the production terminal's pre-rendering and audio mixing process, the audio renderer is configured to determine whether the user is within the sweet spot based on the user's location information and the sweet spot information in the audio optimization metadata.

[0243] If the user's current position is within the sweet spot, the audio renderer is configured to render each audio signal based on the 6DoF basic audio metadata and the user's position information to obtain a rendering signal. Audio mixing is performed on the rendering signal based on audio mixing parameters corresponding to each decoded audio signal. After audio mixing is performed on all audio signals, all audio obtained by audio mixing is mixed. Final audio mixing is performed based on the audio mixing parameters of the mixed signal, and the processed audio signal is sent to an audio device such as a user's headset.

[0244] If the user is not within the sweet spot, the audio renderer is configured to render each audio signal based on the 6DoF basic audio metadata and the user's position information, and directly mix all the rendered audio to generate a final binaural signal for playback.

[0245] The following describes in detail the audio processing method in the embodiment of the present application by using two specific embodiments.

[0246] Embodiment 1 5 is a schematic diagram of a VR concert scene supporting 6DoF according to an embodiment of the present application. A typical VR concert scene supporting 6DoF is used as an example to describe in detail the technical solution in this embodiment of the present application. The concert scene includes two parts, namely, a stage area and an audience area, and there are four types of target sound sources, namely, human voice, violin sound, cello sound, and drum sound. All sound sources are assumed to be static sound sources, and the positions of the sound sources in the VR scene are shown in FIG. 5.

[0247] In this embodiment, an end-to-end process of the VR concert scene is carried out from the production terminal, the sending terminal to the decoding terminal. The specific process of embodiment 1 mainly includes the following steps:

[0248] Steps S01 to S05 are processes of the production terminal in the VR music scene, step S06 is a process of the transmission terminal in the VR music scene, and steps S07 and S08 are processes of the decoding terminal in the VR music scene.

[0249] S01: The production terminal acquires VR video and image data, VR video and image metadata, VR scene audio data, and 6DoF basic audio metadata.

[0250] VR video and image data, VR video and image metadata, VR scene audio data, and 6DoF basic audio metadata are pre-produced in the VR music scene.

[0251] S02: The production terminal acquires sweet spot metadata.

[0252] The production terminal renders a VR scene based on the VR video and image data and the VR video and image metadata, and then the production terminal selects sweet spots in the VR scene and records the center position coordinates and shape information of each sweet spot. The number of sweet spots may be N, and the sweet spots selected by the production terminal should match the listening area of ​​the user's interest. Finally, the production terminal generates sweet spot metadata based on the sweet spot information according to a specific data structure.

[0253] The sweet spot metadata includes the reference coordinate system, center position coordinates, and shape information of the sweet spot. An example of the data structure of the sweet spot metadata is as follows: <sweet spot identifier> <Sweet Spot 1 Identifier> <Reference coordinate system> <Center position coordinates> <Shape information> <Sweet Spot 2 Identifier> <Reference coordinate system> <Center position coordinates> <Shape information> … <Sweet Spot N Identifier> <Reference coordinate system> <Center position coordinates> <Shape information>

[0254] The shape of each sweet spot may be a sphere, a cylinder, an arbitrary shape formed by a triangular mesh, etc. The sweet spot metadata includes the reference coordinate system and the center position coordinates of the sweet spot, and the shape information is the shape information set by default by the production terminal and the decoding terminal. Another example of the data structure of the sweet spot metadata is as follows: <sweet spot identifier> <Sweet Spot 1 Identifier> <Reference coordinate system> <Center position coordinates> <Sweet Spot 2 Identifier> <Reference coordinate system> <Center position coordinates> … <Sweet Spot N Identifier> <Reference coordinate system> <Center position coordinates>

[0255] The data and data structure included in the sweet spot metadata are not limited to the two types described above. For example, the center position information of sweet spot 2 to sweet spot N may be position information relative to sweet spot 1.

[0256] S03: For each sweet spot, the production terminal performs first rendering on the M pieces of audio data (referred to as first audio data) one by one based on the sweet spot metadata, the VR audio data, and the 6DoF basic audio metadata to obtain M pieces of rendered audio data (referred to as second audio data). The audio signal obtained by the first rendering may be a single-channel signal or a binaural rendering signal. The N sweet spots have a total of N*M audio signals obtained by the first rendering. These audio signals are referred to as second audio data. Each first audio signal may be an object-based signal, a multi-channel-based audio signal, or a scene-based audio signal.

[0257] S04: The production terminal performs audio mixing on each rendered audio signal to obtain first audio mixing parameters in each sweet spot, and performs audio mixing on the final mixed signal to obtain second audio mixing parameters.

[0258] The production terminal performs a first audio mixing process on each second audio signal, extracts parameters for each audio mixing step of each audio signal in the audio mixing process, and the parameters are referred to as first audio mixing parameters. The audio data obtained by the audio mixing are referred to as third audio data.

[0259] Optionally, all audio signals in the third audio data are mixed to obtain a fourth audio signal. A second audio mixing is performed on the fourth audio signal, and parameters for each audio mixing step are reserved. The parameters are referred to as second audio mixing parameters. The two audio mixing steps can be completed by a production terminal in a VR music scene.

[0260] Each of the first and second audio mixing parameters includes an audio signal identification number, an equalization parameter, a compressor parameter, and a reverberator parameter. The equalization parameters include a frequency band, a gain value, and a Q value. The Q value is a parameter of an equalization filter and represents the quality factor of the equalization filter, and can be used to describe the bandwidth of the equalization filter. The compressor parameters include a threshold, a compression ratio, a start time, a release time, and a gain compensation value. The reverberator parameters include a reverberation time, a delay time, and a direct-to-reverberant sound ratio.

[0261] Optionally, the audio mixing parameters of important audio mixing steps may be reserved based on a specific application scenario. The types of audio mixing parameters included in the first audio mixing parameter and the second audio mixing parameter may be different.

[0262] S05: The production terminal generates 6DoF audio optimization metadata according to a specific data structure, based on the sweet spot metadata and the audio mixing parameters corresponding to each sweet spot. The sweet spot metadata and the audio mixing parameters in step S04 may be stored and transmitted in the form of independent data structures. The data structure of the sweet spot metadata is shown in step S02. An example of the data structure of the audio mixing parameters is as follows. <Audio mixing metadata identifier> <Sweet spot 1 identifier> <Audio signal 1 identification id> <Equalization parameter identifier> <Frequency band 1> <Gain value> <Q value> … <Frequency band P> <Gain value> <Q value> <Compressor parameter identifier> <Threshold> <Compression ratio> <Start time> <Release time> <Gain compensation value> <Reverb parameter> <Reverb type> <Reverb time> <Delay time> <Direct sound to reverberant sound ratio> <…> … … <Audio signal M identification id> <Equalization parameter identifier> <Frequency band 1> <Gain value> <Q value> … <Frequency band P> <Gain value> <Q value> <Compressor parameter identifier> <Threshold value> <Compression ratio> <Start time> <Release time> <Gain compensation value> <Reverb parameter> <Reverb type> <Reverb time> <Delay time> <Direct-to-reverb ratio> <…> … <Second audio mixing parameter identifier> <Equalization parameter identifier> <Frequency band 1> <Gain value> <Q value> … <Frequency band P> <Gain value> <Q value> <Compressor parameter identifier> <Threshold value> <Compression ratio> <Start time> <Release time> <Gain compensation value> <Reverb parameter> <Reverb type> <Reverb time> <Delay time> <Direct-to-reverb ratio> … <Sweet spot N identifier> …

[0263] Note that the sweet spot N identifier of sweet spot 1 and the audio mixing parameters have the same data structure.

[0264] In the above data structure, the types of audio mixing parameters in sweet spot 1 to sweet spot N are completely the same.

[0265] Optionally, the parameter type stored in sweet spot 1 is the same as the aforementioned data structure, and the audio mixing parameters of sweet spot 2 to sweet spot N are differential parameters relative to the audio mixing parameters of sweet spot 1, thereby reducing the quantity of parameters in the 6DoF audio optimization metadata.

[0266] Optionally, the data structure of the sweet spot metadata and the data structure of the audio mixing parameters are integrated into the same data structure, thereby reducing the number of parameters of the 6DoF audio optimization metadata.

[0267] S06: In addition to encoding and transmitting the VR video and image data, the VR video and image metadata, the audio data, and the 6DoF basic audio metadata, the transmitting terminal further needs to encode and transmit the 6DoF audio optimization metadata.

[0268] Optionally, to improve the transmission efficiency of the 6DoF audio optimization metadata, a sweet spot metadata change identifier and an audio mixing parameter change identifier are added to the encoded 6DoF audio optimization metadata. When the VR music scene is initialized, the initial audio optimization metadata is transmitted. When the VR scene changes and the sweet spot position and shape information changes, the sweet spot metadata change identifier is true and the sweet spot change information is transmitted. When the audio mixing metadata of the current frame changes, the audio mixing parameter change identifier is true and the audio mixing metadata change information is transmitted.

[0269] S07: In the decoding terminal, the user's VR head-mounted device or the like obtains the user's 6DoF position information, and the video renderer renders the video based on the decoded VR video and image data, the VR video and image metadata, and the user's position information. In addition, a sweet spot is identified based on the decoded sweet spot metadata. Optionally, a recommended 6DoF music experience route can be further identified to guide the user to better experience 6DoF music.

[0270] S08: In the decoding terminal, the user's VR head-mounted device acquires the user's 6DoF position information, and the audio decoder decodes the audio bitstream to obtain first decoded audio data, decoded 6DoF basic audio metadata, and decoded 6DoF audio optimization metadata. The audio renderer determines whether the user is located within the sweet spot based on the user's position information and the decoded sweet spot metadata.

[0271] If the user's current location is within the sweet spot, each first decoder-side audio signal is rendered based on the 6DoF basic audio metadata and the user's location information to obtain M rendered audio signals (denoted as second decoder-side audio signals). Audio mixing is performed on each second decoder-side audio signal based on the decoded first audio mixing parameters to obtain M third decoder-side audio signals. The M third decoder-side audio signals are mixed based on the 6DoF basic audio metadata, and if the decoded second audio mixing parameters exist, second audio mixing is performed on the mixed signal to obtain a final music signal. When the user's current location is within the sweet spot, an optimal immersive music experience can be provided to the user.

[0272] If the user's current position is outside the sweet spot, the audio renderer renders each first decoder-side audio signal separately based on the 6DoF basic audio metadata and the user's position information to obtain M rendered audio signals (denoted as second decoder-side audio signals). The M second decoder-side audio signals are mixed to obtain a final music signal.

[0273] Optionally, a transition distance is set at each sweet spot, and a smoothing algorithm is used to ensure that the audible music signal transitions naturally as the user moves freely in and out of the sweet spot. The smoothing algorithm is not limited in this embodiment of the present application.

[0274] For example, an area that is a specific distance (i.e., a transition distance) away from the edge of the sweet spot may be set as a transition area, in which each parameter of the 6DoF audio optimization metadata gradually changes to 0, so that the music effect heard by the user can transition naturally.

[0275] Embodiment 2 FIG. 6 illustrates another 6DoFVR according to an embodiment of the present application. music 1 is an end-to-end flowchart of a scene.

[0276] The main difference between embodiment 2 and embodiment 1 lies in the different production process in the 6DoF music scene. In embodiment 1, during the audio mixing in step S04, audio mixing is performed to extract audio mixing metadata while the produced VR video metadata, VR video data, audio data, and 6DoF basic audio metadata remain unchanged.

[0277] In embodiment 2, in the audio mixing process of step S04, the created VR video metadata, VR video data, audio data, and 6DoF basic audio metadata may be adjusted and optimized, and audio mixing metadata may be extracted simultaneously. For example, in the audio mixing process of step S04, the VR scene creation terminal may adjust the audio frequency response and gain, adjust the position of the target sound source, the acoustic parameters of the room, etc.

[0278] Compared with Embodiment 1, the audio mixing metadata of Embodiment 2 is less than that of Embodiment 1, and the 3D immersive music effect after audio mixing may also be better than that of Embodiment 1. In the production process shown in Embodiment 1, the basic 6DoF audio metadata is not modified, and only new audio optimization metadata is generated. However, in Embodiment 2, the 6DoF basic audio metadata is adjusted. For example, if the sound of a musical instrument directly in front of the user does not sound harmonious, the position of the instrument may be adjusted in the VR video scene, and the sound source position information corresponding to the instrument in the 6DoF basic audio metadata is modified.

[0279] Optionally, the reverberation effect of the music signal heard by the user is adjusted by adjusting the room acoustic parameters in the 6DoF basic audio metadata, and the first audio mixing parameters and the second audio mixing parameters of embodiment 1 may not include reverberator parameters.

[0280] Optionally, to adjust the effect of the music signal finally heard by the user, parameters such as the frequency response of one or more audio signals in the audio data of Fig. 6 or the position and gain of the audio signals in the 6DoF basic audio metadata of Fig. 6 are adjusted. The first audio mixing parameters in embodiment 1 may not include equalization parameters corresponding to these signals.

[0281] By using the exemplary description of the foregoing embodiment, it can be seen that an embodiment of the present application provides a method for creating, transmitting, and rendering a 6DoF virtual music scene. The decoding terminal can guide the user to better experience the 6DoF music scene and effectively convey the user's personal aesthetic sense of the music. The user can hear more complete 3D immersive music in each sweet spot, and the user can have different music experiences in different sweet spots. In addition, this embodiment of the present application proposes that a sweet spot position change identifier and an audio mixing parameter change identifier may be added to the 6DoF audio optimization metadata, thereby effectively improving the transmission efficiency of the 6DoF audio optimization metadata.

[0282] It should be noted that for simplicity of explanation, the above method embodiments are expressed as a series of operations. However, those skilled in the art should understand that the present application is not limited to the described order of operations, as some steps may be performed in other orders or simultaneously. Those skilled in the art should understand that all embodiments described herein belong to exemplary embodiments, and the operations and modules involved are not necessarily required by the present application.

[0283] To better implement the solutions of the embodiments of the present application, related apparatuses for implementing the solutions are further provided below.

[0284] See Figure 7. A decoding terminal 700 provided in an embodiment of the present application may include: a decoding module 701, a rendering module 702, an audio mixing module 703, and a mixing module 704.

[0285] The decoding module is configured to decode the audio bitstream to obtain audio-optimization metadata, basic audio metadata, and M pieces of decoded audio data, where the audio-optimization metadata includes first metadata of a first optimized listening area and first decoded audio mixing parameters corresponding to the first optimized listening area, and M is a positive integer.

[0286] The rendering module performs M operations based on the user's current location and the basic audio metadata to obtain M rendered audio data. Return of The audio processing unit is configured to render the encoded audio data.

[0287] The audio mixing module is configured to perform first audio mixing on the M pieces of rendered audio data based on the first decoded audio mixing parameters to obtain M pieces of first audio mixing data when the current location is within the first optimized listening area.

[0288] The mixing module is configured to mix the M first audio mixing data to obtain mixed audio data corresponding to a first optimized listening area.

[0289] In the above-described embodiment of the present application, the audio optimization metadata of this embodiment of the present application includes first metadata of a first optimized listening area and first audio mixing parameters corresponding to the first optimized listening area, where the first optimized listening area is determined based on the user's current location. Therefore, to perform audio mixing, first decoded audio mixing parameters corresponding to the first optimized listening area can be obtained. Therefore, audio optimization metadata suitable for the user to move freely in the first optimized listening area can be provided, and the listening experience obtained when the user moves freely can be improved.

[0290] See Figure 8. A sending terminal 800 provided in an embodiment of the present application may include: a receiving module 801, an encoding module 802, and a sending module 803.

[0291] The receiving module is configured to receive audio-optimization metadata, basic audio metadata, and M pieces of first audio data, where the audio-optimization metadata includes first metadata of a first optimized listening area and first audio mixing parameters corresponding to the first optimized listening area, and M is a positive integer.

[0292] The encoding module is configured to perform compression encoding on the audio-optimized metadata, the basic audio metadata, and the M pieces of first audio data to obtain an audio bitstream.

[0293] The transmission module is configured to transmit the audio bitstream.

[0294] In the above-described embodiment of the present application, audio optimization metadata is first received from a production terminal, an audio bitstream is generated based on the audio optimization metadata, and the audio bitstream is transmitted to a decoding terminal. The decoding terminal may obtain the audio optimization metadata by using the audio bitstream. The audio optimization metadata includes first metadata of a first optimized listening area and first audio mixing parameters corresponding to the first optimized listening area. Therefore, audio optimization metadata suitable for a user to move freely in the first optimized listening area can be provided, and the listening experience obtained when the user moves freely can be improved.

[0295] See Figure 9. A production terminal 900 provided in an embodiment of the present application may include: an acquisition module 901, a rendering module 902, an audio mixing module 903, and a generation module 904.

[0296] The acquisition module is configured to acquire basic audio metadata and metadata of N optimized listening areas, where N is a positive integer, and the N optimized listening areas include the first optimized listening area.

[0297] The rendering module is configured to render the M pieces of audio data to be processed based on the first optimized listening area and the basic audio metadata to obtain M pieces of rendered audio data corresponding to the first optimized listening area, where M is a positive integer.

[0298] The audio mixing module is configured to perform first audio mixing on the M pieces of rendered audio data to obtain first audio mixing parameters corresponding to the M pieces of first audio mixing data and the first optimized listening area.

[0299] The generation module is configured to generate audio-optimized metadata based on the first metadata and the first audio mixing parameters for the first optimized listening area, where the audio-optimized metadata includes the first metadata and the first audio mixing parameters.

[0300] In the above-described embodiment of the present application, metadata for N optimized listening areas may be obtained, where the N optimized listening areas include a first optimized listening area. Therefore, M first audio data may be rendered and mixed for the first optimized listening area. Finally, audio optimization metadata may be generated, where the audio optimization metadata includes first metadata for the first optimized listening area and first audio mixing parameters corresponding to the first optimized listening area. Therefore, audio optimization metadata suitable for a user to freely move in the first optimized listening area may be provided, and the listening experience obtained when the user freely moves may be improved.

[0301] It should be noted that the contents of information exchange between the modules / units of the device and their execution processes are based on the same idea as the method embodiments of the present application, and produce the same technical effects as the method embodiments of the present application. For specific contents, please refer to the above description of the method embodiments of the present application. The details will not be described again here.

[0302] An embodiment of the present application further provides a computer storage medium, which stores a program, and the program performs some or all of the steps described in the foregoing method embodiments.

[0303] The following describes another decoding terminal according to an embodiment of the present application. Please refer to Figure 10. The decoding terminal 1000 includes: The decoding terminal 1000 includes a receiver 1001, a transmitter 1002, a processor 1003, and a memory 1004 (there may be one or more processors 1003 in the decoding terminal 1000, and one processor is used as an example in FIG. 10). In some embodiments of the present application, the receiver 1001, the transmitter 1002, the processor 1003, and the memory 1004 may be connected via a bus or in another manner. In FIG. 10, the receiver 1001, the transmitter 1002, the processor 1003, and the memory 1004 are used as an example connected via a bus.

[0304] The memory 1004 may include read-only memory and random-access memory and may provide instructions and data to the processor 1003. A portion of the memory 1004 may further include non-volatile random access memory (NVRAM). The memory 1004 stores an operating system and operating instructions, executable modules or data structures, or a subset or extended set thereof. The operating instructions may include various operating instructions for performing various operations. The operating system may include various system programs for performing various basic services and handling hardware-based tasks.

[0305] The processor 1003 controls the operation of the decoding terminal and may also be referred to as a central processing unit (CPU). In particular applications, the components of the decoding terminal are coupled to each other by using a bus system. In addition to a data bus, the bus system may further include a power bus, a control bus, and a status signal bus. However, for clarity of explanation, the various types of buses in the figures are shown as a bus system.

[0306] The methods disclosed in the embodiments of the present application may be applied to or implemented by the processor 1003. The processor 1003 may be an integrated circuit chip and have signal processing capabilities. In the implementation process, the method steps may be implemented by using hardware integrated logic circuitry within the processor 1003 or by using instructions in the form of software. The processor 1003 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or another programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The processor 1003 may implement or perform the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc. The steps of the methods disclosed in the embodiments of the present application may be directly executed and completed by using a hardware decoding processor, or may be executed and completed by using a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium well-established in the art, such as a random access memory, a flash memory, a read-only memory, an electrically erasable programmable read-only memory, or a register. The storage medium is located in the memory 1004, and the processor 1003 reads information in the memory 1004 and completes the steps of the aforementioned method in combination with the hardware of the processor 1003.

[0307] The receiver 1001 may be configured to receive input digital or textual information and generate signal inputs related to relevant settings and function control of the decoding terminal. The transmitter 1002 may include a display device such as a display, and the transmitter 1002 may be configured to output numeric or textual information by using an external interface.

[0308] In this embodiment of the present application, the processor 1003 is configured to perform the method shown in Figures 2A and 2B of the previous embodiment, performed by a decoding terminal.

[0309] The following describes another transmitting terminal provided in one embodiment of the present application. Please refer to Figure 11. The transmitting terminal 1100 includes: The transmitting terminal 1100 includes a receiver 1101, a transmitter 1102, a processor 1103, and a memory 1104 (there may be one or more processors 1103, and one processor is used as an example in FIG. 11). In some embodiments of the present application, the receiver 1101, the transmitter 1102, the processor 1103, and the memory 1104 may be connected via a bus or in another manner. In FIG. 11, the receiver 1101, the transmitter 1102, the processor 1103, and the memory 1104 are used as an example connected via a bus.

[0310] Memory 1104 may include read-only memory and random-access memory and provide instructions and data to processor 1103. A portion of memory 1104 may further include non-volatile random access memory (NVRAM). Memory 1104 stores an operating system and operating instructions, as well as executable modules or data structures, or a subset or extended set thereof, which may include various operating instructions and are used to perform various operations. The operating system may include various system programs for performing various basic services and handling hardware-based tasks.

[0311] The processor 1103 controls the operation of the transmitting terminal and may also be called a central processing unit (CPU). In particular applications, the components of the transmitting terminal are coupled to each other by using a bus system. In addition to a data bus, the bus system may further include a power bus, a control bus, and a status signal bus. However, for clarity of explanation, the various types of buses in the figures are shown as a bus system.

[0312] The methods disclosed in the above embodiments of the present application may be applied to or implemented by the processor 1103. The processor 1103 may be an integrated circuit chip and have signal processing capabilities. In the implementation process, the steps of the above methods may be completed by using instructions in the form of hardware integrated logic circuits or software within the processor 1103. The processor 1103 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or another programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The processor 1103 may implement or perform the methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc. The steps of the methods disclosed in the embodiments of the present application may be directly executed and completed by using a hardware decoding processor, or may be executed and completed by using a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium well-established in the art, such as a random access memory, a flash memory, a read-only memory, an electrically erasable programmable read-only memory, or a register. The storage medium is located in the memory 1104, and the processor 1103 reads information in the memory 1104 and completes the steps of the aforementioned method in combination with the hardware of the processor 1103.

[0313] The receiver 1101 may be configured to receive input digital or textual information and generate signal inputs related to relevant settings and function control of the transmitting terminal. The transmitter 1102 may include a display device such as a display, and the transmitter 1102 may be configured to output numeric or textual information by using an external interface.

[0314] In this embodiment of the present application, the processor 1103 is configured to perform the method shown in Figures 2A and 2B of the previous embodiment, performed by the transmitting terminal.

[0315] The following describes another production terminal provided in one embodiment of the present application. Please refer to Figure 12. The production terminal 1200: The production terminal 1200 includes a receiver 1201, a transmitter 1202, a processor 1203, and a memory 1204 (the production terminal 1200 may have one or more processors 1203, and one processor is used as an example in FIG. 12). In some embodiments of the present application, the receiver 1201, the transmitter 1202, the processor 1203, and the memory 1204 may be connected via a bus or in another manner. In FIG. 12, the receiver 1201, the transmitter 1202, the processor 1203, and the memory 1204 are used as an example connected via a bus.

[0316] Memory 1204 may include read-only memory and random-access memory and provide instructions and data to processor 1203. A portion of memory 1204 may further include NVRAM. Memory 1204 stores an operating system and operating instructions, executable modules or data structures, or a subset or extended set thereof. The operating instructions may include various operating instructions used to perform various operations. The operating system may include various system programs for performing various basic services and handling hardware-based tasks.

[0317] The processor 1203 controls the operation of the production terminal and may also be referred to as a CPU. In particular applications, the components of the production terminal are coupled to each other using a bus system. In addition to a data bus, the bus system may further include a power bus, a control bus, and a status signal bus. However, for clarity of explanation, the various types of buses in the figures are shown as a bus system.

[0318] The methods disclosed in the embodiments of the present application may be applied to or implemented by the processor 1203. The processor 1203 may be an integrated circuit chip and have signal processing capabilities. During implementation, the steps of the aforementioned methods may be completed by using instructions in the form of hardware integrated logic circuits or software in the processor 1203. The processor 1203 may be a general-purpose processor, a DSP, an ASIC, an FPGA or another programmable logic device, a discrete gate or transistor logic device, or a discrete hardware assembly. The processor 1203 may implement or perform the methods, steps, and logical block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc. The steps of the methods disclosed in the embodiments of the present application may be directly executed and completed by using a hardware decoding processor, or may be executed and completed by using a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium that is well-known in the art, such as a random access memory, a flash memory, a read-only memory, an electrically erasable programmable read-only memory, or a register. The storage medium is located in the memory 1204, and the processor 1203 reads information in the memory 1204 and completes the steps of the aforementioned method in combination with the hardware of the processor 1203.

[0319] In this embodiment of the present application, the processor 1203 is configured to perform the audio processing method shown in Figures 2A and 2B of the previous embodiment, performed by a production terminal.

[0320] In another possible design, when the decoding terminal, the transmitting terminal, or the producing terminal is a chip within the terminal, the chip includes a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin, a circuit, etc. The processing unit may execute computer-executable instructions stored in the storage unit, so that the chip within the terminal performs the audio processing method of any one of the first to third aspects. Optionally, the storage unit is a storage unit within the chip, for example, a register or a cache. Alternatively, the storage unit may be a storage unit located outside the chip within the terminal, for example, a read-only memory (ROM) or another type of static storage device capable of storing static information and instructions, or a random access memory (RAM).

[0321] The processor may be a general purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits configured to control program execution of the methods of the first to third aspects.

[0322] It should also be noted that the described device embodiments are merely examples. Units described as separate parts may or may not be physically separate, and parts displayed as units may or may not be physical units, and may be located in one location or distributed over multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the solutions of the embodiments. In addition, in the accompanying drawings of the device embodiments provided in this application, the connection relationships between modules indicate that the modules have communication connections with each other, which may be specifically implemented as one or more communication buses or signal cables.

[0323] Based on the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by software in addition to the necessary general-purpose hardware, or by dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, any function that can be performed by a computer program can be easily implemented by using corresponding hardware. Furthermore, the specific hardware structure used to achieve the same function may be in various forms, such as an analog circuit, a digital circuit, or a dedicated circuit. However, for the present application, a software program implementation is a better implementation in most cases. Based on such understanding, the technical solution of the present application, or a part that contributes to the prior art, may be embodied in the form of a software product. A computer software product is stored in a readable storage medium such as a computer floppy disk, USB flash drive, removable hard disk, ROM, RAM, magnetic disk, or optical disk, and includes several instructions that enable a computer device (which may be a personal computer, a server, a network device, etc.) to perform the methods described in the embodiments of the present application.

[0324] All or part of the above-described embodiments may be implemented using software, hardware, firmware, or any combination thereof. If software is used to implement the embodiments, all or part of the embodiments may be implemented in the form of a computer program product.

[0325] A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the procedures or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or another programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, or digital subscriber line (DSL)) or wireless (e.g., infrared, radio, or microwave) transmission. The computer-readable storage medium may be any available medium accessible by a computer, or a data storage device, such as a server or data center, that integrates one or more available media. The usable medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)). [Explanation of symbols]

[0326] 100 Audio Processing System 101 Production Terminal 102 transmitting terminal 103 Decoding Terminal 700 Decoding Terminal 701 Decryption Module 702 Rendering Module 703 Audio Mixing Module 704 Mixing Module 800 Sending Terminal 801 Receiver Module 802 Encoding Module 803 Transmit Module 900 Production Terminal 901 Acquisition Module 902 Rendering Module 903 Audio Mixing Module 904 Generation Module 1000 Decoding Terminal 1001 Receiver 1002 Transmitter 1003 processor 1004 memory 1100 transmitting terminal 1101 Receiver 1102 Transmitter 1103 processor 1104 Memory 1200 Production Terminal 1201 Receiver 1202 Transmitter 1203 processor 1204 memory

Claims

1. 1. A method of audio processing, comprising: decoding the audio bitstream to obtain audio-optimization metadata, basic audio metadata, and M decoded audio data, wherein the audio-optimization metadata includes first metadata of a first optimized listening area and first decoded audio mixing parameters corresponding to the first optimized listening area, where M is a positive integer; rendering the M decoded audio data based on a user's current location and the basic audio metadata to obtain M rendered audio data; performing first audio mixing on the M rendered audio data based on the first decoded audio mixing parameters to obtain M first audio mixing data when the current location is within the first optimized listening area; mixing the M first audio mixing data to obtain mixed audio data corresponding to the first optimized listening area; 1. An audio processing method comprising:

2. the audio-optimization metadata further includes second decoded audio mixing parameters corresponding to the first optimized listening area; The method comprises: performing second audio mixing on the mixed audio data based on the second decoded audio mixing parameters to obtain second audio mixing data corresponding to the first optimized listening area; 10. The method of claim 1, further comprising:

3. The method of claim 2 , wherein the second decoded audio mixing parameters include at least one of an identifier of the second audio mixing data, an equalization parameter, a compressor parameter, and a reverberator parameter.

4. 3. The method of claim 2, wherein the audio-optimization metadata further includes N-1 difference parameters of N-1 second decoded audio mixing parameters corresponding to N-1 optimized listening areas other than the first optimized listening area within the N optimized listening areas, relative to the second decoded audio mixing parameters corresponding to the first optimized listening area, where N is a positive integer.

5. The method of claim 4 , wherein the first decoded audio mixing parameters include at least one of an identifier of the rendered audio data, an equalization parameter, a compressor parameter, and a reverberator parameter.

6. The method comprises: decoding the video image bitstream to obtain decoded video image data and video image metadata, wherein the video image metadata includes video metadata and image metadata; rendering the decoded video image data based on the video image metadata to obtain rendered video image data; establishing a virtual scene based on the rendered video image data; identifying the first optimized listening area within the virtual scene based on the rendered video image data and the audio optimization metadata; 5. The method of claim 4, further comprising:

7. 5. The method of claim 4, wherein the first metadata includes at least one of a reference coordinate system of the first optimized listening area, a center position coordinate of the first optimized listening area, and a shape of the first optimized listening area.

8. A decoding terminal, a decoding module configured to decode the audio bitstream to obtain audio-optimization metadata, basic audio metadata, and M decoded audio data, wherein the audio-optimization metadata includes first metadata of a first optimized listening area and first decoded audio mixing parameters corresponding to the first optimized listening area, where M is a positive integer; and a rendering module configured to render the M decoded audio data based on a user's current location and the basic audio metadata to obtain M rendered audio data; an audio mixing module configured to perform first audio mixing on the M rendered audio data based on the first decoded audio mixing parameters to obtain M first audio mixing data when the current location is within the first optimized listening area; a mixing module configured to mix the M first audio mixing data to obtain mixed audio data corresponding to the first optimized listening area; and A decoding terminal comprising:

9. 10. A computer-readable storage medium containing instructions that, when executed on a computer, enable the computer to perform the method of claim 1.

10. A computer program comprising instructions, which when run on a computer, enable the computer to perform the method of claim 1.

Citation Information

Patent Citations

  • Concept for generating extended or modified sound field descriptions using multi-point sound field descriptions

    JP2020527746A

  • Concepts for generating extended or modified sound field descriptions using depth-extended DirAC techniques or other techniques

    JP2020527887A

  • Audio playback method and audio playback apparatus in six degrees of freedom environment

    US20200162833A1

  • Three-dimensional audio playing method and playing apparatus

    US20200374646A1

  • Audio rendering with spatial metadata interpolation

    WO2021170900A1