An audio stream processing method, apparatus, and electronic device

By decoding and preprocessing the audio stream, generating and adjusting the audio object metadata information, separate control of voice-over and ambient sounds in TV voice-over are achieved, solving the problem that cannot be controlled separately in the prior art and improving the user experience.

CN119211635BActive Publication Date: 2025-05-27中央广播电视总台
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411306154.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2025-05-27
Estimated Expiration
2044-09-19

AI Technical Summary

Technical Problem

In the prior art, ambient sound and voiceover cannot be controlled separately, resulting in poor user experience.

Method used

By decoding the input audio stream, multiple audio tracks and their metadata information are generated, and when the voiceover identifier is included in the program list, the voiceover audio stream is preprocessed, the second audio object metadata information is generated, the voiceover audio track is muted, and the audio rendering module is called for audio rendering.

Benefits of technology

Separate control of voice-over and ambient sounds in traditional TV accompaniment, such as mute, volume adjustment and sound position change, improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119211635B_ABST
    Figure CN119211635B_ABST
Patent Text Reader

Abstract

The present invention provides an audio stream processing method, apparatus, and electronic device. The method includes: decoding an input audio stream to generate N audio tracks, as well as metadata information and audio signals corresponding to each of the audio tracks; when the received program schedule information includes a voiceover identifier, preprocessing the audio stream to obtain a voiceover audio stream, and generating second audio object metadata information for the voiceover audio stream; setting the target audio track corresponding to the voiceover among the N audio tracks as a muted audio track; calling an audio rendering module to perform audio rendering on the adjusted audio stream and the voiceover audio stream based on the metadata information and the second audio object metadata information to obtain a target audio rendering signal and output it. The audio stream processing solution provided by this application can achieve separate control of the voiceover and ambient sounds in traditional TV audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and in particular, to an audio stream processing method, an apparatus, and an electronic device. Background Art

[0002] TV accompanying sound is an important component of TV programs. Currently, TV programs generally use stereo or surround sound to transmit TV accompanying sound. For programs such as sports events and documentaries, the TV accompanying sound includes two parts, one is the ambient sound (on-site sound recording in coordination with the picture), and the other part is the voice-over such as commentary.

[0003] In order to implement the transmission of TV accompanying sound, the following scheme is mainly adopted at present: The ambient sound and the voice-over are combined into a sound track for transmission, and are decoded at the playback terminal and sent to different speakers for broadcast. In this processing method, the ambient sound and the voice-over are played as a whole, and the separate control of the ambient sound and the voice-over cannot be achieved. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide an audio stream processing method, an apparatus, and an electronic device, which can solve the problem that the ambient sound and the voice-over cannot be separately controlled in the prior art.

[0005] To solve the above technical problems, the present invention provides the following technical solutions:

[0006] The embodiments of the present invention provide an audio stream processing method, and the method includes:

[0007] Decode the input audio stream to generate N sound tracks, as well as the metadata information and audio signals corresponding to each of the sound tracks;

[0008] Determine whether the received program list information contains a voice-over identifier;

[0009] In the case where the program list information contains a voice-over identifier, preprocess the audio stream to obtain a voice-over audio stream, and generate second audio object metadata information for the voice-over audio stream;

[0010] Set the target sound track corresponding to the voice-over in the N sound tracks as a muted sound track;

[0011] Call an audio rendering module, and based on the metadata information and the second audio object metadata information, perform audio rendering on the adjusted audio stream and the voice-over audio stream to obtain a target audio rendering signal and output it.

[0012] Optionally, after the step of determining whether the received program list information contains a voice-over identifier, the method further includes:

[0013] In the case where the voiceover identifier is not included in the program list information, the audio rendering module is called to render the audio stream, and a target audio rendering signal is obtained and output.

[0014] Optionally, the N audio tracks include: a track corresponding to the left main channel, a track corresponding to the right main channel, a center track corresponding to the center channel, a track corresponding to the left surround channel, a track corresponding to the right surround channel, and a track corresponding to the subwoofer channel; wherein, the target track is the center track.

[0015] Optionally, in the case where the voiceover identifier is included in the program list information, the steps of preprocessing the audio stream to obtain a voiceover audio stream and generating second audio object metadata information for the voiceover audio stream include:

[0016] In the case where the voiceover identifier is included in the program list information, copy the voiceover audio stream in the audio stream;

[0017] Generate second audio object metadata information for the copied voiceover audio stream so that the voiceover audio stream becomes a sound object, wherein the second audio object metadata information includes: an audio object name, an identifier indicating whether interaction is supported, an initial volume, and an initial position.

[0018] Optionally, the step of calling the audio rendering module to perform audio rendering on the adjusted audio stream and the voiceover audio stream based on the metadata information and the second audio object metadata information, obtaining a target audio rendering signal and outputting it includes:

[0019] Based on the metadata information and the second audio object metadata information, initialize the parameters of the audio track for user interaction;

[0020] Call the audio rendering module to perform audio rendering on the adjusted audio stream and the voiceover audio stream, obtain a target audio rendering signal matching the specified speaker, and output the target audio rendering signal to the specified speaker.

[0021] Optionally, the method further includes:

[0022] Receive an adjustment operation on the metadata information and / or the second audio object metadata information, wherein the adjustment operation includes at least one of the following: muting, volume adjustment, and sound azimuth adjustment;

[0023] In response to the adjustment operation, modify the corresponding audio object metadata information.

[0024] An embodiment of the present invention further provides an audio stream processing device, wherein the device includes:

[0025] A decoding module, configured to decode an input audio stream to generate N audio tracks, as well as metadata information and audio signals corresponding to each of the audio tracks;

[0026] A judgment module, configured to judge whether the received program list information contains a voiceover identifier;

[0027] A preprocessing module, configured to, when the program list information contains a voiceover identifier, preprocess the audio stream to obtain a voiceover audio stream, and generate second audio object metadata information for the voiceover audio stream;

[0028] An audio track adjustment module, configured to set the target audio track corresponding to the voiceover among the N audio tracks as a muted audio track;

[0029] A first calling module, configured to call an audio rendering module to perform audio rendering on the adjusted audio stream and the voiceover audio stream based on the metadata information and the second audio object metadata information, to obtain a target audio rendering signal and output it.

[0030] Optionally, the apparatus further includes:

[0031] A second calling module, configured to, after the judgment module judges whether the received program list information contains a voiceover identifier, when the program list information does not contain a voiceover identifier, call the audio rendering module to render the audio stream, to obtain a target audio rendering signal and output it.

[0032] Optionally, the N audio tracks include: an audio track corresponding to the left main channel, an audio track corresponding to the right main channel, a center audio track corresponding to the center channel, an audio track corresponding to the left surround channel, an audio track corresponding to the right surround channel, and an audio track corresponding to the subwoofer channel; wherein, the target audio track is the center audio track.

[0033] Optionally, the preprocessing module includes:

[0034] A first sub-module, configured to, when the program list information contains a voiceover identifier, copy the voiceover audio stream in the audio stream;

[0035] A second sub-module, configured to generate second audio object metadata information for the copied voiceover audio stream, so that the voiceover audio stream becomes a sound object, wherein the second audio object metadata information includes: an audio object name, an identifier for indicating whether interaction is supported, an initial volume, and an initial position.

[0036] Optionally, the first calling module includes:

[0037] A third sub-module, configured to initialize parameters of an audio track for user interaction based on the metadata information and the second audio object metadata information;

[0038] A fourth sub-module, configured to call an audio rendering module to perform audio rendering on the adjusted audio stream and the voice-over audio stream, obtain a target audio rendering signal matching a specified speaker, and output the target audio rendering signal to the specified speaker.

[0039] Optionally, the apparatus further includes:

[0040] A receiving module, configured to receive an adjustment operation on the metadata information and / or the second audio object metadata information, where the adjustment operation includes at least one of the following: muting, volume adjustment, and sound azimuth adjustment;

[0041] A response module, configured to modify corresponding audio object metadata information in response to the adjustment operation.

[0042] In another aspect of the implementation of the present invention, an electronic device is further provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; the memory is used to store a computer program; when the processor executes the program stored on the memory, it implements the audio stream processing method described in any one of the above.

[0043] The audio stream processing solution disclosed in the present invention decodes an input audio stream to generate N audio tracks and first audio object metadata information corresponding to each audio track; when a voice-over identifier is included in the program list information, preprocesses the audio stream to obtain a voice-over audio stream, and generates second audio object metadata information for the voice-over audio stream; sets the target audio track corresponding to the voice-over in the N audio tracks as a muted audio track; calls an audio rendering module to perform audio rendering on the adjusted audio stream and the voice-over audio stream based on the metadata information and the second audio object metadata information, obtain a target audio rendering signal and output it. This audio stream processing solution can, without changing the TV broadcast system and the TV channel transmission system, adapt through software at the playback terminal to achieve separate control of the voice-over and environmental sounds in traditional TV audio, such as muting, volume adjustment, and spatial position change of the voice-over, thereby improving the user experience of traditional TV audio. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a flowchart showing the steps of an audio stream processing method according to an embodiment of the present application;

[0045] Figure 2 is a schematic diagram showing the audio stream processing flow according to an embodiment of the present application;

[0046] Figure 3 It is a structural block diagram showing an audio stream processing device according to an embodiment of the present application. Detailed implementation manners

[0047] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0048] To solve the problem in the prior art that the ambient sound and the voice-over cannot be controlled separately, the ambient sound can be used as a sound bed, and the voice-over can be used as a sound object to be encoded and transmitted separately, and then unified rendering is performed at the terminal according to the metadata information to form the required sound signal and transmitted to the speaker for playback. In this case, the voice-over is used as a sound object, and the voice-over can be muted, adjusted in volume, position of occurrence, etc. by using the interactive metadata information, but the television broadcast and transmission system needs to be transformed to adapt to this requirement, with a large investment and more consumption of bandwidth resources. The present invention provides an audio stream processing method for adjusting the audio track and supplementing the adjusted audio metadata in the playback terminal. This method can perform separate playback control on the voice-over under the condition of using the traditional television program broadcast and transmission scheme without transforming the television broadcast and transmission system.

[0049] The following will, with reference to the accompanying drawings, through specific embodiments and their application scenarios, describe in detail the audio stream processing solution provided by the embodiments of the present application.

[0050] As shown in the attached Figure 1 figure, the audio stream processing method of the embodiment of the present application includes the following steps:

[0051] Step 101: Decode the input audio stream to generate N audio tracks and the corresponding metadata information and audio signals for each audio track.

[0052] Exemplarily, the number of N is 6, and the N audio tracks include: the audio track corresponding to the left main channel, the audio track corresponding to the right main channel, the center audio track corresponding to the center channel, the audio track corresponding to the left surround channel, the audio track corresponding to the right surround channel, and the audio track corresponding to the subwoofer channel. It should be noted that the number of N is not limited to 6, and can also be any other feasible positive integer, and the number of N can be adaptively adjusted according to the audio stream type and output source.

[0053] The audio stream processing method provided by the embodiments of the present application can be applied to an electronic device. The electronic device stores a computer program related to audio stream processing, and the processor of the electronic device executes the computer program to perform audio stream processing. During the process of playing a TV, when the TV sound contains a voice-over, such as a live sports game, a documentary, etc., the audio stream processing method provided by the embodiments of the present application can be used to achieve a better listening effect through the interaction with the voice-over audio track.

[0054] Step 102: Determine whether the currently received program schedule information contains a voice-over identifier.

[0055] The voice-over identifier is used to indicate whether the audio stream corresponding to the program schedule information contains a voice-over. In the case where the program schedule information does not contain a voice-over identifier, call the audio rendering module to render the audio stream to obtain a target audio rendering signal and output it. In the case where the program schedule information contains a voice-over identifier, execute Step 103.

[0056] Step 103: In the case where the program schedule information contains a voice-over identifier, preprocess the audio stream to obtain a voice-over audio stream, and generate second audio object metadata information for the voice-over audio stream.

[0057] In an optional embodiment, the method of preprocessing the audio stream to obtain a voice-over audio stream and generating second audio object metadata information for the voice-over audio stream may be as follows:

[0058] In the case where the program schedule information contains a voice-over identifier, copy the voice-over audio stream in the audio stream; generate second audio object metadata information for the copied voice-over audio stream to make the voice-over audio stream a sound object, where the second audio object metadata information includes: audio object name, an identifier indicating whether interaction is supported, initial volume, and initial position.

[0059] The preprocessing of the audio stream can be performed by a preprocessing module set in the playback terminal. Through preprocessing, the voice-over audio stream can be processed into an independent sound object. Subsequently, during playback, the user can flexibly adjust the sound object through the interaction function provided by the playback terminal, such as: muting, increasing volume, decreasing volume, adjusting sound orientation, etc.

[0060] Step 104: Set the target audio track corresponding to the voice-over in the N audio tracks as a muted audio track.

[0061] Among them, the target audio track can be a center audio track. In the actual implementation process, the specific selection of the target audio track needs to be determined according to the audio track corresponding to the voice-over audio, that is, the target audio track is the audio track corresponding to the voice-over audio.

[0062] The voice-over audio stream in the original audio stream is generally transmitted through the center audio track. Therefore, the target audio track can be set as the center audio track.

[0063] Step 105: Call the audio rendering module to perform audio rendering on the adjusted audio stream and the voice-over audio stream based on the metadata information and the second audio object metadata information to obtain a target audio rendering signal and output it.

[0064] An optionally invoked audio rendering module can render the adjusted audio stream and the voiceover audio stream based on metadata information and second audio object metadata information to obtain a target audio rendering signal and output it in the following manner:

[0065] First, initialize the parameters of the audio track for user interaction based on the metadata information and the second audio object metadata information; among them, the parameters of the audio track include information such as the volume size and azimuth of the audio track.

[0066] Second, invoke the audio rendering module to render the adjusted audio stream and the voiceover audio stream to obtain a target audio rendering signal matching the specified speaker, and output the target audio rendering signal to the specified speaker.

[0067] In an optionally implemented embodiment, the audio stream of the voiceover can also be adjusted during playback. The playback terminal receives the user's adjustment operation on the metadata information and / or the second audio object metadata information, and in response to the adjustment operation, modifies the corresponding audio object metadata information.

[0068] Among them, the adjustment operation includes at least one of the following: muting, volume adjustment, and sound azimuth adjustment;

[0069] This optionally way of adjusting the parameters related to the audio object during playback facilitates the user to flexibly adjust the played audio stream according to needs, and can improve the interest and convenience.

[0070] The audio stream processing method provided by the embodiments of the present application decodes the input audio stream to generate N audio tracks and the corresponding metadata information for each audio track; in the case where the voiceover identifier is included in the program list information, preprocess the audio stream to obtain the voiceover audio stream, and generate second audio object metadata information for the voiceover audio stream; set the target audio track corresponding to the voiceover in the N audio tracks as the muted audio track; invoke the audio rendering module to render the adjusted audio stream and the voiceover audio stream based on the metadata information and the second audio object metadata information to obtain a target audio rendering signal and output it. This audio stream processing solution can, without changing the TV broadcast system and the TV channel transmission system, adapt through software on the playback terminal to realize the separate control of the voiceover and environmental sounds in traditional TV sound, such as muting, volume adjustment, and spatial position change of the voiceover, thereby improving the user experience of traditional TV sound.

[0071] Next, refer to the attached Figure 2 Another embodiment is used to illustrate the audio stream processing method of the embodiments of the present application.

[0072] In order to achieve interactive control of a specific audio track in the television audio on the TV channel playback terminal, a method of preprocessing audio track data and metadata by utilizing the rendering capabilities of the playback terminal is designed in this application. In this embodiment, the original TV audio is 5.1 surround sound, and the center audio track is used as the voiceover for illustration. The specific implementation steps are as follows:

[0073] Step 1: Decode the audio - video stream of the TV channel to form 6 audio tracks and corresponding metadata information, forming 6 - channel PCM audio signals (left and right main channels LR, center channel C, left and right surround channels SR / LR, and a sub - woofer channel SW);

[0074] Step 2: Receive the program list information. If there is no special audio description for the current program, the audio signal stream is passed through to the audio rendering output module and broadcast as ordinary 5.1 audio signals. If there is a voiceover identifier in the program list, the audio signal is input into the audio signal pre - processing module;

[0075] Step 3: The pre - processing module pre - processes the input audio stream. The voiceover audio (usually the center audio track) in the original audio signal is copied to form a new audio signal stream, that is, the audio stream, and metadata information (i.e., the second object metadata information) is prepared for this audio signal stream. The metadata information includes the name of the audio object, whether it supports interaction, the initial volume, the initial position, making this audio signal become a sound object (i.e., AudioObject). The audio track (center audio track) in the original audio stream is set to a mute track. Thus, a structure consisting of 6 audio tracks forming a 5.1 sound bed channel plus a voiceover audio object can be formed, along with metadata information describing the audio stream;

[0076] Step 4: Based on the metadata information of the audio stream, initialize information such as the volume size and azimuth of the audio track for interaction; and the rendering module renders the pre - processed sound bed, sound object, and metadata information into an audio signal output with a specified speaker layout.

[0077] Step 5: During the video stream playback, the player provides an interactive function between the user and the playback terminal, and adjusts the corresponding metadata through the metadata processing module to achieve adjustment of the rendering effect. Among them, the adjustments include muting, volume adjustment, sound azimuth adjustment, etc.

[0078] The method of using metadata to control audio rendering to achieve audio track interaction in the embodiments of this application, without changing the TV broadcast system and the TV channel transmission system, makes full use of the existing production resources and transmission resources. Through adaptation on the playback terminal, it realizes control such as muting, volume adjustment, and spatial position change of the voiceover in traditional TV audio, and can improve the user experience of traditional TV audio.

[0079] Figure 3 The structural block diagram of an audio stream processing device according to an embodiment of the present application is shown.

[0080] The audio stream processing device provided by the embodiment of the present application includes the following functional modules:

[0081] A decoding module 301, configured to decode an input audio stream to generate N audio tracks and corresponding metadata information and audio signals for each of the audio tracks;

[0082] A judgment module 302, configured to judge whether the received program list information contains a voice-over identifier;

[0083] A preprocessing module 303, configured to, when the program list information contains a voice-over identifier, preprocess the audio stream to obtain a voice-over audio stream and generate second audio object metadata information for the voice-over audio stream;

[0084] An audio track adjustment module 304, configured to set a target audio track corresponding to the voice-over among the N audio tracks as a muted audio track;

[0085] A first calling module 305, configured to call an audio rendering module to perform audio rendering on the adjusted audio stream and the voice-over audio stream based on the metadata information and the second audio object metadata information to obtain and output a target audio rendering signal.

[0086] Optionally, the device further includes:

[0087] A second calling module, configured to, after the judgment module judges whether the received program list information contains a voice-over identifier, when the program list information does not contain a voice-over identifier, call an audio rendering module to render the audio stream to obtain and output a target audio rendering signal.

[0088] Optionally, the N audio tracks include: an audio track corresponding to the left main channel, an audio track corresponding to the right main channel, a center audio track corresponding to the center channel, an audio track corresponding to the left surround channel, an audio track corresponding to the right surround channel, and an audio track corresponding to the subwoofer channel; wherein, the target audio track is the center audio track.

[0089] Optionally, the preprocessing module includes:

[0090] A first sub-module, configured to copy the voice-over audio stream in the audio stream when the program list information contains a voice-over identifier;

[0091] A second sub-module, configured to generate second audio object metadata information for the copied voiceover audio stream, so that the voiceover audio stream becomes a sound object, where the second audio object metadata information includes: an audio object name, an identifier indicating whether interaction is supported, an initial volume, and an initial position.

[0092] Optionally, the first calling module includes:

[0093] A third sub-module, configured to initialize parameters of a track for user interaction based on the metadata information and the second audio object metadata information;

[0094] A fourth sub-module, configured to call an audio rendering module to perform audio rendering on the adjusted audio stream and the voiceover audio stream, obtain a target audio rendering signal matching a specified speaker, and output the target audio rendering signal to the specified speaker.

[0095] Optionally, the apparatus further includes:

[0096] A receiving module, configured to receive an adjustment operation on the metadata information and / or the second audio object metadata information, where the adjustment operation includes at least one of the following: muting, volume adjustment, and sound azimuth adjustment;

[0097] A response module, configured to modify corresponding audio object metadata information in response to the adjustment operation.

[0098] In another aspect of the implementation of the present invention, there is also provided an electronic device, including a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory complete mutual communication through the communication bus; the memory is used to store a computer program; when the processor executes the program stored on the memory, it implements the audio stream processing method described in any one of the above.

[0099] Provided by the embodiments of the present application Figure 3 The audio stream processing apparatus shown can implement Figure 1 Each process implemented by the method embodiment, and for the sake of brevity, it will not be repeated here.

[0100] The audio stream processing apparatus provided by the embodiments of the present application can, without changing the TV broadcast system and the TV channel transmission system, adapt through software at the playback terminal to achieve separate control of the voiceover and ambient sounds in traditional TV audio, such as muting, volume adjustment, spatial position change, etc. for the voiceover, thereby improving the user experience of traditional TV audio.

[0101] An embodiment of the present invention further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus.

[0102] The memory is used to store a computer program;

[0103] When the processor is used to execute the program stored on the memory, it implements the audio stream processing method shown in the above method embodiment.

[0104] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0105] The communication interface is used for communication between the above terminal and other devices.

[0106] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0107] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0108] In another embodiment provided by the present invention, a computer-readable storage medium is also provided. Instructions are stored in the computer-readable storage medium. When it runs on a computer, it enables the computer to implement any one of the audio stream processing methods in the above embodiments.

[0109] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, which, when running on a computer, enables the computer to implement the audio stream processing method described in any one of the above embodiments.

[0110] It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising that element.

[0111] The above are the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An audio stream processing method, characterized in that: The method comprises: Decoding the input audio stream to generate N audio tracks and metadata information and audio signals corresponding to each of the audio tracks; the audio stream is an audio stream containing television audio; Determine whether the currently received program list information contains a voice-over identifier; In a case where the program list information includes a voice-over identifier, preprocessing the audio stream to obtain a voice-over audio stream, and generating second audio object metadata information for the voice-over audio stream; Setting the target audio track corresponding to the voice-over in the N audio tracks as a silent audio track; Calling an audio rendering module to perform audio rendering on the adjusted audio stream and the voice-over audio stream based on the metadata information and the metadata information of the second audio object, obtain a target audio rendering signal and output it; In the case where the program list information includes a voice-over identifier, the step of preprocessing the audio stream to obtain a voice-over audio stream, and generating second audio object metadata information for the voice-over audio stream includes: In the case where the program list information includes a voice-over identifier, copying the voice-over audio stream in the audio stream; Generate second audio object metadata information for the copied voice-over audio stream to make the voice-over audio stream a sound object, wherein the second audio object metadata information includes: an audio object name, an identifier for indicating whether interaction is supported, an initial volume, and an initial position.

2. The method according to claim 1, characterized in that After the step of determining whether the currently received program list information contains a voice-over identifier, the method further includes: In the case that the program list information does not include a voice-over identifier, an audio rendering module is called to render the audio stream to obtain and output a target audio rendering signal.

3. The method according to claim 1, characterized in that The N audio tracks include: an audio track corresponding to the left main channel, an audio track corresponding to the right main channel, a center audio track corresponding to the center channel, an audio track corresponding to the left surround channel, an audio track corresponding to the right surround channel, and an audio track corresponding to the subwoofer channel; wherein the target audio track is the center audio track.

4. The method according to claim 1, characterized in that: The step of calling the audio rendering module, performing audio rendering on the adjusted audio stream and the voice-over audio stream based on the metadata information and the metadata information of the second audio object, obtaining a target audio rendering signal and outputting the signal includes: Initialize parameters of the audio track for user interaction based on the metadata information and the second audio object metadata information; An audio rendering module is called to perform audio rendering on the adjusted audio stream and the voice-over audio stream to obtain a target audio rendering signal that matches a designated speaker, and the target audio rendering signal is output to the designated speaker.

5. The method according to claim 1, characterized in that The method further comprises: Receiving an adjustment operation on the metadata information and / or the metadata information of the second audio object, wherein the adjustment operation includes at least one of the following: muting, volume adjustment, and sound orientation adjustment; In response to the adjustment operation, the corresponding audio object metadata information is modified.

6. An audio stream processing device, characterized in that: The device comprises: A decoding module, used to decode the input audio stream to generate N audio tracks and metadata information and audio signals corresponding to each audio track; A judging module, used for judging whether the currently received program list information contains a voice-over mark; A preprocessing module, configured to preprocess the audio stream to obtain a voice-over audio stream, and generate second audio object metadata information for the voice-over audio stream when the program list information includes a voice-over identifier; An audio track adjustment module, used for setting a target audio track corresponding to the voice-over in the N audio tracks as a silent audio track; A first calling module is used to call an audio rendering module to perform audio rendering on the adjusted audio stream and the voice-over audio stream based on the metadata information and the metadata information of the second audio object, obtain a target audio rendering signal and output it; The preprocessing module comprises: The first submodule is used for copying the voice-over audio stream in the audio stream when the program list information includes a voice-over identifier; The second submodule is used to generate second audio object metadata information for the copied voice-over audio stream to make the voice-over audio stream a sound object, wherein the second audio object metadata information includes: an audio object name, an identifier for indicating whether interaction is supported, an initial volume, and an initial position.

7. The device according to claim 6, characterized in that The device also includes: The second calling module is used to call the audio rendering module to render the audio stream after the judgment module determines whether the currently received program list information contains a voice-over identifier, if the program list information does not contain a voice-over identifier, to obtain a target audio rendering signal and output it.

8. The device according to claim 6, characterized in that The N audio tracks include: an audio track corresponding to the left main channel, an audio track corresponding to the right main channel, a center audio track corresponding to the center channel, an audio track corresponding to the left surround channel, an audio track corresponding to the right surround channel, and an audio track corresponding to the subwoofer channel; wherein the target audio track is the center audio track.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; The processor is used to implement the audio stream processing method as described in any one of claims 1 to 5 when executing the program stored in the memory.

Citation Information

Patent Citations

  • Music data processing method and device and computer storage medium

    CN112037738A

  • Audio metadata modification at rendering device

    US20180357038A1