Audio rendering system, method and electronic device
By spatially encoding the audio signal and using metadata information, the problem that existing audio rendering systems are difficult to achieve economically feasible and practical immersive audio experience in commercial deployment is solved, and efficient audio rendering in different usage scenarios is achieved.
Patent Information
- Application Number
- CN202280042877.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-06-15
- Filing Date
- 2022-06-15
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-06-15
AI Technical Summary
Existing audio rendering systems are difficult to achieve economically feasible and practical immersive audio experiences in commercial deployment, especially in the production and consumption of user-generated content and professional workers.
By an audio encoding method, an audio signal in a specific audio content format and its associated metadata information are obtained, and the audio signal is spatially encoded based on these metadata information to obtain the encoded audio signal.
It realizes providing an efficient and economical audio rendering system in different usage scenarios, which can meet the different needs of professional workers and ordinary users for immersion.
Smart Images

Figure CN117501362B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of international patent application No. PCT / CN2021 / 100062 filed on June 15, 2021, which is incorporated herein by reference. Technical Field
[0003] The present disclosure relates to the technical field of audio signal processing, and in particular to an audio rendering system, an audio rendering method, an electronic device, and a non-transitory computer-readable storage medium. Background Art
[0004] Audio rendering refers to the appropriate processing of sound signals from a sound source to provide users with a desired listening experience in user application scenarios, especially an immersive experience.
[0005] Generally speaking, an excellent immersive audio system should provide listeners with a feeling of being immersed in a virtual environment. However, immersion itself is not a sufficient condition for the successful commercial deployment of virtual reality multimedia services. In order to be commercially successful, the audio system should also provide content creation tools, content creation workflows, content distribution methods and platforms, and a rendering system that is economically feasible and easy to use for both consumers and creators.
[0006] Whether an audio system is practical and economically feasible for successful commercial deployment depends on the usage scenario and the level of sophistication expected in the content production and consumption process. For example, there will be very different expectations for the entire creation and consumption chain and the content playback experience for user-generated content (UGC) and professional-generated content (PGC). For example, an ordinary user who is looking for leisure purposes and a professional user will have very different requirements for the quality of content and the immersiveness provided during playback, but at the same time, they will also have different playback devices. For example, professional users may build a more sophisticated listening environment. Summary of the invention
[0007] According to some embodiments of the present disclosure, there is provided an audio encoding method for audio rendering, comprising an acquisition step for acquiring an audio signal in a specific audio content format and metadata-related information associated with the audio signal in the specific audio content format; and an encoding step for spatially encoding the audio signal in the specific audio content format based on the metadata-related information associated with the audio signal in the specific audio content format to obtain an encoded audio signal.
[0008] According to some other embodiments of the present disclosure, an audio rendering method is provided, comprising an audio signal encoding step for spatially encoding an audio signal in a specific audio content format using the audio encoding method of any embodiment described in the present disclosure to obtain an encoded audio signal; and an audio signal decoding step for spatially decoding the encoded audio signal to obtain a decoded audio signal for audio rendering.
[0009] According to some other embodiments of the present disclosure, an audio encoder for audio rendering is provided, comprising: an acquisition unit configured to acquire an audio signal in a specific audio content format and metadata-related information associated with the audio signal in the specific audio content format; and an encoding unit configured to spatially encode the audio signal in the specific audio content format based on the metadata-related information associated with the audio signal in the specific audio content format to obtain an encoded audio signal.
[0010] According to some other embodiments of the present disclosure, an audio rendering device is provided, comprising: an audio encoder according to any embodiment described in the present disclosure; and an audio signal decoder, configured to spatially decode an encoded audio signal obtained by the audio encoder to obtain a decoded audio signal for audio rendering.
[0011] According to some other embodiments of the present disclosure, a chip is provided, comprising: at least one processor and an interface, wherein the interface is used to provide computer execution instructions to the at least one processor, and the at least one processor is used to execute the computer execution instructions to implement at least one of the audio encoding method and the audio rendering method of any embodiment described in the present disclosure.
[0012] According to some further embodiments of the present disclosure, a computer program is provided, comprising: instructions, which, when executed by a processor, cause the processor to execute at least one of the audio encoding method and the audio rendering method of any embodiment described in the present disclosure.
[0013] According to some further embodiments of the present disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute at least one of the audio encoding method and the audio rendering method of any embodiment described in the present disclosure based on instructions stored in the memory device.
[0014] According to some further embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, at least one of the audio encoding method and the audio rendering method of any embodiment described in the present disclosure is implemented.
[0015] According to some further embodiments of the present disclosure, a computer program product is provided, comprising instructions, which, when executed by a processor, implement at least one of the audio encoding method and the audio rendering method of any one of the embodiments described in the present disclosure.
[0016] Other features and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of the present application. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation on the present disclosure. In the drawings:
[0018] Figure 1 A schematic diagram illustrating some embodiments of an audio signal processing process;
[0019] Figure 2A and Figure 2B Schematic diagrams showing some embodiments of audio system architectures;
[0020] Figure 3A shows a schematic diagram of a tetrahedral B-format microphone;
[0021] Figure 3B Schematic diagrams showing spherical harmonic functions of N = 0th order (first row) to 3rd order (last row);
[0022] Figure 3C A schematic diagram of a HOA microphone is shown;
[0023] Figure 3D shows a schematic diagram of an XY pair of stereo microphones;
[0024] Figure 4A A block diagram of an audio rendering system according to an embodiment of the present disclosure is shown;
[0025] Figure 4B A schematic conceptual diagram of an audio rendering process according to an embodiment of the present disclosure is shown;
[0026] Figure 4C and 4D A schematic diagram showing a pre-processing operation in an audio rendering system according to an embodiment of the present disclosure is shown;
[0027] Figure 4E shows a block diagram of an audio signal encoding module according to an embodiment of the present disclosure,
[0028] Figure 4F A flowchart of spatial encoding of an audio signal according to an embodiment of the present disclosure is shown;
[0029] Figure 4G A flowchart showing an exemplary implementation of an audio rendering process according to an embodiment of the present disclosure is shown;
[0030] Figure 4H A schematic diagram showing an exemplary implementation of an audio rendering process according to an embodiment of the present disclosure;
[0031] Fig. 4I A flowchart of an audio rendering method according to an embodiment of the present disclosure is shown;
[0032] Figure 5 A block diagram showing some embodiments of the electronic device of the present disclosure;
[0033] Figure 6 A block diagram showing some other embodiments of the electronic device of the present disclosure;
[0034] Figure 7 A block diagram showing some embodiments of the chip of the present disclosure.
[0035] It should be understood that, for ease of description, the sizes of the various parts shown in the drawings are not necessarily drawn according to the actual proportional relationship. The same or similar reference numerals are used in the various drawings to represent the same or similar parts. Therefore, once an item is defined in one drawing, it may not be further discussed in subsequent drawings. DETAILED DESCRIPTION
[0036] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is by no means intended to limit the present disclosure and its application or use. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0037] Unless otherwise specifically stated, the relative arrangement of the parts and steps set forth in these embodiments, numerical expressions and numerical values do not limit the scope of the present disclosure. The techniques, methods and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but in appropriate cases, the techniques, methods and equipment should be considered as part of the authorization specification. In all examples shown and discussed here, any specific value should be interpreted as being merely exemplary, rather than as a limitation. Therefore, other examples of exemplary embodiments may have different values.
[0038] It should be understood that the various steps recorded in the method embodiments of the present disclosure can be performed in different orders, and / or performed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown in the execution. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions and numerical values of the parts and steps set forth in these embodiments should be interpreted as being merely exemplary and do not limit the scope of the present disclosure.
[0039] The term "include" and its variations used in the present disclosure mean an open term that includes at least the following elements / features but does not exclude other elements / features, that is, "including but not limited to". In addition, the term "include" and its variations used in the present disclosure mean an open term that includes at least the following elements / features but does not exclude other elements / features, that is, "including but not limited to". Therefore, include is synonymous with include. The term "based on" means "based at least in part on".
[0040] References throughout this specification to "one embodiment," "some embodiments," or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. For example, the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; and the term "some embodiments" means "at least some embodiments." Furthermore, the appearances of the phrases "in one embodiment," "in some embodiments," or "in an embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment, but may do so.
[0041] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units. Unless otherwise specified, the concepts of "first", "second", etc. are not intended to imply that the objects described in this way must be in a given order in time, space, ranking, or any other manner.
[0042] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0043] Figure 1 Some conceptual diagrams of audio signal processing, especially from acquisition to rendering process / system are shown. Figure 1As shown, in this system, the audio signal is processed or produced after being collected, and the processed / produced audio signal is distributed to the rendering end for rendering, so that it is presented to the user in an appropriate form to meet the user experience. It should be pointed out that such an audio signal processing flow can be applied to various application scenarios, especially virtual reality audio content expression.
[0044] In particular, according to an embodiment of the present disclosure, virtual reality audio content expression broadly involves metadata, a renderer / rendering system, an audio codec, etc., wherein metadata, a renderer / rendering system, and an audio codec can be logically separated from each other. When performing local storage and production, the renderer / rendering system can directly process metadata and audio signals without audio encoding and decoding. In particular, the renderer / rendering system here is used for audio content production. On the other hand, when used for transmission (such as live broadcast or two-way communication), the transmission format of metadata+audio stream can be set, and then the metadata and audio content are transmitted to the renderer / rendering system through an intermediate process including a coding and decoding process for rendering to the user. In some embodiments, such as an exemplary embodiment of virtual reality audio content expression, input audio signals and metadata can be obtained from the acquisition end, wherein the input audio signal includes various appropriate forms, such as channels, objects, HOA, or mixed formats thereof. Metadata may include appropriate types, such as dynamic metadata and static metadata, wherein dynamic metadata may be transmitted together with the input audio signal, for example, in various appropriate ways. As an example, metadata information may be generated according to metadata definitions, wherein dynamic metadata may be transmitted along with the audio stream, and the specific encapsulation format is defined according to the type of transmission protocol adopted by the system layer. Of course, metadata may also be directly transmitted to the playback end without further generating metadata information. For example, static metadata may be directly transmitted to the playback end without going through the encoding and decoding process. During the transmission process, the input audio signal will be audio encoded, then transmitted to the playback end, and then decoded for playback to the user through a playback device, such as a renderer. At the playback end, the renderer renders and outputs the decoded audio file with metadata. Logically, metadata and audio encoding and decoding are independent of each other, and the decoder and renderer are decoupled. The renderer may be configured with an identifier, that is, the renderer has a corresponding identifier, and different renderers have different identifiers. As an example, the renderer adopts a registration system, that is, the playback end is set with multiple IDs, which respectively indicate the multiple renderers / rendering systems that the playback end can support. For example, it may include at least 4 IDs, ID1 indicates a renderer based on binaural output, ID2 indicates a renderer based on speaker output, ID3-ID4 can indicate other types of renderers, and various renderers can indicate the same metadata definition, and of course can also support different metadata definitions. Each renderer can have a corresponding metadata definition. In this case, a specific metadata identifier can be used to indicate a specific metadata definition during the transmission process, so that the renderer can have a corresponding metadata identifier, so that the playback end can select the corresponding renderer according to the metadata identifier to play back the audio signal.
[0045] Figure 2A and 2BAn exemplary implementation of an audio system is shown. Figure 2A A schematic diagram showing an exemplary architecture of an audio system according to some embodiments of the present disclosure is shown. Figure 2A As shown, the audio system may include, but is not limited to, audio acquisition, audio content production, audio storage / distribution, and audio rendering. Figure 2B An exemplary implementation of the various stages of an audio rendering process / system is shown. The production and consumption stages in the audio system are mainly shown, and optionally also include intermediate processing stages, such as compression. The production and consumption stages here can correspond to Figure 2A The intermediate processing stage can be included in the exemplary implementation of the production and rendering stages shown in Figure 2A The distribution phase shown in , of course, can be included in the production phase, rendering phase. Figure 2A and 2B To describe the implementation of each part of the audio system. It should be noted that in addition to the consideration of the complexity of acquisition, production, distribution and rendering, the audio system may also need to meet other requirements for audio scenarios to support communication, such as latency, and such requirements can be met by corresponding processing means, which will not be described in detail here.
[0046] Audio Capture
[0047] In the audio acquisition stage, the audio scene is captured to acquire the audio signal. The audio acquisition can be processed by appropriate audio acquisition means / systems / devices, etc.
[0048] The audio acquisition system may be closely related to the format used in the production of audio content, and the audio content format may include at least one of the following three: scene-based audio representation, channel-based audio representation, and object-based audio representation, and for each audio content format, corresponding or adapted devices and / or methods may be used for capture. As an example, for applications that support scene-based audio representation, a spherical microphone array may be used to capture scene audio signals, while in applications that use channel-based audio and object-based audio representation, one or more specifically optimized microphones may be used to record sound to capture audio signals. Additionally, audio acquisition may also include appropriate post-processing of the captured audio signals. The following will exemplarily describe audio acquisition of various audio content formats.
[0049] Acquisition of scene-based audio representations
[0050] A scene-based audio representation is a scalable, speaker-independent representation of a sound field, for example, an example definition is given in ITU RBS.2266-2. According to some embodiments, scene-based audio may be based on a set of orthogonal basis functions, such as spherical harmonics.
[0051] According to some embodiments, examples of scene-based audio formats used may include B-Format, First Order Ambisonics (FOA), Higher Order Ambisonics (HOA), etc. Ambisonics (Ambisonics) indicates an omnidirectional audio system, i.e., it can include sound sources above and below the listener in addition to the horizontal plane. The auditory scene of Ambisonics can be captured by using first-order or higher-order Ambisonics microphones. As an example, a scene-based audio representation may generally indicate an audio signal including a HOA.
[0052] According to some embodiments, a B-format microphone or a first-order Ambisonics (FOA) format may use the first four low-order spherical harmonics to represent a three-dimensional sound field with four signals W, X, Y, and Z. Among them, W is used to record the sound pressure in all directions, X is used to record the front / rear sound pressure gradient of the collection position, Y is used to record the left / right sound pressure gradient at the collection position, and Z is used to record the upper / lower sound pressure gradient at the collection position. These four signals can be generated by processing the original signals of the so-called "tetrahedron" microphone, which may be composed of four microphones in a configuration of left front upper (LFU), right front lower (RFD), left rear lower (LBD), and right rear upper (RBU), such as Figure 3A shown.
[0053] In some embodiments, a B-format microphone array configuration may be deployed on a portable spherical audio and video capture device, with raw microphone signal components processed in real time to derive W, X, Y, and Z components. According to some examples, horizontal only B-format microphones may be used for auditory scene capture and audio capture. In particular, some configurations may support horizontal only B-format, where only W, X, and Y components are captured, without capturing the Z component. Compared to the 3D audio capabilities of FOA and HOA, the horizontal only B-format gives up the additional immersion provided by the height information.
[0054] In some embodiments, multiple formats for exchanging higher-order Ambisonics data may be included. In the HOA data exchange format, the channel order, normalization method, and polarity should be correctly defined. In some embodiments, for HOA signals, the auditory scene can be captured by a higher-order Ambisonics microphone. In particular, compared to first-order Ambisonics, the spatial resolution and listening area can be greatly enhanced by increasing the number of directional microphones, for example, by second-order, third-order, fourth-order, and higher-order Ambisonics systems (collectively referred to as HOA, Higher Order Ambisonics). An N-order three-dimensional Ambisonics system requires (N+1) 2 The distribution of these microphones can be consistent with the distribution of spherical harmonics of the same order. Figure 3B Spherical harmonic functions of N=0th order (first row) to 3rd order (last row) are shown. Figure 3C The HOA microphone is shown.
[0055] Acquisition of channel-based audio representations
[0056] The acquisition of channel-based audio representations often uses microphones for audio acquisition and may also include channel-based post-processing. As an example, an object-based audio representation may generally indicate an audio signal comprising channels. Such an acquisition system may use multiple microphones to capture sounds from different directions; or use overlapping or spaced microphone arrays. According to some embodiments, different channel-based formats may be created based on the number and spatial arrangement of microphones, for example, Figure 3D The XY pair stereo microphone shown in the figure uses a microphone array to record 8.0 channel content. In addition, the microphone built into the user's device can also achieve channel-based audio format recording, such as using a mobile phone to record stereo.
[0057] Acquisition of object-based audio representations
[0058] According to some embodiments, an object-based audio representation may represent an entire complex audio scene using a collection of a series of single audio elements, each of which includes an audio waveform and a set of related parameters or metadata. The metadata may specify the movement and transformation of each audio element in the sound scene, thereby reproducing the audio scene originally designed by the artist. The experience provided by object-based audio generally exceeds that of general mono audio acquisition, making the audio more likely to meet the artistic intent of the producer. As an example, an object-based audio representation may generally indicate an audio signal including an object.
[0059] According to some embodiments, the spatial accuracy of the object-based audio representation depends on the metadata and the rendering system. It is not directly related to the number of channels the audio contains.
[0060] The collection of object-based audio representations can be captured using appropriate collection equipment, such as speakers, and is appropriately processed. For example, a mono track can be collected and further processed based on metadata to obtain an object-based audio representation. As an example, sound objects typically use mono tracks that have been recorded or generated through sound design. These mono tracks can be further processed as sound elements in tools such as digital audio workstations (DAWs), such as using metadata to specify that the sound elements are on the horizontal plane around the listener, or even at any position in three-dimensional space. Therefore, a "track" in a DAW can correspond to an audio object.
[0061] Additionally, according to the embodiments of the present disclosure, in order to achieve or even further optimize the sense of immersion, the audio acquisition system may generally also consider the following factors and optimize accordingly:
[0062] -Signal-to-Noise Ratio (SNR). Noise sources that are not part of the audio scene tend to reduce realism and immersion, so the audio capture system should have a low enough noise floor that it is appropriately masked by the recorded content but unnoticeable in the reproduction process.
[0063] -Acoustic Overload Point (AOP). The nonlinear behavior of the audio acquisition system may reduce the sense of realism. Therefore, the microphone in the audio acquisition system should have a sufficiently high acoustic overload point to avoid the audio scene of interest exceeding the threshold and causing nonlinear distortion.
[0064] -Microphone frequency response. The microphone should have a flat frequency response across the entire frequency range.
[0065] - Wind noise protection. Wind noise may cause non-linear audio behavior, which reduces realism. Therefore, the audio acquisition system or microphone should be designed to attenuate wind noise, for example, to keep it below a certain threshold.
[0066] - Microphone element configuration, such as spacing, crosstalk, gain, and directivity matching: These aspects ultimately enhance or degrade the spatial accuracy of scene-based audio reproduction. Therefore, the above configuration aspects of the microphone can be optimized while ensuring spatial accuracy.
[0067] - Latency. If two-way communication is required, the mouth to ear latency should be low enough to allow a natural conversational experience. Therefore, the audio acquisition system should be designed to achieve low latency, for example, below a certain latency threshold.
[0068] It should be noted that the above audio acquisition process and various audio representations are merely exemplary and non-restrictive. The audio representation may also be in other suitable forms known or to be known in the future, and may be acquired by using appropriate devices, as long as such audio representation can be acquired from the music scene and can be used to present to the user.
[0069] Audio content production
[0070] After the audio signal is acquired by the audio capture / collection system, the audio signal will be input into the production stage for audio content production.
[0071] In some embodiments, in the audio content production process, the producer's creation function of the audio content needs to be satisfied. For example, for an object-based sound representation system, the creator needs to have the ability to edit sound objects and generate metadata, and the aforementioned metadata generation operation can be performed here. The producer can achieve the creation of audio content in various appropriate ways.
[0072] In one example, if Figure 2B As shown in , in the production stage, input audio data and audio metadata are received, and the audio data and audio metadata are processed, especially authorization and metadata tagging, to obtain a production result. In some embodiments, exemplarily, the input of the audio processing may include, but is not limited to, target-based audio signals, FOA (First-Order Ambisonics, first-order spherical sound field signal), HOA (Higher-Order Ambisonics, higher-order spherical sound field signal), stereo, surround sound, etc., in particular, the input of the audio processing may also include scene information and metadata, etc., which are associated with the input metadata. In some embodiments, the audio data is input to the audio track interface for processing, and the audio metadata is processed via general audio source data (such as ADM extension, etc.). Optionally, standardization processing can also be performed, especially for the results obtained by authorization and metadata tagging.
[0073] In some embodiments, in the audio content production process, the creator also needs to be able to monitor and modify the work in a timely manner. As an example, an audio rendering system can be provided to provide a scene monitoring function. In addition, in order for consumers to obtain the artistic intent that the creator wants to express, the rendering system provided for the creator to monitor should be the same as the rendering system provided to the consumer to ensure a consistent experience.
[0074] Audio production formats
[0075] In or after the audio content production process, audio content with an appropriate audio production format can be obtained. According to an embodiment of the present disclosure, the audio production format can be various appropriate formats. As an example, the audio production format can be specified in ITU-R BS.2266-2. Channel-based, object-based, and scene-based audio representations are specified in ITU-R BS.2266-2, as shown in Table 1 below. For example, all signal types in Table 1 can describe three-dimensional audio whose goal is to bring an immersive experience.
[0076] Table 1: Audio production formats
[0077]
[0078] According to some embodiments, the signal types shown in the table can be combined with audio metadata to control rendering. As an example, the audio metadata includes at least one of the following:
[0079] - Channel configuration.
[0080] -The normalization method and channel order used for scene-based audio representation.
[0081] - The configuration and properties of an object, such as its position in space.
[0082] -Narration, in particular, using head tracking technology to make the narration adapt to the movement of the listener's head, or remain stationary in the scene. For example, for commentary tracks where the speaker is invisible, head tracking is not required and static audio processing is used. For visible commentary tracks, the track is positioned to the speaker in the scene based on the head tracking results.
[0083] It should be noted that the above audio production process and various audio production formats are merely exemplary and non-restrictive. Audio production can also be performed by any other appropriate means, any other appropriate device, and any other appropriate audio production format, as long as the acquired audio signal can be processed for rendering.
[0084] Intermediate processing stage before audio rendering
[0085] According to some embodiments of the present disclosure, after the captured audio signal is produced and before being provided to the audio rendering stage, further intermediate processing may be performed on the audio signal.
[0086] In some embodiments, the intermediate processing of the audio signal may include storage and distribution of the audio signal. For example, the audio signal may be stored and distributed in a suitable format, such as an audio storage format and an audio distribution format, respectively. The audio storage format and the audio distribution format may be in various suitable forms. The following describes, as examples, existing spatial audio formats or spatial audio exchange formats related to audio storage and / or audio distribution.
[0087] An example may be a container format, such as a .mp4 container, that can hold both spatial (scene-based) and non-ambiguous audio. Such a container format may include a Spatial Audio Box (SA3D), which contains information such as the Ambisonics type, order, channel order, and normalization. The container format may also include a Non-Diegetic Audio Box (SAND), which is used to represent audio that should remain unchanged when the listener's head rotates (such as commentary, stereo music, etc.). In implementation, Ambisonic Channel Number (ACN) channel ordering and Schmidt semi-normalization (SN3D) normalization calculation may be used.
[0088] Another example may be based on the Audio Definition Model (ADM), which is an open standard that seeks to be compatible with object-, channel-, and scene-based audio systems through XML. Its purpose is to provide a way to describe audio metadata so that each individual audio track in a file or stream can be correctly rendered, processed, or distributed. The model is divided into a content part and a format part. The content part describes what is contained in the audio, such as the track language (Chinese, English, Japanese, etc.) and loudness. The format part contains technical information required for the audio to be correctly decoded or rendered, such as the position coordinates of the sound object and the order of the HOA components. For example, Recommendation ITU-R BS.2076-0 specifies a series of ADM elements, such as audioTrackFormat (describing what format the data is in), audioTrackUID (uniquely identifying an audio track or asset with an audio scene recording), audioPackFormat (grouping audio channels), etc. AMD can be used for channel-, object-, and scene-based audio.
[0089] Yet another example is AmbiX. AmbiX supports audio content based on HOA scenarios. AmbiX files contain linear PCM data with a word length of 16, 24 or 32 bits of fixed point, or 32 bits of floating point, and can support all valid sampling rates in .caf (Apple's Core Audio Format). AmbiX uses ACN sorting and SN3D normalization, and supports HOA and mixed-order Ambisonics. As a popular format for exchanging Ambisonics content, AmbiX is gaining rapid development.
[0090] As another example, the intermediate processing of the audio signal may also include appropriate compression processing. As an example, the audio content produced may be encoded / decoded to obtain a compression result, which is then provided to the rendering side for rendering. For example, such compression processing may help reduce data transmission overhead and improve data transmission efficiency. The encoding and decoding in the compression may be implemented using any appropriate technology.
[0091] It should be noted that the above-mentioned audio intermediate processing, storage, distribution, etc. formats are merely exemplary and non-restrictive. The audio intermediate processing may also include any other appropriate processing and may also adopt any other appropriate format, as long as the processed audio signal can be effectively transmitted to the audio rendering end for rendering.
[0092] It should be noted that the transmission of metadata is also included in the audio transmission process. The metadata can be in various appropriate forms and can be applicable to all audio renderers / rendering systems, or can be applied to each audio renderer / rendering system respectively. Such metadata can be referred to as rendering-related metadata, for example, it can include basic metadata and extended metadata, and the basic metadata is, for example, ADM basic metadata that conforms to BS.2076. The ADM metadata describing the audio format can be given in XML (Extensible Markup Language) form. In some embodiments, the metadata can be appropriately controlled, such as hierarchical control.
[0093] Metadata is mainly implemented using XML encoding. Metadata in XML format can be included in the "axml" or "bxml" block in the BW64 format audio file for transmission. The "audio package format identifier", "audio track format identifier" and "audio track unique identifier" in the generated metadata can be provided to the BW64 file for linking the metadata with the actual audio track. Metadata basic elements may include but are not limited to at least one of the following: audio program, audio content, audio object, audio package format, audio channel format, audio stream format, audio track format, audio track unique identifier, audio block format, etc. Extended metadata can be encapsulated in various appropriate forms, for example, it can be encapsulated in a manner similar to the aforementioned basic metadata, and can contain appropriate information, identifiers, etc.
[0094] Audio Rendering
[0095] After receiving the audio signal transmitted from the audio production stage, the audio signal is processed at the audio rendering end / playback end to be played back / presented to the user. In particular, the audio signal is rendered and presented to the user with a desired effect.
[0096] In some embodiments, the processing at the audio rendering end may include processing the signal from the audio production stage before rendering, as an example, Figure 2B As shown, according to the processing result of the production side, the metadata is restored and rendered using the audio track interface and the general audio metadata (such as ADM extension, etc.); the result after metadata restoration and rendering is audio rendered, and the obtained result is input into the audio device for consumption by the consumer. As another example, in the case where the audio signal representation is compressed in the intermediate stage, the corresponding decompression processing can also be performed at the audio rendering end.
[0097] According to an embodiment of the present disclosure, the processing of the audio rendering end may include various appropriate types of audio rendering. In particular, corresponding audio rendering processing can be used for each type of audio representation. As an example, the input data of the audio rendering end may be composed of a renderer identifier, metadata, and an audio signal. The audio rendering end may select a corresponding renderer according to the rendered indicator transmitted, and then the selected renderer reads the corresponding metadata information and audio file to perform audio playback. The input data of the audio rendering end may be in various appropriate forms, for example, various appropriate encapsulation formats may be used, such as a layered format, metadata and audio files may be encapsulated in an inner layer, and a renderer identifier may be encapsulated in an outer layer. For example, the metadata and audio files may be in a BW64 file format, and the outermost layer may be encapsulated with a renderer identifier, such as a renderer number, a renderer ID, etc.
[0098] In some embodiments, the audio rendering process may adopt scene-based audio rendering. In particular, for scene-based audio (SBA), the rendering may be independent of the capture or creation of the sound scene, and may be adaptively generated mainly for the application scene.
[0099] In one example, in a speaker presentation scenario, the rendering of the sound scene may be performed on the receiving device, and a real or virtual speaker signal may be generated. The speaker signal may be a speaker array signal S in vector form = [S1 ... S n ] T , where 1, ..., n represents the 1st, ..., nth loudspeakers. As an example, the loudspeaker signal S may be generated by S=D·B, where B is a vector of SBA signals B=[B (0,0) …B (n,m) ] T , the subscripts n and m in the vector represent the order and degree of the spherical harmonics, and D is the rendering matrix (also called decoding matrix) of the target speaker system.
[0100] In one example, in a binaural rendering scenario, the audio scene can be rendered by playing back binaural signals through headphones. The binaural signals can be represented by a binaural impulse response matrix IR of a virtual speaker signal S and a speaker position. BIN The convolution S BIN =(DB)*IR BIN get.
[0101] In one example, in an immersive application, it is desired that the sound field rotates according to the movement of the head. An audio signal suitable for this rotation situation can be realized by multiplying a rotation matrix F with the SBA signal B′=FB.
[0102] In some embodiments, the audio rendering process may employ channel-based audio rendering. In particular, for channel-based audio representation, each channel is associated with and can be presented by a corresponding speaker. The positions of the speakers are standardized in, for example, ITU-R BS.2051 or MPEG CICP.
[0103] In some embodiments, in an immersive audio scenario, each speaker channel is treated as a virtual sound source in the scene and rendered to headphones; that is, the audio signal of each channel is rendered to the correct position of a virtual listening room according to the standard. The most direct approach is to filter the audio signal of each virtual sound source with the response function measured in the reference listening room. The acoustic response function can be measured with microphones placed in the ears of a person or an artificial head. They are called binaural room impulse responses (BRIR). This method can provide high audio quality and accurate positioning, but the disadvantage is high computational complexity, especially for a large number of channels to be rendered and longer BRIRs. Therefore, some alternative methods have been developed to reduce complexity while maintaining audio quality. Typically, these alternative methods involve parametric models of BRIR, for example, by using sparse filters or recursive filters.
[0104] In some embodiments, the audio rendering process may employ object-based audio rendering. In particular, for object-based audio representation, audio rendering may be performed taking into account objects and associated metadata. In particular, in object-based audio rendering, each object sound source is presented independently along with its metadata, which describes the spatial properties of each sound source, such as position, direction, width, etc. Using these properties, the sound sources are rendered individually in the three-dimensional audio space around the listener.
[0105] Rendering can be performed for speaker arrays or headphones. In one example, speaker array rendering uses different types of speaker panning methods (such as VBAP, vector based amplitude panning) to use the sound played by the speaker array to give the listener the feeling that the object sound source is at a specified position. In another example, there are many different ways to render headphones, such as using the HRTF (Head-related transfer function) corresponding to the direction of each sound source to directly filter the sound source signal. An indirect rendering method can also be used to render the sound source to a virtual speaker array, and then perform binaural rendering on each virtual speaker.
[0106] At present, a variety of file formats and metadata that support immersive audio transmission and playback are being used. In particular, in conventional immersive audio systems, there are different audio representation methods, such as scene-based audio representation, channel-based audio representation, and object-based audio representation, and therefore, various types / formats of inputs need to be processed accordingly. In addition, the playback devices of immersive audio are different for consumer usage scenarios. Typical examples include standard speaker arrays, custom speaker arrays, special speaker arrays, headphones (binaural playback), etc., for which various types / formats of outputs need to be generated. However, there is currently no common or public file exchange standard. This will cause trouble for creators, because for different platforms, it is often necessary to repeatedly render works for each platform's definition, and in particular, it is necessary to repeatedly generate audio based on objects, channels, and scenes for each platform, as well as metadata for guiding the correct rendering of all audio elements, which leads to low efficiency and poor compatibility of existing audio systems. Therefore, it is desirable to provide a standard immersive audio rendering system that can be compatible with all the above input and output formats while ensuring rendering effects and efficiency.
[0107] In view of this, the present disclosure conceives a highly compatible and efficient audio rendering, which is compatible with various input audios and various desired audio outputs, while ensuring rendering effects and efficiency. In particular, in the present disclosure, it is possible to obtain an audio signal in a common space format that can be used in a user application scenario based on the received input audio signal, that is, even if the received input audio signal may contain or be an audio representation signal in a different format, such an audio representation signal may be transformed / encoded into an audio signal in a common space format; then the audio signal in the common space format may be decoded and processed according to the type of playback device in the user's listening environment, thereby obtaining output audio that is particularly suitable for the playback device in the user's listening environment, so that it can be well compatible with various input and output formats, and output formats that are particularly suitable for the playback device in the user's listening environment can be obtained for various inputs, thereby realizing an audio rendering system with good compatibility, and then realizing an audio system with good compatibility. As a result, the present disclosure achieves improved audio rendering, especially improved immersive audio rendering.
[0108] Hereinafter, an audio rendering system and method according to embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.
[0109] Figure 4AA block diagram of some embodiments of an audio rendering system according to an embodiment of the present disclosure is shown. The audio rendering system 4 includes an acquisition module 41, which is configured to acquire an audio signal in a specific spatial format based on an input audio signal, and the audio signal in the specific spatial format may be an audio signal in a common spatial format obtained from various possible audio representation signals for use in user application scenarios; and an audio signal decoding module 42, which is configured to be able to spatially decode the encoded audio signal in the specific spatial format to obtain a decoded audio signal for audio rendering, thereby presenting / playing back audio to the user based on the spatially decoded audio signal.
[0110] According to some embodiments of the present disclosure, the audio signal in the specific spatial format may be referred to as an intermediate audio signal in audio rendering, or may be referred to as an intermediate signal medium, which has a common specific spatial format that can be obtained from various input audio signals, for example, it may be any appropriate spatial format, as long as it can be supported by the user application scenario / user playback environment and is suitable for playback in the user playback environment. In particular, the intermediate signal may be a signal that is relatively independent of the sound source, and may be applied to different scenarios / devices for playback according to different decoding methods, thereby improving the universality of the audio rendering system of the present application. As an example, the audio signal in the specific spatial format may be an Ambisonics type audio signal, and more particularly, the audio signal in the specific spatial format is any one or more of FOA (First Order Ambisonics), HOA (Higher Order Ambisonics), and MOA (Mixed-order Ambisonics).
[0111] According to an embodiment of the present disclosure, the audio signal of the specific spatial format can be appropriately obtained based on the format of the input audio signal. In some embodiments, the input audio signal can be a distributed spatial audio exchange format, which can be obtained from various collected audio content formats, thereby performing spatial audio processing on such input audio signal to obtain an audio signal with the specific spatial format. In particular, in some embodiments, the spatial audio processing may include appropriate processing of the input audio, especially including parsing, format conversion, information processing, encoding, etc., to obtain the audio signal of the specific spatial format. In other embodiments, the audio signal of the specific spatial format can be directly obtained from the input audio signal without performing at least some of the spatial audio processing. In some embodiments, the input audio signal may be in other appropriate formats other than the non-spatial audio exchange format. In particular, the input audio signal may contain or be directly a signal in a specific audio content format, such as a specific audio representation signal, or contain or be directly an audio signal in a specific spatial format. In this case, the input audio signal may not need to perform at least some of the spatial audio processing, so that the aforementioned spatial audio processing may not need to be performed, such as not performing parsing, format conversion, information processing, encoding, etc.; or only part of the spatial audio processing may be performed, such as only performing encoding without performing parsing, format conversion, etc., so that an audio signal in a specific spatial format can be obtained.
[0112] According to an embodiment of the present disclosure, the acquisition module 41 may include an audio signal encoding module 413, which is configured to spatially encode the audio signal of the specific audio content format based on metadata related information associated with the audio signal of the specific audio content format to obtain an encoded audio signal. The encoded audio signal may be included in an audio signal of a specific spatial format. According to an embodiment of the present disclosure, the audio signal of the specific audio content format may, for example, include a spatial audio signal of a specific spatial audio representation method, and in particular, the spatial audio signal is at least one of a scene-based audio representation signal, a channel-based audio representation signal, and an object-based audio representation signal. In some embodiments, the audio signal encoding module 413 specifically encodes a specific type of audio signal in the audio signal of the specific audio content format, and the specific type of audio signal is an audio signal that needs or is required to be spatially encoded in an audio rendering system, and may, for example, include at least one of a scene-based audio representation signal, an object-based audio representation signal, and a specific channel signal (for example, a non-narrative channel / track) in the channel-based audio representation signal.
[0113] Optionally, the acquisition module 41 may include an audio signal acquisition module 411, which is configured to acquire an audio signal in a specific audio content format and metadata information associated with the audio signal. In some embodiments, the audio signal acquisition module may obtain an audio signal in a specific audio content format and metadata information associated with the audio signal by parsing the input signal, or receive a directly input audio signal in the specific audio content format and metadata information associated with the audio signal.
[0114] Optionally, the acquisition module 41 may also include an audio information processing module 412, which is configured to extract audio parameters of an audio signal in a specific audio content format based on metadata associated with the audio signal in a specific audio content format, so that the audio signal encoding module may be further configured to perform spatial encoding on the audio signal in the specific audio content format based on metadata associated with the audio signal and at least one of the audio parameters. As an example, the audio information processing module may be referred to as a scene information processor, which may provide the audio parameters extracted based on the metadata to the audio signal encoding module for encoding. The audio information processing module is not required for the audio rendering of the present disclosure, for example, its information processing function may not be performed, or it may be outside the audio rendering system, or the audio information processing module may be included in other modules, such as an audio signal acquisition module or an audio signal encoding module, or its function is implemented by other modules, and is therefore indicated by dotted lines in the accompanying drawings.
[0115] In some embodiments, additionally or optionally, the audio rendering system may include a signal adjustment module 43, which is configured to perform signal processing on the decoded audio signal. The signal processing performed by the signal adjustment module can be referred to as a signal post-processing, especially the post-processing performed on the decoded audio signal before being played back by the playback device. Therefore, the signal adjustment module may also be referred to as a signal post-processing module. In particular, the signal adjustment module 43 may be configured to adjust the decoded audio signal based on the characteristics of the playback device in the user application scenario, so that the adjusted audio signal can present a more appropriate acoustic experience when rendered by the audio rendering device. It should be noted that the audio signal adjustment module is not necessary for the audio rendering of the present disclosure, for example, the signal adjustment function may not be performed, or it may be outside the audio rendering system, or the audio signal adjustment module may be included in other modules, such as in an audio signal decoding module or its function is implemented by a decoding module, so it is indicated by dotted lines in the accompanying drawings.
[0116] Additionally, the audio rendering system 4 may also include or be connected to an audio input port for receiving an input audio signal, which may be distributed and transmitted to the audio rendering system in the audio system, as described above, or directly input by the user at the user end or consumer end, which will be described later. Additionally, the audio rendering system 4 may also include or be connected to an output device, such as an audio rendering device, an audio playback device, which may present the spatially decoded audio signal to the user. According to some embodiments of the present disclosure, the audio rendering device or audio playback device according to an embodiment of the present disclosure may be any appropriate audio device, such as a speaker, a speaker array, headphones, and any other appropriate device capable of presenting an audio signal to the user.
[0117] Figure 4B A schematic conceptual diagram of audio rendering processing according to an embodiment of the present disclosure is shown, showing a process of obtaining an output audio signal suitable for rendering in a user application scenario, especially for presenting / playing back to a user through a device in a playback environment, based on an input audio signal.
[0118] First, an audio signal in a specific spatial format that can be used for playback in a user application scenario is obtained. In particular, appropriate processing is performed depending on the format of the input audio signal to obtain the audio signal in the specific spatial format.
[0119] On the one hand, in the case where the input audio signal includes an audio signal having a spatial audio exchange format distributed to the audio rendering system, the input audio signal can be subjected to spatial audio processing to obtain an audio signal in a specific spatial format. In particular, the spatial audio exchange format can be any known appropriate format of an audio signal in signal transmission, such as the audio distribution format in the audio signal distribution described above, which will not be described in detail here. In some embodiments, spatial audio processing may include at least one of parsing, format conversion, information processing, encoding, etc. of the input audio signal. In particular, audio signals of various audio content formats can be derived from the input audio signal by audio parsing, and then the parsed signal is encoded to obtain an audio signal in a spatial format suitable for rendering in a user application scenario, i.e., a playback environment, for playback. In addition, format conversion and signal information processing can optionally be performed before encoding. Thus, an audio signal having a specific spatial audio representation can be derived from the input audio signal, and an audio signal of the specific spatial format can be obtained based on the audio signal having the specific spatial audio representation.
[0120] As an example, an audio signal having a specific audio representation can be obtained from an input audio signal, such as at least one of a scene-based audio representation signal, an object-based audio representation signal, and a channel-based audio representation signal. For example, when the input audio signal is an audio signal having a spatial audio exchange format, the input audio signal is parsed to obtain a spatial audio signal having a specific spatial audio representation, such as a scene-based audio representation signal, a channel-based audio representation signal, and at least one of an object-based audio representation signal, and metadata information corresponding to the signal, and optionally, the spatial audio signal can be further converted into a predetermined format, such as a format pre-defined and predetermined by an audio rendering system or even an audio system. Of course, such a format conversion is not necessary.
[0121] Further, for the obtained audio signal of the specific audio representation, audio processing is performed based on the audio representation of the audio signal. Specifically, spatial audio encoding is performed on at least one of the narrative channels in the scene-based audio representation signal, the object-based audio representation signal, and the channel-based audio representation signal to obtain an audio signal with a specific spatial format. That is, although the format / representation of the input audio signal may be different, the input audio signal can still be converted into a common audio signal with a specific spatial format for decoding and rendering. The spatial audio encoding process can be performed based on metadata-related information associated with the audio signal, where the metadata-related information may include metadata of the audio signal directly obtained, such as derived from the input audio signal during the parsing process, and / or optionally, it may also include audio parameters corresponding to the spatial audio signal obtained by performing information processing on the metadata information of each signal obtained, and the spatial audio encoding process can be performed based on the audio parameters.
[0122] On the other hand, the input audio signal may be in other appropriate formats other than the spatial audio exchange format, in particular, for example, a specific spatial representation signal or even a specific spatial format signal. In this case, at least some of the aforementioned spatial audio processing may be skipped to obtain an audio signal in a specific spatial format. In some embodiments, when the input audio signal is not a distributed audio signal in a spatial audio exchange format, but a directly input audio signal with a specific spatial audio representation, the aforementioned audio parsing process may not be required, and format conversion and encoding may be performed directly. Even, when the input audio signal has a predetermined format, the aforementioned format conversion may not be required, and encoding may be performed directly. In other embodiments, the input audio signal is directly an audio signal in the specific spatial format, and such an input audio signal may be directly transmitted / transparently transmitted to the audio signal spatial decoder without spatial audio processing, such as parsing, format conversion, information processing, encoding, etc. For example, when the input audio signal is a scene-based spatial audio representation signal, such an input audio signal may be directly transmitted to the spatial decoder as a specific spatial format signal without the aforementioned spatial audio processing. According to some embodiments, when the input audio signal is not a distributed audio signal having a spatial audio exchange format, for example, it may be an audio signal of the aforementioned specific spatial audio representation or an audio signal of a specific spatial format, it may be directly input at the user end / consumer end, for example, it may be directly obtained from an application programming interface (API) directly provided in the rendering system.
[0123] For example, in the case of a signal with a specific representation directly input from the user end / consumer end, such as one of the three audio representations mentioned above, the aforementioned parsing process may not be required, and it may be directly converted into a format specified by the system. For another example, when the input audio signal is already in a format specified by the system and a representation that the system can process, it may be directly transmitted to the spatial encoding processing module without the aforementioned parsing and code conversion. For another example, if the input audio signal is a non-narrative channel signal, a binaural signal after reverberation processing, etc., the input audio signal may be directly transmitted to the spatial decoding module for decoding without performing the aforementioned spatial audio encoding process. In this case, a judgment unit / module may exist in the system to determine whether the input audio signal meets the above conditions.
[0124] Then, spatial decoding can be performed on the obtained audio signal with a specific spatial format. In particular, the obtained audio signal with a specific spatial format can be referred to as an audio signal to be decoded, and the audio signal spatial decoding is intended to convert the audio signal to be decoded into a format suitable for playback by a user application scenario, such as an audio playback environment, a playback device in an audio rendering environment, and a rendering device. According to an embodiment of the present disclosure, decoding can be performed according to an audio signal playback mode, and the playback mode can be indicated in various appropriate ways, such as by an identifier, and can be informed of a decoding module in various appropriate ways, such as informing the decoding module together with the input audio signal, or can be input by other input devices and informed of the decoding module. As an example, the renderer ID mentioned above can be used as an identifier to inform whether the playback mode is binaural playback or speaker playback, etc. In some embodiments, audio signal decoding can use a decoding method corresponding to the playback device in the user application scenario, especially a decoding matrix, to decode the audio signal of the specific spatial format, and transform the audio signal to be decoded into an audio of a suitable format. In other embodiments, audio signal decoding can also be performed in other appropriate ways, such as virtual signal decoding.
[0125] Optionally, after the audio signal is decoded, the decoded output may be post-processed, in particular signal adjustment may be performed to adjust the spatially decoded audio signal for a specific playback device in the user application scenario, especially to adjust the audio signal characteristics, so that the adjusted audio signal can present a more appropriate acoustic experience when rendered by an audio rendering device.
[0126] Thus, the decoded audio signal or the adjusted audio signal can be presented to the user in a user application scenario, for example, through an audio rendering device / audio playback device in an audio playback environment, to meet the needs of the user.
[0127] It should be noted that the processing of audio data and / or metadata in the above rendering process can be performed in various appropriate formats. According to some embodiments, audio signal processing can be performed in blocks, and the block size can be set. For example, the block size can be pre-set and not changed during the processing. For example, the block size can be set when the audio rendering system is initialized. In some embodiments, metadata can be parsed in blocks and then the scene information can be adjusted for the metadata. This operation can be included in the operation of the scene information processing module according to an embodiment of the present disclosure.
[0128] Various processing / module operations in the audio rendering process / system according to an embodiment of the present disclosure will be described in further detail below with reference to the accompanying drawings.
[0129] Input signal acquisition
[0130] The signal suitable for rendering processing by the audio rendering system can be obtained by various appropriate methods. According to an embodiment of the present disclosure, the signal suitable for rendering processing by the audio rendering system can be an audio signal in a specific audio content format. In some embodiments, the audio signal in a specific audio content format can be directly input into the audio rendering system, that is, the audio signal in a specific audio content format can be directly input as an input signal, so that it can be directly obtained. In other embodiments, the audio signal in a specific audio content format can be obtained from the audio signal input to the audio rendering system. As an example, the input audio signal may be an audio signal in other formats, such as a specific combination signal containing an audio signal in a specific audio content format, a signal in other formats, in which case the audio signal in a specific audio content format can be obtained by parsing the input audio signal. In this case, the input signal acquisition module can be referred to as an audio signal parsing module, and the signal processing performed by it can be referred to as a signal pre-processing, especially the processing before the audio signal is encoded.
[0131] Audio signal analysis
[0132] Figure 4C and 4D An exemplary process of an audio signal parsing module according to an embodiment of the present disclosure is shown.
[0133] According to some embodiments of the present disclosure, considering different application scenarios, the audio signal may be input in different input formats. Therefore, audio signal parsing can be performed before the audio rendering process to be compatible with inputs of different formats. Such audio signal parsing processing can be considered as a kind of pre-processing / preprocessing. In some embodiments, the audio signal parsing module can be configured to obtain an audio signal having an audio content format compatible with an audio rendering system and metadata information associated with the audio signal from the input audio signal. In particular, any input spatial audio exchange format signal can be parsed to obtain an audio signal having an audio content format compatible with an audio rendering system, which may include at least one of an object-based audio representation signal, a scene-based audio representation signal, and a channel-based audio representation signal, as well as associated metadata information. Figure 4C The parsing process for an arbitrary Spatial Audio Interchange Format signal input is shown.
[0134] Further, in some embodiments, the audio signal parsing module can further convert the acquired audio signal having an audio content format compatible with the audio rendering system so that the audio signal has a predetermined format, in particular, a predetermined format of the audio rendering system, for example, converting the signal into a format agreed upon by the audio rendering system according to the signal format type. In particular, the predetermined format can correspond to a predetermined configuration parameter of an audio signal in a specific audio content format, so that in the audio signal parsing operation, the audio signal in the specific audio content format can be further converted into the predetermined configuration parameter. In some embodiments, in the case where the audio signal having an audio content format compatible with the audio rendering system is a scene-based audio representation signal, the signal parsing module is configured to convert a scene-based audio signal having different channel orderings and normalization coefficients into channel orderings and normalization coefficients agreed upon by the audio rendering system.
[0135] As an example, for any spatial audio exchange format signal for distribution, whether it is a non-streaming or streaming signal, such signals can be divided into three types of signals according to the signal representation method of spatial audio through an input signal parser, namely, at least one of a scene-based audio representation signal, a channel-based audio representation signal, and an object-based audio representation signal, and metadata corresponding to such signals. On the other hand, the signal can also be converted into a system-constrained format according to the format type in pre-processing. For example, for the scene-based spatial audio representation signal HOA, different channel orderings (such as ACN, Ambisonic Channel Number, FuMa, Furse-Malham and SID, Single index designation) and different normalization coefficients (N3D, SN3D, FuMa) are used in different data exchange formats. In this step, they can be converted into a certain agreed channel ordering and normalization coefficient, such as (ACN+SN3D).
[0136] In some embodiments, when the input audio signal is not a distributed spatial audio exchange format signal, it may not be necessary to perform at least some of the spatial audio processing on the input audio signal. As an example, the input specific audio signal may be directly at least one of the three signal representations described above, thereby eliminating the signal parsing process described above, and the audio signal and its associated metadata may be directly passed to the audio signal encoding module. Figure 4D The processing of specific audio signal input according to other embodiments of the present disclosure is shown. In other embodiments, the input audio signal may even be an audio signal in the specific spatial format described above, and such an input audio signal may be directly transmitted / transparently transmitted to the audio signal decoding module without performing the aforementioned spatial audio processing including parsing, format conversion, audio encoding, etc.
[0137] In some embodiments, for such an input audio signal, the audio rendering system may further include a specific audio input device, which is used to directly receive the input audio signal and directly transmit / transparently transmit it to the audio signal encoding module or the audio signal decoding module. It should be noted that such a specific input device may be, for example, an application program interface (API), and the format of the input audio signal that it can receive has been pre-set, for example, corresponding to the specific spatial format described above, for example, it may be at least one of the three signal representation methods mentioned above, and so on, so that when the input device receives the input audio signal, the input audio signal can be directly transmitted / transparently transmitted without performing at least some of the spatial audio processing. It should be noted that such a specific input device may also be part of the audio signal acquisition operation / module, and even be included in the audio signal parsing module.
[0138] It should be noted that the implementation of the aforementioned audio signal parsing module and the specific audio input device is merely exemplary and not restrictive. According to some embodiments of the present disclosure, the audio signal parsing module may be implemented in various appropriate ways. In some embodiments, the audio signal parsing module may include a parsing submodule and a direct transmission submodule, the parsing submodule may only receive an audio signal in a spatial exchange format for audio parsing, and the direct transmission submodule may receive an audio signal in a specific audio content format or a specific audio representation signal for direct transmission. In this way, the audio rendering system may be configured so that the audio signal parsing module receives two inputs, namely an audio signal in a spatial exchange format and an audio signal in a specific audio content format or a specific audio representation signal. In other embodiments, the audio signal parsing module may include a judgment submodule, a parsing submodule, and a direct transmission submodule, so that the audio signal parsing module can receive any type of input signal and perform appropriate processing. Among them, the judgment submodule can determine what format / type the input audio signal is, and in the case of determining that the input audio signal is an audio signal in a spatial audio exchange format, turn to the parsing submodule to perform the above-mentioned parsing operation, otherwise the audio signal can be directly transmitted / transparently transmitted to the format conversion, audio encoding, audio decoding and other stages by the direct transmission submodule, as described above. Of course, the judgment submodule may also be outside the audio signal analysis module. Audio signal judgment may be implemented in various known appropriate ways, which will not be described in detail here.
[0139] Audio Information Processing
[0140] In some embodiments, the audio rendering system may include an audio information processing module configured to obtain audio parameters of an audio signal of a specific audio content format based on metadata associated with an audio signal of a specific audio content format, and in particular, to obtain audio parameters based on metadata associated with the specific type of audio signal, as metadata information that can be used for encoding. According to an embodiment of the present disclosure, the audio information processing module may be referred to as a scene information processing module / processor, and the audio parameters obtained by the audio information processing module may be input to an audio signal encoding module, whereby the audio signal encoding module may be further configured to perform spatial encoding on the specific type of audio signal based on the audio parameters. Here, the specific type of audio signal may include the aforementioned audio signal of an audio content format compatible with the audio rendering system obtained from the input audio signal, such as at least one of the aforementioned scene-based audio representation signal, object-based audio representation signal, and channel-based audio representation signal, and in particular, at least one of the specific type of channel signals in the object-based audio representation signal, scene-based audio representation signal, and channel-based audio representation signal. As an example, the specific type of channel signal may be referred to as a first specific type of channel signal, which may include a non-narrative channel / track in a channel-based audio representation signal. In another example, the specific type of channel signal may also include a narrative channel / audio track that does not need to be spatially encoded according to an application scenario.
[0141] In some embodiments, the audio information processing module is further configured to obtain audio parameters of the specific type of audio signal based on the audio content format of the specific type of audio signal, and in particular to obtain audio parameters based on the audio content format of an audio signal having an audio content format compatible with an audio rendering system obtained from an input audio signal. For example, the audio parameters may be parameters of specific types respectively corresponding to the audio content formats, as described above.
[0142] According to some embodiments of the present disclosure, the audio signal is an object-based audio representation signal, and the audio information processing module is configured to obtain spatial attribute information of the object-based audio representation signal as an audio parameter that can be used for spatial audio encoding processing. In some embodiments, the spatial attribute information of the audio signal includes the orientation information of each audio element in a coordinate system, or the relative orientation information of a sound source related to the audio signal relative to a listener. In some embodiments, the spatial attribute information of the audio signal further includes the distance information of each sound element of the audio signal in a coordinate system. As an example, in metadata processing based on object-based audio representation, the orientation information of each sound element in a coordinate system, such as azimuth and elevation, can be obtained, and optionally, distance information can also be obtained, or the relative orientation information of each sound source relative to the listener's head can be obtained.
[0143] According to some embodiments of the present disclosure, the audio signal is a scene-based audio representation signal, and the audio information processing module is configured to obtain rotation information related to the audio signal based on metadata information associated with the audio signal for spatial audio coding processing. In some embodiments, the rotation information related to the audio signal includes at least one of the rotation information of the audio signal and the rotation information of the listener of the audio signal. As an example, in the metadata processing of the scene-based audio representation, the rotation information of the scene audio and the rotation information of the listener are read from the metadata.
[0144] According to some embodiments of the present disclosure, the audio signal is a channel-based audio signal, and the audio information processing module is configured to obtain audio parameters based on the channel track type of the audio signal. In particular, the audio encoding process will mainly target specific types of channel-based audio signals that require spatial encoding, especially narrative channel tracks of channel-based audio signals, and the audio information processing module may be configured to split the audio representation of the channel into audio elements by channel to convert into metadata as audio parameters. It should be noted that narrative channel tracks of channel-based audio signals may not perform spatial audio encoding. For example, spatial audio encoding may not be performed depending on the specific application scenario. Such tracks may be directly transmitted to the decoding stage, or may be further processed depending on the playback method.
[0145] As an example, in metadata processing based on channel-based audio representation, for narrative channel audio tracks, the audio representation of the channel can be split into audio elements according to the standard definition of the channel, and converted into metadata for processing. Depending on the needs of the application scenario, spatial audio processing may not be performed, and mixing may be performed for different playback modes in subsequent links. For non-narrative channel audio tracks, since dynamic spatialization processing is not required, mixing may be performed for different playback modes in subsequent links. In other words, non-narrative channel audio tracks will not be processed by the audio information processing module, that is, spatial audio processing will not be performed, but can bypass the audio information processing module and be directly transmitted / transparently transmitted.
[0146] Audio signal encoding
[0147] The following will refer to Figure 4E and 4F The audio signal encoding module according to an embodiment of the present disclosure is described. Figure 4EA block diagram of some embodiments of an audio signal encoding module is shown, wherein the audio signal encoding module may be configured to spatially encode an audio signal in a specific audio content format based on metadata related information associated with the audio signal in the specific audio content format to obtain an encoded audio signal. Additionally, the audio signal encoding module may also be configured to obtain an audio signal in a specific audio content format and associated metadata related information. In one example, the audio signal encoding module may receive the audio signal and metadata related information, such as the audio signal and metadata related information generated by the aforementioned audio signal parsing module and the audio signal processing module, such as by means of an input port / input device. In another example, the audio signal encoding module may implement the operation of the aforementioned audio signal acquisition module and / or audio signal processing module, for example, the aforementioned audio signal acquisition module and / or audio signal processing module may be included to acquire the audio signal and metadata. Here, the audio signal encoding module may also be referred to as an audio signal spatial encoding module / encoder. Figure 4F A flowchart showing some embodiments of audio signal encoding operations is provided, wherein an audio signal in a specific audio content format and metadata-related information associated with the audio signal are obtained; and for the audio signal in the specific audio content format, the audio signal in the specific audio content format is spatially encoded based on the metadata-related information associated with the audio signal in the specific audio content format to obtain an encoded audio signal.
[0148] According to an embodiment of the present disclosure, the acquired audio signal of a specific audio content format may be referred to as an audio signal to be encoded. As an example, the acquired audio signal may be a non-direct / transparent audio signal, and may have various audio content formats or audio representations, such as at least one of the three types of audio signals represented as described above, or other suitable audio signals. As an example, such an audio signal may be, for example, an object-based audio representation signal, a scene-based audio representation signal, or a narrative channel track in a channel-based audio representation signal that may have been pre-specified to be encoded for a specific application scenario, such as the aforementioned. In particular, the acquired audio signal may be directly input, such as a signal without the need for signal parsing as described above, or may be an audio signal extracted / parsed from an input audio signal, such as an audio signal obtained by the signal parsing module described above without the need for audio encoding, such as a specific type of channel signal in a channel-based audio representation signal, which may be referred to herein as a second specific type of channel signal, such as a narrative channel audio track that is not specified to require encoding as described above, or a non-narrative channel audio track that does not require encoding itself, which will not be input into the audio signal encoding module, for example, but will be directly transmitted to a subsequent decoding module.
[0149] According to an embodiment of the present disclosure, the specific spatial format may be a spatial format that can be supported by the audio rendering system, for example, it can be played back to the user in different user application scenarios, such as different audio playback environments. In a sense, the encoded audio signal of the specific spatial format can be used as an intermediate signal medium, that is, it indicates that an intermediate signal of a common format is obtained by encoding an input audio signal that may contain various spatial representations, and decoding and processing are performed from the intermediate signal for rendering. The encoded audio signal of the specific spatial format may be an audio signal of a specific spatial format as described above, such as FOA, HOA, MOA, etc., which will not be described in detail here. Thus, for an audio signal that may have at least one of a plurality of different spatial representations, it can be spatially encoded to obtain an encoded audio signal of a specific spatial format that can be used for playback in a user application scenario, that is, even if the audio signal may contain different content formats / audio representations, an audio signal of a common or common spatial format can still be obtained by encoding. In some embodiments, the encoded audio signal may be added to the intermediate signal, for example, encoded into an intermediate signal. In another embodiment, the encoded audio signal may also be directly transmitted / transparently transmitted to the spatial decoder without being added to the intermediate signal. In this way, the audio signal encoding module can be compatible with various types of input signals to obtain an encoded audio signal in a common spatial format, thereby enabling the audio rendering process to be performed efficiently.
[0150] According to an embodiment of the present disclosure, the audio signal encoding module can be implemented in various appropriate ways, for example, it can include an acquisition unit and an encoding unit that respectively implement the above-mentioned acquisition and encoding operations. Such a spatial encoder, acquisition unit, and encoding unit can be various appropriate implementation forms, such as software, hardware, firmware, etc. or any combination. In some embodiments, the audio signal encoding module can be implemented to only receive the audio signal to be encoded, such as the audio signal to be encoded that is directly input or obtained from the audio signal parsing module. In other words, the signal input to the audio signal encoding module must be encoded. As an example, in this case, the acquisition unit can be implemented as a signal input interface, which can directly receive the audio signal to be encoded. In other embodiments, the audio signal encoding module can be implemented to receive audio signals or audio representation signals in various audio content formats. Thus, in addition to the acquisition unit and the encoding unit, the audio signal encoding module may further include a determination unit, which may determine whether the audio signal received by the audio signal encoding module is an audio signal that needs to be encoded, and transmit the audio signal to the acquisition unit and the encoding unit when it is determined to be an audio signal that needs to be encoded; and when it is determined to be an audio signal that does not need to be encoded, the audio signal is directly transmitted to the decoding module without audio encoding. In some embodiments, the determination may be performed in various appropriate ways, for example, the audio content format of the audio or the audio signal representation method may be referred to for comparison, and when the format or representation method of the input audio signal matches the format or representation method of the audio signal that needs to be encoded, it is determined that the input audio signal needs to be encoded. For example, the determination unit may also receive other reference information, such as application scenario information, rules pre-defined for specific application scenarios, etc., and may make a determination based on the reference information. As described above, when the rules pre-defined for specific application scenarios are known, the audio signal that needs to be encoded in the audio signal may be selected according to the rules. For another example, the determination unit may also obtain an identifier related to the signal type, and determine whether the signal needs to be encoded according to the identifier related to the signal type. The identifier may be in various appropriate forms, such as a signal type identifier, and any other appropriate indication information capable of indicating the signal type.
[0151] According to some embodiments of the present disclosure, metadata-related information associated with an audio signal may include metadata in an appropriate form and may depend on the signal type of the audio signal. In particular, the metadata information may correspond to the signal representation of the signal. For example, for an object-based signal representation, the metadata information may be related to the properties of the audio object, especially the spatial properties; for a scene-based signal representation, the metadata information may be related to the properties of the scene; for a channel-based signal representation, the metadata information may be related to the properties of the channel. In some embodiments of the present disclosure, it can be referred to as encoding the audio signal according to the type of the audio signal. In particular, the audio signal can be encoded based on the metadata-related information corresponding to the type of the audio signal.
[0152] According to an embodiment of the present disclosure, metadata-related information associated with an audio signal may include metadata associated with the audio signal and at least one of an audio parameter of the audio signal obtained based on the metadata. In some embodiments, metadata-related information may include metadata associated with the audio signal, such as metadata obtained together with the audio signal, such as directly input or obtained by signal parsing. In other embodiments, metadata-related information may also include audio parameters of the audio signal obtained based on the metadata, as described above for the operation of the information processing module.
[0153] According to an embodiment of the present disclosure, metadata-related information can be obtained in various appropriate ways. In particular, metadata information can be obtained through signal parsing processing, or directly input, or obtained through specific processing. In some embodiments, metadata-related information can be obtained by the signal parsing process as described above when parsing the input signal having the spatial audio exchange format distributed, and metadata associated with the specific audio representation signal is obtained. In some embodiments, metadata-related information can be directly input when the audio signal is input, for example, when the input audio signal can be directly input through the API without the need for the aforementioned audio signal parsing, metadata-related information can be input together with the audio signal when the audio signal is input, or input separately from the audio signal. In other embodiments, the metadata of the parsed audio signal or the metadata directly input can be further processed, such as information processing, thereby obtaining appropriate audio parameters / information as metadata information for audio encoding. According to an embodiment of the present disclosure, the information processing can be referred to as scene information processing, and in the information processing, processing can be performed based on the metadata associated with the audio signal to obtain appropriate audio parameters / information. In some embodiments, for example, signals of different formats can be extracted based on metadata and corresponding audio parameters can be calculated, and as an example, the audio parameters can be related to the rendering application scene. In other embodiments, for example, scene information may be adjusted based on metadata.
[0154] According to an embodiment of the present disclosure, for an audio signal to be encoded, encoding will be performed based on metadata related information associated with the audio signal. In particular, the audio signal to be encoded may include a specific type of audio signal in the aforementioned audio signal of a specific audio content format, and for such an audio signal, spatial encoding will be performed on the specific type of audio signal based on metadata related information associated with the specific type of audio signal to obtain an encoded audio signal of a specific spatial format. Such encoding may be referred to as spatial encoding.
[0155] According to some embodiments, the audio signal encoding module may be configured to weight the audio signal based on metadata information. In particular, the audio signal encoding module may be configured to weight according to the weight in the metadata. The metadata may be associated with the audio signal to be encoded acquired by the audio signal encoding module, for example, associated with various audio content format signals / audio representation signals, as described above. In particular, in some embodiments, the audio signal encoding module may also be configured to weight the audio signal acquired, especially the audio signal with a specific audio content format, based on the metadata associated with the audio signal. In other embodiments, the audio signal encoding module may also be configured to further perform additional processing on the encoded audio signal, such as weighting, rotation, etc. In particular, the audio signal encoding module may be configured to convert the audio signal of a specific audio content format into an audio signal with a specific spatial format, and then weight the obtained audio signal with a specific spatial format based on metadata, thereby obtaining as an intermediate signal. In some embodiments, the audio signal encoding module may be configured to further process the audio signal with a specific spatial format obtained by conversion based on metadata, such as format conversion, rotation, etc. In some embodiments, the audio signal encoding module can be configured to convert the audio signal of a specific spatial format obtained by encoding or directly input to meet the format supported and constrained by the current system. For example, the channel arrangement method, regularization method, etc. can be converted to meet the requirements of the system.
[0156] According to some embodiments of the present disclosure, the audio signal of the specific audio content format is an object-based audio representation signal, and the audio signal encoding module is configured to spatially encode the object-based audio representation signal based on the spatial attribute information of the object's audio representation signal. In particular, the encoding can be performed by matrix multiplication. In some embodiments, the spatial attribute information of the object-based audio representation signal may include information related to the spatial propagation of sound objects based on the audio signal, and in particular, information related to the spatial propagation path from the sound object to the listener. In some embodiments, the information related to the spatial propagation path from the sound object to the listener includes at least one of the propagation duration, propagation distance, orientation information, path intensity energy, and nodes along the way of each spatial propagation path from the sound object to the listener.
[0157] In some embodiments, the audio signal encoding module is configured to spatially encode the object-based audio signal according to at least one of a filter function and a spherical harmonic function, wherein the filter function may be a filter function for filtering the audio signal based on the path energy intensity of the spatial propagation path from the sound object in the audio signal to the listener, and the spherical harmonic function may be a spherical harmonic function based on the azimuth information of the spatial propagation path. In some embodiments, the audio signal encoding may be performed based on a combination of the filter function and the spherical harmonic function. As an example, the audio signal encoding may be performed based on the product of the filter function and the spherical harmonic function.
[0158] In some embodiments, the spatial audio encoding of the object-based audio signal can be further based on the delay of the sound object in the spatial propagation, for example, based on the propagation duration of the spatial propagation path. In this case, the filter function for filtering the audio signal based on the path energy intensity is a filter function for filtering the audio signal of the sound object before it propagates along the spatial propagation path, based on the path intensity energy of the path. In some embodiments, the audio signal of the sound object before it propagates along the spatial propagation path refers to the audio signal at a moment before the time required for the sound object to reach the listener along the spatial propagation path, for example, the audio signal of the sound object before the propagation duration.
[0159] In some embodiments, the azimuth information of the spatial propagation path may include the azimuth angle of the spatial propagation path to the listener or the azimuth angle of the spatial propagation path relative to the coordinate system. In some embodiments, the spherical harmonic function based on the azimuth angle of the spatial propagation path may be any suitable form of spherical harmonic function.
[0160] In some embodiments, the spatial audio coding of the object-based audio signal may further be based on the length of the spatial propagation path from the sound object in the audio signal to the listener, and at least one of the near-field compensation function and the diffusion function may be used to encode the audio signal. For example, depending on the length of the spatial propagation path, at least one of the near-field compensation function and the diffusion function may be applied to the audio signal of the sound object for the propagation path to perform appropriate audio signal compensation and enhance the effect.
[0161] In some embodiments, spatial encoding of object-based audio signals (such as the spatial encoding of object-based audio signals described above) can be performed for one or more spatial propagation paths from the sound object to the listener. In particular, when there is one spatial propagation path from the sound object to the listener, spatial encoding of the object-based audio signal is performed for the spatial propagation path, and when there are multiple spatial propagation paths from the sound object to the listener, it can be performed for at least one of the multiple spatial propagation paths, or even all of the spatial propagation paths. Specifically, the relevant information of each spatial propagation path from the sound object to the listener can be considered separately, and the audio signal corresponding to the spatial propagation path can be encoded accordingly, and then the encoding results of each spatial propagation path can be combined to obtain the encoding result for the sound object. The spatial propagation path from the sound object to the listener can be determined by various appropriate methods, and in particular, it can be determined by the information processing module described above by obtaining spatial attribute information.
[0162] In some embodiments, spatial encoding of the object-based audio signal may be performed for each of the one or more sound objects contained in the audio signal, and the encoding process for each sound object may be performed as described above. In some embodiments, the audio signal encoding module is further configured to weightedly combine the encoded signals of each object-based audio representation signal based on the weight of the sound object defined in the metadata. In particular, in the case where the audio signal contains multiple sound objects, for each sound object in the audio signal, the object-based audio representation signal may be spatially encoded based on the spatial propagation related information of the sound object of the audio signal, for example, after the audio representation signal is spatially encoded for the spatial propagation path of each sound object as described above, the weight of each sound object contained in the metadata associated with the audio representation signal is used to weightedly combine the encoded audio signals of each sound object.
[0163] As an example, in the spatial encoding process of object-based audio representation, for each audio object, the audio signal will be written into a delay device in consideration of the delay of sound propagation in space. It can be known from the metadata information associated with the audio representation signal, especially the audio parameters obtained by the audio information processing module, that each sound object will have one or more propagation paths to reach the listener. According to the length of each path, the time t1 required for the sound object to reach the listener is calculated. Therefore, the audio signal s of the sound object before time t1 can be obtained from the delay device of the audio object, and the audio signal is filtered by the filter function E based on the path energy intensity. Furthermore, the azimuth information of the path, such as the direction angle θ of the path to the listener, can be obtained from the metadata information associated with the audio representation signal, especially the audio parameters obtained by the audio information processing module, and a specific function based on the azimuth angle, such as the spherical harmonics Y of the corresponding channel, can be used to encode the audio signal into an encoded signal based on the two, such as the HOA signal S. Let N be the number of channels of the HOA signal, then the HOA signal S obtained by the audio encoding process is N It can be expressed as follows:
[0164] s N =E(s(t-t1))Y N (θ)
[0165] Alternatively or optionally, for the orientation information of the path, the direction of the path relative to the coordinate system may be used instead of the direction to the listener, so that the target sound field signal may be obtained as the encoded audio signal by multiplying with the rotation matrix in a subsequent step. For example, in the case where the path orientation information is the direction of the path relative to the coordinate system, the rotation matrix may be further multiplied on the basis of the above formula to obtain the encoded HOA signal.
[0166] In some embodiments of the present disclosure, the encoding operation can be performed in the time domain or the frequency domain. Furthermore, encoding can also be performed based on the distance of the spatial propagation path from the sound object to the listener. In particular, at least one of the near-field compensation function and the source spread function can be further applied according to the distance of the path to enhance the effect. For example, the near-field compensation function and / or the source spread function can be further applied on the basis of the aforementioned encoded HOA signal. In particular, it can be considered that if the distance of the path is less than a threshold, the near-field compensation function is applied, and if it is greater than the threshold, the source spread function is applied, and vice versa, to further optimize the aforementioned encoded HOA signal.
[0167] Finally, the HOA signals obtained after the signal conversion of each sound object are weighted and superimposed according to the weight of the sound object defined in the metadata, so that the weighted sum signal of all object-based audio signals can be obtained as the encoded signal, which can be used as the intermediate signal.
[0168] In some embodiments, the audio signal of the object-based audio signal is spatially encoded and the audio signal can also be encoded based on the reverberation information, so that the encoded signal obtained can be directly transmitted to the spatial decoder for decoding, or can be added to the intermediate signal output by the encoder. In some embodiments, the audio signal encoding module is further configured to obtain reverberation parameter information, and perform reverberation processing on the audio signal to obtain a reverberation-related signal of the audio signal. In particular, the spatial reverberation response of the scene can be obtained, and the audio signal is convolved based on the spatial reverberation response to obtain the reverberation-related signal of the audio signal. The reverberation parameter information can be obtained in various appropriate ways, such as obtained from metadata information, obtained from the aforementioned information processing module, obtained by a user or other input device, and so on.
[0169] As an example, for a more advanced information processor, the spatial house reverberation response of the user application scenario may be generated, including but not limited to RIR (Room Impulse Response), ARIR (Ambisonics Room Impulse Response), BRIR (Binaural Room Impulse Response), MO-BRIR (Multi orientation Binaural Room Impulse Response). When such information is obtained, a convolution device may be added to the encoding module to process the audio signal. Depending on the type of reverberation, the processing result may be an intermediate signal (ARIR), an omnidirectional signal (RIR) or a binaural signal (BRIR, MO-BRIR), and the processing result may be added to the intermediate signal or transmitted to the next step for corresponding playback decoding. Optionally, the information processor may also provide reverberation parameter information such as reverberation duration, and an artificial reverberation generator (for example, a feedback delay network) may be added to the encoding module to process the artificial reverberation, and the result may be output to the intermediate signal or transmitted to the decoder for processing.
[0170] In some embodiments, the audio signal of the specific audio content format is a scene-based audio representation signal, and the audio signal encoding module is further configured to weight the scene-based audio representation signal based on weight information indicated or included in metadata associated with the audio representation signal. In this way, the weighted signal can be used as an encoded audio signal for spatial decoding. In some embodiments, the audio signal of the specific audio content format is a scene-based audio representation signal, and the audio signal encoding module is further configured to perform a sound field rotation operation on the scene-based audio representation signal based on spatial rotation information indicated or included in metadata associated with the audio representation signal. In this way, the rotated audio signal can be used as an encoded audio signal for spatial decoding.
[0171] As an example, for a scene audio signal, it is itself a FOA, HOA or MOA signal, so it can be directly weighted according to the weight information in the metadata, that is, the intermediate signal that is desired to be obtained. In addition, if the metadata indicates that the sound field needs to be rotated, the sound field rotation process can be performed in the encoding module according to different implementations. For example, the scene audio signal can be multiplied by a parameter indicating the sound field rotation characteristics, such as a vector, a matrix, etc., so that the audio signal can be further processed. It should be noted that this sound field rotation operation can also be performed in the decoding stage. In some implementations, the sound field rotation operation can be performed in one of the encoding and decoding stages, or in both.
[0172] In some embodiments, the audio signal of the specific audio content format is a channel-based audio representation signal, and the audio signal encoding module is further configured to convert the channel-based audio representation signal that needs to be converted into an object-based audio representation signal and encode it when the channel-based audio representation signal needs to be converted. The encoding operation here can be performed as described above for encoding the object-based audio representation signal. In some embodiments, the channel-based audio representation signal that needs to be converted may include a narrative channel track of the channel-based audio representation signal, and the audio signal encoding module is further configured to convert the audio representation signal converted from the narrative channel track into an object-based audio representation signal and encode it, as described above. In other embodiments, for the narrative channel track of the channel-based audio representation signal, the audio representation signal corresponding to the narrative channel track can be split into audio elements by channel and converted into metadata for encoding.
[0173] In some embodiments, the audio signal of a specific audio content format is a channel-based audio representation signal, and the channel-based audio representation information may not be subjected to spatial audio processing, especially spatial audio encoding. Such a channel-based audio representation signal will be directly transmitted to the audio decoding module and processed in an appropriate manner for playback / rendering. In particular, in some embodiments, when the narrative channel audio track of the channel-based audio representation signal does not require spatial audio processing according to scene requirements, for example, it is pre-specified that the narrative channel audio track does not need to be encoded, and the narrative channel audio track can be directly transmitted to the decoding step. In other embodiments, the non-narrative channel audio track of the channel-based audio representation signal itself does not require spatial audio processing, and therefore can be directly transmitted to the decoding step.
[0174] As an example, the spatial encoding processing of the channel-based audio representation signal can be performed based on a predetermined rule, which can be provided in a suitable manner, and in particular can be specified in the information processing module. For example, it can be specified that the channel-based audio representation signal, especially the narrative channel track in the channel-based audio representation signal, needs to be processed by audio encoding. Therefore, the audio encoding can be performed in a suitable manner according to the regulations. The audio encoding method can be converted into an object-based audio representation and processed as described above, or it can be any other encoding method, such as a pre-agreed encoding method for channel-based audio signals. On the other hand, in the case where it has been specified that the channel-based audio representation signal, especially the narrative channel track therein, does not need to be converted, or in the case of a non-narrative channel track in the channel-based audio representation signal, the audio representation signal can be directly transmitted to the decoding module / stage, so that it can be processed for different playback modes.
[0175] Audio signal decoding
[0176] According to an embodiment of the present disclosure, after the audio signal is audio encoded or directly transmitted / transparently transmitted as described above, such encoded audio signal or directly transmitted / transparently transmitted audio signal will be subjected to audio decoding processing to obtain an audio signal suitable for playback / rendering in the user application scenario. In particular, such an encoded audio signal or directly transmitted / transparently transmitted audio signal may be referred to as a signal to be decoded, which may correspond to an audio signal of a specific spatial format as described above, or an intermediate signal. As an example, the audio signal of the specific spatial format may be the aforementioned intermediate signal, or may be an audio signal directly transmitted / transparently transmitted to a spatial decoder, including an unencoded audio signal, or an encoded audio signal that is spatially encoded but not included in an intermediate signal, such as a non-narrative channel signal or a binaural signal after reverberation processing. The audio decoding process may be performed by an audio signal decoding module.
[0177] According to an embodiment of the present disclosure, the audio signal decoding module can decode the intermediate signal and the transparent signal to the playback / playing device according to the playback mode. Thus, the audio signal to be decoded can be converted into a format suitable for playback by a playback device in a user application scenario, such as an audio playback environment, an audio rendering environment. According to an embodiment of the present disclosure, the playback mode can be related to the configuration of the playback device in the user application scenario. In particular, depending on the configuration information of the playback device in the user application scenario, such as the identifier, type, arrangement, etc. of the playback device, a corresponding decoding method can be adopted. In this way, the decoded audio signal can be suitable for a specific type of playback environment, especially for the playback device in the playback environment, so that compatibility with various types of playback environments can be achieved. As an example, the audio signal decoder can be decoded according to information related to the type of the user application scenario, which can be a type indicator of the user application scenario, such as a type indicator of the rendering device / playback device in the user application scenario, such as a renderer ID, so that a decoding process corresponding to the renderer ID can be performed to obtain an audio signal suitable for playback by the renderer. As an example, the renderer ID may be as described above, each renderer ID may correspond to a specific renderer arrangement / playback scene / playback device arrangement, etc., so that an audio signal suitable for playback of the renderer arrangement / playback scene / playback device arrangement, etc. corresponding to the renderer ID may be decoded. In some embodiments, the playback mode, such as the renderer ID, may be pre-specified, transmitted to the rendering end, or input through an input port. In some embodiments, the audio signal decoder decodes the audio signal of a specific spatial format using a decoding method corresponding to the playback device in the user application scenario.
[0178] In some embodiments, the playback device in the user application scenario may include a speaker array, which may correspond to a speaker playback / rendering scenario. In this case, the audio signal decoder may use a decoding matrix corresponding to the speaker array in the user application scenario to decode the audio signal of the specific spatial format. As an example, such a user application scenario may correspond to a specific renderer ID, such as the aforementioned renderer ID2. In particular, for example, corresponding identifiers may be set separately according to the type of speaker array to more accurately indicate the user application scenario. For example, corresponding identifiers may be set separately for standard speaker arrays, custom speaker arrays, etc.
[0179] The decoding matrix may be determined depending on the configuration information of the speaker array, such as the type and arrangement of the speaker array. In some embodiments, when the playback device in the user application scenario is a predetermined speaker array, the decoding matrix is a decoding matrix corresponding to the predetermined speaker array built into the audio signal decoder or received from the outside. In particular, the decoding matrix may be a preset decoding matrix, which may be pre-stored in the decoding module, for example, may be stored in a database in association with / corresponding to the type of the speaker array, or provided to the decoding module in other ways. Thus, the decoding module may call the corresponding decoding matrix according to the known predetermined speaker array type for decoding processing. The decoding matrix may be in various suitable forms, for example, it may include a gain, such as a gain value from an HOA track / channel to a speaker, so that the gain may be directly applied to the HOA signal to generate an output audio channel so as to render the HOA signal into the speaker array.
[0180] As an example, for a standard speaker array defined in the standard, such as 5.1, the decoder will have built-in decoding matrix coefficients, and the playback signal L can be obtained by multiplying the intermediate signal with the decoding matrix.
[0181] L=DS N ,
[0182] Where L is the speaker array signal, D is the decoding matrix, S N is an intermediate signal, obtained as described above. On the other hand, for the direct / transparent audio signal, the signal can be converted into the speaker array according to the definition of the standard speaker, for example, it can be multiplied by the decoding matrix as described above, and other suitable methods can also be used, such as vector-based amplitude translation (Vector-base amplitude panning, VBAP), etc. As another example, in the case of special speaker array spatial decoding, for Sound Bar or some more special speaker arrays, the speaker manufacturer is required to provide a correspondingly designed decoding matrix. The system provides a decoding matrix setting interface to receive decoding matrix related parameters corresponding to the special speaker array, so that the received decoding matrix can be used for decoding processing, as described above.
[0183] In other embodiments, when the playback device in the user application scenario is a custom speaker array, the decoding matrix is a decoding matrix calculated according to the arrangement of the custom speaker array. As an example, the decoding matrix is calculated according to the azimuth and pitch angle of each speaker in the speaker array or the three-dimensional coordinate value of the speaker. As an example, in the custom speaker array spatial decoding, in the scenario of the custom speaker array, such speakers usually have a spherical, hemispherical design or a rectangle, which can surround or semi-surround the listener. The decoding module can calculate the decoding matrix according to the arrangement of the custom speakers, and the required input is the azimuth and pitch angle of each speaker, or the three-dimensional coordinate value of the speaker. The calculation method of the speaker decoding matrix can be SAD (Sampling Ambisonic Decoder), MMD (Mode Matching Decoder), EPAD (Energy preserved Ambisonic Decoder), AllRAD (All Round Ambisonic Decoder), etc.
[0184] According to some embodiments of the present disclosure, when the playback device in the user application scenario is a headset, it may correspond to scenarios such as headset rendering / playback, binaural rendering / playback, etc., and the audio signal decoder is configured to directly decode the audio signal to be decoded into a binaural signal as a decoded audio signal, or obtain the decoded signal as a decoded audio signal through speaker virtualization. As an example, such a user application scenario may correspond to a specific renderer ID, such as the aforementioned renderer ID1. As an example, there may be a variety of appropriate decoding methods for the playback environment of the headset. In some embodiments, for example, the signal to be decoded, such as the aforementioned intermediate signal, may be directly decoded into a binaural signal. In particular, the signal to be decoded may be directly decoded, for example, the HOA signal may be converted by determining a rotation matrix according to the posture of the listener, and then the HOA channel / track may be adjusted, such as convolution (for example, convolution using a gain matrix, a harmonic function, HRIR (head-related impulse response), spherical harmonic HRIR, etc., such as frequency domain convolution), so that a binaural signal may be obtained. In other words, such a process can also be regarded as directly multiplying the HOA signal by the decoding matrix, which can include a rotation matrix, a gain matrix, a harmonic function, etc. As an example, typical methods include LS (least squares), Magnitude LS, SPR (Spatial resampling), etc. For the transparent signal, which is usually a binaural signal, it is directly played back. As another example, indirect rendering can also be performed, that is, first using a speaker array, and then performing HRTF convolution according to the position of the speaker to virtualize the speaker, so as to obtain a decoded signal.
[0185] In some embodiments, in the audio decoding process, the audio signal to be decoded can also be processed based on the metadata information associated with the audio signal to be decoded. In particular, the audio signal to be decoded can be spatially transformed according to the spatial transformation information in the metadata information. For example, when the metadata information indicates that rotation is required, the sound field rotation operation can be performed on the audio representation signal to be decoded based on the rotation information indicated in the metadata. As an example, first, according to the processing method of the previous module and the rotation information in the metadata, the intermediate signal is multiplied by the rotation matrix as needed to obtain the rotated intermediate signal, so that the rotated intermediate signal can be decoded. It should be noted that the spatial transformation here, such as spatial rotation, can be performed alternatively with the spatial encoding in the spatial encoding process described above, such as spatial rotation.
[0186] Audio signal post-processing
[0187] According to an embodiment of the present disclosure, optionally or additionally, the spatially decoded audio signal can be adjusted for a specific playback device in a user application scenario, so that the adjusted audio signal can present a more appropriate acoustic experience when rendered by an audio rendering device. In particular, the audio signal adjustment can be mainly intended to eliminate the inconsistencies that may exist between different playback types, or different playback modes, etc., so that the adjusted audio signal can maintain a consistent playback experience when played back in the application scenario, thereby improving the user's experience. In the context of the present disclosure, the audio signal adjustment process can be referred to as a post-processing, which refers to post-processing the output signal obtained by audio decoding, which can be referred to as output signal post-processing. In some embodiments, the signal post-processing module is configured to perform at least one of frequency response compensation and dynamic control range on the decoded audio signal for a specific playback device.
[0188] As an example, the post-processing module takes into account the inconsistency of different playback methods. Different playback devices have different frequency response curves and gains. In order to present a consistent acoustic experience, the output signal is post-processed and adjusted. The post-processing operation includes but is not limited to frequency response compensation (EQ, Equalization) and dynamic range control (DRC, Dynamic range control) for specific devices.
[0189] In the audio rendering system disclosed in the present invention, the audio information processing module, audio signal encoding module, signal space decoder and output signal post-processing mentioned above can constitute the core rendering module of the system, which is responsible for processing the signals in three audio representation formats and their metadata obtained after pre-processing and playing them back through the playback device in the user application environment.
[0190] It should be noted that the various modules of the audio rendering system described above are only logical modules divided according to the specific functions they implement, and are not used to limit the specific implementation method. For example, they can be implemented in software, hardware, or a combination of software and hardware. In actual implementation, the above-mentioned modules can be implemented as independent physical entities, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). For example, encoders, decoders, etc. can use chips (such as integrated circuit modules including a single chip), hardware components or complete products. In addition, the above-mentioned modules are shown with dotted lines in the accompanying drawings to indicate that these units may not actually exist, and the operations / functions they implement can be implemented by other modules containing the module or the system or device itself. For example, Figure 4A At least one of the audio signal parsing module 411, the information processing module 412, and the audio signal encoding module 413 shown in the figure may be located outside the acquisition module 41 and exist in the audio rendering system 4, for example, between the acquisition module 41 and the decoder 42, and sequentially process the input audio signal to obtain the audio signal to be processed by the decoder. It may even be located outside the audio rendering system.
[0191] In addition, although not shown, the audio rendering system 4 may also include a memory, which may store various information generated by the various modules included in the system and the device during operation, programs and data for operation, data to be sent by the communication unit, etc. The memory may be a volatile memory and / or a non-volatile memory. For example, the memory may include, but is not limited to, a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a read-only memory (ROM), and a flash memory. Of course, the memory may also be located outside the device.
[0192] In addition, optionally, the audio rendering system 4 may also include other components not shown, such as an interface, a communication unit, etc. As an example, an interface and / or a communication unit may be used to receive an input audio signal to be rendered, and the audio signal finally generated may also be output to a playback device in a playback environment for playback. In one example, the communication unit may be implemented in a suitable manner known in the art, such as including communication components such as an antenna array and / or a radio frequency link, various types of interfaces, communication units, etc. This will not be described in detail. In addition, the device may also include other components not shown, such as a radio frequency link, a baseband processing unit, a network interface, a processor, a controller, etc. This will not be described in detail.
[0193] The following will describe an exemplary implementation of audio rendering according to an embodiment of the present disclosure in conjunction with the accompanying drawings, wherein Figure 4G and 4HA flowchart of an exemplary implementation of an audio rendering process according to an embodiment of the present disclosure is shown. As an example, the audio rendering system mainly includes a rendering metadata system and a core rendering system. The metadata system contains control information describing audio content and rendering technology, such as whether the audio input form is single-channel, dual-channel, multi-channel, or object (object) or sound field HOA, as well as dynamic sound sources and listener position information, and rendered acoustic environment information such as house shape, size, wall quality, etc. The core rendering system renders the corresponding playback device and environment based on different audio signal representation forms and metadata parsed from the metadata system.
[0194] First, an input audio signal is received, and parsed or directly transmitted according to the format of the input audio signal. On the one hand, when the input audio signal is an input signal with any spatial audio exchange format, the input audio signal can be parsed to obtain an audio signal with a specific spatial audio representation, such as an object-based spatial audio representation signal, a scene-based spatial audio representation signal, a channel-based spatial audio representation signal, and associated metadata, and then the parsing result is passed to the subsequent processing stage. On the other hand, when the input audio signal is directly an audio signal with a specific spatial audio representation, it does not need to be parsed and is directly passed to the subsequent processing stage. For example, such an audio signal can be directly transmitted to the audio encoding stage, such as an object-based audio representation signal, a scene-based audio representation signal, or a narrative channel track that needs to be encoded in a channel-based audio representation signal. Even in the case where the audio signal of the specific spatial representation is of a type / format that does not require encoding, it can be directly transmitted to the audio decoding stage, such as a non-narrative channel track in the parsed channel-based audio representation, or a narrative channel track that does not need to be encoded.
[0195] Then, information processing can be performed based on the obtained metadata to extract and obtain audio parameters related to each audio signal, and such audio parameters can be used as metadata information. The information processing here can be performed on either the parsed audio signal or the directly transmitted audio signal. Of course, as mentioned above, such information processing is optional and does not have to be performed.
[0196] Next, the audio signal of the specific spatial audio representation is encoded. On the one hand, the audio signal of the specific spatial audio representation can be encoded based on the metadata information, and the obtained encoded audio signal is either directly transmitted to the subsequent audio decoding stage, or an intermediate signal is obtained and then transmitted to the subsequent audio decoding stage. On the other hand, in the case where the audio signal of the specific spatial audio representation does not need to be encoded, such an audio signal can be directly transmitted to the audio decoding stage.
[0197] Then, in the audio decoding stage, the received audio signal can be decoded to obtain an audio signal suitable for playback in the user application scenario as an output signal. Such an output signal can be presented to the user through the user application scenario, such as an audio playback device in an audio playback environment.
[0198] Fig. 4I Flowcharts of some embodiments of the audio rendering method according to the present disclosure are shown. Fig. 4I As shown, in method 400, in step S430 (also referred to as an audio signal encoding step), for the audio signal in the specific audio content format, based on metadata information associated with the audio signal in the specific audio content format, the audio signal in the specific audio content format is spatially encoded to obtain an encoded audio signal; and in step S440 (also referred to as an audio signal decoding step), the encoded audio signal in the specific spatial format may be spatially decoded to obtain a decoded audio signal for audio rendering.
[0199] In some embodiments of the present disclosure, the method 400 may further include step S410 (also referred to as an audio signal acquisition step), in which an audio signal in a specific audio content format and metadata information associated with the audio signal are acquired. In the audio signal acquisition step, the step may further include parsing the input audio signal to obtain an audio signal in accordance with a specific spatial audio representation method, and performing format conversion on the audio signal in accordance with the specific spatial audio representation method to obtain an audio signal in the specific audio content format.
[0200] In some embodiments of the present disclosure, the method 400 may further include step S420 (also referred to as an information processing step), in which the audio parameters of the specific type of audio signal may be extracted based on the metadata information associated with the specific type of audio signal. In particular, in the audio information processing step, the audio parameters of the specific type of audio signal may be further extracted based on the audio content format of the specific type of audio signal. Thus, in the audio signal encoding step, the step may further include spatially encoding the specific type of audio signal based on the audio parameters.
[0201] In some embodiments of the present disclosure, in the audio signal decoding step, the audio signal in the specific spatial format may be further decoded based on the playback mode. In particular, the decoding may be performed using a decoding method corresponding to a playback device in a user application scenario.
[0202] In some embodiments of the present disclosure, method 400 may further include a signal input step, in which an input audio signal is received, and when the input audio signal is an audio signal of a specific type in an audio signal in a specific audio content format, the input audio signal is directly transmitted to the audio signal encoding step, or when the input audio signal is an input audio signal in a specific audio content format and is not an audio signal of the specific type, the input audio signal is directly transmitted to the audio signal decoding step.
[0203] In some embodiments of the present disclosure, method 400 may further include step S450 (also referred to as a signal post-processing step), in which the decoded audio signal may be post-processed, in particular, the post-processing may be performed based on the characteristics of the playback device in the user application scenario.
[0204] It should be noted that the above-mentioned signal acquisition step, information processing step, signal input step, and signal post-processing step are not necessarily included in the rendering method according to the present disclosure, that is, even if the step is not included, the method according to the present disclosure is still complete and can effectively solve the problems of the present disclosure and achieve favorable effects. For example, these steps may be implemented outside the method according to the present disclosure, and the results of the steps are provided to the method of the present disclosure, or the result signal of the method of the present disclosure is received. In addition, in the exemplary line of sight, these steps may also be combined with other steps of the present disclosure, for example, the signal acquisition step may be included in the signal encoding step, for example, the information processing step, the signal input step may be included in the signal acquisition step, or the information processing step may be included in the signal encoding step, or the signal post-processing step may be included in the signal decoding step. Therefore, these steps are shown with dotted lines in the accompanying drawings.
[0205] Although not shown, the audio rendering method according to the present disclosure may also include other steps to implement the processing / operations in the pre-processing, audio information processing, audio signal spatial coding, etc. described above, which will not be described in detail here. It should be noted that the audio rendering method according to the present disclosure and the steps therein can be performed by any appropriate device, such as a processor, an integrated circuit, a chip, etc., for example, it can be performed by the aforementioned audio rendering system and each module therein, and the method can also be embodied in a computer program, an instruction, a computer program medium, a computer program product, etc.
[0206] Figure 5 1 shows a block diagram of an electronic device according to some embodiments of the present disclosure. Figure 5As shown, the electronic device 5 of this embodiment includes: a memory 51 and a processor 52 coupled to the memory 51, and the processor 52 is configured to execute the reverberation duration estimation method or the audio signal rendering method in any embodiment of the present disclosure based on the instructions stored in the memory 51.
[0207] The memory 51 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory may store, for example, an operating system, an application program, a boot loader, a database, and other programs.
[0208] Reference below Figure 6 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. The electronic device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0209] Figure 6 A block diagram showing some other embodiments of the electronic device of the present disclosure.
[0210] like Figure 6 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In RAM 603, various programs and data required for the operation of the electronic device are also stored. The processing device 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0211] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0212] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0213] In some embodiments, a chip is also provided, including: at least one processor and an interface, the interface being used to provide computer execution instructions to the at least one processor, and the at least one processor being used to execute the computer execution instructions to implement the reverberation duration estimation method or audio signal rendering method of any of the above embodiments.
[0214] Figure 7 1 shows a block diagram of a chip capable of implementing some embodiments of the present disclosure. Figure 7 As shown, the processor 70 of the chip is mounted on the host CPU as a coprocessor, and the host CPU assigns tasks. The core part of the processor 70 is the operation circuit, and the controller 704 controls the operation circuit 703 to extract data from the memory (weight memory or input memory) and perform operations.
[0215] In some embodiments, the operation circuit 703 includes a plurality of processing units (Process Engine, PE) therein. In some embodiments, the operation circuit 703 is a two-dimensional systolic array. The operation circuit 703 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some embodiments, the operation circuit 703 is a general-purpose matrix processor.
[0216] For example, suppose there is an input matrix A, a weight matrix B, and an output matrix C. The operation circuit takes the corresponding data of matrix B from the weight memory 702 and caches it on each PE in the operation circuit. The operation circuit takes the matrix A data from the input memory 701 and performs matrix operation with matrix B, and the partial result or final result of the matrix is stored in the accumulator 708.
[0217] The vector calculation unit 707 can further process the output of the operation circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc.
[0218] In some embodiments, the vector calculation unit 707 can store the processed output vector to the unified buffer 706. For example, the vector calculation unit 707 can apply a nonlinear function to the output of the operation circuit 703, such as a vector of accumulated values, to generate an activation value. In some embodiments, the vector calculation unit 707 generates a normalized value, a merged value, or both. In some embodiments, the processed output vector can be used as an activation input to the operation circuit 703, for example, for use in a subsequent layer in a neural network.
[0219] The unified memory 706 is used to store input data and output data.
[0220] The memory unit access controller 705 (Direct Memory Access Controller, DMAC) moves the input data in the external memory to the input memory 701 and / or the unified memory 706, stores the weight data in the external memory into the weight memory 702, and stores the data in the unified memory 706 into the external memory.
[0221] The bus interface unit (BIU) 510 is used to implement the interaction between the main CPU, DMAC and instruction fetch memory 709 through the bus.
[0222] An instruction fetch buffer 709 connected to the controller 704 and used to store instructions used by the controller 704;
[0223] The controller 704 is used to call the instructions cached in the memory 709 to control the working process of the computing accelerator.
[0224] Generally, the unified memory 706, the input memory 701, the weight memory 702 and the instruction fetch memory 709 are all on-chip memories, and the external memory is a memory outside the NPU, which can be a double data rate synchronous dynamic random access memory (DDR SDRAM), a high bandwidth memory (HBM) or other readable and writable memory.
[0225] In some embodiments, a computer program is further provided, comprising: instructions, which, when executed by a processor, cause the processor to perform the audio signal processing of any of the above embodiments, especially any processing in the audio signal rendering process.
[0226] Those skilled in the art will appreciate that the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When a computer instruction or computer program is loaded or executed on a computer, a process or function according to an embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable non-transient storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0227] Although some specific embodiments of the present disclosure have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present disclosure. It should be understood by those skilled in the art that the above embodiments may be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. An audio processing method for audio rendering, comprising: an acquisition step, for acquiring an audio signal in a specific audio content format and related parameters of the audio signal in the specific audio content format, wherein the related parameters are acquired based on metadata associated with the audio signal in the specific audio content format; and A spatial encoding step, for spatially encoding the audio signal of the specific audio content format based on relevant parameters of the audio signal of the specific audio content format to obtain a spatially encoded audio signal with a common spatial format, wherein the spatially encoded audio signal is an Ambisonics type audio signal; Wherein, in the case where the audio signal in the specific audio content format includes an object-based audio representation signal, the spatial encoding step includes spatially encoding the object-based audio representation signal based on spatial attribute information in relevant parameters of the object-based audio representation signal, and the spatial attribute information includes relevant information of the spatial propagation path from the sound object of the audio representation signal to the listener, which includes at least one of the propagation duration, propagation distance, orientation information, path energy intensity, and nodes along the way of the spatial propagation path from the sound object to the listener.
2. The method according to claim 1, wherein: The Ambisonics type audio signal can include at least one of FOA (First Order Ambisonics), HOA (Higher Order Ambisonics), and MOA (Mixed-order Ambisonics).
3. The method according to claim 1, wherein: The spatial encoding step further includes spatially encoding the audio representation signal based on at least one of a filter function that filters the audio signal based on the path energy intensity of the spatial propagation path from the sound object in the audio representation signal to the listener and a spherical harmonic function based on the azimuth information of the spatial propagation path.
4. The method according to claim 1, wherein: The spatial encoding step further includes using at least one of a near-field compensation function and a diffusion function to perform spatial encoding of the audio representation signal based on the length of a spatial propagation path from a sound object in the audio representation signal to a listener.
5. The method according to claim 1, wherein: The spatial encoding step further comprises, in the case where the audio representation signal contains a plurality of sound objects, For each sound object in the audio representation signal, spatially encode the audio signal based on information about the spatial propagation path of the sound object in the audio representation signal to a listener, and Based on the weights of the sound objects defined in the metadata, the spatially encoded signals of the audio representation signals of the sound objects are weightedly superimposed.
6. The method according to claim 1, wherein: The spatial encoding step further comprises: In case that the audio signal in the specific audio content format comprises an object-based audio representation signal, a reverberation-related signal of the object-based audio representation signal is obtained based on a reverberation parameter among the related parameters of the object-based audio representation signal.
7. The method according to claim 1, wherein: The spatial encoding step further comprises weighting the scene-based audio representation signal based on weight information in relevant parameters of the scene-based audio representation signal when the audio signal of the specific audio content format comprises a scene-based audio representation signal.
8. The method according to claim 1, wherein: The spatial encoding step further includes performing a sound field rotation operation on the scene-based audio representation signal based on rotation information indicated in relevant parameters of the scene-based audio representation signal when the audio signal in the specific audio content format includes a scene-based audio representation signal.
9. The method according to claim 1, wherein: The spatial encoding step further comprises, when the audio signal in the specific audio content format comprises a specific type of channel signal in a channel-based audio representation signal, converting the specific type of channel signal into an object-based audio representation signal and performing spatial encoding on the signal.
10. The method according to claim 1, wherein: The spatial encoding step further includes, when the audio signal in the specific audio content format includes a specific type of channel signal in a channel-based audio representation signal, splitting the specific type of channel signal into audio elements by channel and converting the audio elements into metadata for spatial encoding.
11. The method according to claim 1, wherein: The spatial attribute information of the object-based audio representation signal also includes at least one of the position information of each audio element in the audio representation signal in the coordinate system, the distance information of each audio element, or the relative position information of the sound source related to the audio signal relative to the listener.
12. The method according to any one of claims 1 to 11, wherein: In the case that the audio signal in the specific audio content format comprises a scene-based audio representation signal, the relevant parameters comprise rotation information related to the audio signal.
13. The method according to claim 12, wherein: The rotation information related to the audio signal includes at least one of rotation information of the audio signal and rotation information of a listener of the audio signal.
14. The method according to any one of claims 1 to 11, wherein: In the case where the audio signal in the specific audio content format includes a specific type of channel signal in a channel-based audio representation signal, the relevant parameters include metadata converted by splitting the audio representation of the specific type of channel signal into audio elements by channel.
15. The method according to any one of claims 1 to 11, wherein: The audio signal in the specific audio content format is obtained by parsing an input audio signal in a spatial audio exchange format.
16. An audio rendering method, comprising: an audio signal encoding step, for spatially encoding an audio signal in a specific audio content format to obtain an encoded audio signal using the method according to any one of claims 1 to 15; and The audio signal decoding step is used to spatially decode the encoded audio signal to obtain a decoded audio signal for audio rendering.
17. The audio rendering method according to claim 16, wherein: The spatial decoding step further includes spatially decoding an audio signal that has not been spatially encoded, wherein the audio signal that has not been spatially encoded includes at least one of a scene-based audio representation signal, a specific type of channel signal in a channel-based audio representation signal, and an audio signal processed with reverberation.
18. The audio rendering method according to claim 16, wherein: The spatial decoding step further includes spatially decoding the audio signal based on a playback mode, wherein the playback mode is indicated by at least one of a playback type, a playback environment, a playback device type, and a playback device identifier.
19. The audio rendering method according to any one of claims 16 to 18, further comprising a signal post-processing step for post-processing the decoded audio signal.
20. An audio processing device for audio rendering, comprising: an acquiring unit configured to acquire an audio signal in a specific audio content format and related parameters of the audio signal in the specific audio content format, wherein the related parameters are acquired based on metadata associated with the audio signal in the specific audio content format; and a spatial encoding unit configured to spatially encode the audio signal of the specific audio content format based on relevant parameters of the audio signal of the specific audio content format to obtain a spatially encoded audio signal having a common spatial format, wherein the spatially encoded audio signal is an Ambisonics type audio signal; Wherein, in the case where the audio signal in the specific audio content format includes an object-based audio representation signal, the spatial encoding unit is configured to spatially encode the object-based audio representation signal based on spatial attribute information in relevant parameters of the object-based audio representation signal, and the spatial attribute information includes relevant information of the spatial propagation path from the sound object of the audio representation signal to the listener, which includes at least one of the propagation duration, propagation distance, orientation information, path energy intensity, and nodes along the way of the spatial propagation path from the sound object to the listener.
21. The apparatus according to claim 20, wherein: The Ambisonics type audio signal can include at least one of FOA (First Order Ambisonics), HOA (Higher Order Ambisonics), and MOA (Mixed-order Ambisonics).
22. The apparatus of claim 20, wherein: The spatial encoding unit is configured to perform spatial encoding of the audio representation signal based on at least one of a filter function that filters the audio signal based on the path energy intensity of the spatial propagation path from the sound object in the audio representation signal to the listener and a spherical harmonic function based on the azimuth information of the spatial propagation path.
23. The apparatus of claim 20, wherein: The spatial encoding unit is further configured to perform spatial encoding of the audio representation signal using at least one of a near field compensation function and a diffusion function based on the length of a spatial propagation path from a sound object in the audio representation signal to a listener.
24. The apparatus of claim 20, wherein: The spatial encoding unit is configured to: when the audio representation signal contains a plurality of sound objects, For each sound object in the audio representation signal, spatially encode the audio signal based on information about the spatial propagation path of the sound object in the audio representation signal to a listener, and Based on the weights of the sound objects defined in the metadata, the spatially encoded signals of the audio representation signals of the sound objects are weightedly superimposed.
25. The apparatus of claim 20, wherein: The spatial coding unit is further configured as follows: In case that the audio signal in the specific audio content format comprises an object-based audio representation signal, a reverberation-related signal of the object-based audio representation signal is obtained based on a reverberation parameter among the related parameters of the object-based audio representation signal.
26. The apparatus of claim 20, wherein: The spatial encoding unit is further configured to, when the audio signal in the specific audio content format includes a scene-based audio representation signal, weight the scene-based audio representation signal based on weight information in relevant parameters of the scene-based audio representation signal.
27. The apparatus of claim 20, wherein: The spatial encoding unit is further configured to perform a sound field rotation operation on the scene-based audio representation signal based on rotation information indicated in relevant parameters of the scene-based audio representation signal when the audio signal in the specific audio content format includes a scene-based audio representation signal.
28. The apparatus of claim 20, wherein: The spatial encoding unit is further configured to, when the audio signal in the specific audio content format includes a specific type of channel signal in a channel-based audio representation signal, convert the specific type of channel signal into an object-based audio representation signal and perform spatial encoding on the signal.
29. The apparatus of claim 20, wherein: The spatial encoding unit is further configured to, when the audio signal in the specific audio content format includes a specific type of channel signal in a channel-based audio representation signal, split the specific type of channel signal into audio elements by channel and convert the audio elements into metadata for spatial encoding.
30. The apparatus of claim 20, wherein: The spatial attribute information of the object-based audio representation signal also includes at least one of the position information of each audio element in the audio representation signal in the coordinate system, the distance information of each audio element, or the relative position information of the sound source related to the audio signal relative to the listener.
31. The apparatus according to any one of claims 20 to 30, wherein: In the case that the audio signal in the specific audio content format comprises a scene-based audio representation signal, the relevant parameters comprise rotation information related to the audio signal.
32. The apparatus of claim 31, wherein: The rotation information related to the audio signal includes at least one of rotation information of the audio signal and rotation information of a listener of the audio signal.
33. The apparatus according to any one of claims 20 to 30, wherein: In the case where the audio signal in the specific audio content format includes a specific type of channel signal in a channel-based audio signal, the relevant parameters include metadata converted by splitting the audio representation of the specific type of channel signal into audio elements according to channels.
34. The apparatus according to any one of claims 20 to 30, wherein: The audio signal in the specific audio content format is obtained by parsing an input audio signal in a spatial audio exchange format.
35. An audio rendering device, comprising: The audio processing device according to any one of claims 20 to 34; as well as The audio signal decoder is used to spatially decode the encoded audio signal obtained by the audio processing device to obtain a decoded audio signal for audio rendering.
36. The audio rendering device according to claim 35, wherein: The audio signal decoder is further configured to perform spatial decoding on an audio signal that has not been spatially encoded, wherein the audio signal that has not been spatially encoded includes at least one of a scene-based audio representation signal, a specific type of channel signal in a channel-based audio representation signal, and an audio signal processed with reverberation.
37. The audio rendering device according to claim 35, wherein: The audio signal decoder is further configured to spatially decode the audio signal based on a playback mode, wherein the playback mode is indicated by at least one of a playback type, a playback environment, a playback device type, and a playback device identifier.
38. The audio rendering device according to any one of claims 35-37, further comprising a signal post-processor for post-processing the decoded audio signal.
39. A chip, comprising: At least one processor and an interface, wherein the interface is used to provide computer-executable instructions to the at least one processor, and the at least one processor is used to execute the computer-executable instructions to implement the method according to any one of claims 1-19.
40. An electronic device comprising: Memory; and A processor coupled to the memory, the processor being configured to execute the method according to any one of claims 1-19 based on instructions stored in the memory device.
41. A non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of claims 1 to 19.
42. A computer program product comprising instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 1-19.
Citation Information
Patent Citations
Method, apparatus and system for pre-rendered signal for audio rendering
CN111955020A
Apparatus and method for audio encoding
WO2021074007A1