Audio coding method and device and electronic equipment

The sound source direction information of the scene audio signal is predicted through the sound source direction prediction model, and the problem of mismatch extraction of HOA signal direction information in the prior art is solved, and the encoding efficiency and quality are improved.

CN119993172APending Publication Date: 2025-05-13HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311503652.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art does not match the signal when extracting the direction information of the HOA signal, resulting in a decrease in encoding efficiency.

Method used

By obtaining the characteristic information of the scene audio signal and inputting it into the sound source direction prediction model, the sound source direction prediction result is output, and the sound source direction information is determined based on the result and encoding.

Benefits of technology

The encoding efficiency is improved, so that the encoding quality at the encoding end is higher, and the sound source direction information that conforms to the characteristics of the scene audio signal can be output more flexibly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993172A_ABST
    Figure CN119993172A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an audio coding method and device and electronic equipment. The audio coding method comprises the following steps: acquiring feature information of a scene audio signal to be coded; inputting the feature information of the scene audio signal into a sound source direction prediction model, and outputting a sound source direction prediction result through the sound source direction prediction model; determining sound source direction information of the scene audio signal according to the sound source direction prediction result; and encoding the scene audio signal according to the sound source direction information to obtain a first code stream. According to the embodiment of the invention, the sound source direction prediction result of the scene audio signal can be adaptively and flexibly output according to the real-time change of the scene audio signal, so that the sound field of the position where a listener is located in the space is close to the original sound field when the scene audio signal is recorded as much as possible, the coding quality of a coding end is ensured, and the coding efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of audio coding technology, and more particularly to an audio coding method and device and an electronic device. Background Art

[0002] Three-dimensional audio technology is an audio technology that uses computers and signal processing to acquire, process, transmit, render and play back sound events and three-dimensional sound field information in the real world. Three-dimensional audio gives sound a strong sense of space, envelopment and immersion, giving people an extraordinary auditory experience of "being there". Among them, higher order ambisonics (HOA) technology has the property of being independent of speaker layout during recording, encoding and playback, as well as the rotatable playback characteristics of HOA format data. It has higher flexibility when playing back three-dimensional audio, and has therefore received more extensive attention and research.

[0003] For an N-order HOA signal, it is necessary to extract the directional information of the HOA signal and convert the HOA signal into a virtual speaker signal according to the directional information to achieve the playback of three-dimensional audio. However, the directional information extracted from the HOA signal in the prior art does not match the HOA signal, which reduces the signal coding efficiency. Summary of the invention

[0004] The embodiments of the present application provide an audio encoding method and apparatus and an electronic device for improving encoding efficiency.

[0005] To solve the above technical problems, the present application provides the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides an audio encoding method, comprising: first obtaining feature information of a scene audio signal to be encoded; then inputting the feature information of the scene audio signal into a sound source direction prediction model, and outputting a sound source direction prediction result through the sound source direction prediction model; next determining the sound source direction information of the scene audio signal according to the sound source direction prediction result; and finally encoding the scene audio signal according to the sound source direction information to obtain a first bit stream.

[0007] In the above scheme, the characteristic information of the scene audio signal is input into the sound source direction prediction model. Since the characteristic information of the scene audio signal changes in real time with the scene audio signal, the sound source direction is predicted by combining the characteristic information of the scene audio signal through the sound source direction prediction model, so that the output sound source direction prediction result can indicate the sound source direction that conforms to the scene audio signal. The embodiment of the present application can adaptively and flexibly output the sound source direction prediction result of the scene audio signal according to the real-time changes of the scene audio signal, and then generate the sound source direction information according to the sound source direction prediction result. The scene audio signal is encoded according to the sound source direction information, so that the sound field at the position of the listener in the space is as close as possible to the original sound field when the scene audio signal is recorded, thereby ensuring the encoding quality of the encoding end and improving the encoding efficiency.

[0008] In a possible implementation, the sound source direction prediction result includes: a sound source direction angle and a confidence level corresponding to the sound source direction angle. In the above scheme, the encoder predicts the sound source direction angle and the corresponding confidence level through a sound source direction prediction model, so that the predicted sound source direction angle can be selected through the confidence level to determine the sound source direction information that meets the characteristics of the scene audio signal.

[0009] In a possible implementation, the number of the sound source direction angles is N1, and the number of the confidence levels is N1, there is a one-to-one correspondence between the N1 sound source direction angles and the N1 confidence levels, and N1 is a positive integer. In the above scheme, there is no limitation on the value of N1, there is a correspondence between the N1 sound source direction angles and the N1 confidence levels, for example, there is a one-to-one correspondence between the N1 sound source direction angles and the N1 confidence levels, and the confidence level can be used to represent the confidence level corresponding to the sound source direction angle predicted by the sound source direction prediction model. In the embodiment of the present application, the encoding end predicts the sound source direction angle and the corresponding confidence level through the sound source direction prediction model, so that the predicted sound source direction angle can be selected through the confidence level to determine the sound source direction information that meets the characteristics of the scene audio signal.

[0010] In a possible implementation, the N1 sound source direction angles include: N1 sound source direction angle pairs, wherein the sound source direction angle pairs include: a horizontal angle of the sound source direction and a pitch angle of the sound source direction. In the above scheme, the sound source direction of the scene audio signal described in the embodiment of the present application can be a direction angle in a spherical coordinate system. The encoding end uses N1 pairs of sound source direction angles, wherein a sound source direction angle pair includes: a horizontal angle of the sound source direction and a pitch angle of the sound source direction, that is, the sound source direction of the scene audio signal can be described by the horizontal angle of the sound source direction and the pitch angle of the sound source direction.

[0011] In a possible implementation, the sound source direction prediction result includes: N2 confidences, there is a one-to-one correspondence between the N2 confidences and the N2 sound source spherical direction angles, and N2 is a positive integer. In the above scheme, the encoding end outputs a number of confidences through the sound source direction prediction model. Specifically, N2 represents the number of confidences predicted by the sound source direction prediction model. Compared with the above embodiment, the output of the sound source direction prediction model in the embodiment of the present application is N2 confidences, which simplifies the output content of the sound source direction prediction model and reduces the size of the sound source direction prediction result. The encoding end pre-configures the correspondence between the confidence and the sound source spherical direction angle, that is, the encoding end can pre-set the sound source spherical direction angle, and the encoding end can determine the N2 sound source spherical direction angles through the N2 confidences output by the sound source direction prediction model, so that the predicted sound source direction angle can be selected through the N2 sound source spherical direction angles to determine the sound source direction information that meets the characteristics of the scene audio signal.

[0012] In a possible implementation, there is a one-to-one correspondence between the N2 confidences and the N2 index values; the N2 index values ​​are index values ​​corresponding to the N2 spherical direction angles of the sound source. In the above scheme, the encoding end outputs a number of confidences corresponding to the index values ​​through the sound source direction prediction model, and N2 represents the number of confidences predicted by the sound source direction prediction model. Compared with the previous embodiment, the output of the sound source direction prediction model in the embodiment of the present application is N2 confidences, and there is a one-to-one correspondence between the N2 confidences and the N2 index values, which simplifies the output content of the sound source direction prediction model and reduces the size of the sound source direction prediction result.

[0013] In a possible implementation, the sound source direction prediction result includes: the Cartesian coordinates of the sound source direction and the confidence corresponding to the Cartesian coordinates of the sound source direction. In the above scheme, the encoder predicts the Cartesian coordinates of the sound source direction and the corresponding confidence through the sound source direction prediction model, so that the predicted Cartesian coordinates of the sound source direction can be selected according to the confidence to determine the sound source direction information that meets the characteristics of the scene audio signal.

[0014] In a possible implementation, the determining the sound source direction information of the scene audio signal according to the sound source direction prediction result includes: obtaining sound source direction selection configuration information; determining the sound source direction information of the scene audio signal according to the sound source direction prediction result and the sound source direction selection configuration information. In the above scheme, the sound source direction information of the scene audio signal can be output more accurately by using the sound source direction selection configuration information and the sound source direction prediction model. The sound source direction information output after being selected by the sound source direction selection configuration information can improve the flexibility of scene audio signal encoding and improve encoding efficiency.

[0015] In a possible implementation, the sound source direction selection configuration information includes at least one of the following: a sound source direction number threshold for signal encoding, a sound source direction confidence threshold, and a configurable rule for selecting a sound source direction. In the above scheme, the sound source direction selection configuration information may include the sound source direction number threshold for signal encoding, or the sound source direction selection configuration information may include the sound source direction confidence threshold, or the sound source direction selection configuration information may include a configurable rule for selecting a sound source direction, or the sound source direction selection configuration information may include any combination of the above information.

[0016] In a possible implementation, the sound source direction prediction result includes: at least one confidence of a candidate; the sound source direction information of the scene audio signal is determined according to the sound source direction prediction result and the sound source direction selection configuration information, including: selecting M1 confidences from the at least one confidence of the candidate, wherein M1 represents the threshold of the number of sound source directions for signal encoding, and M1 is a positive integer; and determining that the sound source direction information includes M1 sound source direction angles corresponding to the M1 confidences. In the above scheme, the M1 sound source direction angles of the scene audio signal are determined according to the sound source direction prediction result and the sound source direction selection configuration information. The embodiment of the present application can flexibly and adaptively output the sound source direction information of the scene audio signal according to the real-time changes of the scene audio signal, which can improve the flexibility of the scene audio signal encoding and improve the encoding efficiency.

[0017] In a possible implementation, the sound source direction prediction result includes: at least one confidence of a candidate; the determining the sound source direction information of the scene audio signal based on the sound source direction prediction result and the sound source direction selection configuration information includes: selecting M2 confidences greater than or equal to the sound source direction confidence threshold from the at least one confidence of the candidate, wherein M2 is a positive integer; determining that the sound source direction information includes M2 sound source direction angles corresponding to the M2 confidences. In the above scheme, the M2 sound source direction angles of the scene audio signal are determined based on the sound source direction prediction result and the sound source direction selection configuration information. The embodiment of the present application can flexibly and adaptively output the sound source direction information of the scene audio signal according to the real-time changes of the scene audio signal, which can improve the flexibility of the scene audio signal encoding and improve the encoding efficiency.

[0018] In a possible implementation, the sound source direction prediction result includes: at least one confidence level of a candidate, where N is a positive integer; determining the sound source direction information of the scene audio signal based on the sound source direction prediction result and the sound source direction selection configuration information includes: selecting M3 confidence levels that meet the configurable rule for selecting the sound source direction from the at least one confidence level of the candidate, where M3 is a positive integer; determining that the sound source direction information includes M3 sound source direction angles corresponding to the M3 confidence levels. In the above scheme, the M3 sound source direction angles of the scene audio signal are determined based on the sound source direction prediction result and the sound source direction selection configuration information. The embodiment of the present application can flexibly and adaptively output the sound source direction information of the scene audio signal according to the real-time changes of the scene audio signal, which can improve the flexibility of scene audio signal encoding and improve encoding efficiency.

[0019] In a possible implementation, determining the sound source direction information of the scene audio signal according to the sound source direction prediction result includes: obtaining sound source direction selection configuration information; determining the sound source direction information of the scene audio signal according to the N2 confidences and the sound source direction selection configuration information. In the above scheme, the sound source direction prediction result includes N2 confidences, and the encoding end can determine the sound source direction information of the scene audio signal according to the N2 confidences and the sound source direction selection configuration information, and screen the N2 confidences included in the sound source direction prediction result according to the specific configuration requirements of the sound source direction selection configuration information to obtain the final sound source direction information of the scene audio signal. In the embodiment of the present application, the sound source direction selection configuration information and the sound source direction prediction model are used to more accurately output the sound source direction information of the scene audio signal, and the sound source direction information output after being selected by the sound source direction selection configuration information can improve the flexibility of the scene audio signal encoding and improve the encoding efficiency.

[0020] In a possible implementation, the sound source direction prediction result includes: at least one confidence of a candidate; the determining the sound source direction information of the scene audio signal according to the sound source direction prediction result and the sound source direction selection configuration information includes: selecting M4 confidences from at least one confidence of the candidate according to the sound source direction selection configuration information, wherein M4 is a positive integer; determining the sound source direction information includes M4 sound source direction Cartesian coordinates corresponding to the M4 confidences. In the above scheme, the M4 sound source direction Cartesian coordinates of the scene audio signal are determined according to the sound source direction prediction result and the sound source direction selection configuration information. The embodiment of the present application can flexibly and adaptively output the sound source direction information of the scene audio signal according to the real-time changes of the scene audio signal, which can improve the flexibility of the scene audio signal encoding and improve the encoding efficiency.

[0021] In a possible implementation, the sound source direction prediction result includes: at least one confidence of a candidate; the determining the sound source direction information of the scene audio signal according to the sound source direction prediction result and the sound source direction selection configuration information includes: selecting M5 confidences from at least one confidence of the candidate according to the sound source direction selection configuration information, wherein M5 is a positive integer; determining M5 direction indexes corresponding to the M5 confidences according to the correspondence between the confidence and the direction index; determining that the sound source direction information includes M5 sound source direction angles corresponding to the M5 direction indexes. In the above scheme, the M5 sound source direction angles of the scene audio signal are determined according to the sound source direction prediction result and the sound source direction selection configuration information. The embodiment of the present application can flexibly and adaptively output the sound source direction information of the scene audio signal according to the real-time changes of the scene audio signal, which can improve the flexibility of the scene audio signal encoding and improve the encoding efficiency.

[0022] In a possible implementation manner, the feature information of the scene audio signal includes at least one of the following: a time domain feature of the scene audio signal, a transform domain feature of the scene audio signal, or a feature obtained by using the time domain feature or the transform domain feature.

[0023] In one possible implementation, the transform domain features include at least one of the following: Fourier transform features, discrete cosine transform features, and discrete sine transform features; the features obtained through the time domain features or the transform domain features include at least one of the following: relative harmonic coefficient features, relative modal coherence features, and acoustic density features.

[0024] In a second aspect, an embodiment of the present application provides an audio encoding device, comprising: a feature acquisition module, used to acquire feature information of a scene audio signal to be encoded; a direction prediction module, used to input the feature information of the scene audio signal into a sound source direction prediction model, and output a sound source direction prediction result through the sound source direction prediction model; a direction determination module, used to determine the sound source direction information of the scene audio signal according to the sound source direction prediction result; and an encoding module, used to encode the scene audio signal according to the sound source direction information to obtain a first code stream.

[0025] The second aspect and any implementation of the second aspect correspond to the first aspect and any implementation of the first aspect respectively. The technical effects corresponding to the second aspect and any implementation of the second aspect can refer to the technical effects corresponding to the above-mentioned first aspect and any implementation of the first aspect, which will not be repeated here.

[0026] In a third aspect, an embodiment of the present application provides a method for generating a bitstream, which can generate a bitstream according to the first aspect and any one of the implementation methods of the first aspect.

[0027] The third aspect and any implementation of the third aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the third aspect and any implementation of the third aspect can refer to the technical effects corresponding to the first aspect and any implementation of the first aspect, which will not be repeated here.

[0028] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising: a memory and a processor, the memory being coupled to the processor; the memory storing program instructions, and when the program instructions are executed by the processor, the electronic device executes the audio encoding method in the first aspect or any possible implementation of the first aspect.

[0029] The fourth aspect and any implementation of the fourth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the fourth aspect and any implementation of the fourth aspect can refer to the technical effects corresponding to the above-mentioned first aspect and any implementation of the first aspect, which will not be repeated here.

[0030] In a fifth aspect, an embodiment of the present application provides a chip comprising one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from a memory of an electronic device and send a signal to the processor, the signal comprising a computer instruction stored in the memory; when the processor executes the computer instruction, the electronic device executes the audio encoding method in the first aspect or any possible implementation of the first aspect.

[0031] The fifth aspect and any implementation of the fifth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the fifth aspect and any implementation of the fifth aspect can refer to the technical effects corresponding to the first aspect and any implementation of the first aspect, which will not be repeated here.

[0032] In a sixth aspect, an embodiment of the present application provides an apparatus for generating a code stream, the apparatus comprising: a processor and at least one storage medium, the at least one storage medium being used to store the code stream, the processor executing the first aspect and any one implementation method of the first aspect to generate a first code stream.

[0033] The sixth aspect and any implementation of the sixth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the sixth aspect and any implementation of the sixth aspect can refer to the technical effects corresponding to the first aspect and any implementation of the first aspect, which will not be repeated here.

[0034] In the seventh aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a computer or a processor, the computer or the processor executes the audio encoding method in the first aspect or any possible implementation of the first aspect.

[0035] The seventh aspect and any implementation of the seventh aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the seventh aspect and any implementation of the seventh aspect can refer to the technical effects corresponding to the above-mentioned first aspect and any implementation of the first aspect, which will not be repeated here.

[0036] In an eighth aspect, an embodiment of the present application provides a computer program product, which includes a software program. When the software program is executed by a computer or a processor, the computer or the processor executes the audio encoding method in the first aspect or any possible implementation of the first aspect.

[0037] The eighth aspect and any implementation of the eighth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the eighth aspect and any implementation of the eighth aspect can refer to the technical effects corresponding to the first aspect and any implementation of the first aspect, which will not be repeated here.

[0038] In a ninth aspect, an embodiment of the present application provides a device for transmitting a code stream, the device comprising: a transmitter and at least one storage medium, the at least one storage medium is used to store the code stream, the code stream is generated according to the first aspect and any one of the implementation methods of the first aspect; the transmitter is used to obtain the code stream from the storage medium and send the code stream to the terminal side device through the transmission medium.

[0039] The ninth aspect and any implementation of the ninth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the ninth aspect and any implementation of the ninth aspect can refer to the technical effects corresponding to the first aspect and any implementation of the first aspect, which will not be repeated here.

[0040] In a tenth aspect, an embodiment of the present application provides a system for distributing code streams, the system comprising: at least one storage medium for storing at least one code stream, the at least one code stream being generated according to the first aspect and any one of the implementation methods of the first aspect, a streaming media device for obtaining a target code stream from the at least one storage medium and sending the target code stream to an end-side device, wherein the streaming media device comprises a content server or a content distribution server.

[0041] The tenth aspect and any implementation of the tenth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the tenth aspect and any implementation of the tenth aspect can refer to the technical effects corresponding to the above-mentioned first aspect and any implementation of the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 A schematic diagram of the structure of the audio processing system provided in the embodiment of the present application;

[0043] Figure 2a A schematic diagram of an audio encoder and an audio decoder provided in an embodiment of the present application being applied to a terminal device;

[0044] Figure 2b A schematic diagram of an audio encoder provided in an embodiment of the present application being applied to a wireless device or a core network device;

[0045] Figure 2c A schematic diagram of an audio decoder provided in an embodiment of the present application being applied to a wireless device or a core network device;

[0046] Figure 3a A schematic diagram of a multi-channel encoder and a multi-channel decoder provided in an embodiment of the present application applied to a terminal device;

[0047] Figure 3b A schematic diagram of a multi-channel encoder provided in an embodiment of the present application applied to a wireless device or a core network device;

[0048] Figure 3c A schematic diagram of a multi-channel decoder provided in an embodiment of the present application applied to a wireless device or a core network device;

[0049] Figure 4 A flowchart of an audio encoding method provided in an embodiment of the present application;

[0050] Figure 5 An application scenario architecture diagram of an encoding end for an audio encoding method provided in an embodiment of the present application;

[0051] Figure 6An application scenario architecture diagram of an encoding end for an audio encoding method application is provided for an embodiment of the present application;

[0052] Figure 7 A schematic diagram of the neural network model provided in an embodiment of the present application outputting N sound source direction angles and the confidence levels corresponding to the sound source direction angles;

[0053] Figure 8 A schematic diagram of a sound source direction selection module provided in an embodiment of the present application performing sound source direction selection;

[0054] Fig. 9 An application scenario architecture diagram of an encoding end for another audio encoding method application is provided for an embodiment of the present application;

[0055] Fig.10 A schematic diagram of a neural network model outputting N confidence levels provided in an embodiment of the present application;

[0056] Fig.11 A schematic diagram of a sound source direction selection module provided in an embodiment of the present application performing sound source direction selection;

[0057] Fig.12 An application scenario architecture diagram of an encoding end for another audio encoding method application is provided for an embodiment of the present application;

[0058] Fig.13 A schematic diagram of the neural network model provided in an embodiment of the present application outputting N sound source direction angles and the confidence levels corresponding to the sound source direction angles;

[0059] Fig.14 A schematic diagram of a neural network model outputting N confidence levels provided in an embodiment of the present application;

[0060] Fig.15 A schematic diagram of a sound source direction selection module provided in an embodiment of the present application performing sound source direction selection;

[0061] Fig.16 A schematic diagram of the structure of an audio encoding device provided in an embodiment of the present application;

[0062] Fig.17 A schematic diagram of the composition structure of another audio encoding device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0063] The embodiments of the present application provide an audio encoding method and apparatus and an electronic device for improving encoding efficiency.

[0064] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0065] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0066] The terms "first" and "second" in the description and claims of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order of objects. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, which is merely a way of distinguishing objects with the same attributes when describing the embodiments of the present application. For example, the first target object and the second target object are used to distinguish different target objects, rather than to describe a specific order of target objects.

[0067] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0068] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "multiple" refers to two or more than two. For example, multiple processing units refer to two or more processing units; multiple systems refer to two or more systems.

[0069] Furthermore, the terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover a non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a list of elements is not necessarily limited to those elements but may include other elements not expressly listed or inherent to such process, method, product, or apparatus.

[0070] In order to make the description of the following embodiments clear and concise, a brief introduction to the related technology is first given.

[0071] Sound is a continuous wave generated by the vibration of an object. The object that vibrates and emits sound waves is called a sound source. When sound waves propagate through a medium (such as air, solid or liquid), the auditory organs of humans or animals can perceive the sound.

[0072] The characteristics of sound waves include pitch, intensity and timbre. Pitch refers to the high or low pitch of a sound. Intensity refers to the size of a sound. Intensity can also be called loudness or volume. The unit of intensity is decibel (dB). Timbre is also called timbre quality.

[0073] The frequency of sound waves determines the pitch of the sound. The higher the frequency, the higher the pitch. The number of times an object vibrates in one second is called frequency, and the unit of frequency is hertz (Hz). The frequency of sound that the human ear can recognize is between 20Hz and 20,000Hz.

[0074] The amplitude of the sound wave determines the intensity of the sound. The greater the amplitude, the greater the sound intensity. The closer to the sound source, the greater the sound intensity.

[0075] The waveform of the sound wave determines the timbre, and the waveforms of the sound wave include square wave, sawtooth wave, sine wave and pulse wave.

[0076] According to the characteristics of sound waves, sounds can be divided into regular sounds and irregular sounds. Irregular sounds refer to sounds produced by irregular vibrations of the sound source. Irregular sounds are, for example, noises that affect people's work, study, and rest. Regular sounds refer to sounds produced by regular vibrations of the sound source. Regular sounds include speech and music. When sound is represented by electricity, regular sounds are analog signals that continuously change in the time-frequency domain. This analog signal can be called an audio signal. An audio signal is an information carrier that carries speech, music, and sound effects.

[0077] Since human hearing has the ability to distinguish the positional distribution of sound sources in space, when listeners hear sounds in space, in addition to being able to feel the pitch, volume and timbre of the sounds, they can also feel the direction of the sounds.

[0078] As people pay more and more attention to the experience of auditory systems and demand more quality, three-dimensional audio technology has emerged to enhance the depth, presence and spatial sense of sound. Therefore, listeners can not only feel the sound from the front, back, left and right sound sources, but also feel the space they are in is surrounded by the spatial sound field (referred to as "sound field") generated by these sound sources, and the sound spreads around, creating an "immersive" sound effect that makes listeners feel like they are in a theater or concert hall.

[0079] The scene audio signal involved in the embodiments of the present application may refer to a signal used to describe a sound field; wherein the scene audio signal may include: an HOA signal (wherein the HOA signal may include a three-dimensional HOA signal and a two-dimensional HOA signal (also referred to as a planar HOA signal)) and a three-dimensional audio signal; the three-dimensional audio signal may refer to other audio signals in the scene audio signal except the HOA signal, and the HOA signal is used as an example for explanation below.

[0080] The scene audio signal is an information carrier that carries the spatial position information of the sound source in the sound field, and describes the sound field of the listener in space. Formula (4) shows that the sound field can be expanded on the sphere according to spherical harmonics, that is, the sound field can be decomposed into the superposition of multiple plane waves. The scene audio signal can also be called a three-dimensional stereo signal, that is, an Ambisonics signal. For example, the scene audio signal can include an HOA signal. Therefore, the sound field described by the HOA signal can be expressed by the superposition of multiple plane waves, and the sound field can be reconstructed by the HOA coefficients.

[0081] The HOA signal to be encoded involved in the embodiments of the present application may refer to an N-order HOA signal, which may be represented by an HOA coefficient or an Ambisonic (stereo reverberation) coefficient, where N is an integer greater than or equal to 1 (wherein, when N is equal to, the first-order HOA signal may be referred to as a FOA (First Order Ambisonic, first-order stereo reverberation) signal). The N-order HOA signal includes (N+1) 2 channels of audio signal.

[0082] The technical solution of the embodiment of the present application can be applied to various audio processing systems, such as Figure 1 As shown, it is a schematic diagram of the composition structure of the audio processing system provided by the embodiment of the present application. The audio processing system 100 may include: an audio encoding device 101 and an audio decoding device 102. Among them, the audio encoding device 101 can be used to generate an audio encoding stream, and then the audio encoding stream can be transmitted to the audio decoding device 102 through an audio transmission channel. The audio decoding device 102 can receive the audio encoding stream, and then perform the audio decoding function of the audio decoding device 102, and finally obtain a reconstructed signal.

[0083] In an embodiment of the present application, the audio encoding device can be applied to various terminal devices with audio communication needs, wireless devices and core network devices with transcoding needs, for example, the audio encoding device can be an audio encoder of the above-mentioned terminal device or wireless device or core network device. Similarly, the audio decoding device can be applied to various terminal devices with audio communication needs, wireless devices and core network devices with transcoding needs, for example, the audio decoding device can be an audio decoder of the above-mentioned terminal device or wireless device or core network device. For example, the audio encoder can include a wireless access network, a media gateway of the core network, a transcoding device, a media resource server, a mobile terminal, a fixed network terminal, etc. The audio encoder can also be an audio codec used in virtual reality (VR) streaming media services.

[0084] In the application embodiment, taking the audio encoding and decoding module (audio encoding and audio decoding) suitable for virtual reality streaming media (VR streaming) service as an example, the end-to-end processing flow of the audio signal includes: the audio signal A passes through the acquisition module (acquisition) and then performs a preprocessing operation (audio preprocessing), the preprocessing operation includes filtering out the low-frequency part of the signal, which can be based on 20Hz or 50Hz as the dividing point, extracting the orientation information in the signal, and then performing encoding processing (audio encoding) and packaging (file / segment encapsulation) and sending (delivery) to the decoding end, the decoding end first unpacks (file / segment decapsulation), and then decodes (audio decoding), and performs binaural rendering (audio rendering) on ​​the decoded signal. The rendered signal is mapped to the listener's headphones (headphones), which can be independent headphones or headphones on eyewear devices.

[0085] like Figure 2a As shown, it is a schematic diagram of the audio encoder and audio decoder provided by the embodiment of the present application being applied to a terminal device. For each terminal device, it can include: an audio encoder, a channel encoder, an audio decoder, and a channel decoder. Specifically, the channel encoder is used to perform channel encoding on the audio signal, and the channel decoder is used to perform channel decoding on the audio signal. For example, in the first terminal device 20, it can include: a first audio encoder 201, a first channel encoder 202, a first audio decoder 203, and a first channel decoder 204. In the second terminal device 21, it can include: a second audio decoder 211, a second channel decoder 212, a second audio encoder 213, and a second channel encoder 214. The first terminal device 20 is connected to a wireless or wired first network communication device 22, and the first network communication device 22 and the wireless or wired second network communication device 23 are connected through a digital channel, and the second terminal device 21 is connected to the wireless or wired second network communication device 23. Among them, the above-mentioned wireless or wired network communication device can generally refer to a signal transmission device, such as a communication base station, a data exchange device, etc.

[0086] In audio communication, the terminal device at the sending end first collects audio, encodes the collected audio signal, and then transmits it in a digital channel through a wireless network or core network after channel encoding. The terminal device at the receiving end performs channel decoding based on the received signal to obtain a bit stream, and then restores the audio signal through audio decoding, which is then played back by the terminal device at the receiving end.

[0087] like Figure 2b As shown, it is a schematic diagram of the audio encoder provided in the embodiment of the present application applied to a wireless device or a core network device. Among them, the wireless device or core network device 25 includes: a channel decoder 251, other audio decoders 252, an audio encoder 253 provided in the embodiment of the present application, and a channel encoder 254, wherein other audio decoders 252 refer to other audio decoders other than the audio decoder. In the wireless device or core network device 25, the signal entering the device is first channel-decoded by the channel decoder 251, and then the other audio decoder 252 is used to perform audio decoding, and then the audio encoder 253 provided in the embodiment of the present application is used to perform audio encoding, and finally the channel encoder 254 is used to perform channel encoding on the audio signal, and then the channel encoding is completed before it is transmitted. Among them, the other audio decoder 252 performs audio decoding on the code stream decoded by the channel decoder 251.

[0088] like Figure 2c As shown, it is a schematic diagram of the audio decoder provided in the embodiment of the present application being applied to a wireless device or a core network device. Among them, the wireless device or core network device 25 includes: a channel decoder 251, an audio decoder 255 provided in the embodiment of the present application, other audio encoders 256, and a channel encoder 254, wherein other audio encoders 256 refer to other audio encoders other than the audio encoder. In the wireless device or core network device 25, the signal entering the device is first channel-decoded by the channel decoder 251, and then the received audio coding stream is decoded using the audio decoder 255, and then the other audio encoder 256 is used for audio encoding, and finally the channel encoder 254 is used to channel-encode the audio signal, and then the channel encoding is completed before it is transmitted. In the wireless device or core network device, if transcoding needs to be implemented, the corresponding audio codec processing needs to be performed. Among them, the wireless device refers to the radio frequency-related device in the communication, and the core network device refers to the core network-related device in the communication.

[0089] In some embodiments of the present application, the audio encoding device can be applied to various terminal devices with audio communication needs, wireless devices and core network devices with transcoding needs, for example, the audio encoding device can be a multi-channel encoder of the above terminal device or wireless device or core network device. Similarly, the audio decoding device can be applied to various terminal devices with audio communication needs, wireless devices and core network devices with transcoding needs, for example, the audio decoding device can be a multi-channel decoder of the above terminal device or wireless device or core network device.

[0090] like Figure 3aAs shown, it is a schematic diagram of a multi-channel encoder and a multi-channel decoder provided in an embodiment of the present application applied to a terminal device, and each terminal device may include: a multi-channel encoder, a channel encoder, a multi-channel decoder, and a channel decoder. The multi-channel encoder can execute the audio encoding method provided in an embodiment of the present application, and the multi-channel decoder can execute the audio decoding method provided in an embodiment of the present application. Specifically, the channel encoder is used to perform channel encoding on a multi-channel signal, and the channel decoder is used to perform channel decoding on a multi-channel signal. For example, in the first terminal device 30, it may include: a first multi-channel encoder 301, a first channel encoder 302, a first multi-channel decoder 303, and a first channel decoder 304. In the second terminal device 31, it may include: a second multi-channel decoder 311, a second channel decoder 312, a second multi-channel encoder 313, and a second channel encoder 314. The first terminal device 30 is connected to a wireless or wired first network communication device 32, and the first network communication device 32 and the wireless or wired second network communication device 33 are connected through a digital channel, and the second terminal device 31 is connected to a wireless or wired second network communication device 33. The wireless or wired network communication equipment mentioned above may generally refer to signal transmission equipment, such as communication base stations, data exchange equipment, etc. In audio communication, the terminal device as the transmitter performs multi-channel encoding on the collected multi-channel signal, and then performs channel encoding, and then transmits it in a digital channel through a wireless network or a core network. The terminal device as the receiving end performs channel decoding according to the received signal to obtain a multi-channel signal encoding stream, and then recovers the multi-channel signal through multi-channel decoding, and then plays it back by the terminal device as the receiving end.

[0091] like Figure 3b As shown, it is a schematic diagram of a multi-channel encoder provided in an embodiment of the present application applied to a wireless device or a core network device, wherein the wireless device or the core network device 35 includes: a channel decoder 351, other audio decoders 352, a multi-channel encoder 353, and a channel encoder 354, which are the same as the aforementioned Figure 2b Similar, no further description is given here.

[0092] like Figure 3c As shown, it is a schematic diagram of a multi-channel decoder provided in an embodiment of the present application applied to a wireless device or a core network device, wherein the wireless device or the core network device 35 includes: a channel decoder 351, a multi-channel decoder 355, other audio encoders 356, and a channel encoder 354, which are the same as the aforementioned Figure 2c Similar, no further description is given here.

[0093] Among them, the audio encoding process can be a part of a multi-channel encoder, and the audio decoding process can be a part of a multi-channel decoder. For example, multi-channel encoding of the collected multi-channel signal can be to obtain an audio signal after processing the collected multi-channel signal, and then encode the obtained audio signal according to the method provided in the embodiment of the present application; the decoding end encodes the code stream according to the multi-channel signal, decodes the audio signal, and restores the multi-channel signal after upmixing. Therefore, the embodiment of the present application can also be applied to multi-channel encoders and multi-channel decoders in terminal devices, wireless devices, and core network devices. In wireless or core network devices, if transcoding needs to be implemented, corresponding multi-channel encoding and decoding processing is required.

[0094] In order to achieve better audio auditory effects, Ambisonics technology requires a large amount of data to record more detailed information about the sound scene. As the order of the Ambisonics signal increases, more data will be generated. A large amount of data causes difficulties in transmission and storage, so it is necessary to encode and decode the Ambisonics signal. The coding technology for the Ambisonics signal needs to consider the information redundancy between channels, and the Ambisonics signal can be effectively compressed and encoded. By analyzing the data of each channel of the Ambisonics signal, the directional signal (or foreground signal) and its directional information and background signal and auxiliary side information are extracted, and the Ambisonics signal can be efficiently compressed and encoded. In the Ambisonics signal coding technology, the Ambisonics signal is converted into an actual speaker signal for playback, or the Ambisonics signal is converted into a virtual speaker signal and then mapped to a binaural signal for playback. In this process, an important step is to convert the Ambisonics signal into a (virtual) speaker signal, wherein extracting the directional information of the Ambisonics signal is a core link in such technology. The embodiment of the present application provides an Ambisonics signal coding scheme based on extracting signal directional information using a machine learning model, which is described in detail below.

[0095] The audio encoding and decoding method provided in the embodiment of the present application may include: an audio encoding method and an audio decoding method, wherein the audio encoding method is performed by an audio encoding device, and the audio decoding method is performed by an audio decoding device, and the audio encoding device and the audio decoding device can communicate with each other. Figure 4 As shown in FIG. 1 , it is a schematic diagram of an audio encoding method performed by an audio encoding device in an embodiment of the present application, wherein the following steps 401 to 404 can be performed by an audio encoding device (hereinafter referred to as an encoding end), and the audio decoding method performed by an audio decoding device (hereinafter referred to as a decoding end) corresponds to the audio encoding method. Figure 4As shown, the audio encoding method performed by the encoding end mainly includes the following processes:

[0096] 401. Obtain feature information of a scene audio signal to be encoded.

[0097] The encoding end obtains a scene audio signal to be encoded, where the scene audio signal refers to an audio signal obtained by collecting the sound field at the position of the microphone in the space, and the scene audio signal can also be called an original scene audio signal. For example, the scene audio signal can be an Ambisonics signal. Specifically, the scene audio signal can be an audio signal obtained by using a higher order ambisonics (HOA) technology.

[0098] The encoding end obtains feature information of the scene audio signal, which can also be called signal feature information. The feature information of the scene audio signal can include various types of features of the scene audio signal. The embodiment of the present application does not limit the feature extraction method of the scene audio signal.

[0099] In some embodiments of the present application, the feature information of the scene audio signal includes at least one of the following: a time domain feature of the scene audio signal, a transform domain feature of the scene audio signal, or a feature obtained through the time domain feature or the transform domain feature.

[0100] Among them, the time domain features of the scene audio signal refer to the features obtained by the encoder extracting the features of the scene audio signal in the time domain. The transform domain features of the scene audio signal refer to the features obtained by transforming the time domain features of the scene audio signal according to a feature transformation method. The feature transformation methods may include multiple methods, such as Fourier transform, discrete cosine transform, etc. Through the above feature transformation methods, the time domain features can be transformed into transform domain features. In addition, in the embodiment of the present application, the feature information of the scene audio signal may also include other features that can characterize the scene audio signal in addition to the time domain features or frequency domain features. For example, the feature information of the scene audio signal may also include features obtained through time domain features or transform domain features. The specific method of transforming the time domain features or transform domain features is not limited in the embodiment of the present application. In the embodiment of the present application, the feature information of the scene audio signal can be accurately obtained through multiple implementation methods of the feature information.

[0101] Furthermore, in some embodiments of the present application, the transform domain feature includes at least one of the following: Fourier transform feature, discrete cosine transform feature, and discrete sine transform feature.

[0102] Among them, the encoding end can transform the time domain characteristics of the scene audio signal through any one of Fourier transform, discrete cosine transform, and discrete sine transform, so as to obtain the corresponding Fourier transform characteristics, discrete cosine transform characteristics, and discrete sine transform characteristics. The calculation methods of Fourier transform, discrete cosine transform, and discrete sine transform are not described in detail here.

[0103] The features obtained through time domain features or transform domain features include at least one of the following: relative harmonic coefficients (RHC) features, relative modal coherence (RMC) features, and acoustic intensity features.

[0104] Specifically, after obtaining the transform domain features of the scene audio signal, the encoder can derive one or more features from the transform domain features as feature information of the scene audio signal. For example, the features derived from the transform domain features by the encoder may include at least one of the following: relative harmonic coefficient features, relative modal coherence features, and acoustic density features. In the embodiment of the present application, the feature information of the scene audio signal can be accurately obtained through multiple implementation methods of the feature information.

[0105] 402. Input feature information of the scene audio signal into a sound source direction prediction model, and output a sound source direction prediction result through the sound source direction prediction model.

[0106] After the encoder obtains the characteristic information of the scene audio signal, the encoder inputs the characteristic information of the scene audio signal into the sound source direction prediction model. The sound source direction prediction model can be a pre-trained machine learning model. In the embodiment of the present application, the training process of the sound source direction prediction model, the adopted machine learning algorithm and the training samples are not described in detail. In the subsequent embodiments, the sound source direction prediction model is specifically illustrated as a neural network model.

[0107] The sound source direction prediction model after pre-training can predict the sound source direction of the scene audio signal, so that the sound source direction prediction model can output the sound source direction prediction result, and the sound source direction prediction result can include one or more sound source directions of the scene audio signal predicted by the sound source direction prediction model. The encoding end provided in the embodiment of the present application can pre-configure the sound source direction prediction model, and the encoding end inputs the characteristic information of the scene audio signal into the sound source direction prediction model. Since the characteristic information of the scene audio signal changes in real time with the scene audio signal, that is, the characteristic information corresponding to different scene audio signals is different, the sound source direction is predicted by combining the characteristic information of the scene audio signal with the sound source direction prediction model, so that the output sound source direction prediction result can indicate the sound source direction that conforms to the scene audio signal. The embodiment of the present application can adaptively and flexibly predict the sound source direction of the scene audio signal according to the real-time changes of the scene audio signal, which can improve the flexibility of scene audio signal encoding and improve encoding efficiency.

[0108] In some embodiments of the present application, the sound source direction prediction result includes: the sound source direction angle and the confidence level corresponding to the sound source direction angle.

[0109] Among them, the encoding end outputs the sound source direction prediction result through the sound source direction prediction model. The encoding end can predict the sound source direction angle of the scene audio signal through the sound source direction prediction model, and output the confidence corresponding to the predicted sound source direction angle, and the confidence indicates the credibility range of the sound source direction prediction model for the predicted sound source direction angle that meets the characteristics of the scene audio signal. In the embodiment of the present application, the encoding end predicts the sound source direction angle and the corresponding confidence through the sound source direction prediction model, so that the predicted sound source direction angle can be selected through the confidence to determine the sound source direction information that meets the characteristics of the scene audio signal.

[0110] In some embodiments of the present application, the encoding end outputs a number of sound source direction angles and confidence levels corresponding to the sound source direction angles through a sound source direction prediction model, where N represents the number of sound source direction angles predicted by the sound source direction prediction model, the number of sound source direction angles is N, and the confidence level is N. The value of N in different application scenarios is not limited. For example, in one scenario, the value of N is N1, and in another scenario, the value of N is N2, and so on. N1 and N2 respectively refer to the number of sound source direction angles determined in different scenarios. In addition, each sound source direction angle output by the sound source direction prediction model corresponds to a confidence level, and the confidence level ranges from [0, 1]. There are N confidence levels, and the value of N is not limited in different application scenarios. For example, in one scenario, the value of N is N1, and N1 sound source direction angles correspond to N1 confidence levels. In another scenario, the value of N is N2, and N2 sound source direction angles correspond to N2 confidence levels. N1 and N2 respectively refer to the number of confidence levels determined in different scenarios, and there is no limitation on the specific value of the confidence level predicted by the sound source direction prediction model. It is understandable that the specific values ​​of N1 and N2 are not limited, and N1 can be equal to N2, or N1 is not equal to N2, depending on the application scenario.

[0111] In some embodiments of the present application, the number of sound source direction angles is N1, and the number of confidence levels is N1.

[0112] There is a one-to-one correspondence between N1 sound source direction angles and N1 confidence levels, where N1 is a positive integer.

[0113] There is no limitation on the value of N1, and there is a corresponding relationship between the N1 sound source direction angles and the N1 confidences, for example, there is a one-to-one correspondence between the N1 sound source direction angles and the N1 confidences, and the confidence can be used to represent the confidence corresponding to the sound source direction angle predicted by the sound source direction prediction model. In the embodiment of the present application, the encoding end predicts the sound source direction angle and the corresponding confidence by the sound source direction prediction model, so that the predicted sound source direction angle can be selected by the confidence to determine the sound source direction information that meets the characteristics of the scene audio signal.

[0114] Further, in some embodiments of the present application, the N1 sound source direction angles include: N1 sound source direction angle pairs,

[0115] in,

[0116] The sound source direction angle pair includes: the sound source direction horizontal angle and the sound source direction elevation angle.

[0117] Among them, the sound source direction of the scene audio signal described in the embodiment of the present application can be a direction angle in a spherical coordinate system. The encoding end uses N1 pairs of sound source direction angles, wherein a sound source direction angle pair includes: a horizontal angle of the sound source direction and a pitch angle of the sound source direction, that is, the sound source direction of the scene audio signal can be described by the horizontal angle of the sound source direction and the pitch angle of the sound source direction. For example, the range of the horizontal angle of the sound source direction is [0, 360], and the range of the pitch angle of the sound source direction is [0, 180]. There is no limitation on the specific values ​​of the horizontal angle of the sound source direction and the pitch angle of the sound source direction predicted by the sound source direction prediction model.

[0118] In some embodiments of the present application, the sound source direction prediction result includes: N2 confidence levels, and there is a one-to-one correspondence between the N2 confidence levels and the N2 sound source spherical direction angles.

[0119] Among them, the encoding end outputs a number of confidences through the sound source direction prediction model. Specifically, N2 represents the number of confidences predicted by the sound source direction prediction model. Compared with the previous embodiment, the output of the sound source direction prediction model in the embodiment of the present application is N2 confidences, which simplifies the output content of the sound source direction prediction model and reduces the size of the sound source direction prediction result. The encoding end pre-configures the corresponding relationship between the confidence and the spherical direction angle of the sound source, that is, the encoding end can pre-set the spherical direction angle of the sound source. The encoding end can determine the N2 spherical direction angles of the sound source through the N2 confidences output by the sound source direction prediction model, so that the predicted sound source direction angle can be selected through the N2 sound source spherical direction angles to determine the sound source direction information that meets the characteristics of the scene audio signal.

[0120] Further, in some embodiments of the present application, there is a one-to-one correspondence between the N2 confidences and the N2 index values;

[0121] The N2 index values ​​are index values ​​corresponding to the spherical direction angles of the N2 sound sources.

[0122] Among them, the encoding end outputs a number of confidences corresponding to the index values ​​through the sound source direction prediction model, and N2 represents the number of confidences predicted by the sound source direction prediction model. Compared with the previous embodiment, the output of the sound source direction prediction model in the embodiment of the present application is N2 confidences, and there is a one-to-one correspondence between the N2 confidences and the N2 index values, which simplifies the output content of the sound source direction prediction model and reduces the size of the sound source direction prediction result. The N2 confidences output by the sound source direction prediction model correspond to N index values, and these N2 index values ​​are index values ​​of the spherical direction angles of the N2 sound sources.

[0123] An example is given below. The spherical direction angle preset by the decoding end represents the uniform or non-uniform extraction of N2 sound source positions on the sphere, where N2≥1, and each sound source position corresponds to a set of preset sound source directions and an index corresponding thereto. The value of the index is [1, N2]. Each index corresponds to a preset sound source direction, and the encoding end outputs the confidence corresponding to the N2 indexes through the sound source direction prediction module.

[0124] Furthermore, in some other embodiments of the present application, the sound source direction prediction result includes: Cartesian coordinates of the sound source direction and confidence levels corresponding to the Cartesian coordinates of the sound source direction.

[0125] Among them, the encoding end outputs the sound source direction prediction result through the sound source direction prediction model. The encoding end can predict the Cartesian coordinates of the sound source direction of the scene audio signal through the sound source direction prediction model, and output the confidence corresponding to the predicted Cartesian coordinates of the sound source direction. For example, the Cartesian coordinates of the sound source direction output by the sound source direction prediction model are (X, Y, Z), and the value range of (X, Y, Z) is X∈[-1, 1], Y∈[-1, 1], Z∈[-1, 1]. The confidence represents the range of credibility of the sound source direction prediction model for the predicted Cartesian coordinates of the sound source direction that conforms to the characteristics of the scene audio signal. In the embodiment of the present application, the encoding end predicts the Cartesian coordinates of the sound source direction and the corresponding confidence through the sound source direction prediction model, so that the predicted Cartesian coordinates of the sound source direction can be selected according to the confidence to determine the sound source direction information that conforms to the characteristics of the scene audio signal.

[0126] 403. Determine sound source direction information of the scene audio signal according to the sound source direction prediction result.

[0127] After the encoding end outputs the sound source direction prediction result through the sound source direction prediction model, the encoding end parses the sound source direction prediction result and determines the sound source direction information of the scene audio signal from the sound source direction prediction result. The sound source direction information is the sound source direction used when encoding the scene audio signal. The sound source direction information is obtained through the sound source direction prediction of the sound source direction prediction model, and the sound source direction information matches the scene audio signal.

[0128] For example, the encoding end may include a sound source direction selection module, which selects the sound source direction for the scene audio signal to determine the sound source direction information of the scene audio signal. In the embodiment of the present application, the sound source direction of the scene audio signal is not fixed, but the sound source direction is predicted according to the sound source direction prediction model combined with the characteristic information of the scene audio signal, so as to determine the sound source direction information of the scene audio signal according to the sound source direction prediction result. The embodiment of the present application can flexibly and adaptively output the sound source direction information of the scene audio signal according to the real-time changes of the scene audio signal, which can improve the flexibility of the scene audio signal encoding and improve the encoding efficiency.

[0129] In some embodiments of the present application, step 403 determines the sound source direction information of the scene audio signal according to the sound source direction prediction result, including:

[0130] Get the sound source direction selection configuration information;

[0131] The sound source direction information of the scene audio signal is determined according to the sound source direction prediction result and the sound source direction selection configuration information.

[0132] Among them, the encoding end can obtain the sound source direction selection configuration information, and the sound source direction selection configuration information is used to determine the sound source direction for the scene audio signal in combination with the sound source direction prediction result output by the sound source direction selection module. The sound source direction selection configuration information can be configuration information pre-stored in the encoding end, and the sound source direction selection configuration information may include predetermined sound source direction selection parameters and sound source direction selection rules. The encoding end can determine the sound source direction information of the scene audio signal based on the sound source direction prediction result and the sound source direction selection configuration information, and screen the sound source direction prediction result according to the specific configuration requirements of the sound source direction selection configuration information to obtain the final sound source direction information of the scene audio signal. In the embodiment of the present application, the sound source direction selection configuration information and the sound source direction prediction model are used to more accurately output the sound source direction information of the scene audio signal. The sound source direction information output after the sound source direction selection configuration information is selected can improve the flexibility of the scene audio signal encoding and improve the encoding efficiency.

[0133] Furthermore, in some embodiments of the present application, the sound source direction selection configuration information includes at least one of the following: a sound source direction number threshold for signal encoding, a sound source direction confidence threshold, and a configurable rule for selecting a sound source direction.

[0134] Among them, the sound source direction selection configuration information may include a threshold value for the number of sound source directions used for signal encoding, or the sound source direction selection configuration information may include a sound source direction confidence threshold, or the sound source direction selection configuration information may include a configurable rule for selecting the sound source direction, or the sound source direction selection configuration information may include any combination of the above information.

[0135] The sound source direction number threshold may be used to screen the number of sound source directions of the scene audio signal. For example, the sound source direction number threshold refers to the number of sound source direction information that can be processed by the signal encoding module included in the encoding end.

[0136] The sound source direction confidence threshold can be used to screen the confidence corresponding to the sound source direction of the scene audio signal. For example, assuming that the value of the confidence threshold T is set to 0.8, the sound source direction angle with a confidence greater than or equal to 0.8 is finally selected.

[0137] The configurable rule for selecting the sound source direction may include a flexibly configurable rule for selecting the sound source direction input to the encoding end. For example, the configurable rule for selecting the sound source direction includes configuration information capable of selecting the sound source direction, for example, the user sets the number of output sound source directions, and this number is less than or equal to the number of sound source direction angles output by the sound source direction prediction model.

[0138] In some embodiments of the present application, the sound source direction prediction result includes: at least one confidence level of the candidate;

[0139] Determining the sound source direction information of the scene audio signal according to the sound source direction prediction result and the sound source direction selection configuration information includes:

[0140] Selecting M1 confidence levels from at least one candidate confidence level, where M1 represents a threshold value of the number of sound source directions used for signal coding, and M1 is a positive integer;

[0141] The sound source direction information is determined to include M1 sound source direction angles corresponding to M1 confidence levels.

[0142] Among them, the sound source direction prediction result includes at least one candidate confidence, and the at least one candidate confidence is a candidate confidence predicted by the sound source direction prediction model. For example, the at least one candidate confidence may include N1 confidences in the aforementioned embodiment, or the at least one candidate confidence may include N2 confidences in the aforementioned embodiment. For example, the encoding end determines that the sound source direction angles are N1, and the N1 sound source direction angles correspond to N1 confidences, and the at least one candidate confidence includes N1 confidences. For another example, the encoding end determines that the sound source direction angles are N2, and the N2 sound source direction angles correspond to N2 confidences, and the at least one candidate confidence includes N2 confidences.

[0143] The sound source direction selection configuration information may include a sound source direction number threshold for signal encoding, for example, M1 represents the sound source direction number threshold for signal encoding, the encoding end selects M1 confidences from at least one candidate confidence, for example, these M1 confidences include: the first M1 confidences selected from at least one candidate confidence from high to low, there is a one-to-one correspondence between the confidence and the sound source direction angle, and the M1 sound source direction angles are determined according to the M1 confidences. In the embodiment of the present application, the M1 sound source direction angles of the scene audio signal are determined according to the sound source direction prediction result and the sound source direction selection configuration information. The embodiment of the present application can flexibly and adaptively output the sound source direction information of the scene audio signal according to the real-time changes of the scene audio signal, which can improve the flexibility of the scene audio signal encoding and improve the encoding efficiency.

[0144] An example is given as follows: the encoding end includes a signal encoding module and a sound source direction selection module. When the sound source direction selection configuration information is set to a threshold M1 of the number of sound source direction information that the signal encoding module can process, the sound source direction selection module selects the sound source direction angles corresponding to the first M1 maximum confidences based on the confidence of the sound source direction.

[0145] In some embodiments of the present application, the sound source direction prediction result includes: at least one confidence level of the candidate;

[0146] Determining the sound source direction information of the scene audio signal according to the sound source direction prediction result and the sound source direction selection configuration information includes:

[0147] Selecting M2 confidences greater than or equal to the sound source direction confidence threshold from at least one candidate confidence, where M2 is a positive integer;

[0148] The sound source direction information is determined to include M2 ​​sound source direction angles corresponding to M2 confidence levels.

[0149] Among them, the encoding end outputs at least one confidence of the candidate through the sound source direction prediction model, the sound source direction selection configuration information may include a sound source direction confidence threshold, the encoding end selects M2 confidences greater than or equal to the sound source direction confidence threshold from the at least one confidence of the candidate, there is a one-to-one correspondence between the confidence and the sound source direction angle, and the M2 sound source direction angles are determined according to the M2 confidences. In the embodiment of the present application, the M2 sound source direction angles of the scene audio signal are determined according to the sound source direction prediction result and the sound source direction selection configuration information. The embodiment of the present application can flexibly and adaptively output the sound source direction information of the scene audio signal according to the real-time changes of the scene audio signal, which can improve the flexibility of the scene audio signal encoding and improve the encoding efficiency.

[0150] An example is given as follows: the encoding end includes a sound source direction selection module. When the sound source direction selection configuration information is set to the sound source direction confidence threshold T, the sound source direction selection module selects the sound source direction angle whose confidence is greater than or equal to T.

[0151] In some embodiments of the present application, the sound source direction prediction result includes: at least one confidence level of the candidate;

[0152] Determining the sound source direction information of the scene audio signal according to the sound source direction prediction result and the sound source direction selection configuration information includes:

[0153] Selecting M3 confidences that meet a configurable rule for selecting a sound source direction from at least one candidate confidence, where M3 is a positive integer;

[0154] The sound source direction information is determined to include M3 sound source direction angles corresponding to M3 confidence levels.

[0155] Among them, the encoding end outputs at least one confidence of the candidate through the sound source direction prediction model, the sound source direction selection configuration information may include a configurable rule for selecting the sound source direction, the encoding end selects M3 confidences that meet the configurable rule for selecting the sound source direction from the at least one confidence of the candidate, there is a one-to-one correspondence between the confidence and the sound source direction angle, and the M3 sound source direction angles are determined according to the M3 confidences. In the embodiment of the present application, the M3 sound source direction angles of the scene audio signal are determined according to the sound source direction prediction result and the sound source direction selection configuration information. The embodiment of the present application can flexibly and adaptively output the sound source direction information of the scene audio signal according to the real-time changes of the scene audio signal, which can improve the flexibility of the scene audio signal encoding and improve the encoding efficiency.

[0156] An example is given as follows: the encoding end includes a signal encoding module and a sound source direction selection module. When the sound source direction selection configuration information is set to the sound source direction information number threshold M3 that the signal encoding module can process and the sound source direction confidence threshold T, the sound source direction selection module selects the sound source direction angle corresponding to the confidence of the sound source direction greater than or equal to T. When the number of selected sound source direction angles is greater than M3, the sound source direction angles corresponding to the first M3 maximum confidences are output.

[0157] An example is given below. When the sound source direction selection configuration information is set to one information or a combination of multiple information in the sound source direction selection configuration information, the sound source direction selection module selects the corresponding sound source direction information according to the specified rules. An example is given below. When a combination of multiple information in the sound source direction selection configuration information is used, the number of sound source direction information finally selected is specified according to the configuration information that can select the minimum number of sound source direction information. For example, the sound source direction selection configuration information is the sound source direction confidence threshold T, and the user sets the number of sound source direction information selected to be P1. According to the sound source direction confidence threshold T, the number of sound source direction information selected is Q1. When P1 is less than or equal to Q1, P1 sound source direction angles with the maximum confidence values ​​are finally selected from the Q1 sound source direction information.

[0158] In some embodiments of the present application, the sound source direction prediction result includes: N2 confidence levels, where the N2 confidence levels correspond to N2 spherical direction angles of the sound source.

[0159] Step 403 determines the sound source direction information of the scene audio signal according to the sound source direction prediction result, including:

[0160] Get the sound source direction selection configuration information;

[0161] The sound source direction information of the scene audio signal is determined according to the N2 confidence levels and the sound source direction selection configuration information.

[0162] Among them, the encoding end can obtain the sound source direction selection configuration information, and the sound source direction selection configuration information is used to determine the sound source direction for the scene audio signal in combination with the sound source direction prediction result output by the sound source direction selection module. The sound source direction selection configuration information can be configuration information pre-stored in the encoding end, and the sound source direction selection configuration information can include predetermined sound source direction selection parameters and sound source direction selection rules. The sound source direction prediction result includes N2 confidences, and the encoding end can determine the sound source direction information of the scene audio signal according to the N2 confidences and the sound source direction selection configuration information, and screen the N2 confidences included in the sound source direction prediction result according to the specific configuration requirements of the sound source direction selection configuration information to obtain the final sound source direction information of the scene audio signal. In the embodiment of the present application, the sound source direction selection configuration information and the sound source direction prediction model are used to more accurately output the sound source direction information of the scene audio signal, and the sound source direction information output after the sound source direction selection configuration information is selected can improve the flexibility of the scene audio signal encoding and improve the encoding efficiency.

[0163] In some embodiments of the present application, the sound source direction prediction result includes: at least one confidence level of the candidate. For example, the at least one confidence level of the candidate includes N2 confidence levels, and the N2 confidence levels correspond one-to-one to the N2 Cartesian coordinates of the sound source directions;

[0164] Determining the sound source direction information of the scene audio signal according to the sound source direction prediction result and the sound source direction selection configuration information includes:

[0165] Selecting M4 confidence levels from at least one candidate confidence level according to the sound source direction selection configuration information, where M4 is a positive integer;

[0166] The sound source direction information is determined to include M4 Cartesian coordinates of the sound source directions corresponding to M4 confidence levels.

[0167] Among them, in the implementation scenario where the sound source direction prediction result includes N2 sound source direction Cartesian coordinates and N2 confidences, the encoding end selects M4 confidences from at least one candidate confidence according to the sound source direction selection configuration information, and there is a one-to-one correspondence between the confidence and the sound source direction angle, and the M4 sound source direction angles are determined according to the M4 confidences. The content of the sound source direction selection configuration information is detailed in the aforementioned embodiment description. In the embodiment of the present application, the M4 sound source direction Cartesian coordinates of the scene audio signal are determined according to the sound source direction prediction result and the sound source direction selection configuration information. The embodiment of the present application can flexibly and adaptively output the sound source direction information of the scene audio signal according to the real-time changes of the scene audio signal, which can improve the flexibility of the scene audio signal encoding and improve the encoding efficiency.

[0168] In some embodiments of the present application, the sound source direction prediction result includes: at least one confidence level of the candidate;

[0169] Determining the sound source direction information of the scene audio signal according to the sound source direction prediction result and the sound source direction selection configuration information includes:

[0170] Selecting M5 confidence levels from at least one candidate confidence level according to the sound source direction selection configuration information, where M5 is a positive integer;

[0171] Determine M5 direction indexes corresponding to M5 confidence levels according to the corresponding relationship between the confidence level and the direction index;

[0172] The sound source direction information is determined to include M5 sound source direction angles corresponding to M5 direction indexes.

[0173] Among them, the encoding end outputs at least one confidence of the candidate through the sound source direction prediction model, and the encoding end selects M5 confidences from at least one confidence of the candidate according to the sound source direction selection configuration information. There is a corresponding relationship between the confidence and the direction index. The encoding end determines the M5 direction indexes corresponding to the M5 confidences according to the corresponding relationship between the confidence and the direction index, and then determines the M5 sound source direction angles according to the corresponding relationship between the direction index and the sound source direction angle. In the embodiment of the present application, the M5 sound source direction angles of the scene audio signal are determined according to the sound source direction prediction result and the sound source direction selection configuration information. The embodiment of the present application can flexibly and adaptively output the sound source direction information of the scene audio signal according to the real-time changes of the scene audio signal, which can improve the flexibility of the scene audio signal encoding and improve the encoding efficiency.

[0174] It can be understood that the specific values ​​of M1 to M5 are not limited. The values ​​of M1 to M5 can be unequal, or at least two values ​​of M1 to M5 can be equal, or at least three values ​​of M1 to M5 can be equal, or at least four values ​​of M1 to M5 can be equal, depending on the application scenario.

[0175] 404. Encode the scene audio signal according to the sound source direction information to obtain a first bit stream.

[0176] In an embodiment of the present application, after determining the sound source direction information, the encoding end can encode the scene audio signal according to the sound source direction information to obtain a first code stream, and the first code stream is an audio encoding code stream generated by the encoding end. For example, the encoding end can specifically be a core encoder, and the core encoder encodes the scene audio signal according to the sound source direction information to obtain a first code stream. The code stream can also be called an audio signal encoding code stream. The encoding end of the embodiment of the present application encodes the scene audio signal according to the sound source direction information. Since the sound source direction information can be flexibly output in combination with the real-time transformation of the scene audio signal, the sound field at the position of the listener in the space is as close as possible to the original sound field when the scene audio signal was recorded, thereby ensuring the encoding quality of the encoding end and improving the encoding efficiency.

[0177] For example, the encoding end includes a sound source direction selection module and a signal encoding module. After the sound source direction selection module determines the sound source direction information of the scene audio signal according to the sound source direction prediction result, the sound source direction selection module sends the sound source direction information to the signal encoding module. The signal encoding module encodes the scene audio signal according to the sound source direction information, thereby obtaining a first bit stream. The encoding end then sends the first bit stream to the decoding end through an audio transmission channel. The decoding end receives the first bit stream, obtains the sound source direction information from the first bit stream, and then obtains the target scene audio signal from the first bit stream according to the sound source direction information, converts the target scene audio signal into an actual speaker signal and then plays it back, or converts the target scene audio signal into a virtual speaker signal and then maps it to a binaural signal for playback.

[0178] It can be seen from the examples of the aforementioned embodiments that in the embodiments of the present application, the characteristic information of the scene audio signal is input into the sound source direction prediction model. Since the characteristic information of the scene audio signal changes in real time with the scene audio signal, the sound source direction is predicted by combining the characteristic information of the scene audio signal through the sound source direction prediction model, so that the output sound source direction prediction result can indicate the sound source direction that conforms to the scene audio signal. The embodiments of the present application can adaptively and flexibly output the sound source direction prediction result of the scene audio signal according to the real-time changes of the scene audio signal, and then generate the sound source direction information according to the sound source direction prediction result. The scene audio signal is encoded according to the sound source direction information, so that the sound field at the position of the listener in the space is as close as possible to the original sound field when the scene audio signal is recorded, thereby ensuring the encoding quality of the encoding end and improving the encoding efficiency.

[0179] In order to better understand and implement the above-mentioned solutions in the embodiments of the present application, the following examples are given for specific explanation by using corresponding application scenarios.

[0180] In the embodiment of the present application, the sound source direction prediction model is specifically a neural network, and the scene audio signal is specifically an Ambisonics signal. The embodiment of the present application proposes an Ambisonics signal encoding scheme that extracts signal direction information through a neural network model, and uses the neural network model to flexibly and efficiently extract the direction information of the Ambisonics signal, thereby improving the signal encoding efficiency.

[0181] like Figure 5 As shown, it is an application scenario architecture diagram of the encoding end of an audio encoding method provided in an embodiment of the present application, and the encoding end includes: a pre-trained neural network model, a sound source direction selection module and an Ambisonics signal encoding module.

[0182] The characteristic information representing the Ambisonics signal is input into a pre-trained neural network model, and the neural network model outputs the directional information related to the sound source of the Ambisonics signal.

[0183] The sound source direction selection module further determines the sound source direction information of the Ambisonics signal according to the sound source direction information output by the neural network model.

[0184] In an optional implementation, the sound source direction selection configuration information is input into the sound source direction selection module, and the sound source direction selection module further determines the sound source direction information of the Ambisonics signal according to the sound source direction information output by the neural network model and the sound source direction selection configuration information.

[0185] The Ambisonics signal encoding module encodes the Ambisonics signal using the determined sound source direction information and outputs a code stream, which includes the encoding information.

[0186] From the above description of the encoding end, it can be seen that the Ambisonics signal encoding module can flexibly use the sound source direction information to complete the efficient encoding of the Ambisonics signal.

[0187] Next, the present application is described in detail through multiple embodiments.

[0188] like Figure 6 As shown, an application scenario architecture diagram of an encoding end of an audio encoding method applied in an embodiment of the present application is provided. An embodiment of the present application provides a neural network model to output the sound source direction angle of an Ambisonics signal and the confidence corresponding to each sound source direction angle, and then adaptively selects the sound source direction angle of the Ambisonics signal according to the confidence value, and inputs the sound source direction angle into the Ambisonics signal encoding module to complete the encoding process.

[0189] After the characteristics of the Ambisonics signal are input into the neural network model, a number of sound source direction angles and the confidence corresponding to the sound source direction angles are output. The neural network model provided in the embodiment of the present application is generated by pre-training and can be any model of a neural network structure that can complete the functions of this embodiment, and the embodiment of the present application does not make specific limitations.

[0190] The features of the Ambisonics signal may be the time domain features, transform domain features, or features derived using the time domain features and transform domain features of the Ambisonics signal. Specifically, the features input to the neural network model in the embodiment of the present application are acoustic intensity features of the Ambisonics signal derived from the short-time Fourier transform (STFT) domain features. The process of deriving the acoustic density features is as follows:

[0191]

[0192]

[0193] In the above formulas (1) and (2), W(t, f) represents the STFT coefficient spectrum of the 0th order channel of the Ambisonics signal, X(t, f), Y(t, f), and Z(t, f) represent the STFT coefficient spectrum of the channels corresponding to the X, Y, and Z axis directions of the 1st order channel of the Ambisonics signal, * represents the conjugate transposed matrix, t represents the sampling time, and f represents the frequency point corresponding to the sampling time. a (t, f) represents the real part of the acoustic density feature, I r (t, f) represents the imaginary part of the acoustic density feature.

[0194] The acoustic density feature of the Ambisonics signal input to the neural network model in the embodiment of the present application can be expressed as follows:

[0195]

[0196] In the above formula (3), C(t, f) is expressed as follows:

[0197]

[0198] like Figure 7 As shown, it is a schematic diagram of the neural network model provided in an embodiment of the present application outputting N sound source direction angles and the confidence corresponding to the sound source direction angles. The encoder inputs the characteristics of the Ambisonics signal into the neural network model and outputs (horizontal angle 1, pitch angle 1), ..., (horizontal angle N , pitch angle N ), confidence 1, ..., confidence N .

[0199] The following is an example of the input and output of the neural network model in the embodiment of the present application: Figure 7As shown, the neural network model outputs N direction angle pairs, including horizontal angles and elevation angles and confidences corresponding to the direction angle pairs, where N≥1. The horizontal angle range is [0, 360], the elevation angle range is [0, 180], and the confidence value range is [0, 1]. For example, N can specifically be N1 in the aforementioned embodiment.

[0200] The sound source direction selection module adaptively selects the sound source direction angle according to the input direction angle pair including the horizontal angle and the pitch angle and the corresponding confidence of the direction angle pair, and combines the sound source direction selection configuration information, such as Figure 8 , which is a schematic diagram of sound source direction selection performed by the sound source direction selection module provided in an embodiment of the present application.

[0201] Specifically, the neural network model converts (horizontal angle 1, pitch angle 1), ..., (horizontal angle N , pitch angle N ), confidence 1, ..., confidence N Input to the sound source direction selection module. In the embodiment of the present application, the sound source direction selection configuration information is the confidence threshold T. In the embodiment of the present application, T∈(0,1). The sound source direction angle selection module selects the sound source horizontal angle and pitch angle pair with a confidence greater than or equal to T according to the set confidence threshold T. The sound source direction selection module outputs (horizontal angle P , pitch angle P ),...,(horizontal angle Q , pitch angle Q ).

[0202] Where P∈[1,N], Q∈[1,N], in the embodiment of the present application, Q≥P, and,

[0203] QP≤MAX process (5)

[0204] Among them, MAX process Indicates the maximum number of sound source direction angle pairs that the Ambisonics signal encoding module can process.

[0205] The sound source direction selection module will (horizontal angle P , pitch angle P ),...,(horizontal angle Q , pitch angle Q ) is input into the Ambisonics signal encoding module, which refers to a module that uses the sound source direction information to encode the Ambisonics signal. The Ambisonics signal encoding module is not specifically limited in the embodiment of the present application. In this embodiment, the Ambisonics signal encoding module outputs the Ambisonics signal component information.

[0206] As an example, the Ambisonics signal encoding module uses the input sound source direction information to process the input Ambisonics signal and output encoding information.

[0207] The output of the Ambisonics signal module includes one of the following: signal components, side information, or signal components and side information.

[0208] The signal component may be, for example, a Fourier transform coefficient of an Ambisonics signal. Side information is information describing the Ambisonics signal itself or the relationship between signals. For example, the side information may include signal gain, or the side information may include signal correlation.

[0209] For example, the signal components include one or more of the following signal components: the foreground signal component may include a signal component related to the direction of the sound source, and the background signal component may include a signal component unrelated to the direction of the sound source.

[0210] For example, the side information includes one or more of the following information: gain information, direction information, dispersion information, correlation information, and other parameter information that can describe the Ambisonics signal.

[0211] Through the examples of the above embodiments, it can be seen that the embodiments of the present application use a pre-trained neural network model to more accurately output the sound source direction angle of the Ambisonics signal. The sound source direction angle output after the sound source direction angle selection can improve the flexibility of Ambisonics signal encoding and improve the encoding efficiency. Different from the prior art, the embodiments of the present application can adaptively calculate the sound source direction angle according to the real-time changes of the Ambisonics signal, thereby improving the encoding efficiency of the Ambisonics signal.

[0212] like Fig. 9 As shown, an application scenario architecture diagram of an encoding end of another audio encoding method applied in an embodiment of the present application is provided. The embodiment of the present application provides a neural network model to output the sound source direction angle of the Ambisonics signal and the confidence corresponding to each sound source direction angle, and then adaptively selects the sound source direction angle of the Ambisonics signal according to the confidence value, and inputs the sound source direction angle into the Ambisonics signal encoding module to complete the encoding process.

[0213] The embodiments of the present application Figure 6 The difference between the illustrated embodiment and the embodiment shown is that the output of the neural network model is the confidence level corresponding to the preset spherical direction angle.

[0214] In the embodiment of the present application, the preset spherical direction angle represents the uniform or non-uniform extraction of N sound source positions on the spherical surface, where N ≥ 1, and each sound source position corresponds to a set of preset sound source directions and an index corresponding thereto, and the value of the index is [1, N]. Figure 6 The illustrated embodiments are the same, and the directions of the embodiments of the present application are represented by a pair of horizontal angles and elevation angles.

[0215] In the embodiment of the present application, the neural network model outputs N confidence values, and the range of confidence in the embodiment of the present application is [0, 1]. The index of each confidence value is consistent with the index of the preset spherical direction angle. For example, the value of N can specifically be N2 in the aforementioned embodiment.

[0216] like Fig.10 As shown, it is a schematic diagram of the output of N confidences by the neural network model provided in the embodiment of the present application. The encoding end inputs the characteristics of the Ambisonics signal into the neural network model, and the neural network model outputs N confidences corresponding to N groups of preset spherical direction angles. The output N confidences include confidence 1, ..., confidence N .

[0217] like Fig.11 , which is a schematic diagram of the sound source direction selection module provided in the embodiment of the present application for selecting the sound source direction. Specifically, the sound source direction selection configuration information of the sound source direction selection module in the embodiment of the present application is Figure 6 The same as the embodiment shown, is set to the confidence threshold T. The selection process of the sound source direction of this embodiment is the same as Figure 6 The embodiment shown is the same, except that the selected sound source direction angle of the Ambisonics signal corresponds to a preset spherical direction angle.

[0218] It can be seen from the examples of the aforementioned embodiments that the embodiments of the present application can improve the flexibility of Ambisonics signal encoding and improve the encoding efficiency.

[0219] like Fig.12 As shown, an application scenario architecture diagram of an encoding end of another audio encoding method applied in an embodiment of the present application is provided. The embodiment of the present application provides a neural network model to output the sound source direction angle of the Ambisonics signal and the confidence corresponding to each sound source direction angle, and then adaptively selects the sound source direction angle of the Ambisonics signal according to the confidence value, and inputs the sound source direction angle into the Ambisonics signal encoding module to complete the encoding process.

[0220] This embodiment illustrates that the output of the pre-trained neural network model is the Cartesian coordinate representation of the sound source direction of the Ambisonics signal. Fig.12 As shown, what is illustrated in the embodiment of the present application is the Cartesian coordinate representation of the sound source direction output by the neural network model.

[0221] like Fig.13 As shown in FIG. 1 , a schematic diagram of the neural network model provided in an embodiment of the present application outputting N sound source direction angles and the confidences corresponding to the sound source direction angles is shown. The encoder inputs the features of the Ambisonics signal into the neural network model and outputs (X1, Y1, Z1), ..., (X N ,Y N ,Z N ), confidence 1, ..., confidence N .

[0222] In this embodiment, the Cartesian coordinates have a value range of X∈[-1, 1], Y∈[-1, 1], and Z∈[-1, 1]. The other processes of this embodiment are the same as Figure 6 The embodiments shown are identical.

[0223] It can be seen from the examples of the aforementioned embodiments that the embodiments of the present application can improve the flexibility of Ambisonics signal encoding and improve the encoding efficiency.

[0224] like Fig.14 As shown, it is a schematic diagram of the output of N confidences by the neural network model provided in the embodiment of the present application. The encoding end inputs the characteristics of the Ambisonics signal into the neural network model, and the neural network model outputs N confidences corresponding to N groups of preset spherical direction angles. The output N confidences include confidence 1, ..., confidence N .

[0225] and Figure 6 The difference between the embodiment shown in FIG. 1 and FIG. 2 is that the sound source direction selection module in this embodiment outputs the Cartesian coordinates of the sound source direction. Fig.12 The same embodiment as shown, other limiting conditions and processes are the same as Figure 6 The embodiments shown are identical.

[0226] It can be seen from the examples of the aforementioned embodiments that the embodiments of the present application can improve the flexibility of Ambisonics signal encoding and improve the encoding efficiency.

[0227] like Fig.15 , which is a schematic diagram of the sound source direction selection module provided in the embodiment of the present application for selecting the sound source direction. Specifically, the sound source direction selection configuration information of the sound source direction selection module in the embodiment of the present application is Figure 6 The same embodiment as shown Figure 8The difference between the embodiment shown is that the sound source direction selection module in this embodiment outputs the direction index corresponding to the preset sound source direction. The process of the sound source direction selection module outputting the direction index and other limiting conditions are the same as Fig. 9 The embodiments shown are identical.

[0228] It can be seen from the examples of the aforementioned embodiments that in the embodiments of the present application, the characteristic information characterizing the Ambisonics signal is input into a pre-trained neural network model to obtain the sound source direction information of the Ambisonics signal. The Ambisonics signal direction information is obtained by using the neural network model, which has the advantage of being more flexible. The Ambisonics signal encoding module can flexibly use the sound source direction information to complete the efficient encoding of the Ambisonics signal.

[0229] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0230] In order to better implement the above-mentioned solution of the embodiment of the present application, relevant devices for implementing the above-mentioned solution are also provided below.

[0231] See also Fig.16 As shown, an audio encoding device 1600 provided in an embodiment of the present application is shown. Fig.16 The audio encoding device in the embodiment can be used to execute the audio encoding method of the aforementioned embodiment. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding method provided above, which will not be repeated here. The audio encoding device 1600 may include: a feature acquisition module 1601, a direction prediction module 1602, a direction determination module 1603 and an encoding module 1604, wherein:

[0232] A feature acquisition module, used to acquire feature information of a scene audio signal to be encoded;

[0233] A direction prediction module, used for inputting the characteristic information of the scene audio signal into a sound source direction prediction model, and outputting a sound source direction prediction result through the sound source direction prediction model;

[0234] A direction determination module, used to determine the sound source direction information of the scene audio signal according to the sound source direction prediction result;

[0235] The encoding module is used to encode the scene audio signal according to the sound source direction information to obtain a first code stream.

[0236] In some embodiments of the present application, the sound source direction prediction result includes: a sound source direction angle and a confidence level corresponding to the sound source direction angle.

[0237] In some embodiments of the present application, the number of the sound source direction angles is N1, and the number of the confidence levels is N1.

[0238] There is a one-to-one correspondence between the N1 sound source direction angles and the N1 confidence levels, and N1 is a positive integer.

[0239] In some embodiments of the present application, the N1 sound source direction angles include: N1 sound source direction angle pairs,

[0240] in,

[0241] The sound source direction angle pair includes: a sound source direction horizontal angle and a sound source direction elevation angle.

[0242] In some embodiments of the present application, the sound source direction prediction result includes: N2 confidence levels, there is a one-to-one correspondence between the N2 confidence levels and the N2 sound source spherical direction angles, and N2 is a positive integer.

[0243] In some embodiments of the present application, there is a one-to-one correspondence between the N2 confidences and the N2 index values;

[0244] The N2 index values ​​are index values ​​corresponding to the spherical direction angles of the N2 sound sources respectively.

[0245] In some embodiments of the present application, the sound source direction prediction result includes: Cartesian coordinates of the sound source direction and confidence levels corresponding to the Cartesian coordinates of the sound source direction.

[0246] In some embodiments of the present application, the direction determination module is specifically used to obtain sound source direction selection configuration information; and determine the sound source direction information of the scene audio signal according to the sound source direction prediction result and the sound source direction selection configuration information.

[0247] In some embodiments of the present application, the sound source direction selection configuration information includes at least one of the following: a sound source direction number threshold for signal encoding, a sound source direction confidence threshold, and a configurable rule for selecting a sound source direction.

[0248] In some embodiments of the present application, the sound source direction prediction result includes: at least one confidence level of the candidate;

[0249] The direction determination module is specifically used to select M1 confidences from at least one of the candidate confidences, where M1 represents the threshold of the number of sound source directions for signal encoding, and M1 is a positive integer; and determine that the sound source direction information includes M1 sound source direction angles corresponding to the M1 confidences.

[0250] In some embodiments of the present application, the sound source direction prediction result includes: at least one confidence level of the candidate;

[0251] The direction determination module is specifically used to select M2 confidences greater than or equal to the sound source direction confidence threshold from at least one confidence of the candidate, where M2 is a positive integer; and determine that the sound source direction information includes M2 sound source direction angles corresponding to the M2 confidences.

[0252] In some embodiments of the present application, the sound source direction prediction result includes: at least one confidence level of the candidate, where N is a positive integer;

[0253] The direction determination module is specifically used to select M3 confidences that meet the configurable rules for selecting the sound source direction from at least one of the candidate confidences, where M3 is a positive integer; and determine that the sound source direction information includes M3 sound source direction angles corresponding to the M3 confidences.

[0254] In some embodiments of the present application, the direction determination module is specifically used to obtain sound source direction selection configuration information; and determine the sound source direction information of the scene audio signal according to the N2 confidence levels and the sound source direction selection configuration information.

[0255] In some embodiments of the present application, the sound source direction prediction result includes: at least one confidence level of the candidate;

[0256] The direction determination module is specifically used to select M4 confidence levels from at least one of the candidate confidence levels according to the sound source direction selection configuration information, where M4 is a positive integer; and determine that the sound source direction information includes M4 sound source direction Cartesian coordinates corresponding to the M4 confidence levels.

[0257] In some embodiments of the present application, the sound source direction prediction result includes: at least one confidence level of the candidate;

[0258] The direction determination module is specifically used to select M5 confidences from at least one of the candidate confidences according to the sound source direction selection configuration information, where M5 is a positive integer; determine the M5 direction indexes corresponding to the M5 confidences according to the correspondence between the confidence and the direction index; and determine that the sound source direction information includes M5 sound source direction angles corresponding to the M5 direction indexes.

[0259] In some embodiments of the present application, the feature information of the scene audio signal includes at least one of the following: a time domain feature of the scene audio signal, a transform domain feature of the scene audio signal, or a feature obtained by using the time domain feature or the transform domain feature.

[0260] In some embodiments of the present application, at least one of the following is included: a Fourier transform feature, a discrete cosine transform feature, and a discrete sine transform feature;

[0261] The features obtained through the time domain features or the transform domain features include at least one of the following: relative harmonic coefficient features, relative modal coherence features, and acoustic density features.

[0262] It should be noted that the information interaction, execution process, etc. between the modules / units of the above-mentioned device are based on the same concept as the method embodiment of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the present application, and will not be repeated here.

[0263] An embodiment of the present application also provides a computer storage medium, wherein the computer storage medium stores a program, and the program executes some or all of the steps recorded in the above method embodiment.

[0264] Next, another audio encoding device provided by the embodiment of the present application is introduced. Fig.17 As shown, the audio encoding device 1700 includes:

[0265] The receiver 1701, the transmitter 1702, the processor 1703 and the memory 1704 (wherein the number of the processor 1703 in the audio encoding device 1700 can be one or more, Fig.17 In some embodiments of the present application, the receiver 1701, the transmitter 1702, the processor 1703 and the memory 1704 may be connected via a bus or other means, wherein: Fig.17 The example of connecting through bus is taken in the following.

[0266] The memory 1704 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1703. A portion of the memory 1704 may also include a non-volatile random access memory (NVRAM). The memory 1704 stores an operating system and operating instructions, executable modules or data structures, or a subset thereof, or an extended set thereof, wherein the operating instructions may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic services and processing hardware-based tasks.

[0267] The processor 1703 controls the operation of the audio encoding device, and the processor 1703 may also be referred to as a central processing unit (CPU). In a specific application, the various components of the audio encoding device are coupled together through a bus system, wherein the bus system may include a power bus, a control bus, and a status signal bus in addition to a data bus. However, for the sake of clarity, various buses are referred to as bus systems in the figure.

[0268] The method disclosed in the above embodiment of the present application can be applied to the processor 1703, or implemented by the processor 1703. The processor 1703 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 1703. The above processor 1703 can be a general processor, a digital signal processor (digital signal processing, DSP), an application specific integrated circuit (application specific integrated circuit, ASIC), a field programmable gate array (field-programmable gate array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1704, and the processor 1703 reads the information in the memory 1704 and completes the steps of the above method in combination with its hardware.

[0269] Receiver 1701 can be used to receive input digital or character information, and generate signal input related to the relevant settings and function control of the audio encoding device. Transmitter 1702 may include a display device such as a display screen. Transmitter 1702 can be used to output digital or character information through an external interface.

[0270] In the embodiment of the present application, the processor 1703 is used to execute the above-mentioned embodiment. Figure 4 The audio encoding method shown is performed by the audio encoding device.

[0271] In another possible design, when the audio encoding device is a chip in a terminal, the chip includes: a processing unit and a communication unit, the processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin or a circuit, etc. The processing unit may execute computer-executable instructions stored in the storage unit so that the chip in the terminal executes the audio encoding method of any one of the above-mentioned first aspects. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc., and the storage unit may also be a storage unit in the terminal located outside the chip, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0272] The processor mentioned in any of the above places may be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the above-mentioned first aspect or second aspect method.

[0273] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.

[0274] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0275] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0276] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, a computer, a server, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.

Claims

1. An audio encoding method, characterized in that: The method comprises: Acquiring characteristic information of a scene audio signal to be encoded; Inputting feature information of the scene audio signal into a sound source direction prediction model, and outputting a sound source direction prediction result through the sound source direction prediction model; Determining the sound source direction information of the scene audio signal according to the sound source direction prediction result; The scene audio signal is encoded according to the sound source direction information to obtain a first code stream.

2. The method according to claim 1, characterized in that The sound source direction prediction result includes: a sound source direction angle and a confidence level corresponding to the sound source direction angle.

3. The method according to claim 2, characterized in that The number of sound source direction angles is N1, and the number of confidence levels is N1, There is a one-to-one correspondence between the N1 sound source direction angles and the N1 confidence levels, and N1 is a positive integer.

4. The method according to claim 3, characterized in that The N1 sound source direction angles include: N1 sound source direction angle pairs, in, The sound source direction angle pair includes: a sound source direction horizontal angle and a sound source direction elevation angle.

5. The method according to claim 1, characterized in that The sound source direction prediction result includes: N2 confidence levels, there is a one-to-one correspondence between the N2 confidence levels and the N2 sound source spherical direction angles, and N2 is a positive integer.

6. The method according to claim 5, characterized in that There is a one-to-one correspondence between the N2 confidences and the N2 index values; The N2 index values ​​are index values ​​corresponding to the spherical direction angles of the N2 sound sources respectively.

7. The method according to claim 1, characterized in that The sound source direction prediction result includes: Cartesian coordinates of the sound source direction and confidence levels corresponding to the Cartesian coordinates of the sound source direction.

8. The method according to any one of claims 1 to 7, characterized in that The determining the sound source direction information of the scene audio signal according to the sound source direction prediction result includes: Get the sound source direction selection configuration information; The sound source direction information of the scene audio signal is determined according to the sound source direction prediction result and the sound source direction selection configuration information.

9. The method according to claim 8, characterized in that The sound source direction selection configuration information includes at least one of the following: a sound source direction number threshold for signal encoding, a sound source direction confidence threshold, and a configurable rule for selecting a sound source direction.

10. The method according to claim 9, characterized in that The sound source direction prediction result includes: at least one confidence level of the candidate; The determining the sound source direction information of the scene audio signal according to the sound source direction prediction result and the sound source direction selection configuration information includes: Selecting M1 confidence levels from at least one of the candidate confidence levels, where M1 represents the threshold of the number of sound source directions for signal coding, and M1 is a positive integer; Determine that the sound source direction information includes M1 sound source direction angles corresponding to the M1 confidence levels.

11. The method according to claim 9 or 10, characterized in that: The sound source direction prediction result includes: at least one confidence level of the candidate; The determining the sound source direction information of the scene audio signal according to the sound source direction prediction result and the sound source direction selection configuration information includes: Selecting M2 confidences greater than or equal to the sound source direction confidence threshold from at least one of the candidate confidences, where M2 is a positive integer; Determine that the sound source direction information includes M2 sound source direction angles corresponding to the M2 confidence levels.

12. The method according to any one of claims 9 to 11, characterized in that The sound source direction prediction result includes: at least one confidence level of the candidate, where N is a positive integer; The determining the sound source direction information of the scene audio signal according to the sound source direction prediction result and the sound source direction selection configuration information includes: Selecting M3 confidences that meet the configurable rule for selecting the direction of the sound source from the at least one confidence of the candidate, where M3 is a positive integer; Determine that the sound source direction information includes M3 sound source direction angles corresponding to the M3 confidence levels.

13. The method according to claim 5, characterized in that The determining the sound source direction information of the scene audio signal according to the sound source direction prediction result includes: Get the sound source direction selection configuration information; The sound source direction information of the scene audio signal is determined according to the N2 confidence levels and the sound source direction selection configuration information.

14. The method according to any one of claims 8 to 12, characterized in that The sound source direction prediction result includes: at least one confidence level of the candidate; The determining the sound source direction information of the scene audio signal according to the sound source direction prediction result and the sound source direction selection configuration information includes: Selecting M4 confidence levels from at least one of the candidate confidence levels according to the sound source direction selection configuration information, where M4 is a positive integer; Determining the sound source direction information includes M4 sound source direction Cartesian coordinates corresponding to the M4 confidence levels.

15. The method according to any one of claims 8 to 12, characterized in that The sound source direction prediction result includes: at least one confidence level of the candidate; The determining the sound source direction information of the scene audio signal according to the sound source direction prediction result and the sound source direction selection configuration information includes: Selecting M5 confidence levels from at least one of the candidate confidence levels according to the sound source direction selection configuration information, where M5 is a positive integer; Determine M5 direction indexes corresponding to the M5 confidence levels according to the corresponding relationship between the confidence levels and the direction indexes; Determine that the sound source direction information includes M5 sound source direction angles corresponding to the M5 direction indexes.

16. The method according to any one of claims 1 to 15, characterized in that The feature information of the scene audio signal includes at least one of the following: a time domain feature of the scene audio signal, a transform domain feature of the scene audio signal, or a feature obtained by using the time domain feature or the transform domain feature.

17. The method according to claim 16, characterized in that The transform domain feature includes at least one of the following: Fourier transform feature, discrete cosine transform feature, discrete sine transform feature; The features obtained through the time domain features or the transform domain features include at least one of the following: relative harmonic coefficient features, relative modal coherence features, and acoustic density features.

18. An audio encoding device, characterized in that: The device comprises: A feature acquisition module, used to acquire feature information of a scene audio signal to be encoded; A direction prediction module, used for inputting the characteristic information of the scene audio signal into a sound source direction prediction model, and outputting a sound source direction prediction result through the sound source direction prediction model; A direction determination module, used to determine the sound source direction information of the scene audio signal according to the sound source direction prediction result; The encoding module is used to encode the scene audio signal according to the sound source direction information to obtain a first code stream.

19. An electronic device, characterized in that: include: a memory and a processor, the memory being coupled to the processor; The memory stores program instructions, and when the program instructions are executed by the processor, the electronic device executes the audio encoding method according to any one of claims 1 to 17.

20. A chip, characterized in that: It comprises one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from a memory of an electronic device and send the signal to the processor, the signal comprising a computer instruction stored in the memory; when the processor executes the computer instruction, the electronic device executes the audio encoding method described in any one of claims 1 to 17.

21. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program runs on a computer or a processor, the computer or the processor executes the audio encoding method as described in any one of claims 1 to 17.

22. A computer-readable storage medium, characterized in that: Comprising a code stream generated by the audio encoding method according to any one of claims 1 to 17.

23. A computer program product, characterized in that The computer program product comprises a software program, and when the software program is executed by a computer or a processor, the steps of the audio encoding method according to any one of claims 1 to 17 are performed.