Method and device for allocating bits of audio objects
By pre-rendering and perceived importance parameter evaluation of audio objects in audio frames, the number of bits is accurately allocated, which solves the problem of low bit allocation efficiency of audio objects in the prior art, and improves the reconstruction quality and encoding efficiency of audio objects.
Patent Information
- Application Number
- CN202110083715.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-01-21
AI Technical Summary
The prior art causes the overall quality of reconstructed audio objects to be low and the encoding efficiency is low when the number of bits is allocated between audio objects.
By pre-rendering multiple audio objects in the audio frame to be encoded, the respective perceived importance parameter values are obtained and the number of bits is allocated based on these parameter values to ensure higher quality audio object reconstruction.
The overall quality and encoding efficiency of reconstructed audio objects are improved, and the immersion of three-dimensional audio scenes is enhanced through more accurate bit number allocation.
Smart Images

Figure CN114822564B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio coding and decoding, and in particular to a method and device for allocating bits of an audio object. Background Art
[0002] 3D audio gives sound a strong sense of space, envelopment and immersion, giving people an extraordinary auditory experience of "sound immersion". In recent years, people have paid more and more attention to the development of audio technology.
[0003] Object-based audio technology is an important way to achieve three-dimensional audio. Through rendering technology, relatively independent audio objects can be represented as audio scenes with a sense of space and a more realistic auditory experience. The number of bits used by the encoder to encode the audio object is an important factor affecting the quality of the reconstructed audio object at the decoder. Therefore, at a fixed bit rate, how to allocate the number of bits between audio objects so that the rendered three-dimensional audio scene has high quality is one of the important directions of current audio coding research.
[0004] At present, a commonly used method for allocating bits of audio objects is to evenly allocate the total number of bits to multiple audio objects in an audio frame, which results in low overall quality of the reconstructed audio objects and low coding efficiency. Summary of the invention
[0005] The embodiments of the present application provide a method and device for allocating bits of an audio object, which helps to improve the overall quality and coding efficiency of the reconstructed audio object.
[0006] In order to achieve the above objectives, this application provides the following technical solutions:
[0007] In a first aspect, a method for allocating bits of an audio object is provided, comprising: pre-rendering a plurality of audio objects to be pre-rendered in an audio frame to be encoded, respectively, to obtain a plurality of pre-rendered audio objects. Then, obtaining a perceptual importance parameter value of each of the plurality of pre-rendered audio objects. The perceptual importance parameter value of the current pre-rendered audio object in the plurality of pre-rendered audio objects is used to indicate the perceptual importance of the current rendered audio object in the plurality of pre-rendered audio objects. The current pre-rendered audio object may be any one of the plurality of pre-rendered audio objects. Then, based on the perceptual importance parameter values of the plurality of pre-rendered audio objects, obtaining a bit allocation parameter value of the current audio object to be pre-rendered in the plurality of audio objects to be pre-rendered. The current audio object to be pre-rendered may be any one of the plurality of audio objects to be pre-rendered. Finally, determining a target number of bits to be allocated for the current audio object to be pre-rendered based on the bit allocation parameter value of the current audio object to be pre-rendered and the total number of bits to be allocated corresponding to the plurality of audio objects to be pre-rendered. For example, the total number of bits to be allocated may be used to encode the plurality of audio objects to be pre-rendered. The target number of bits may be used to encode the current audio object to be pre-rendered.
[0008] This technical solution takes into account the differences in the perceptual characteristics of different pre-rendered audio objects at the rendering playback end when allocating bits to the audio objects to be pre-rendered. Compared with the technical solution of using the same number of bits to encode different audio objects in the traditional technology, it helps to improve the overall quality of the reconstructed audio objects. For example, the higher the degree of perceptual importance indicated by the perceptual importance parameter value of a pre-rendered audio object, the more bits the encoder can allocate to the audio object to be pre-rendered corresponding to the pre-rendered audio object (that is, the audio object of the pre-rendered audio object before pre-rendering), and the number of bits can be used to encode the audio object to be pre-rendered. At this time, the quality of the audio object reconstructed by the decoder will be higher. In this way, it helps to improve the overall quality of the reconstructed audio frames containing multiple audio objects. At the same time, the coding efficiency can be improved.
[0009] In one possible design, the degree of perceptual importance includes at least one of an energy intensity and a spectrum variation degree.
[0010] In one possible design, the perceptual importance parameter includes an energy importance parameter, wherein the energy importance parameter of the current pre-rendered audio object is calculated based on the energy of the current pre-rendered audio object, and is used to indicate a ratio between the energy of the current pre-rendered audio object and the sum of the energies of the multiple pre-rendered audio objects.
[0011] In a possible design, the perceptual importance parameter includes a perceptual intensity importance parameter, wherein the perceptual intensity importance parameter of the current pre-rendered audio object is calculated by combining the human hearing curve and the energy of the current pre-rendered audio object, and is used to indicate the ratio between the sum of the energies of a preset number of frequency bands with the largest energies among multiple frequency bands of the current pre-rendered audio object and the sum of the energies of a preset number of frequency bands with the largest energies among multiple frequency bands of the multiple pre-rendered audio objects.
[0012] In one possible design, the perceptual importance parameter includes a spectrum flatness parameter. The spectrum flatness parameter of the current pre-rendered audio object is used to indicate the spectrum flatness of the current pre-rendered audio object among the multiple pre-rendered audio objects.
[0013] In a possible design, the current pre-rendered audio object is an audio object obtained by pre-rendering the current audio object to be pre-rendered. The bit allocation parameter value of the current audio object to be pre-rendered includes a first ratio, or a parameter value determined according to the first ratio. The first ratio is a ratio between the perceptual importance parameter value of the current pre-rendered audio object and the sum of the perceptual importance parameter values of the multiple pre-rendered audio objects. This possible design provides a specific implementation method for obtaining the bit allocation parameter value of the current audio object to be pre-rendered, and the implementation method is simple.
[0014] In a possible design, the method further includes: obtaining the content importance parameter value of each of the multiple audio objects to be pre-rendered. Among them, the content importance parameter value of the current audio object to be pre-rendered is used to indicate the importance of the sound type represented by the content of the current audio object to be pre-rendered in the sound type represented by the content of the multiple audio objects to be pre-rendered. In this case, based on the perceptual importance parameter values of each of the multiple pre-rendered audio objects, the bit allocation parameter value of the current audio object to be pre-rendered in the multiple audio objects to be pre-rendered is obtained, including: based on the perceptual importance parameter values of each of the multiple pre-rendered audio objects and the content importance parameter values of each of the multiple audio objects to be pre-rendered, the bit allocation parameter value of the current audio object to be pre-rendered is obtained. This possible design, when allocating the number of bits for the audio objects to be pre-rendered, also takes into account the differences in content characteristics of different audio objects to be pre-rendered. Therefore, compared with the technical solution of encoding different audio objects using the same number of bits in the traditional technology, the overall quality and coding efficiency of the reconstructed audio objects can be further improved.
[0015] In one possible design, the current pre-rendered audio object is an audio object obtained by pre-rendering the current audio object to be pre-rendered. The bit allocation parameter value of the current audio object to be pre-rendered includes a second ratio, or a parameter value determined according to the second ratio. Among them, the second ratio is the ratio between the first value of the current audio object to be pre-rendered and the sum of the first values of the multiple audio objects to be pre-rendered. The first value of the current audio object to be pre-rendered is the product of the content importance parameter value of the current audio object to be pre-rendered and the perceptual importance parameter value of the current pre-rendered audio object, or a parameter value determined according to "the product of the content importance parameter value of the current audio object to be pre-rendered and the perceptual importance parameter value of the current pre-rendered audio object". This possible design provides another specific implementation method for obtaining the bit allocation parameter value of the current audio object to be pre-rendered, which is simple to implement.
[0016] In one possible design, the sound type includes at least one of the following: speech, music, sound effect, ambient sound or noise.
[0017] In a possible design, the ratio of the target number of bits allocated to the current audio object to be pre-rendered to the total number of bits to be allocated is equal to a third ratio, or equal to a parameter value determined according to the third ratio. The third ratio is the ratio between the bit allocation parameter value of the current audio object to be pre-rendered and the sum of the bit allocation parameter values of the multiple audio objects to be pre-rendered. This possible design provides a specific implementation method for determining the target number of bits allocated to the current audio object to be pre-rendered. In this possible design, audio objects to be pre-rendered with different bit allocation parameter values can correspond to different target numbers of bits.
[0018] In a possible design, based on the bit allocation parameter value of the current audio object to be pre-rendered and the total number of bits to be allocated, determining the target number of bits to be allocated for the current audio object to be pre-rendered includes: determining the priority level of the current audio object to be pre-rendered based on the correspondence between multiple bit allocation parameter values and multiple priority levels, and the bit allocation parameter value of the current audio object to be pre-rendered; then, based on the priority level of the current audio object to be pre-rendered and the total number of bits to be allocated, determining the target number of bits to be allocated for the current audio object to be pre-rendered. This possible design provides another specific implementation method for determining the target number of bits to be allocated for the current audio object to be pre-rendered, in which audio objects to be pre-rendered with different bit allocation parameter values can correspond to the same target number of bits or different target numbers of bits.
[0019] In one possible design, a ratio of a target number of bits allocated to the current audio object to be pre-rendered to a total number of bits to be allocated is equal to a fourth ratio, or is equal to a parameter value determined according to the fourth ratio, wherein the fourth ratio is a ratio of a priority level of the current audio object to be pre-rendered to a sum of priority levels of the plurality of audio objects to be pre-rendered.
[0020] In one possible design, based on the bit allocation parameter value of the current audio object to be pre-rendered and the total number of bits to be allocated corresponding to the multiple audio objects to be pre-rendered, the method includes: obtaining an initial number of bits allocated to the current audio object to be pre-rendered; adjusting the bit allocation parameter value of the current audio object to be pre-rendered based on the initial number of bits; and determining a target number of bits allocated to the current audio object to be pre-rendered based on the total number of bits to be allocated and the adjusted bit allocation parameter value of the current audio object to be pre-rendered. This possible design provides another implementation method for determining the target number of bits allocated to the current audio object to be pre-rendered.
[0021] This possible design adjusts the bit allocation parameter value of the current audio object to be pre-rendered by allocating the initial number of bits to the current audio object to be pre-rendered, which helps to further improve the overall quality and coding efficiency of the reconstructed audio object. In addition, the initial number of bits can be obtained according to traditional techniques, that is, this possible design provides a solution combining traditional techniques with the techniques provided in the embodiments of the present application. Alternatively, the initial number of bits can be obtained according to one of the technical solutions provided in the embodiments of the present application, that is, this possible design provides a solution combining multiple techniques provided in the embodiments of the present application.
[0022] In a possible design, the adjusted bit allocation parameter value of the current audio object to be pre-rendered includes: a fifth ratio or a parameter value determined according to the fifth ratio. The fifth ratio is the ratio between the second value of the current audio object to be pre-rendered and the sum of the second values of multiple audio objects to be pre-rendered. The second value of the current audio object to be pre-rendered is the product of the initial number of bits allocated to the current audio object to be pre-rendered and the bit allocation parameter value of the current audio object to be pre-rendered, or is a parameter value determined according to "the product of the initial number of bits allocated to the current audio object to be pre-rendered and the bit allocation parameter value of the current audio object to be pre-rendered". This possible design provides a specific implementation method for adjusting the bit allocation parameter value.
[0023] In one possible design, a ratio of a target number of bits allocated to a current audio object to be pre-rendered to a total number of bits to be allocated is equal to an adjusted bit allocation parameter value of the current audio object to be pre-rendered, or is equal to a parameter value determined according to the adjusted bit allocation parameter value of the current audio object to be pre-rendered. This possible design provides a specific implementation method for determining a target number of bits allocated to a current audio object to be pre-rendered.
[0024] In a possible design, the method further includes: sending ratio information between target bit numbers respectively allocated to the multiple audio objects to be pre-rendered, and the ratio information is used to reconstruct the multiple audio objects to be pre-rendered.
[0025] In a second aspect, a bit allocation device for an audio object is provided. The bit allocation device for an audio object may be an encoder or an encoding device including an encoder. For example, the encoder may be a stereo encoder or a multi-channel encoder. For example, the encoding device may be a terminal such as a mobile terminal, a fixed network terminal, etc. Alternatively, the encoding device may be a network device such as a media gateway, a transcoding device, a media resource server, etc. in a wireless access network or a core network.
[0026] In one possible design, the bit allocation device of the audio object is used to execute any one of the methods provided in the first aspect above. The present application can divide the bit allocation device of the audio object into functional modules according to the method provided in the first aspect above. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. Exemplarily, the present application can divide the bit allocation device of the audio object into a pre-rendering module, an acquisition module, a determination module, etc. according to the function. The description of the possible technical solutions and beneficial effects executed by the above-mentioned divided functional modules can refer to the corresponding technical solutions in the first aspect above, and will not be repeated here.
[0027] In another possible design, the bit allocation device of the audio object includes: a processor, which is used to implement any of the methods described in the first aspect above. The device may also include a memory, which is coupled to the processor. When the processor executes the instructions stored in the memory, any of the methods described in the first aspect above can be implemented. The device may also include a communication interface, which is used for the device to communicate with other devices. Exemplarily, the communication port can be a transceiver, circuit, bus, module or other type of communication interface. In the present application, the instructions in the memory can be pre-stored or stored after being downloaded from the Internet when the device is used. The present application does not make a unique limitation on the source of the instructions in the memory. The coupling in the embodiment of the present application is an indirect coupling or connection between units or modules, which can be electrical, mechanical or other forms, and is used for information exchange between units or modules.
[0028] In a third aspect, a computer-readable storage medium is provided, such as a non-transitory computer-readable storage medium, on which a computer program (or instruction) is stored, and when the computer program (or instruction) is executed on a computer, the computer executes any one of the methods provided in the first aspect.
[0029] According to a fourth aspect, a computer program product is provided, which enables any one of the methods provided in the first aspect to be executed when the computer program product is run on the computer.
[0030] In a fifth aspect, an audio system is provided, comprising: an encoding device and a decoding device. The encoding device is used to execute any one of the methods provided in the first aspect. The decoding device is used to receive information sent by the encoding device and perform a decoding process. For example, the encoding device can be an encoder (such as a stereo encoder or a multi-channel encoder) or an encoding device including an encoder (such as a terminal or a network device). Correspondingly, the decoding device can be a decoder (such as a stereo decoder or a multi-channel decoder) or a decoding device including a decoder (such as a terminal or a network device).
[0031] It can be understood that any of the bit allocation devices, computer storage media, computer program products or audio systems of the audio objects provided above can be applied to the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods and will not be repeated here.
[0032] In this application, the name of the above-mentioned audio object bit allocation device does not limit the device or functional module itself. In actual implementation, these devices or functional modules may appear with other names. As long as the functions of each device or functional module are similar to those of this application, they belong to the scope of the claims of this application and their equivalent technologies.
[0033] These and other aspects of the present application will become more apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1A A structural diagram of an audio system applicable to the technical solution provided in an embodiment of the present application;
[0035] Figure 1B A schematic diagram of the structure of an audio system to which the technical solution provided in the embodiment of the present application is applicable Figure 2 ;
[0036] Figure 2 A third structural diagram of an audio system to which the technical solution provided in the embodiment of the present application is applicable;
[0037] Figure 3AA schematic diagram of the structure of an audio system to which the technical solution provided in the embodiment of the present application is applicable Figure 4 ;
[0038] Figure 3B A schematic diagram of the structure of an audio system to which the technical solution provided in the embodiment of the present application is applicable Figure 5 ;
[0039] Figure 4 A schematic diagram of the structure of an audio system to which the technical solution provided in the embodiment of the present application is applicable Figure 6 ;
[0040] Figure 5 A schematic diagram of the hardware structure of a computer device provided in an embodiment of the present application;
[0041] Figure 6 Schematic diagram 1 of a flow chart of a method for allocating bits of an audio object provided in an embodiment of the present application;
[0042] Figure 7 A schematic diagram of a process for determining a target number of bits provided in an embodiment of the present application;
[0043] Figure 8 Schematic diagram of the process of the bit allocation method for audio objects provided in the embodiment of the present application Figure 2 ;
[0044] Fig. 9 A schematic diagram of a method for calculating a content importance parameter value provided in an embodiment of the present application;
[0045] Fig.10 Flowchart 3 of the method for allocating bits of an audio object provided in an embodiment of the present application;
[0046] Fig.11 Schematic diagram of the process of the bit allocation method for audio objects provided in the embodiment of the present application Figure 4 ;
[0047] Fig.12 A schematic diagram of the structure of a bit allocation device for an audio object provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] The following describes some of the terms and technologies involved in this application:
[0049] 1) Audio frame
[0050] Audio data is streamed. In practical applications, in order to facilitate audio processing and transmission, the amount of audio data within a certain period of time is usually taken as a frame of audio, that is, an audio frame. This period of time is called "sampling time", which can be determined according to the requirements of the codec and specific applications. For example, the period of time is 2.5ms to 60ms, where ms is milliseconds.
[0051] 2) Audio object
[0052] An important way to realize three-dimensional audio is object-based audio technology. In object-based audio technology, each audio frame can contain multiple audio objects. During encoding and decoding, the multiple audio objects are encoded and decoded separately.
[0053] In some scenarios, an audio object may also be referred to as an object audio signal or an audio signal.
[0054] 3) Metadata
[0055] Metadata, also known as intermediary data or relay data, is data about data. It is mainly used to describe data properties and supports functions such as indicating storage location, historical data, resource search, and file records. Metadata is information about the organization of data, data domains, and their relationships.
[0056] 4) Other terms
[0057] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0058] In the embodiments of the present application, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, "plurality" means two or more.
[0059] In the present application, the term "at least one" means one or more, and the term "multiple" means two or more. For example, multiple second messages refer to two or more second messages.
[0060] It should be understood that the terms used in the description of the various examples herein are only for describing specific examples and are not intended to be limiting. As used in the description of the various examples and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0061] It should also be understood that the term "and / or" used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. The term "and / or" is a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this application generally indicates that the associated objects before and after are in an "or" relationship.
[0062] It should also be understood that in the various embodiments of the present application, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0063] It should be understood that determining B based on A does not mean determining B only based on A. B can also be determined based on A and / or other information.
[0064] It should also be understood that the term “comprise” (also known as “includes,” “including,” “comprises” and / or “comprising”) when used in this specification specifies the presence of stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0065] It should also be understood that the term "if" may be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting." Similarly, the phrase "if it is determined that ..." or "if [a stated condition or event] is detected" may be interpreted to mean "upon determining that ..." or "in response to determining that ..." or "upon detecting [a stated condition or event]" or "in response to detecting [a stated condition or event]," depending on the context.
[0066] It should be understood that the references to "one embodiment", "some embodiments", or "a possible implementation" throughout the specification mean that specific features, structures, or characteristics related to the embodiment or implementation are included in at least one embodiment of the present application. Therefore, the references to "in one embodiment" or "in some embodiments", or "a possible implementation" appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0067] In some embodiments, the bit allocation method for audio objects provided in the embodiments of the present application can be applied to a stereo encoder of a terminal. For example, the terminal can be a mobile terminal, a fixed network terminal, etc.
[0068] like Figure 1A FIG. 1 is a schematic diagram of the structure of an audio system 1 applicable to the technical solution provided in the embodiment of the present application. The audio system 1 includes a first terminal 11 and a second terminal 12 .
[0069] The first terminal 11 includes an audio collection module 111 , a stereo encoder 112 and a channel encoder 113 . The second terminal 12 includes a channel decoder 121 , a stereo decoder 122 and an audio playback module 123 .
[0070] based on Figure 1A In the first terminal 11, the audio acquisition module 111 is used to acquire a stereo signal, and the stereo encoder 112 is used to perform stereo encoding on the stereo signal. The channel encoder 113 is used to perform channel encoding on the stereo encoded signal. Optionally, the channel encoded signal is transmitted in a digital channel after being processed by the first communication device 13. After being transmitted to the second terminal 12 through the second communication device 14, any one of the first communication device 13 and the second communication device 14 can be a wireless network communication device or a wired network communication device.
[0071] based on Figure 1A In the second terminal 12, the channel decoder 121 is used to perform channel decoding on the received signal. The stereo decoder 122 is used to perform stereo decoding on the channel-decoded signal. The audio playback module 123 is used to play back the stereo-decoded signal.
[0072] Figure 1AThe first terminal 11 and the first communication device 13 are transmitting side devices, and the second terminal 12 and the second communication device 14 are receiving side devices. In some scenarios, the first terminal 11 and the first communication device 13 can also be used as receiving side devices, and correspondingly, the second terminal 12 and the second communication device 14 are used as transmitting side devices. In this case, the first terminal 11 can also include a channel decoder 121, a stereo decoder 122 and an audio playback module 123, and the second terminal 12 can also include an audio acquisition module 111, a stereo encoder 112 and a channel encoder 113, as shown in FIG. Figure 1B The functions of each module can be found in the above text and will not be described here.
[0073] In some embodiments, the bit allocation method for audio objects provided in the embodiments of the present application can be applied to a stereo encoder of a network device (including a wireless network device or a core network device). For example, the network device can be a media gateway, a transcoding device, a media resource server, etc. in a wireless access network or a core network.
[0074] like Figure 2 FIG. 2 is a schematic diagram of the structure of an audio system 2 applicable to the technical solution provided in the embodiment of the present application. The audio system 2 includes: a first network device 21 and a second network device 22.
[0075] The first network device 21 includes: a first channel decoder 211, other audio decoders 212, a stereo encoder 213 and a first channel encoder 214. The second network device 22 includes: a second channel decoder 221, a stereo decoder 222, other audio encoders 223 and a second channel decoder 224.
[0076] In the first network device 21, the first channel decoder 211 is used to perform channel decoding on the received signal. The other audio decoder 212 is used to transcode the channel-decoded signal. The stereo encoder 213 is used to perform stereo encoding on the transcoded signal. The first channel encoder 214 is used to perform channel encoding on the stereo-encoded signal.
[0077] In the second network device 22, the second channel decoder 221 is used to perform channel decoding on the received signal. The stereo decoder 222 is used to perform stereo decoding on the channel-decoded signal. The other audio encoder 223 is used to transcode the stereo-decoded signal. The second channel decoder 224 is used to perform channel coding on the transcoded signal.
[0078] It should be noted that the stereo encoding and decoding process can be part of a multi-channel codec. For example, the encoding end performs multi-channel encoding on the collected multi-channel signal, which may include: the encoding end performs downmixing on the collected multi-channel signal to obtain a stereo signal, and encodes the stereo signal. The decoding end decodes the bit stream according to the multi-channel signal to obtain a stereo signal, and performs upmixing on the stereo signal to restore the multi-channel signal.
[0079] Based on this, the bit allocation method for the audio object provided in the embodiment of the present application can also be applied to the multi-channel encoder of the terminal. The audio system where the multi-channel encoder is located can refer to Figure 3A or Figure 3B Alternatively, the bit allocation method for audio objects provided in the embodiment of the present application can also be applied to a multi-channel encoder of a network device (including a wireless network device or a core network device). The audio system where the multi-channel encoder is located can refer to Figure 4 .
[0080] like Figure 3A FIG. 1 is a schematic diagram of the structure of an audio system 3 applicable to the technical solution provided in the embodiment of the present application. Figure 3A is based on Figure 1A To draw, specifically Figure 1A The stereo encoder 112 in is replaced by a multi-channel encoder 114 , and the stereo decoder 122 is replaced by a multi-channel decoder 124 .
[0081] based on Figure 3A In the first terminal 11, the audio acquisition module 111 is used to acquire a multi-channel signal. The multi-channel encoder 114 is used to perform multi-channel encoding on the multi-channel signal, including stereo encoding. The channel encoder 113 is used to perform channel encoding on the multi-channel encoded signal. The channel-encoded signal is processed by the first communication device 13 and transmitted in a digital channel. After being transmitted to the second terminal 12 via the second communication device 14.
[0082] based on Figure 3A In the second terminal 12, the channel decoder 121 is used to perform channel decoding on the received signal. The multi-channel decoder 124 is used to perform multi-channel decoding on the channel-decoded signal, including stereo decoding. The audio playback module 123 is used to play back the multi-channel decoded signal.
[0083] like Figure 3B As shown, it is a structural schematic diagram of another audio system 3 applicable to the technical solution provided in the embodiment of the present application. Figure 3B is based on Figure 1B and Figure 3A The interpretation of the relevant content can be based on Figure 1B and Figure 3A As well as the above Figure 1B and Figure 3A The text description is inferred and will not be repeated here.
[0084] like Figure 4 As shown, it is a structural diagram of an audio system 4 applicable to the technical solution provided in the embodiment of the present application. Figure 4 is based on Figure 2 To draw, specifically Figure 2 The stereo encoder 213 in the audio system is replaced by a multi-channel encoder 215, and the stereo decoder 223 is replaced by a multi-channel decoder 225. The multi-channel encoder 215 is used to perform multi-channel encoding on the signal transcoded by the other audio decoder 212, including stereo encoding. The first channel encoder 214 is used to perform channel encoding on the multi-channel encoded signal. The multi-channel decoder 225 is used to perform multi-channel decoding on the signal channel-decoded by the second channel decoder 221, including stereo decoding. The other audio encoders 223 are used to transcode the multi-channel decoded signal. The functions of other modules / devices can refer to the above description. Figure 2 The description of the functions of the corresponding modules in will not be repeated here.
[0085] In some embodiments, the bit allocation method of the audio object provided in the embodiment of the present application can be applied to the audio encoder (audio encoding) in the virtual reality (VR) streaming service. In this scenario, the end-to-end processing flow of the audio object includes: the audio object A is preprocessed (audio preprocessing) after passing through the acquisition module (acquisition), and the preprocessing operation may include filtering out the low-frequency part of the signal, usually extracting the orientation information in the signal with 20Hz (hertz) or 50Hz as the demarcation point, and then encoding (audio encoding) and packaging (file / segment encapsulation) through the audio encoder. The signal packaged by the encoding process is sent (delivered) to the decoding end. The decoding end unpacks the received signal (file / segment decapsulation), and decodes it (audio decoding) through the audio decoder, and then performs binaural rendering (audio rendering) on the decoded signal, and the rendered signal is mapped to the listener's headphones (headphones). The headphones can be independent headphones or headphones on glasses devices such as HTC VIVE.
[0086] The modules / devices in any of the above audio systems are distinguished from the perspective of logical functions. Part or all of the above modules / devices can be implemented by software, hardware, or a combination of software and hardware.
[0087] like Figure 5 FIG. 5 is a schematic diagram of the hardware structure of a computer device 5 provided in an embodiment of the present application. The computer device 5 can be used to execute the bit allocation method for audio objects provided in an embodiment of the present application.
[0088] Optionally, the computer device 5 can be used to implement the above Figure 1A , Figure 1B or Figure 2 The function of the stereo encoder in , or for implementing Figure 3A , Figure 3B or Figure 4 The function of the multi-channel encoder in .
[0089] Optionally, the computer device 5 can be used to implement Figure 1A The function of the first terminal in Figure 1B The function of the first terminal or the second terminal in Figure 2 The function of the first network device in Figure 3A The function of the first terminal in Figure 3B The function of the first terminal or the second terminal in Figure 4 The function of the first network device.
[0090] like Figure 5 As shown, the computer device 5 may include a processor 51, a memory 52, a communication interface 53, and a bus 54. The processor 51, the memory 52, and the communication interface 53 may be connected via the bus 54.
[0091] The processor 51 is the control center of the computer device 5, and may be a general-purpose central processing unit (CPU) or other general-purpose processors, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0092] As an example, the processor 51 may include one or more CPUs, such as Figure 5 CPU0 and CPU1 are shown in the figure.
[0093] The memory 52 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.
[0094] In a possible implementation, the memory 52 may exist independently of the processor 51. The memory 52 may be connected to the processor 51 via a bus 54 and used to store data, instructions or program codes. When the processor 51 calls and executes the instructions or program codes stored in the memory 52, the bit allocation method for the audio object provided in the embodiment of the present application can be implemented.
[0095] In another possible implementation, the memory 52 may also be integrated with the processor 51 .
[0096] The communication interface 53 is used for connecting the computer device 5 with other devices through a communication network, and the communication network may be Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 53 may include a receiving unit for receiving data and a sending unit for sending data.
[0097] The bus 54 may be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0098] It should be pointed out that Figure 5 The structure shown in the figure does not constitute a limitation on the computer equipment, except Figure 5In addition to the components shown, the computer device 5 may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0099] The following describes the bit allocation method for audio objects provided by the embodiment of the present application in conjunction with the accompanying drawings. The method can be applied to an encoder. For example, the encoder can be Figure 1A , Figure 1B or Figure 2 The stereo encoder in Figure 3A , Figure 3B or Figure 4 It can also be a multi-channel encoder in VR streaming services.
[0100] like Figure 6 , which is a flow chart of a method for allocating bits of an audio object provided in an embodiment of the present application. Figure 6 The method shown may include the following steps:
[0101] S101: The encoder pre-renders a plurality of audio objects to be pre-rendered in an audio frame to be encoded, respectively, to obtain a plurality of pre-rendering audio objects (prerendering object audio), wherein the audio objects to be pre-rendered correspond to the pre-rendering audio objects one by one.
[0102] The audio frame to be encoded may be any three-dimensional audio frame with encoding requirements. In order to distinguish the audio objects before pre-rendering from the audio objects after pre-rendering, in an embodiment of the present application, the audio objects before pre-rendering are referred to as audio objects to be pre-rendered, and the audio objects after pre-rendering are referred to as pre-rendered audio objects. The number of audio objects to be pre-rendered included in the audio frame to be encoded may be predefined. The "multiple audio objects to be pre-rendered" in S101 may be part or all of the audio objects included in the audio frame to be encoded. It can be understood that if the multiple audio objects to be pre-rendered are part of the audio objects included in the audio frame to be encoded, the bit allocation method for the other part of the audio objects may refer to the prior art.
[0103] The embodiments of the present application do not limit the specific implementation method of pre-rendering. For example, the pre-rendering method can be a method used when actually rendering an audio object, such as a method based on a head related transfer function (HRTF), or a low-complexity rendering method that can obtain a result having similar characteristics to the result of actually rendering an audio object.
[0104] Optionally, the metadata information used in pre-rendering is consistent with the metadata information used in actual rendering (i.e., the same or not much different). In this way, it is helpful to make the perceptual importance parameter values of multiple pre-rendered audio objects subsequently obtained by the encoder closer to the perceptual importance parameter values of multiple audio objects obtained by the actual rendering of the decoder, thereby helping to improve the overall quality and coding efficiency of the reconstructed audio objects after bit allocation using the technical solution.
[0105] S102: The encoder obtains a perceptual importance parameter value of each of the plurality of pre-rendered audio objects.
[0106] The perceptual importance parameter value of the current pre-rendered audio object is used to indicate the perceptual importance of the current pre-rendered audio object among the multiple pre-rendered audio objects. The perceptual importance may include: energy intensity and / or spectrum variation. The current pre-rendered audio object may be any pre-rendered object among the multiple pre-rendered audio objects.
[0107] The perceptual importance parameter of the current pre-rendered audio object may include: a parameter indicating the energy intensity and / or spectrum variation degree of the current pre-rendered audio object in the multiple pre-rendered audio objects within a period of time.
[0108] The degree of perceived importance may be measured by one perceived importance parameter or by a combination of multiple perceived importance parameters.
[0109] The embodiment of the present application does not limit what kind of parameters the perceived importance parameters are. For example, the perceived importance parameters may include one or more of the following parameters 1)-3):
[0110] 1) Energy importance parameter. Among them, the energy importance parameter of the current pre-rendered audio object is calculated based on the energy of the current pre-rendered audio object, and is used to indicate the ratio of the energy of the current pre-rendered audio object to the sum of the energies of the multiple pre-rendered audio objects. Optionally, the energy importance parameter of the current pre-rendered audio object may be the ratio, or a value obtained by mapping the ratio according to a preset algorithm. For example, mapping values corresponding to different ratios may be preset, for example, the value of the energy importance parameter in the interval [0.8, 0.9] may be mapped to 0.85 or 0.8, etc. Of course, other mapping methods may also be used, and the specific mapping method is not limited in the embodiments of the present invention.
[0111] 2) Perceptual intensity importance parameter. Among them, the perceptual intensity importance parameter of the current pre-rendered audio object is calculated by combining the human ear auditory curve and the energy of the current pre-rendered audio object, and is used to indicate the ratio between the sum of the energies of a preset number of frequency bands with the largest energy in multiple frequency bands of the current pre-rendered audio object and the sum of the energies of a preset number of frequency bands with the largest energy in multiple frequency bands of the multiple pre-rendered audio objects. Optionally, the perceptual intensity importance parameter of the current pre-rendered audio object can be the ratio, or a value obtained by mapping the ratio. The specific mapping method can refer to the method described in the energy importance parameter part.
[0112] Among them, the preset number of frequency bands with the largest energy in multiple frequency bands can be: the first preset number of frequency bands in the sequence obtained after sorting the multiple frequency bands in order of energy from large to small, or: the last preset number of frequency bands in the sequence obtained after sorting the multiple frequency bands in order of energy from small to large.
[0113] 3) Spectral flatness parameter: The spectral flatness parameter of the current pre-rendered audio object is used to indicate the spectral flatness of the current pre-rendered audio object among the multiple pre-rendered audio objects.
[0114] Optionally, the perceptual importance parameter value of the current pre-rendered audio object can be obtained based on the features of the multiple pre-rendered audio objects, or can be obtained based on the features of the multiple pre-rendered audio objects after shaping. The feature can be a time domain feature, a frequency domain feature, or a combination of a time domain feature and a frequency domain feature. The following description is given by taking the method of obtaining the feature based on the multiple pre-rendered audio objects as an example.
[0115] The following is an exemplary description of how to obtain the energy importance parameter, the perception intensity importance parameter, and the spectrum flatness parameter:
[0116] 1) Energy importance parameters
[0117] Optionally, the energy importance parameter value of the current pre-rendered audio object may include: a ratio of the energy value of the current pre-rendered audio object to the sum of the energy values of the multiple pre-rendered audio objects, or a parameter value determined according to the ratio of the energy value of the current pre-rendered audio object to the sum of the energy values of the multiple pre-rendered audio objects. The parameter value may be considered to be a value obtained after processing (e.g., mapping) the ratio, and the specific processing method is not limited in the embodiment of the present application.
[0118] For example, the energy importance parameter value E_imp of the i-th pre-rendered audio object i Satisfies the following formula 1:
[0119]
[0120] Among them, E i represents the energy value of the i-th pre-rendered audio object. 1≤i≤N, where N represents the number of pre-rendered audio objects in S102. Indicates the total energy value of N pre-rendered audio objects. E_imp i ∈[0,1].
[0121] 2) Perception intensity importance parameters
[0122] Optionally, the perceived intensity importance parameter value of the current pre-rendered audio object can be obtained based on the band perceived intensity parameter values of some or all frequency bands of the current pre-rendered audio object. The band perceived intensity parameter value of a frequency band is calculated by combining the human ear hearing curve and the energy of the frequency band, and is used to indicate the energy strength of the frequency band in the current pre-rendered audio object.
[0123] For example, the perceptual intensity importance parameter value Intensity_imp of the i-th pre-rendered audio object i It can be obtained by following the steps below:
[0124] a) The encoder calculates a frequency band perceptual intensity parameter value for each frequency band of the i-th pre-rendered audio object.
[0125] Specifically: the encoder divides the frequency domain resources of the i-th pre-rendered audio object into multiple frequency bands, and then obtains the frequency band perception intensity parameter values of each of the multiple frequency bands. The embodiment of the present application does not limit how to divide the frequency bands. For example, the frequency band perception intensity parameter value p of the frequency band b in the multiple frequency bands is i (b) Satisfy the following formula 2:
[0126] Formula 2: p i (b) = E i (b)-T(b).
[0127] Among them, E i (b) represents the energy value of frequency band b of the i-th pre-rendered audio object, and T(b) is a constant factor calculated in frequency band b according to the human hearing curve, and its value can be summarized based on experimental experience. For example: Among them, b f Indicates the center frequency value corresponding to the center frequency point of the frequency band b.
[0128] Based on Formula 2, the encoder can obtain a frequency band perceptual intensity parameter value of each frequency band of the i-th pre-rendered audio object.
[0129] b) The encoder sorts the band perception strength parameter values of each band obtained in a) in descending order to obtain the set P shown in formula 3: i (b)
[0130] Formula 3: P i (b)≡{p i (b 1 ), p i (b 2 ), ..., p i (b L )}.
[0131] Among them, P i (b) represents the set of frequency band perceptual intensity parameter values of the i-th pre-rendered audio object after sorting, L represents the number of frequency bands into which the i-th pre-rendered audio object is divided.
[0132] c) The encoder is based on the set P i (b) Obtain the perceptual intensity importance parameter value of the i-th pre-rendered audio object.
[0133] Exemplarily, the encoder takes the set P i (b) The first l values, these l values and Intensity_imp i Satisfies the following formula 4:
[0134]
[0135] Where l≤L, Intensity_imp i ∈[0,1].
[0136] 3) Spectral flatness parameter
[0137] For example, the spectrum flatness parameter value Flatness_imp of the i-th pre-rendered audio object i Satisfies the following formula 5:
[0138] Formula 5:
[0139] Among them, R i (k) = E i (k) / E i , E i (k) represents the energy value of the kth frequency band of the i-th pre-rendered audio object, and B is the number of frequency bands of the i-th pre-rendered audio object. Flatness_imp i ∈[0,1].
[0140] S103: The encoder obtains a bit allocation parameter value for each of the multiple audio objects to be pre-rendered based on the perceptual importance parameter values for each of the multiple pre-rendered audio objects.
[0141] The bit allocation parameter of the current audio object to be pre-rendered is used to indicate the target number of bits allocated to the current audio object to be pre-rendered. The current audio object to be pre-rendered may be any one of the multiple audio objects to be pre-rendered. That is, the encoder may obtain the bit allocation parameter value of each of the multiple audio objects to be pre-rendered in the same manner as obtaining the bit allocation parameter value of the current audio object to be pre-rendered.
[0142] Optionally, the perceptual importance parameter value of the current pre-rendered audio object and the bit allocation parameter value of the current audio object to be pre-rendered satisfy a certain predefined rule, and the rule may be represented by a function or may not be represented by a function. The current pre-rendered audio object is obtained after pre-rendering the current audio object to be pre-rendered. The encoder may determine the bit allocation parameter value of the current audio object to be pre-rendered based on the rule and the perceptual importance parameter value of the current pre-rendered audio object.
[0143] The following uses a function to represent the rule as an example to illustrate how to obtain the bit allocation parameter value of the i-th audio object to be pre-rendered:
[0144] In some implementations, when the perceptual importance parameter of the i-th pre-rendered audio object includes multiple parameters, the encoder may introduce a parameter Important_P in the process of calculating the bit allocation parameter value of the i-th audio object to be pre-rendered. i , parameter Important_P i It is used to indicate the overall perceived importance of the i-th pre-rendered audio object among the N audio objects to be pre-rendered. In contrast, different perceived importance parameter values of the i-th pre-rendered audio object are used to indicate the perceived importance of the i-th pre-rendered audio object at different angles among the N audio objects to be pre-rendered.
[0145] Optional, Important_P i The value of can be obtained by a certain operation of the multiple perceived importance parameter values, such as Important_P i The value of can satisfy the following formula 6:
[0146] Formula 6: Important_P i =f(parm_p i_1 ,parm_p i_2 ,…,parm_p i_m ).
[0147] Among them, parm_p i_j represents the jth perceptual importance parameter value of the i-th pre-rendered audio object. 1≤j≤m, where m is the number of perceptual importance parameters of the i-th pre-rendered audio object.
[0148] Optionally, the functional relationship represented by Formula 6 may be linear or nonlinear.
[0149] When the perceptual importance parameter includes an energy importance parameter, a perceptual intensity importance parameter, and a spectrum flatness parameter, the above formula 6 can be specifically expressed as the following formula 7:
[0150] Formula 7: Important_P i =f(E_imp i ,Intensity_imp i ,Flatness_imp i ).
[0151] Optionally, the functional relationship represented by Formula 7 may be linear or nonlinear.
[0152] For example, the above formula 7 can be specifically expressed as the following formula 8:
[0153] Formula 8: Important_P i =a 1 ·E impi +a 2 Intensity impi +a 3 Flatness impi .
[0154] Among them, a 1 、a 2 and a 3 is a constant whose value can be obtained through experimental experience.
[0155] Optional, a 1 、a 2 and a 3 The following formula 9 is satisfied:
[0156] Formula 9: a 1 +a 2 +a 3 =1, a 1 ,a 2 ,a 3 ∈[0,1].
[0157] Optional, bit allocation parameter value Important_Bit of the i-th audio object to be pre-rendered iThe following formula 10 is satisfied:
[0158] Formula 10: Important_Bit i =f(Important_P i ).
[0159] Optionally, the functional relationship represented by Formula 10 may be linear or nonlinear.
[0160] Optionally, the bit allocation parameter value of the current audio object to be pre-rendered may include: a first ratio, or a parameter value determined according to the first ratio, wherein the first ratio is a ratio between the perceptual importance parameter value of the current pre-rendered audio object and the sum of the perceptual importance parameter values of the multiple pre-rendered audio objects.
[0161] Specifically, S103 may include: the encoder first uses the ratio between the perceptual importance parameter value of the current pre-rendered audio object and the sum of the perceptual importance parameter values of the multiple pre-rendered audio objects as a first ratio; then, uses the first ratio as the bit allocation parameter value of the current audio object to be pre-rendered, or determines a parameter value according to the first ratio, and uses the parameter value as the bit allocation parameter value of the current audio object to be pre-rendered. The parameter value can be considered as a value obtained after processing the first ratio, and the embodiment of the present application does not limit the specific processing method.
[0162] For example, Formula 10 can be further expressed as the following Formula 11:
[0163]
[0164] It can be seen that in this example, Important_Bit i ∈[0,1].
[0165] In some other implementations, the encoder obtains the bit allocation parameter value Important_Bit of the i-th audio object to be pre-rendered. i The parameter Important_P can be omitted in the process i For example, the above formula 10 can be replaced by the following formula 12:
[0166] Formula 12: Important_Bit i =f(parm_p i_1 ,parm_p i_2 ,…,parm_p i_m ).
[0167] Optionally, the functional relationship represented by Formula 12 may be linear or nonlinear.
[0168] When the perceptual importance parameter includes an energy importance parameter, a perceptual intensity importance parameter, and a spectrum flatness parameter, the above formula 12 can be specifically expressed as the following formula 13:
[0169] Formula 13: Important_Bit i =f(E_imp i ,Intensity_imp i ,Flatness_imp i ).
[0170] The specific embodiment of Formula 13 is not limited in this embodiment of the present application.
[0171] S104: The encoder obtains the total number of bits to be allocated corresponding to the multiple audio objects to be pre-rendered.
[0172] The total number of bits to be allocated is the total number of bits allocated for the multiple audio objects to be pre-rendered. The total number of bits to be allocated is known in advance by the encoder, and the specific implementation method can refer to the prior art. The embodiment of the present application does not limit how the encoder knows in advance, for example, it can be indicated by a user, or it can be predefined.
[0173] S105: The encoder determines target numbers of bits to be allocated respectively for the multiple audio objects to be pre-rendered based on the total number of bits to be allocated and the bit allocation parameter values of the multiple audio objects to be pre-rendered.
[0174] Specifically, the encoder determines the target number of bits allocated to the current audio object to be pre-rendered based on the total number of bits to be allocated and the bit allocation parameter value of the current audio object to be pre-rendered.
[0175] In some implementations, the ratio of the target number of bits allocated to the current audio object to be pre-rendered to the total number of bits to be allocated is equal to the third ratio, or equal to the parameter value determined according to the third ratio. The third ratio is the ratio between the bit allocation parameter value of the current audio object to be pre-rendered and the sum of the bit allocation parameter values of the plurality of audio objects to be pre-rendered. The parameter value can be considered as a value obtained after processing the third ratio, and the specific processing method is not limited in the embodiment of the present application.
[0176] Specifically, the encoder first uses the ratio between the bit allocation parameter value of the current audio object to be pre-rendered and the sum of the bit allocation parameter values of the multiple audio objects to be pre-rendered as the third ratio; then, the product of the third ratio and the total number of bits to be allocated is used as the target number of bits allocated to the current audio object to be pre-rendered, or obtains a parameter value according to the third ratio, and uses the product of the parameter value and the total number of bits to be allocated as the target number of bits allocated to the current audio object to be pre-rendered. The parameter value may be a value obtained by the encoder after processing the third ratio, and the embodiment of the present application does not limit the processing method.
[0177] For example, if "the ratio of the target number of bits allocated to the current audio object to be pre-rendered to the total number of bits to be allocated is equal to the third ratio", then the target number of bits allocated to the current audio object to be pre-rendered can be determined by the product of the bit allocation parameter value of the current audio object to be pre-rendered and the total number of bits to be allocated.
[0178] For example, taking the current audio object to be pre-rendered as the i-th audio object to be pre-rendered as an example, the target number of bits allocated to the i-th audio object to be pre-rendered is Bits_object i It can be obtained by the following formula 14:
[0179] Formula 14: Bits_object i =Important_Bit i *Bits_available.
[0180] Among them, Bits_available represents the total number of bits to be allocated.
[0181] In other implementations, such as Figure 7 As shown, S105 may include the following S105A-S105B:
[0182] S105A: The encoder determines the priority level of each of the multiple audio objects to be pre-rendered based on the correspondence between the multiple bit allocation parameter values and the multiple priority levels and the bit allocation parameter values of the multiple audio objects to be pre-rendered.
[0183] Specifically, the encoder determines the priority level of the current audio object to be pre-rendered based on the correspondence between multiple bit allocation parameter values and multiple priority levels and the bit allocation parameter value of the current audio object to be pre-rendered.
[0184] In one implementation, the encoder determines the priority level of each of the multiple audio objects to be pre-rendered based on the correspondence between the intervals of multiple bit allocation parameter values and the multiple priority levels, and the bit allocation parameter values of each of the multiple audio objects to be pre-rendered.
[0185] Specifically, the encoder determines the priority level of the current audio object to be pre-rendered based on the correspondence between the intervals where the multiple bit allocation parameter values are located and the multiple priority levels, and the bit allocation parameter value of the current audio object to be pre-rendered.
[0186] The correspondence between the intervals where the multiple bit allocation parameter values are located and the multiple priority levels can be predefined. The embodiment of the present application does not limit the number of priority levels and the intervals where the bit allocation parameter values corresponding to each priority level are located, and can be determined according to actual needs.
[0187] Optionally, a higher priority level corresponds to a greater number of target bits. For example, as shown in Table 1, it is an example of a corresponding relationship between intervals where multiple bit allocation parameter values are located and multiple priority levels.
[0188] Table 1
[0189] The interval in which the bit allocation parameter value lies Priority Level [0.9,1] 10 [0.8,0.9) 9 [0.7,0.8) 8 [0.6,0.7) 7 [0.5,0.6) 6 [0.4,0.5) 5 [0.3,0.4) 4 [0.2,0.3) 3 [0.1,0.2) 2 [0,0.1) 1
[0190] Optionally, a higher priority level corresponds to a smaller target number of bits. For example, the priority level 10-1 in Table 1 can be replaced by priority 1-10.
[0191] In this implementation, different bit allocation parameter values belonging to the same interval correspond to the same priority level number.
[0192] In another implementation, the encoder approximates the bit allocation parameter values of the multiple audio objects to be pre-rendered to corresponding preset values based on one or more processing methods such as truncation, truncation, rounding, etc.; and then determines the priority level of the multiple audio objects to be pre-rendered based on the corresponding relationship between the multiple preset values and the multiple priority levels.
[0193] Specifically, the encoder approximates the bit allocation parameter value of the current audio object to be pre-rendered to a preset value; and then determines the priority level of the current audio object to be pre-rendered based on the corresponding relationship between multiple preset values and multiple priority levels.
[0194] In this implementation, different bit allocation parameter values corresponding to the same preset value may correspond to the same priority level.
[0195] S105B: The encoder determines target numbers of bits to be allocated respectively to the multiple audio objects to be pre-rendered based on the total number of bits to be allocated and the priority levels of the multiple audio objects to be pre-rendered.
[0196] Specifically, the encoder determines a target number of bits to be allocated for the current audio object to be pre-rendered based on the total number of bits to be allocated and the priority levels of the multiple audio objects to be pre-rendered.
[0197] Optionally, the ratio of the target number of bits allocated to the current audio object to be pre-rendered to the total number of bits to be allocated is equal to the fourth ratio, or equal to the parameter value determined according to the fourth ratio. The fourth ratio is the ratio between the priority level of the current audio object to be pre-rendered and the sum of the priority levels of the plurality of audio objects to be pre-rendered. The parameter value can be considered as a value obtained after processing the fourth ratio, and the specific processing method is not limited in the embodiment of the present application.
[0198] Specifically, the encoder first uses the ratio of the priority level of the current audio object to be pre-rendered to the sum of the priority levels of the multiple audio objects to be pre-rendered as the fourth ratio; then, the product of the fourth ratio and the total number of bits to be allocated is used as the target number of bits allocated to the current audio object to be pre-rendered, or, determines a parameter value according to the fourth ratio, and uses the product of the parameter value and the total number of bits to be allocated as the target number of bits allocated to the current audio object to be pre-rendered.
[0199] For example, assuming that the multiple audio objects to be pre-rendered are audio objects 1-3 to be pre-rendered, and the bit allocation parameter values of the audio objects 1-3 to be pre-rendered are 0.6, 0.25, and 0.15, respectively. Based on Table 1, the priority levels of the audio objects 1-3 to be pre-rendered are 7, 3, and 2, respectively, and the total number of bits to be allocated corresponding to the three audio objects to be pre-rendered is Bits_available. Then: the proportions of the target number of bits allocated to the audio objects 1-3 to be pre-rendered in Bits_available are respectively: It can be seen that the target number of bits allocated to the audio objects 1-3 to be pre-rendered are:
[0200] The bit allocation method for audio objects provided in this embodiment takes into account the differences in the perceptual characteristics of different pre-rendered audio objects at the rendering playback end when allocating bits for the audio objects to be pre-rendered. Compared with the technical solution of encoding different audio objects with the same number of bits in the traditional technology, it helps to improve the overall quality of the reconstructed audio objects. For example, the higher the degree of perceptual importance indicated by the perceptual importance parameter value of a pre-rendered audio object, the more bits the encoder can allocate to the audio object to be pre-rendered corresponding to the pre-rendered audio object (i.e., the audio object of the pre-rendered audio object before pre-rendering), and the number of bits can be used to encode the audio object to be pre-rendered. At this time, the quality of the audio object reconstructed by the decoder will be higher. In this way, it helps to improve the overall quality of the reconstructed audio frames containing multiple audio objects. At the same time, the coding efficiency can be improved.
[0201] like Figure 8 FIG. 1 is a flow chart of another method for allocating bits of an audio object provided in an embodiment of the present application. The explanation of the relevant terms in this embodiment can be referred to Figure 6 The embodiment shown. Figure 8 The method shown may include the following steps:
[0202] S201: The encoder obtains content importance parameter values of respective audio objects to be pre-rendered in an audio frame to be encoded.
[0203] The content importance parameter value of the current audio object to be pre-rendered is used to indicate the importance of the sound type represented by the content of the current audio object to be pre-rendered among the sound types represented by the contents of the multiple audio objects to be pre-rendered.
[0204] It should be noted that the sound type represented by the content of the current audio object remains unchanged before and after pre-rendering. Therefore, the content importance parameter of the current audio object to be pre-rendered is equivalent to: the content importance parameter of the current pre-rendered audio object. The content importance parameter value of the current pre-rendered audio object is used to indicate the importance of the sound type represented by the content of the current pre-rendered audio object among the sound types represented by the contents of the multiple pre-rendered audio objects.
[0205] Optionally, the sound type may include at least one of the following: voice, music, sound effect, ambient sound, noise, etc. Of course, in actual implementation, the sound type may be divided in other ways.
[0206] Among them, which type of sound is more important and which type of sound is less important is relative, and the embodiment of the present application does not limit the determination method thereof, and the specific determination can be based on actual needs. For example, it can be defined that the importance of sound types from high to low is: voice, music, sound effect, ambient sound, and noise.
[0207] The content importance parameter value of the current audio object to be pre-rendered may be obtained from metadata of the audio frame to be encoded, or from features of the current audio object to be pre-rendered, or from features of an audio object obtained after shaping the current audio object to be pre-rendered.
[0208] In some implementations, the content importance parameter value of each of the plurality of audio objects to be pre-rendered obtained from the metadata of the audio frame to be encoded can be expressed as the following formula 15:
[0209] Formula 15: Important_C≡{I_C 1 ,I_C 2 ,…,I_C N}.
[0210] Among them, {I_C 1 ,I_C 2 ,…,I_C N} is to obtain the content importance parameter values of N audio objects to be pre-rendered from the metadata of the audio frame to be encoded, and all of them are constants. 1 ,I_C 2 ,…,I_C N} belongs to (0,1].
[0211] In other implementations, such as Fig. 9 As shown, assuming that the importance of the predefined sound types is from high to low: speech, music, sound effects, ambient sound, noise, the encoder can use a known audio classifier to obtain the confidence score that the sound type represented by the content of each of the multiple audio objects to be pre-rendered is speech, that is, to obtain multiple confidence scores corresponding to the multiple audio objects to be pre-rendered, wherein one audio object to be pre-rendered corresponds to one confidence score. Then, for each audio object to be pre-rendered, the encoder calculates the content importance parameter value of the audio object to be pre-rendered based on the corresponding relationship between the confidence score corresponding to the audio object to be pre-rendered and the content importance parameter value.
[0212] It can be understood that the implementation method can be summarized as: using the confidence score that the sound type represented by the content of an audio object to be pre-rendered is speech, to distinguish (or reflect) whether the sound represented by the content of the audio object to be pre-rendered is speech, music, sound effect, ambient sound or noise, etc., thereby determining the content importance parameter value of the audio object to be pre-rendered.
[0213] For example, the content importance parameter value of the i-th audio object to be pre-rendered is Important_C i The following formula 16 can be satisfied:
[0214]
[0215] Among them, P_C i Indicates the confidence score that the i-th audio object to be pre-rendered is speech, P_C i ∈(0,1]. A and B are constant factors used to make Important_C i ∈(0,1].
[0216] S202: The encoder obtains a bit allocation parameter value for each of the plurality of audio objects to be pre-rendered based on the content importance parameter value for each of the plurality of audio objects to be pre-rendered.
[0217] For example, the bit allocation parameter value Important_Bit of the i-th audio object to be pre-rendered i The following formula 17 is satisfied:
[0218] Formula 17: Important_Bit i =f(Important_C i ).
[0219] Optionally, the functional relationship represented by Formula 17 may be linear or nonlinear.
[0220] Optionally, the bit allocation parameter value of the current audio object to be pre-rendered includes a ratio, or a parameter value determined according to a ratio. The ratio is a ratio between a perceptual importance parameter value of the current pre-rendered audio object and a sum of perceptual importance parameter values of the plurality of pre-rendered audio objects. The parameter value can be considered to be a value obtained after processing the ratio, and the specific processing method is not limited in the embodiment of the present application.
[0221] For example, Formula 17 can be further expressed as the following Formula 18:
[0222]
[0223] It can be seen that in this example, Important_Bit i ∈[0,1].
[0224] S203: The encoder obtains the total number of bits to be allocated corresponding to the multiple audio objects to be pre-rendered.
[0225] The relevant explanation and examples of S203 can refer to the above S104, which will not be repeated here.
[0226] S204: The encoder determines target numbers of bits to be allocated respectively for the multiple audio objects to be pre-rendered based on the total number of bits to be allocated and the bit allocation parameter values of the multiple audio objects to be pre-rendered.
[0227] The relevant explanation and examples of S204 can refer to the above S105 and will not be repeated here.
[0228] The bit allocation method for audio objects provided in this embodiment takes into account the differences in content features of different audio objects to be pre-rendered when allocating the number of bits for the audio objects to be pre-rendered. Compared with the technical solution of encoding different audio objects with the same number of bits in the traditional technology, it helps to improve the overall quality of the reconstructed audio objects. For example, the higher the degree of content importance indicated by the content importance parameter of an audio object to be pre-rendered, the more bits the encoder can allocate to the audio object to be pre-rendered, and the number of bits can be used for encoding. At this time, the quality of the audio object reconstructed by the decoder will be higher. In this way, it helps to improve the overall quality of the reconstructed audio frames containing multiple audio objects. At the same time, it can improve the coding efficiency.
[0229] like Fig.10 FIG. 1 is a flow chart of another method for allocating bits of an audio object provided in an embodiment of the present application. Figure 6 and Figure 8 The embodiment shown. Fig.10 The method shown may include the following steps:
[0230] S301: The encoder pre-renders a plurality of audio objects to be pre-rendered in an audio frame to be encoded, respectively, to obtain a plurality of pre-rendered audio objects. The audio objects correspond to the pre-rendered audio objects one by one.
[0231] S302: The encoder obtains a perceptual importance parameter value of each of the plurality of pre-rendered audio objects.
[0232] Among them, the relevant explanations and examples of S301-S302 can refer to the above S101-S102, which will not be repeated here.
[0233] S303: The encoder obtains content importance parameter values of each of the multiple audio objects to be pre-rendered.
[0234] For the explanation and examples of S303, please refer to the above S201.
[0235] The embodiment of the present application does not limit the execution order of S301-S302 and S303. For example, S301-S302 may be executed first and then S303, or S303 may be executed first and then S301-S302, or S301-S302 and S303 may be executed simultaneously.
[0236] S304: The encoder obtains a bit allocation parameter value for each of the multiple audio objects to be pre-rendered based on the perceptual importance parameter values for each of the multiple pre-rendered audio objects and the content importance parameter values for each of the multiple audio objects to be pre-rendered.
[0237] Specifically, the encoder obtains a bit allocation parameter value of the current audio object to be pre-rendered based on the perceptual importance parameter value of the current pre-rendered audio object and the content importance parameter value of the current audio object to be pre-rendered.
[0238] For example, the bit allocation parameter value Important_Bit of the i-th audio object to be pre-rendered i The following formula 17 is satisfied:
[0239] Formula 19: Important_Bit i =f(Important_P i ,Important_C i ).
[0240] Optionally, the functional relationship represented by Formula 19 may be linear or nonlinear.
[0241] Optionally, the bit allocation parameter value of the current audio object to be pre-rendered includes a second ratio, or a parameter value determined according to the second ratio. The second ratio is the ratio between the first value of the current audio object to be pre-rendered and the sum of the first values of the multiple audio objects to be pre-rendered. Among them, the parameter value can be considered as the value obtained after processing the second ratio, and the embodiment of the present application does not limit the specific processing method. The first value of the current audio object to be pre-rendered is the product of the content importance parameter value of the current audio object to be pre-rendered and the perceptual importance parameter value of the current pre-rendered audio object; or, the first value of the current audio object to be pre-rendered is a parameter value determined based on "the product of the content importance parameter value of the current audio object to be pre-rendered and the perceptual importance parameter value of the current pre-rendered audio object". Among them, the parameter value can be considered as the value obtained after processing the product, and the embodiment of the present application does not limit the specific processing method.
[0242] Specifically, S304 may include: the encoder first uses a ratio between a first value of the current audio object to be pre-rendered and a sum of the first values of the multiple audio objects to be pre-rendered as a second ratio; then, uses the second ratio as a bit allocation parameter value of the current audio object to be pre-rendered, or determines a parameter value according to the second ratio, and uses the parameter value as the bit allocation parameter value of the current audio object to be pre-rendered.
[0243] For example, Formula 19 can be further expressed as the following Formula 20:
[0244]
[0245] It can be seen that in this example, Important_Bit i ∈[0,1].
[0246] S305: The encoder obtains the total number of bits to be allocated corresponding to the multiple audio objects to be pre-rendered.
[0247] The relevant explanation and examples of S305 can refer to the above S104, which will not be repeated here.
[0248] S306: The encoder determines target numbers of bits to be allocated respectively for the multiple audio objects to be pre-rendered based on the total number of bits to be allocated and the bit allocation parameter values of the multiple audio objects to be pre-rendered.
[0249] The relevant explanation and examples of S306 can refer to the above S105 and will not be repeated here.
[0250] The method for allocating bits to audio objects provided in this embodiment takes into account the differences in the perceptual characteristics of different pre-rendered audio objects at the rendering and playback end, as well as the differences in the contents of different audio objects to be pre-rendered, when allocating bits to pre-rendered audio objects. Compared with the technical solution of encoding different audio objects with the same number of bits in the traditional technology, this method helps to improve the overall quality of the reconstructed audio objects. At the same time, it can improve the coding efficiency.
[0251] like Fig.11 , which is a flow chart of another method for allocating bits of an audio object provided in an embodiment of the present application. Fig.11 The method shown may include the following steps:
[0252] S401: The encoder obtains initial numbers of bits respectively allocated to a plurality of audio objects to be pre-rendered of an audio frame to be encoded, and respective bit allocation parameter values of the plurality of audio objects to be pre-rendered.
[0253] For example, the relationship between the initial number of bits allocated to the plurality of audio objects to be pre-rendered obtained by the encoder can be expressed as the following formula 21:
[0254] Formula 21: Bit 1 +Bit 2 +...+Bit N =Bits_available.
[0255] Among them, Bit 1 、Bit 2 , ..., Bit N Respectively represent the initial number of bits allocated to the first, second, ..., Nth audio object to be pre-rendered among the N audio objects to be pre-rendered, obtained by a known method. Bits_available The total number of bits to be allocated.
[0256] The embodiment of the present application does not limit how the encoder obtains the initial number of bits allocated to the multiple audio objects to be pre-rendered. For example, the encoder can divide the total number of bits to be allocated for the multiple audio objects to be pre-rendered equally to obtain the initial number of bits corresponding to each of the multiple audio objects to be pre-rendered. For another example, the encoder can determine the initial number of bits allocated to the multiple audio objects to be pre-rendered based on the energy of each of the multiple audio objects to be pre-rendered. For another example, the initial number of bits allocated to the multiple audio objects to be pre-rendered can be predefined.
[0257] Optionally, the relevant explanation of the bit allocation parameter value in S401 is Figure 6 , Figure 8 or Fig.10 The explanation of the bit allocation parameter value in the illustrated embodiment will not be repeated here.
[0258] In addition, the encoder can obtain the initial number of bits allocated to the multiple audio objects to be pre-rendered based on the content importance parameter values of the multiple audio objects to be pre-rendered. In this case, the encoder can use Figure 8 or Fig.10 The method in the illustrated embodiment obtains the bit allocation parameter value of each of the multiple audio objects to be pre-rendered.
[0259] S402: The encoder adjusts the bit allocation parameter values of the multiple audio objects to be pre-rendered respectively based on the initial number of bits allocated to the multiple audio objects to be pre-rendered respectively, to obtain the adjusted bit allocation parameter values of the multiple audio objects to be pre-rendered respectively.
[0260] Specifically, the encoder adjusts the bit allocation parameter value of the current audio object to be pre-rendered based on the initial number of bits respectively allocated to the current audio object to be pre-rendered, so as to obtain the adjusted bit allocation parameter value of the current audio object to be pre-rendered.
[0261] Optional, the bit allocation parameter value Adjust after modulation of the i-th audio object to be pre-rendered i , bit allocation parameter value Adjust_info of the i-th audio object to be pre-rendered i and the initial number of bits Bit allocated for the i-th audio object to be pre-rendered i The following formula 22 can be satisfied:
[0262] Formula 22: Adjust i =f(Adjust_info i ,Bit i ).
[0263] In other words, Adjust i By Adjust_infoi and Bit i It is obtained through a functional relationship, which can be linear or nonlinear.
[0264] Further optionally, the adjusted bit allocation parameter value of the current audio object to be pre-rendered includes: a fifth ratio or a parameter value determined according to the fifth ratio. The fifth ratio is the ratio between the second value of the current audio object to be pre-rendered and the sum of the second values of each of the multiple audio objects to be pre-rendered. The parameter value can be considered as a value obtained after processing the fifth ratio, and the embodiment of the present application does not limit the specific processing method. The second value of the current audio object to be pre-rendered is the product of the initial number of bits allocated to the current audio object to be pre-rendered and the bit allocation parameter value of the current audio object to be pre-rendered, or is a parameter value determined according to "the product of the initial number of bits allocated to the current audio object to be pre-rendered and the bit allocation parameter value of the current audio object to be pre-rendered". The parameter value can be considered as a value obtained after processing the product, and the embodiment of the present application does not limit the specific processing method.
[0265] Specifically, the encoder first uses the ratio of the second value of the current audio object to be pre-rendered to the sum of the second values of the multiple audio objects to be pre-rendered as the fifth ratio; then, uses the fifth ratio as the adjusted bit allocation parameter value of the current audio object to be pre-rendered, or determines a parameter value according to the fifth ratio, and uses the parameter value as the adjusted bit allocation parameter value of the current audio object to be pre-rendered.
[0266] For example, Formula 22 can be further expressed as the following Formula 23:
[0267]
[0268] S403: The encoder obtains a total number of bits to be allocated corresponding to encoding the multiple audio objects to be pre-rendered.
[0269] The relevant explanation and examples of S403 can refer to the above S104, which will not be repeated here.
[0270] S404: The encoder determines target numbers of bits to be allocated respectively for the plurality of audio objects to be pre-rendered based on the total number of bits to be allocated and the adjusted bit allocation parameter values of the plurality of audio objects to be pre-rendered.
[0271] Optionally, the ratio of the target number of bits allocated to the current audio object to be pre-rendered to the total number of bits to be allocated is equal to the adjusted bit allocation parameter value of the current audio object to be pre-rendered, or is equal to a parameter value determined according to the adjusted bit allocation parameter value of the current audio object to be pre-rendered. The parameter value can be considered to be a value obtained after processing the adjusted bit allocation parameter value of the current audio object to be pre-rendered, and the specific processing method is not limited in the embodiment of the present application.
[0272] For example, the target number of bits to be allocated for the i-th audio object to be pre-rendered is Adjust_Bit i The following formula 24 is satisfied:
[0273] Formula 24: Adjust_Bit i =Adjust i Bits_available.
[0274] Optionally, the bit allocation parameter values of the plurality of audio objects to be pre-rendered obtained in S501 are determined based on the perceptual importance parameter value. Based on this, the above formula 22 can be specifically expressed as the following formula 25:
[0275] Formula 25: Adjust i =f(Important_P i ,Bit i ).
[0276] The above formula 23 can be specifically expressed as the following formula 26:
[0277]
[0278] The method for allocating audio objects bit provided by the present embodiment adjusts the bit allocation parameter values of the multiple audio objects to be pre-rendered based on the initial bit numbers allocated to the multiple audio objects to be pre-rendered, and determines the target bit numbers allocated to the multiple audio objects to be pre-rendered based on the adjusted bit allocation parameter values of the multiple audio objects to be pre-rendered. This helps to further improve the overall quality of the reconstructed audio objects and improve the coding efficiency.
[0279] It should be noted that, in the absence of conflict, some or all of the features in any of the above embodiments may be combined to form a new embodiment.
[0280] Optionally, based on the bit allocation method for audio objects provided by any of the embodiments provided above, the encoder may further send to the decoder ratio information between the target numbers of bits respectively allocated to the multiple audio objects to be pre-rendered, wherein the ratio information is used by the decoder to reconstruct the multiple audio objects to be pre-rendered.
[0281] The specific implementation method of the ratio information is not limited in the embodiment of the present application. For example, the ratio information may be the ratio between the target number of bits respectively allocated to the multiple audio objects to be pre-rendered. For another example, the ratio information may be the target number of bits respectively allocated to the multiple audio objects to be pre-rendered.
[0282] After receiving the ratio information, the decoder can determine which bits in the bit streams corresponding to the multiple audio objects to be pre-rendered (that is, the bit streams obtained after the multiple audio objects to be pre-rendered are encoded) are for which audio object to be pre-rendered according to the total number of bits to be allocated corresponding to the multiple audio objects to be pre-rendered and the ratio information, so as to further use the bits for the specific audio object to be pre-rendered to reconstruct the specific audio object to be pre-rendered.
[0283] For example, assuming that the bit stream corresponding to the multiple audio objects to be pre-rendered sent by the encoder to the decoder contains 100 bits, the audio frame to be encoded contains audio objects 1-3 to be pre-rendered, and the ratio information sent by the encoder to the decoder is 3:3:4, and "3:3:4" represents the ratio between the target number of bits allocated to the audio objects 1-3 to be pre-rendered, then, based on the 100 bits and "3:3:4", the decoder can determine bits 1-30, bits 31-60, and bits 61-100 in the 100 bits (respectively marked as bits 1-100), which are the bits allocated to the audio objects 1-3 to be pre-rendered, respectively. Then, use bits 1-30 to reconstruct audio object 1 to be pre-rendered, use bits 31-60 to reconstruct audio object 2 to be pre-rendered, and use bits 61-100 to reconstruct audio object 3 to be pre-rendered. The reconstruction process can refer to the prior art and will not be repeated here.
[0284] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. In order to realize the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0285] The embodiment of the present application can divide the bit allocation device (such as an encoder or encoding device) of the audio object into functional modules according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation.
[0286] like Fig.12 As shown, Fig.12 FIG. 1 is a schematic diagram showing the structure of a bit allocation device 120 for an audio object provided in an embodiment of the present application. The bit allocation device 120 for an audio object is used to execute the above-mentioned bit allocation method for an audio object, for example, to execute Figure 6 , Figure 8 , Fig.10 or Fig.11 The bit allocation method of the audio object shown in the figure. For example, the bit allocation device 120 for the audio object includes: a pre-rendering module 1201 , an acquisition module 1202 and a determination module 1203 .
[0287] The pre-rendering module 1201 is used to pre-render multiple audio objects to be pre-rendered in the audio frame to be encoded respectively to obtain multiple pre-rendered audio objects. The acquisition module 1202 is used to obtain the perceptual importance parameter value of each of the multiple pre-rendered audio objects; wherein the perceptual importance parameter value of the current pre-rendered audio object in the multiple pre-rendered audio objects is used to indicate the perceptual importance of the current pre-rendered audio object in the multiple pre-rendered audio objects; based on the perceptual importance parameter value of each of the multiple pre-rendered audio objects, the bit allocation parameter value of the current pre-rendered audio object in the pre-rendered audio objects is obtained. The determination module 1203 is used to determine the target number of bits allocated to the current pre-rendered audio object based on the bit allocation parameter value of the current pre-rendered audio object and the total number of bits to be allocated corresponding to the multiple pre-rendered audio objects.
[0288] For example, combined with Figure 6 The pre-rendering module 1201 may be used to execute S101, the acquiring module 1202 may be used to execute S102-S104, and the determining module 1203 may be used to execute S105.
[0289] Optionally, the degree of perceived importance includes at least one of energy intensity and spectrum variation.
[0290] Optionally, the perceptual importance parameter includes an energy importance parameter, wherein the energy importance parameter of the current pre-rendered audio object is calculated based on the energy of the current pre-rendered audio object, and is used to indicate a ratio between the energy of the current pre-rendered audio object and the sum of the energies of the multiple pre-rendered audio objects.
[0291] Optionally, the perceptual importance parameter includes a perceptual intensity importance parameter, wherein the perceptual intensity importance parameter of the current pre-rendered audio object is calculated by combining the human hearing curve and the energy of the current pre-rendered audio object, and is used to indicate the ratio between the sum of the energies of a preset number of frequency bands with the largest energies among multiple frequency bands of the current pre-rendered audio object and the sum of the energies of a preset number of frequency bands with the largest energies among multiple frequency bands of the multiple pre-rendered audio objects.
[0292] Optionally, the perceptual importance parameter includes a spectrum flatness parameter, wherein the spectrum flatness parameter of the current pre-rendered audio object is used to indicate the spectrum flatness of the current pre-rendered audio object among the multiple pre-rendered audio objects.
[0293] Optionally, the current pre-rendered audio object is an audio object obtained by pre-rendering the current audio object to be pre-rendered. The bit allocation parameter value of the current audio object to be pre-rendered includes a first ratio, or a parameter value determined according to the first ratio. The first ratio is a ratio between a perceptual importance parameter value of the current pre-rendered audio object and a sum of perceptual importance parameter values of the multiple pre-rendered audio objects.
[0294] Optionally, the acquisition module 1202 is further used to: acquire the content importance parameter value of each of the multiple audio objects to be pre-rendered; wherein the content importance parameter value of the current audio object to be pre-rendered is used to indicate the importance of the sound type represented by the content of the current audio object to be pre-rendered in the sound types represented by the contents of the multiple audio objects to be pre-rendered. In the aspect of acquiring the bit allocation parameter value of the current audio object to be pre-rendered based on the perceptual importance parameter values of each of the multiple pre-rendered audio objects, the acquisition module is specifically used to: acquire the bit allocation parameter value of the current audio object to be pre-rendered based on the perceptual importance parameter values of each of the multiple pre-rendered audio objects and the content importance parameter values of each of the multiple audio objects to be pre-rendered. For example, in combination Fig.10 , the acquisition module 1202 can be used to execute S303 and S304.
[0295] Optionally, the current pre-rendered audio object is an audio object obtained by pre-rendering the current audio object to be pre-rendered. The bit allocation parameter value of the current audio object to be pre-rendered includes a second ratio, or a parameter value determined according to the second ratio. The second ratio is the ratio between the first value of the current audio object to be pre-rendered and the sum of the first values of the multiple audio objects to be pre-rendered; the first value of the current audio object to be pre-rendered is the product of the content importance parameter value of the current audio object to be pre-rendered and the perceptual importance parameter value of the current pre-rendered audio object, or the first value of the current audio object to be pre-rendered is a parameter value determined according to the product of the content importance parameter value of the current audio object to be pre-rendered and the perceptual importance parameter value of the current pre-rendered audio object.
[0296] Optionally, the sound type includes at least one of the following: voice, music, sound effect, ambient sound or noise.
[0297] Optionally, a ratio of a target number of bits allocated to the current audio object to be pre-rendered to a total number of bits to be allocated is equal to a third ratio, or is equal to a parameter value determined according to the third ratio. The third ratio is a ratio of a bit allocation parameter value of the current audio object to be pre-rendered to a sum of bit allocation parameter values of the plurality of audio objects to be pre-rendered.
[0298] Optionally, the determination module 1203 is specifically configured to: determine the priority level of the current audio object to be pre-rendered based on the correspondence between multiple bit allocation parameter values and multiple priority levels, and the bit allocation parameter value of the current audio object to be pre-rendered. Then, based on the priority level of the current audio object to be pre-rendered and the total number of bits to be allocated, determine the target number of bits allocated to the current audio object to be pre-rendered.
[0299] For example, combined with Figure 7 , the determination module 1203 can be used to execute S105A-S105B.
[0300] Optionally, a ratio of a target number of bits allocated to the current audio object to be pre-rendered to a total number of bits to be allocated is equal to a fourth ratio, or is equal to a parameter value determined according to the fourth ratio, wherein the fourth ratio is a ratio of a priority level of the current audio object to be pre-rendered to a sum of priority levels of the plurality of audio objects to be pre-rendered.
[0301] Optionally, the acquisition module 1202 is further configured to acquire an initial number of bits allocated to the current audio object to be pre-rendered. In this case, the determination module 1203 is specifically configured to: adjust the bit allocation parameter value of the current audio object to be pre-rendered based on the initial number of bits; and then determine the target number of bits allocated to the current audio object to be pre-rendered based on the total number of bits to be allocated and the adjusted bit allocation parameter value of the current audio object to be pre-rendered.
[0302] For example, combined with Fig.11 The acquisition module 1202 can be used to execute the step of acquiring the initial number of bits in S401. The determination module 1203 can be used to execute S402 and S404.
[0303] Optionally, the adjusted bit allocation parameter value of the current audio object to be pre-rendered includes: a fifth ratio or a parameter value determined according to the fifth ratio. The fifth ratio is a ratio between the second value of the current audio object to be pre-rendered and the sum of the second values of the plurality of audio objects to be pre-rendered. The second value of the current audio object to be pre-rendered is the product of the initial number of bits and the bit allocation parameter value of the current audio object to be pre-rendered, or the second value of the current audio object to be pre-rendered is a parameter value determined according to the product of the initial number of bits and the bit allocation parameter value of the current audio object to be pre-rendered.
[0304] Optionally, the ratio of the target number of bits allocated to the current audio object to be pre-rendered to the total number of bits to be allocated is equal to the adjusted bit allocation parameter value of the current audio object to be pre-rendered, or is equal to a parameter value determined according to the adjusted bit allocation parameter value of the current audio object to be pre-rendered.
[0305] Optional, such as Fig.12 As shown, the audio object bit allocation device 120 further includes: a sending module 1204, which is used to send ratio information between target bit numbers respectively allocated to the multiple audio objects to be pre-rendered, wherein the ratio information is used to reconstruct the multiple audio objects to be pre-rendered.
[0306] For the detailed description of the above optional manner, please refer to the above method embodiment, which will not be repeated here. In addition, the explanation and description of the beneficial effects of any of the above audio object bit allocation devices 120 can refer to the above corresponding method embodiment, which will not be repeated here.
[0307] As an example, combining Figure 1A or Figure 1B , the bit allocation device 120 for the audio object may be the stereo encoder 112. Figure 2 , the bit allocation device 120 for the audio object may be a stereo encoder 213. Figure 3A or Figure 3B , the bit allocation device 120 for the audio object may be the multi-channel encoder 114. Figure 4 , the bit allocation device 120 of the audio object may be a multi-channel encoder 215 .
[0308] As an example, combining Figure 1A or Figure 3A , the bit allocation device 120 of the audio object may be the first terminal 11. Figure 1B or Figure 3B , the audio object bit allocation device 120 may be the first terminal 11 or the second terminal 12. Figure 2 or Figure 4 , the audio object bit allocation device 120 may be the first network device 21 .
[0309] As an example, combining Figure 5 The functions implemented in part or in whole in the above-mentioned pre-rendering module 1201, acquisition module 1202 and determination module 1203 can be Figure 5 Processor 51 in the execution Figure 2 The sending module 1204 can be implemented by the program code in the memory 52 in the memory 52. Figure 5 The receiving unit in the communication interface 53 is implemented.
[0310] The embodiment of the present application also provides an audio system, including an encoding device and a decoding device. The encoding device can be any of the bit allocation devices 120 for audio objects provided above. The decoding device is used to receive information sent by the encoding device and perform a decoding process (including a reconstruction process of the audio object).
[0311] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed on a computer, the computer is enabled to execute the method executed by any one of the encoders provided above.
[0312] For explanations of the relevant contents and descriptions of the beneficial effects of any of the audio systems and computer-readable storage media provided above, reference may be made to the corresponding embodiments described above, and no further details will be given here.
[0313] The embodiment of the present application also provides a chip. The chip integrates a control circuit and one or more ports for implementing the functions of the bit allocation device 120 of the above-mentioned audio object. Optionally, the functions supported by the chip can be referred to above and will not be repeated here. A person of ordinary skill in the art can understand that all or part of the steps of implementing the above-mentioned embodiment can be completed by instructing the relevant hardware through a program. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a random access memory, etc. The above-mentioned processing unit or processor can be a central processing unit, a general-purpose processor, an application specific integrated circuit (ASIC), a microprocessor (digital signal processor, DSP), a field programmable gate array (field programmable gate array, FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof.
[0314] The embodiment of the present application also provides a computer program product including instructions, when the instructions are run on a computer, the computer executes any one of the methods in the above embodiments. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a server, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means to another website site, computer, server, or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more servers that can be integrated with the medium. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., an SSD), etc.
[0315] It should be noted that the above-mentioned devices for storing computer instructions or computer programs provided in the embodiments of the present application, such as but not limited to the above-mentioned memories, computer-readable storage media and communication chips, etc., are all non-transitory.
[0316] In the process of implementing the claimed application, those skilled in the art can understand and implement other changes to the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "one" or "an" does not exclude multiple situations. A single processor or other unit can implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results. Although the present application is described in conjunction with specific features and embodiments thereof, various modifications and combinations may be made to it without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are merely exemplary illustrations of the present application as defined by the appended claims, and are deemed to have covered any and all modifications, changes, combinations or equivalents within the scope of the present application.
Claims
1. A method for allocating bits of an audio object, It is characterized in that include: Pre-rendering the plurality of audio objects to be pre-rendered in the audio frame to be encoded respectively to obtain a plurality of pre-rendered audio objects; Obtaining a perceptual importance parameter value of each of the multiple pre-rendered audio objects and a content importance parameter value of each of the multiple pre-rendered audio objects; wherein the perceptual importance parameter value of the current pre-rendered audio object among the multiple pre-rendered audio objects is used to indicate the perceptual importance of the current pre-rendered audio object among the multiple pre-rendered audio objects; and the content importance parameter value of the current audio object to be pre-rendered is used to indicate the importance of the sound type represented by the content of the current audio object to be pre-rendered among the sound types represented by the contents of the multiple audio objects to be pre-rendered; Based on the perceptual importance parameter values of the plurality of pre-rendered audio objects and the content importance parameter values of the plurality of audio objects to be pre-rendered, obtaining a bit allocation parameter value of a current audio object to be pre-rendered among the plurality of audio objects to be pre-rendered; Based on the bit allocation parameter value of the current audio object to be pre-rendered and the total number of bits to be allocated corresponding to the multiple audio objects to be pre-rendered, a target number of bits to be allocated for the current audio object to be pre-rendered is determined.
2. The method according to claim 1, It is characterized in that The perceptual importance parameter includes at least one of the following: an energy importance parameter, a perceptual intensity importance parameter or a spectrum flatness parameter; wherein: The energy importance parameter of the current pre-rendered audio object is calculated based on the energy of the current pre-rendered audio object, and is used to indicate a ratio between the energy of the current pre-rendered audio object and the sum of the energies of the plurality of pre-rendered audio objects; The perceived intensity importance parameter of the current pre-rendered audio object is calculated based on a human hearing curve and the energy of the current pre-rendered audio object, and is used to indicate a ratio between the sum of energies of a preset number of frequency bands with the largest energies among multiple frequency bands of the current pre-rendered audio object and the sum of energies of a preset number of frequency bands with the largest energies among multiple frequency bands of each of the multiple pre-rendered audio objects; The spectrum flatness parameter of the current pre-rendered audio object is used to indicate the spectrum flatness of the current pre-rendered audio object among the multiple pre-rendered audio objects.
3. The method according to claim 1 or 2, It is characterized in that The current pre-rendered audio object is an audio object obtained by pre-rendering the current audio object to be pre-rendered, and the bit allocation parameter value of the current audio object to be pre-rendered includes a first ratio, or a parameter value determined according to the first ratio; The first ratio is a ratio between the perceptual importance parameter value of the current pre-rendered audio object and the sum of the perceptual importance parameter values of the plurality of pre-rendered audio objects.
4. The method according to claim 1 or 2, It is characterized in that The current pre-rendered audio object is an audio object obtained by pre-rendering the current audio object to be pre-rendered, and the bit allocation parameter value of the current audio object to be pre-rendered includes the second ratio, or a parameter value determined according to the second ratio; The second ratio is a ratio between the first value of the current audio object to be pre-rendered and the sum of the first values of the plurality of audio objects to be pre-rendered; The first value of the current audio object to be pre-rendered is the product of a content importance parameter value of the current audio object to be pre-rendered and a perceptual importance parameter value of the current pre-rendered audio object, or the first value of the current audio object to be pre-rendered is a parameter value determined according to the product of a content importance parameter value of the current audio object to be pre-rendered and a perceptual importance parameter value of the current pre-rendered audio object.
5. The method according to claim 1 or 2, It is characterized in that The sound type includes at least one of the following: voice, music, sound effect, ambient sound or noise.
6. The method according to claim 1 or 2, It is characterized in that The ratio of the target number of bits allocated to the current audio object to be pre-rendered to the total number of bits to be allocated is equal to a third ratio, or equal to a parameter value determined according to the third ratio; The third ratio is a ratio between the bit allocation parameter value of the current audio object to be pre-rendered and the sum of the bit allocation parameter values of the plurality of audio objects to be pre-rendered.
7. The method according to claim 1 or 2, It is characterized in that The determining, based on the bit allocation parameter value of the current audio object to be pre-rendered and the total number of bits to be allocated corresponding to the plurality of audio objects to be pre-rendered, a target number of bits to be allocated for the current audio object to be pre-rendered includes: Determining the priority level of the current audio object to be pre-rendered based on the correspondence between multiple bit allocation parameter values and multiple priority levels, and the bit allocation parameter value of the current audio object to be pre-rendered; Based on the priority level of the current audio object to be pre-rendered and the total number of bits to be allocated, a target number of bits to be allocated for the current audio object to be pre-rendered is determined.
8. The method according to claim 7, It is characterized in that A ratio of a target number of bits allocated to the current audio object to be pre-rendered to the total number of bits to be allocated is equal to a fourth ratio, or is equal to a parameter value determined according to the fourth ratio; The fourth ratio is a ratio between the priority level of the current audio object to be pre-rendered and the sum of the priority levels of the plurality of audio objects to be pre-rendered.
9. The method according to claim 1 or 2, It is characterized in that The determining, based on the bit allocation parameter value of the current audio object to be pre-rendered and the total number of bits to be allocated corresponding to the plurality of audio objects to be pre-rendered, a target number of bits to be allocated for the current audio object to be pre-rendered includes: Obtaining an initial number of bits allocated to the current audio object to be pre-rendered; Based on the initial number of bits, adjusting a bit allocation parameter value of the current audio object to be pre-rendered; Based on the total number of bits to be allocated and the adjusted bit allocation parameter value of the current audio object to be pre-rendered, a target number of bits to be allocated for the current pre-rendered audio object is determined.
10. The method according to claim 9, It is characterized in that The adjusted bit allocation parameter value of the current audio object to be pre-rendered includes: the fifth ratio or a parameter value determined according to the fifth ratio; The fifth ratio is a ratio between the second value of the current audio object to be pre-rendered and the sum of the second values of each of the multiple audio objects to be pre-rendered; wherein the second value of the current audio object to be pre-rendered is the product of the initial number of bits and the bit allocation parameter value of the current audio object to be pre-rendered, or the second value of the current audio object to be pre-rendered is a parameter value determined according to the product of the initial number of bits and the bit allocation parameter value of the current audio object to be pre-rendered.
11. The method according to claim 10, It is characterized in that A ratio of a target number of bits used for the current audio object to be pre-rendered to the total number of bits to be allocated is equal to an adjusted bit allocation parameter value of the current audio object to be pre-rendered, or is equal to a parameter value determined according to the adjusted bit allocation parameter value of the current audio object to be pre-rendered.
12. The method according to claim 1 or 2, It is characterized in that The method further comprises: Sending ratio information between target numbers of bits respectively allocated to the plurality of audio objects to be pre-rendered; wherein the ratio information is used to reconstruct the plurality of audio objects to be pre-rendered.
13. A bit allocation device for an audio object, It is characterized in that include: A pre-rendering module, used for pre-rendering a plurality of audio objects to be pre-rendered in the audio frame to be encoded, respectively, to obtain a plurality of pre-rendered audio objects; an acquisition module, configured to acquire a perceptual importance parameter value of each of the plurality of pre-rendered audio objects and a content importance parameter value of each of the plurality of audio objects to be pre-rendered; wherein the perceptual importance parameter value of a current pre-rendered audio object among the plurality of pre-rendered audio objects is used to indicate a perceptual importance degree of the current pre-rendered audio object among the plurality of pre-rendered audio objects; and the content importance parameter value of the current audio object to be pre-rendered is used to indicate a degree of importance of a sound type represented by a content of the current audio object to be pre-rendered among the sound types represented by the content of the plurality of audio objects to be pre-rendered; based on the perceptual importance parameter values of each of the plurality of pre-rendered audio objects and the content importance parameter values of each of the plurality of audio objects to be pre-rendered, acquire a bit allocation parameter value of a current audio object to be pre-rendered among the plurality of audio objects to be pre-rendered; The determination module is configured to determine a target number of bits to be allocated to the current audio object to be pre-rendered based on the bit allocation parameter value of the current audio object to be pre-rendered and the total number of bits to be allocated corresponding to the multiple audio objects to be pre-rendered.
14. The device according to claim 13, It is characterized in that The perceptual importance parameter includes at least one of the following: an energy importance parameter, a perceptual intensity importance parameter or a spectrum flatness parameter; wherein: The energy importance parameter of the current pre-rendered audio object is calculated based on the energy of the current pre-rendered audio object, and is used to indicate a ratio between the energy of the current pre-rendered audio object and the sum of the energies of the plurality of pre-rendered audio objects; The perceived intensity importance parameter of the current pre-rendered audio object is calculated based on a human hearing curve and the energy of the current pre-rendered audio object, and is used to indicate a ratio between the sum of energies of a preset number of frequency bands with the largest energies among multiple frequency bands of the current pre-rendered audio object and the sum of energies of a preset number of frequency bands with the largest energies among multiple frequency bands of each of the multiple pre-rendered audio objects; The spectrum flatness parameter of the current pre-rendered audio object is used to indicate the spectrum flatness of the current pre-rendered audio object among the multiple pre-rendered audio objects.
15. The device according to claim 13 or 14, It is characterized in that The current pre-rendered audio object is an audio object obtained by pre-rendering the current audio object to be pre-rendered, and the bit allocation parameter value of the current audio object to be pre-rendered includes a first ratio, or a parameter value determined according to the first ratio; The first ratio is a ratio between the perceptual importance parameter value of the current pre-rendered audio object and the sum of the perceptual importance parameter values of the plurality of pre-rendered audio objects.
16. The device according to claim 13 or 14, It is characterized in that The current pre-rendered audio object is an audio object obtained by pre-rendering the current audio object to be pre-rendered, and the bit allocation parameter value of the current audio object to be pre-rendered includes the second ratio, or a parameter value determined according to the second ratio; The second ratio is a ratio between the first value of the current audio object to be pre-rendered and the sum of the first values of the plurality of audio objects to be pre-rendered; The first value of the current audio object to be pre-rendered is the product of a content importance parameter value of the current audio object to be pre-rendered and a perceptual importance parameter value of the current pre-rendered audio object, or the first value of the current audio object to be pre-rendered is a parameter value determined according to the product of a content importance parameter value of the current audio object to be pre-rendered and a perceptual importance parameter value of the current pre-rendered audio object.
17. The device according to claim 13 or 14, It is characterized in that The sound type includes at least one of the following: voice, music, sound effect, ambient sound or noise.
18. The device according to claim 13 or 14, It is characterized in that The ratio of the target number of bits allocated to the current audio object to be pre-rendered to the total number of bits to be allocated is equal to a third ratio, or equal to a parameter value determined according to the third ratio; The third ratio is a ratio between the bit allocation parameter value of the current audio object to be pre-rendered and the sum of the bit allocation parameter values of the plurality of audio objects to be pre-rendered.
19. The device according to claim 13 or 14, It is characterized in that The determination module is specifically used for: Determining the priority level of the current audio object to be pre-rendered based on the correspondence between multiple bit allocation parameter values and multiple priority levels, and the bit allocation parameter value of the current audio object to be pre-rendered; Based on the priority level of the current audio object to be pre-rendered and the total number of bits to be allocated, a target number of bits to be allocated for the current audio object to be pre-rendered is determined.
20. The device according to claim 19, It is characterized in that A ratio of a target number of bits allocated to the current audio object to be pre-rendered to the total number of bits to be allocated is equal to a fourth ratio, or is equal to a parameter value determined according to the fourth ratio; The fourth ratio is a ratio between the priority level of the current audio object to be pre-rendered and the sum of the priority levels of the plurality of audio objects to be pre-rendered.
21. The device according to claim 13 or 14, It is characterized in that The determination module is specifically used for: Obtaining an initial number of bits allocated to the current audio object to be pre-rendered; Based on the initial number of bits, adjusting a bit allocation parameter value of the current audio object to be pre-rendered; Based on the total number of bits to be allocated and the adjusted bit allocation parameter value of the current audio object to be pre-rendered, a target number of bits to be allocated for the current audio object to be pre-rendered is determined.
22. The device according to claim 21, It is characterized in that The adjusted bit allocation parameter value of the current audio object to be pre-rendered includes: the fifth ratio or a parameter value determined according to the fifth ratio; The fifth ratio is a ratio between the second value of the current audio object to be pre-rendered and the sum of the second values of each of the multiple audio objects to be pre-rendered; wherein the second value of the current audio object to be pre-rendered is the product of the initial number of bits and the bit allocation parameter value of the current audio object to be pre-rendered, or the second value of the current audio object to be pre-rendered is a parameter value determined according to the product of the initial number of bits and the bit allocation parameter value of the current audio object to be pre-rendered.
23. The device according to claim 22, It is characterized in that A ratio of a target number of bits allocated to the current audio object to be pre-rendered to the total number of bits to be allocated is equal to an adjusted bit allocation parameter value of the current audio object to be pre-rendered, or is equal to a parameter value determined according to the adjusted bit allocation parameter value of the current audio object to be pre-rendered.
24. The device according to claim 13 or 14, It is characterized in that The device also includes: The sending module is used to send ratio information between target bit numbers respectively allocated to the multiple audio objects to be pre-rendered; wherein the ratio information is used to reconstruct the multiple audio objects to be pre-rendered.
25. The device according to claim 13 or 14, It is characterized in that The device is an encoder, or the device is an encoding apparatus including an encoder.
26. The device according to claim 25, It is characterized in that The encoder is a stereo encoder or a multi-channel encoder.
27. A bit allocation device for an audio object, It is characterized in that include: A memory and a processor, the memory being used to store a computer program, and the processor being used to call the computer program to execute the method according to any one of claims 1 to 12.
28. A computer readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed on a computer, the computer is enabled to execute the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Bit distribution method and apparatus with progressively fine spacing parameter
CN101499279A
Cited By
Bit allocation method and apparatus for audio object
WO2022156556A1