Signal processing method and apparatus thereof

EP4745961A4Pending Publication Date: 2026-05-20BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2023-07-14
Publication Date
2026-05-20

AI Technical Summary

Technical Problem

Existing encoding and decoding methods for combined format audio signals fail to reconstruct the most suitable and desired sound field due to energy parameter-based bit allocation, leading to deteriorated reconstruction effects.

Method used

Perform bit allocation for multiple sound channels based on first parameters, including energy parameters and sound field analysis parameters, to adjust bit allocation closer to actual sound field conditions, enabling reconstruction of the desired sound field.

Benefits of technology

Improves the reconstruction effect of combined format audio signals by aligning bit allocation with specific application scenarios, ensuring the desired sound field is accurately recreated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGAF001_ABST
    Figure IMGAF001_ABST
Patent Text Reader

Abstract

A signal processing method and an apparatus thereof, the method comprising: acquiring a mixed audio signal (S3101), the mixed audio signal comprising multiple sound channels; determining a first parameter of the mixed audio signal and / or energy parameters of one or more sound channels among the multiple sound channels (S3103); on the basis of the first parameter of the mixed audio signal and / or the energy parameters of the multiple sound channels, performing bit allocation on the multiple sound channels, to obtain a bit allocation parameter (S3104); on the basis of the bit allocation parameter, encoding the multiple sound channels to obtain a sound channel signal encoding parameter, and writing the bit allocation parameter and the sound channel signal encoding parameter into a code stream (S3105); sending the code stream (S3106). The present invention solves the problem in the related art of an encoding and decoding method not being able to reconstruct a desired and most suitable sound field.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The disclosure relates to the field of communication technology, and in particular to a method and an apparatus for processing a signal.BACKGROUND

[0002] With the increase of transmission bandwidth, the upgrade of signal acquisition / collection equipment of a terminal, the improvement of signal processor performance, and the upgrade of playback equipment of the terminal, three signal formats, namely sound channel-based audio, object-based audio, and scene-based audio, may be collected and may provide three-dimensional (3D) audio services.

[0003] In an application scenario of 3D audio, the 3D audio usually contains signals in multiple audio signal formats, that is, a combined format audio signal. In the related art, an encoding and decoding method of a combined format audio signal usually performs unified encoding and decoding processing on the combined format audio signal. For example, the encoder performs bit allocation for an available bit budget based on energy parameters of the combined format audio signal, and each sound channel uses a corresponding encoding core to perform encoding processing based on allocated bits to output an encoding parameter, and writes the encoding parameters in a bitstream.

[0004] However, the encoder in the related art adopts energy parameter-based bit allocation for the inputted combined format audio signal. This bit allocation method is relatively simple and one-sided, which makes it impossible to reconstruct the most suitable and / or desired sound field for a specific application scenario, and causes the reconstruction effect of the combined format audio to deteriorate.SUMMARY

[0005] Embodiments of the disclosure provide methods for processing a signal and devices thereof.

[0006] According to a first aspect of embodiments of the disclosure, a method for processing a signal is provided. The method includes: obtaining a combined format audio signal, in which the combined format audio signal is a multi-channel audio signal including a plurality of sound channels; determining a first parameter of the combined format audio signal and / or one or more energy parameters of one or more sound channels among the plurality of sound channels; obtaining bit allocation parameters by performing bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels; encoding the plurality of sound channels based on the bit allocation parameters to obtain sound channel signal encoding parameters, and writing the bit allocation parameters and the sound channel signal encoding parameters in a bitstream; and sending the bitstream.

[0007] According to a second aspect of embodiments of the disclosure, a method for processing a signal is provided. The method includes: receiving a bitstream; parsing the bitstream to obtain a bit allocation parameter and a sound channel signal encoding parameter for each of a plurality of sound channels of a combined format audio signal; reconstructing the combined format audio signal based on the bit allocation parameter and the sound channel signal encoding parameter for each of the plurality of sound channels.

[0008] According to a third aspect of embodiments of the disclosure, a first communication device is provided. The first communication device includes: a processing module, configured to obtain a combined format audio signal, in which the combined format audio signal is a multi-channel audio signal including a plurality of sound channels; determine a first parameter of the combined format audio signal and / or one or more energy parameters of one or more sound channels among the plurality of sound channels; obtain bit allocation parameters by performing bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels; encode the plurality of sound channels based on the bit allocation parameters to obtain sound channel signal encoding parameters, and write the bit allocation parameters and the sound channel signal encoding parameters in a bitstream; and a transceiver module, configured to send the bitstream.

[0009] According to a fourth aspect of embodiments of the disclosure, a second communication device is provided. The second communication device includes: a transceiver module, configured to receive a bitstream; and a processing module, configured to parse the bitstream to obtain a bit allocation parameter and a sound channel signal encoding parameter for each of a plurality of sound channels of a combined format audio signal; and reconstruct the combined format audio signal based on the bit allocation parameter and the sound channel signal encoding parameter for each of the plurality of sound channels.

[0010] According to a fifth aspect of embodiments of the disclosure, a communication system is provided. The communication system includes: an encoding-end device, configured to perform an optional implementation according to the aforementioned first aspect; and a decoding-end device, configured to perform an optional implementation according to the aforementioned second aspect.

[0011] According to a sixth aspect of embodiments of the disclosure, a communication device is provided. The communication device includes one or more processors.

[0012] The one or more processors are configured to call instructions to enable the communication device to perform an optional implementation according to the aforementioned first and second aspects.

[0013] According to a seventh aspect of embodiments of the disclosure, a storage medium is provided. The storage medium has instructions stored thereon. When the instructions are executed on a communication device, the communication device is caused to perform an optional implementation according to the aforementioned first and second aspects.

[0014] With the technical solution according to the disclosure, the problem that the encoding and decoding method in the related art cannot reconstruct the desired and most suitable sound field may be solved, and the reconstruction effect of the combined format audio signal is improved.BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in embodiments of the disclosure, the drawings required for describing embodiments are introduced below. The following drawings are only some embodiments of the disclosure and do not impose specific limitations on the protection scope of the disclosure. FIG. 1 is a schematic diagram illustrating an architecture of a communication system according to an embodiment of the disclosure. FIG. 2A is a schematic diagram illustrating an interaction of a method for processing a signal according to an embodiment of the disclosure. FIG. 2B is a schematic diagram illustrating an interaction of a method for processing a signal according to an embodiment of the disclosure. FIG. 2C is a schematic diagram illustrating an interaction of a method for processing a signal according to an embodiment of the disclosure. FIG. 2D is a schematic diagram illustrating an interaction of a method for processing a signal according to an embodiment of the disclosure. FIG. 3 is a flowchart illustrating a method for processing a signal according to an embodiment of the disclosure. FIG. 4 is a flowchart illustrating another method for processing a signal according to an embodiment of the disclosure. FIG. 5 is a schematic diagram illustrating an interaction of a method for processing a signal according to an embodiment of the disclosure. FIG. 6A is a diagram illustrating an example of a method for processing a signal according to an embodiment of the disclosure. FIG. 6B is a diagram illustrating another example of a method for processing a signal according to an embodiment of the disclosure. FIG. 7A is a schematic diagram illustrating a structure of an encoding-end device according to an embodiment of the disclosure. FIG. 7B is a schematic diagram illustrating a structure of a decoding-end device according to an embodiment of the disclosure. FIG. 8A is a schematic diagram illustrating a structure of a communication device 8100 according to an embodiment of the disclosure. FIG. 8B is a schematic diagram illustrating a structure of a chip 8200 according to an embodiment of the disclosure. DETAILED DESCRIPTION

[0016] Embodiments of the disclosure provide methods for processing a signal and devices thereof.

[0017] In a first aspect, an embodiment of the disclosure provides a method for processing a signal. The method includes: obtaining a combined format audio signal, where the combined format audio signal is a multi-channel audio signal including a plurality of sound channels; determining a first parameter of the combined format audio signal and / or one or more energy parameters of one or more sound channels among the plurality of sound channels; obtaining bit allocation parameters by performing bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels; encoding the plurality of sound channels based on the bit allocation parameters to obtain sound channel signal encoding parameters, and writing the bit allocation parameters and the sound channel signal encoding parameters in a bitstream; and sending the bitstream.

[0018] In the above embodiment, by adopting the first parameter of the energy-based audio signal and / or the combined format audio signal, the bit allocation is performed for the plurality of sound channels of the combined format audio signal. The bit allocation parameters may be closer to actual sound field conditions of a specific application scenario, and the desired and most suitable sound field may be reconstructed, which solves the problem that the encoding and decoding method in the related art cannot reconstruct the desired and most suitable sound field, and the reconstruction effect of the combined format audio signal is improved.

[0019] In combination with some embodiments of the first aspect, in some embodiments, the first parameter includes at least one of: a sound field analysis parameter; a control parameter, where the control parameter includes a second parameter and / or a third parameter, the second parameter is used for describing an importance ranking of the plurality of sound channels in the combined format audio signal, and the third parameter is used for describing whether each of the plurality of sound channels is a diegetic sound or a non-diegetic sound.

[0020] In the above embodiment, in performing the bit allocation for the plurality of sound channels of the combined format audio signal, considering the sound field analysis parameter and / or the control parameter may make the bit allocation parameters closer to the actual sound field conditions of the specific application scenario, the desired and most suitable sound field may be reconstructed, and the reconstruction effect of the combined format audio signal may be improved.

[0021] In combination with some embodiments of the first aspect, in some embodiments, the combined format audio signal includes an audio signal in at least one of following format: a sound channel-based audio signal; an object-based audio signal; a scene-based audio signal; or a metadata-based three-dimensional audio signal.

[0022] In the above embodiment, the combined format audio signal including audio signals in various formats may be coded and decoded, and a 3D audio communication service may be implemented between an audio collection end and an audio playback end.

[0023] In combination with some embodiments of the first aspect, in some embodiments, determining the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels includes: determining the one or more energy parameters of the one or more sound channels; and determining a sound field analysis parameter of the combined format audio signal.

[0024] In combination with some embodiments of the first aspect, in some embodiments, obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels includes: obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the one or more energy parameters of the one or more sound channels and the sound field analysis parameter of the combined format audio signal.

[0025] In the above embodiment, in performing the bit allocation for the plurality of sound channels of the combined format audio signal, in addition to considering the energy parameter(s), the sound field analysis parameter is also considered, and the bit allocation parameters may be closer to the actual sound field conditions of the specific application scenario. The desired and most suitable sound field may be reconstructed and the reconstruction effect of the combined format audio signal may be improved.

[0026] In combination with some embodiments of the first aspect, in some embodiments, determining the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels includes: determining the one or more energy parameters of the one or more sound channels; and determining a control parameter of the combined format audio signal.

[0027] In combination with some embodiments of the first aspect, in some embodiments, obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels includes: obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the one or more energy parameters of the one or more sound channels and the control parameter of the combined format audio signal.

[0028] In the above embodiment, in performing the bit allocation for the plurality of sound channels of the combined format audio signal, in addition to the energy parameter(s), an external control parameter is considered, the allocated bits may be adjusted based on the external control parameter, and the bit allocation parameters may be closer to the actual sound field condition of the specific application scenario. The desired and most suitable sound field may be reconstructed, and the reconstruction effect of the combined format audio signal may be improved.

[0029] In combination with some embodiments of the first aspect, in some embodiments, determining the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels includes: determining the sound field analysis parameter and / or the control parameter of the combined format audio signal.

[0030] In combination with some embodiments of the first aspect, in some embodiments, obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels includes: obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the sound field analysis parameter and / or the control parameter of the combined format audio signal.

[0031] In the above embodiment, in performing the bit allocation for the plurality of sound channels of the combined format audio signal, the sound field analysis parameter and / or the external control parameter are considered, and the bit allocation parameters may be closer to the actual sound field condition of the specific application scenario. The desired and most suitable sound field may be reconstructed, and the reconstruction effect of the combined format audio signal may be improved.

[0032] In combination with some embodiments of the first aspect, in some embodiments, determining the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels includes: determining the one or more energy parameters of the one or more sound channels among the plurality of sound channels; and determining the sound field analysis parameter and the control parameter of the combined format audio signal.

[0033] In combination with some embodiments of the first aspect, in some embodiments, obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels includes: obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the one or more energy parameters of the one or more sound channels among the plurality of sound channels, the sound field analysis parameter and the control parameter of the combined format audio signal.

[0034] In the above embodiment, in performing the bit allocation for the plurality of sound channels of the combined format audio signal, in addition to considering the energy parameter(s), the external control parameter and the sound field analysis parameter are also considered, and the bit allocation may be further close to the actual sound field condition of the specific application scenario. The desired and most suitable sound field may be reconstructed, and the reconstruction effect of the combined format audio signal may be improved.

[0035] In combination with some embodiments of the first aspect, in some embodiments, determining the one or more energy parameters of the one or more sound channels among the plurality of sound channels includes: filtering the plurality of sound channels, and determining cross-correlation coefficients between a plurality of filtered sound channels; performing sound channel combination operation based on the cross-correlation coefficients, and determining inter-channel parameters for sound channels in a same sound channel combination; aligning sound channels and downmixing aligned sound channels with the inter-channel parameters, and performing energy calculation on downmixed sound channels to obtain energy parameters of the plurality of sound channels respectively.

[0036] In the above embodiment, the filtering processing is performed on the sound channels, and the energy calculation is performed on the filtered sound channels. The accuracy of the calculation result may be improved and noise interference may be avoided.

[0037] In combination with some embodiments of the first aspect, in some embodiments, performing the energy calculation includes performing any of following processing on sample points of a frame of a downmixed sound channel: calculating a sum of levels of the sample points; calculating a sum of squared levels of the sample points; or calculating a root mean square of levels of the sample points.

[0038] In the above embodiment, by using any one of the sum of levels of the sample points, the sum of squared levels of the sample points, and the root mean square of levels of the sample points to calculate the energy, the accuracy of the calculation result may be ensured.

[0039] In combination with some embodiments of the first aspect, in some embodiments, determining the sound field analysis parameter of the combined format audio signal includes: obtaining the sound field analysis parameter of the combined format audio signal by performing sound field analysis on the combined format audio signal based on a direction-of-arrival (DOA) estimation and / or a multi-channel cross-correlation coefficient (MCCC) method.

[0040] In combination with some embodiments of the first aspect, in some embodiments, determining the control parameter of the combined format audio signal includes at least one of: determining the control parameter of the combined format audio signal based on functional categories of the plurality of sound channels determined in collecting the combined format audio signal by an audio collection device, where the audio collection device is configured to collect the combined format audio signal; and determining the control parameter of the combined format audio signal based on a setting parameter required for reconstructing a desired sound field by a decoding-end device.

[0041] In the above embodiment, by determining the above control parameter based on the functional categories of the sound channels in collecting the combined format audio signal by the audio collection device and / or the setting parameter required for reconstructing the desired sound field by the decoding-end device, the above control parameter may be made closer to the actual situation of the application scenario. In this way, by using the control parameter to adjust the allocated bits, the adjusted bit allocation parameters may be further close to the actual sound field situation of the specific application scenario. The reconstruction of the desired combined format audio signal at the decoding end may be achieved.

[0042] In combination with some embodiments of the first aspect, in some embodiments, the sound field analysis parameter includes at least one of: the total number of sound images; relative priority of sound images; or a background sound.

[0043] In combination with some embodiments of the first aspect, in some embodiments, performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels includes: performing the bit allocation on the combined format audio signal consisting of an object-based audio signal and a scene-based audio signal based on the sound field analysis parameter of the combined format audio signal; where the number of bits allocated to any sound channel of the object-based audio signal is greater than the number of bits allocated to any sound channel of the scene-based audio signal.

[0044] In combination with some embodiments of the first aspect, in some embodiments, performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels includes: performing the bit allocation on the combined format audio signal consisting of the object-based audio signal and the sound channel-based audio signal based on the sound field analysis parameter of the combined format audio signal; where the number of bits allocated to any sound channel of the object-based audio signal is greater than the number of bits allocated to any sound channel of the sound channel-based audio signal.

[0045] In combination with some embodiments of the first aspect, in some embodiments, performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels includes: performing the bit allocation on the combined format audio signal consisting of the sound channel-based audio signal and the scene-based audio signal based on the sound field analysis parameter of the combined format audio signal; where the number of bits allocated to any sound channel of the sound channel-based audio signal is greater than the number of bits allocated to any sound channel of the scene-based audio signal.

[0046] In combination with some embodiments of the first aspect, in some embodiments, performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels includes: performing the bit allocation on the combined format audio signal consisting of the object-based audio signal and the metadata-based three-dimensional audio signal based on the sound field analysis parameter of the combined format audio signal; where the number of bits allocated to any sound channel of the object-based audio signal is greater than the number of bits allocated to any sound channel of the metadata-based three-dimensional audio signal.

[0047] In combination with some embodiments of the first aspect, in some embodiments, performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels includes: performing the bit allocation on the combined format audio signal consisting of the sound channel-based audio signal and the metadata-based three-dimensional audio signal based on the sound field analysis parameter of the combined format audio signal; where the number of bits allocated to any sound channel of the sound channel-based audio signal is greater than the number of bits allocated to any sound channel of the metadata-based three-dimensional audio signal.

[0048] In combination with some embodiments of the first aspect, in some embodiments, the number of bits allocated to a sound channel corresponding to a primary sound image of the object-based audio signal is greater than the number of bits allocated to a sound channel corresponding to a secondary sound image of the object-based audio signal.

[0049] In a second aspect, the disclosure provides a method for processing a signal. The method includes: receiving a bitstream; parsing the bitstream to obtain bit allocation parameters and sound channel signal encoding parameters for a plurality of sound channels of a combined format audio signal; and reconstructing the combined format audio signal based on the bit allocation parameters and the sound channel signal encoding parameters for the plurality of sound channels.

[0050] In combination with some embodiments of the second aspect, in some embodiments, reconstructing the combined format audio signal based on the bit allocation parameters and the sound channel signal encoding parameters for the plurality of sound channels includes: decoding the sound channel signal encoding parameters based on the bit allocation parameters for the plurality of sound channels; and reconstructing the combined format audio signal using decoded sound channels.

[0051] In a third aspect, embodiments of the disclosure provide an encoding-end device. The coding-end device includes at least one of a transceiver module or a processing module. The encoding-end device is configured to perform an optional implementation method according to the first aspect.

[0052] In a fourth aspect, embodiments of the disclosure provide a decoding-end device. The decoding-end device includes at least one of a transceiver module or a processing module. The decoding-end device is configured to perform an optional implementation method according to the second aspect.

[0053] In a fifth aspect, embodiments of the disclosure provide a communication system. The communication system includes an encoding-end device and a decoding-end device.

[0054] The encoding-end device is configured to perform an optional implementation according to the aforementioned first aspect.

[0055] The decoding-end device is configured to perform an optional implementation according to the aforementioned second aspect.

[0056] In a sixth aspect, embodiments of the disclosure provide a communication device. The communication device includes one or more processors. The one or more processors are configured to call instructions to cause the communication device to perform optional implementations according to the aforementioned first aspect.

[0057] In a seventh aspect, embodiments of the disclosure provide a communication device. The communication device includes one or more processors. The one or more processors are configured to call instructions to enable the communication device to perform optional implementations according to the aforementioned second aspect.

[0058] In an eighth aspect, embodiments of the disclosure provide a storage medium. The storage medium has instructions stored thereon. When the instructions are executed on a communication device, the communication device is caused to perform optional implementations according to the aforementioned first and second aspects.

[0059] In a ninth aspect, embodiments of the disclosure provide a program product. When the program product is executed by a communication device, the communication device is caused to perform the method described in optional implementations according to the first aspect and the second aspect.

[0060] In a tenth aspect, embodiments of the disclosure provide a computer program. When the computer program is executed on a computer, the computer is enabled to perform the method described in optional implementations according to the first and second aspects.

[0061] In an eleventh aspect, embodiments of the disclosure provide a chip or a chip system. The chip or the chip system includes a processing circuit configured to perform the method described in optional implementations according to the first aspect and the second aspect.

[0062] It is understandable that the above encoding-end device, the decoding-end device, the communication system, the storage medium, the program product, the computer program, the chip or the chip system are all used to perform the method provided in embodiments of the disclosure. Therefore, for the beneficial effects that may be achieved, reference may be made to the beneficial effects in the corresponding methods, which will not be repeated here.

[0063] Embodiments of the disclosure provide methods for processing a signal and devices thereof. In some embodiments, the terms such as "method for processing a signal", "communication method", "method for encoding and decoding a signal", "method for encoding a signal", "method for decoding a signal" or the like may be used interchangeably, the terms such as "device for processing a signal", "communication device" or the like may be used interchangeably, and the terms such as "system for processing a signal", "system for processing information", "communication system", or the like may be used interchangeably.

[0064] Embodiments of the disclosure are not exhaustive, but are only illustrative of some embodiments, and are not intended to be a specific limitation on the scope of protection of the disclosure. In the absence of contradiction, each step in a certain embodiment may be implemented as an independent embodiment, and the steps may be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment may also be implemented as an independent embodiment, and the order of the steps in a certain embodiment may be arbitrarily exchanged. In addition, the optional implementations in a certain embodiment may be arbitrarily combined. Furthermore, embodiments may be arbitrarily combined. For example, some or all of the steps of different embodiments may be arbitrarily combined, and a certain embodiment may be arbitrarily combined with the optional implementations of other embodiments.

[0065] In each embodiment of the disclosure, unless otherwise specified or there is a logical conflict, the terms and / or descriptions between embodiments are consistent and may be referenced to each other, and the technical features in different embodiments may be combined to form a new embodiment based on their internal logical relationships.

[0066] The terms used in embodiments of the disclosure are only for the purpose of describing specific embodiments and are not intended to limit the disclosure.

[0067] In embodiments of the disclosure, unless otherwise specified, elements expressed in the singular form, such as "a", "an", "the", "above", "said", "aforementioned", "this", or the like may mean "one and only one", "one or more", "at least one", or the like. For example, when using articles such as "a", "an", "the" in English in translation, the noun after the article may be understood as a singular expression or a plural expression.

[0068] In embodiments of the disclosure, the term "a plurality of' refers to two or more.

[0069] In some embodiments, the terms "at least one", "one or more", "a plurality of", "multiple", or the like may be used interchangeably.

[0070] In some embodiments, "at least one of A and B", "A and / or B", "A, in one case, and B, in another case", "in response to one case, A, and in response to another case, B", or the like may include following technical solutions according to the situation: in some embodiments, A (A is performed independently of B); in some embodiments, B (B is performed independently of A); in some embodiments, one is selected from A and B for execution (A and B are selectively performed); or in some embodiments, A and B (both A and B are performed). When there are more branches such as A, B, C, etc., the above is also applied thereto.

[0071] In some embodiments, the term "A or B" may include following technical solutions according to the situation: in some embodiments, A (A is performed independently of B); in some embodiments, B (B is performed independently of A); or in some embodiments, one is selected from A and B for execution (A and B are selectively performed). When there are more branches such as A, B, C, etc., the above is also applied thereto.

[0072] The prefixes such as "first" and "second" in embodiments of the disclosure are only used to distinguish different described objects, and do not constitute restrictions on the position, order, priority, quantity or content of the described objects. For the described object, reference may be made to the description in the context of the claims or embodiments, and should not constitute unnecessary restrictions due to the use of prefixes. For example, in a case where the described object is "field", the ordinal number before the "field" in the "first field" and the "second field" does not limit the position or order between the "fields", and the "first" and "second" do not limit whether the "fields" they define are in the same message, nor do they limit the order of the "first field" and the "second field". As another example, in a case where the described object is a "level", the ordinal number before the "level" in the "first level" and the "second level" does not limit the priority between the "levels". As another example, the number of described objects is not limited by the ordinal number, and may be one or more. Taking the "first device" as an example, the number of "devices" may be one or more. In addition, the objects defined by different prefixes may be the same or different. For example, in a case where the described object is "device", the "first device" and the "second device" may be the same device or different devices, and their types may be the same or different. As another example, in a case where the described object is "information", the "first information" and the "second information" may be the same information or different information, and their contents may be the same or different.

[0073] In some embodiments, "including A", "containing A", "for indicating A", and "carrying A" may be interpreted as directly carrying A or indirectly indicating A.

[0074] In some embodiments, terms such as "in response to ...", "in response to determining ...", "in the case of ...", "at the time of ...", "when ...", "if ...", or the like may be used interchangeably.

[0075] In some embodiments, terms such as "greater than", "greater than or equal to", "not smaller than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not lower than", and "above" may be used interchangeably, and terms such as "smaller than", "smaller than or equal to", "not greater than", "less than", "less than or equal to", "no more than", "lower than", "lower than or equal to", "not higher than", and "below" may be used interchangeably.

[0076] In some embodiments, devices and apparatuses may be interpreted as physical or virtual, and their names are not limited to the names used in embodiments. In some cases, they may also be understood as "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "subject", or the like.

[0077] In some embodiments, "network" may be interpreted as devices included in the network, such as access network device, core network device, or the like.

[0078] In some embodiments, "access network device (AN device)" may also be referred to as "radio access network device (RAN device)", "base station (BS)", "radio base station", "fixed station", and in some embodiments may also be understood as "node", "access point", "transmission point (TP)", "reception point (RP)", "transmission / reception point (TRP)", "panel", "antenna panel", "antenna array" "cell", "macro cell", "small cell", "femto cell", "pico cell", "sector", "cell group", "serving cell", "carrier", "component carrier", "bandwidth part (BWP)" and the like.

[0079] In some embodiments, "terminal" or "terminal device" may be referred to as "user equipment (UE)", "user terminal" "mobile station (MS)", "mobile terminal (MT)", "subscriber station", "mobile unit", "subscriber unit", "wireless unit", "remote unit", "mobile device", "wireless device", "wireless communication device", "remote device", "mobile subscriber station", "access terminal", "mobile terminal", "wireless terminal", "remote terminal", "handset", "user agent", "mobile client", "client", or the like.

[0080] In some embodiments, collection / acquisition of data, information, or the like may comply with the laws and regulations of the country where the data is obtained.

[0081] In some embodiments, data, information, or the like may be obtained with the user's consent.

[0082] FIG. 1 is a schematic diagram illustrating an architecture of a communication system according to an embodiment of the disclosure. The communication system may include, but is not limited to, an encoding-end device and a decoding-end device. The number and form of devices illustrated in FIG. 1 are only examples and do not constitute a limitation on embodiments of the disclosure. In actual applications, two or more encoding-end devices and two or more decoding-end devices may be included. The communication system 100 illustrated in FIG. 1 including one encoding-end device 101 and one decoding-end device 102 is taken as an example.

[0083] In some embodiments, the encoding-end device 101 may be a terminal with an encoding function, for example, the terminal includes an encoder. The terminal may be understood as a collection / acquisition end for collecting or acquiring a combined format audio signal. In some embodiments, the terminal may be an entity on the user side for receiving or transmitting signals, such as a mobile phone. It may also be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), or the like. The terminal may be a vehicle with communication function, a smart vehicle, a mobile phone, a wearable device, a tablet computer (Pad), a computer with wireless transceiver function, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, or the like. Embodiments of the disclosure do not limit the specific technology and specific device form adopted by the terminal.

[0084] In some embodiments, the encoding-end device 101 may be an encoder on a network device. For example, an encoder is deployed on the network device, and the encoder on the network device encodes the inputted combined format audio signal. For example, the collection end sends the combined format audio signal to the network device. The encoder on the network device encodes the combined format audio signal and sends the bitstream obtained with the encoding process to the decoding-end device 102.

[0085] In some embodiments, the decoding-end device 102 may be a terminal with a decoding function, for example, the terminal includes an encoder. The terminal may be understood as a playback end of the combined format audio signal. In some embodiments, the terminal may be an entity on the user side for receiving or transmitting signals, such as a mobile phone. It may also be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. The terminal may be a vehicle with communication function, a smart vehicle, a mobile phone, a wearable device, a tablet computer (Pad), a computer with wireless transceiver function, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, or the like. Embodiments of the disclosure do not limit the specific technology and specific device form adopted by the terminal.

[0086] In some embodiments, the decoding-end device 102 may be a decoder on a network device. For example, a decoder is deployed on the network device, and the decoder on the network device decodes the received bitstream, reconstructs a combined format audio signal, and sends the reconstructed combined format audio signal to the playback end.

[0087] In some embodiments, the network device may be an access network device. In some embodiments, the access network device is, for example, a node or a device that connects a terminal to a wireless network, and the access network device may include, but is not limited to, at least one of an evolved NodeB (eNB), a next generation eNB (ng-eNB), a next generation NodeB (gNB), a node B (NB), a home node B (HNB), a home evolved nodeB (HeNB), a wireless backhaul device, a radio network controller (RNC), a base station controller (BSC), a base transceiver station (BTS), a base band unit (BBU), a mobile switching center, a base station in a sixth generation (6G) communication system, an open RAN, a cloud RAN, a base station in other communication systems, or an access node in a wireless fidelity (Wi-Fi) system.

[0088] In some embodiments, the technical solution according to the disclosure may be applicable to an Open RAN architecture. In this case, the interfaces between the access network devices or within the access network device involved in embodiments of the disclosure may become internal interfaces of the Open RAN, and the processes and information interactions between these internal interfaces may be implemented through software or programs.

[0089] In some embodiments, the access network device may be composed of a central unit (CU) and a (DU). The CU may also be referred to as control unit. The CU-DU structure may be used to split the protocol layer of the access network device, with some functions of the protocol layer being centrally controlled by the CU, the remaining part or all of the functions of the protocol layer being distributed in the DU, and the DU being centrally controlled by the CU, but the disclosure is not limited thereto.

[0090] It may be understood that the communication system according to embodiments of the disclosure is for the purpose of more clearly illustrating the technical solution according to embodiments of the disclosure, and does not constitute a limitation on the technical solution according to embodiments of the disclosure. A person of ordinary skill in the art may know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution according to embodiments of the disclosure is also applicable to similar technical problems.

[0091] The following embodiments of the disclosure may be applied to the communication system 100 or part of subjects illustrated in FIG. 1, but the disclosure is not limited thereto. The subjects illustrated in FIG. 1 are examples, and the communication system may include all or part of the subjects in FIG. 1, or may include other subjects that are not illustrated in FIG. 1. The number and form of the subjects are arbitrary. The subjects may be physical or virtual. The connection relationship between the subjects is an example. The subjects may be connected or disconnected. The connection may be in any manner, such as a direct connection, an indirect connection, a wired connection, or a wireless connection.

[0092] Embodiments of the disclosure may be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G new radio (NR), future radio access (FRA), new-radio access technology (RAT), new radio (NR), new radio access (NX), future generation radio access (FX), Global System for Mobile communications (GSM ®< ), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi ®< ), IEEE 802.16 (WiMAX ®< ), IEEE 802.20, Ultra-WideBand (UWB), Bluetooth ®< , Public Land Mobile Network (PLMN) network, Device-to-Device (D2D) system, Machine-to-Machine (M2M) system, Internet of Things (IoT) system, Vehicle-to-Everything (V2X), systems using other communication methods, next-generation systems based on them, or the like. In addition, multiple systems may also be combined (for example, a combination of LTE or LTE-A and 5G) for application.

[0093] It should be noted that the first generation mobile communication technology (1G) is the first generation wireless cellular technology, which belongs to analog mobile communication network. When 1G was upgraded to 2G, the mobile phone was transferred from analog communication to digital communication, mainly using the Global System for Mobile Communications (GSM) network standard, and the voice encoder uses Adaptive Multi-Rate (AMR), Enhanced Full Rate (EFR), Full Rate (FR), Half Rate (HR) to provide mono narrowband voice services. The 3G mobile communication system was proposed by the International Telecommunication Union (ITU) for international mobile communications in 2000. China Mobile uses TD-SCDMA, China Telecom uses CDMA2000, and China Unicom uses WCDMA. Its voice encoder uses Adaptive Multi-Rate Wideband (AMR-WB) to provide mono broadband voice services. 4G is a better improvement over 3G technology. Both data and voice use full Internet protocol (IP) mode, providing real-time HD / HD+Voice services for voice and audio. The EVS codec used may take into account high-quality compression and reconstruction of voice and audio.

[0094] The voice and audio communication services provided above have expanded from narrowband signals to ultra-wideband and even full-band services, but they are still mono services. People's demand for high-quality audio is increasing. Compared with mono audio, stereo audio has a sense of orientation and distribution for each sound source, and may improve clarity. With the increase of transmission bandwidth and the upgrade of audio collection / acquisition equipment of the terminal, signal processor performance is improved and the terminal playback equipment is upgraded.

[0095] In some embodiments, signals in three formats, such as a sound channel-based audio signal, an object-based audio signal, and a scene-based audio signals, may provide three-dimensional audio services. The Immersive Voice and Audio Services (IVAS) codec being standardized by 3GPP SA4 may support the encoding and decoding requirements of the above signals in three formats.

[0096] In some embodiments, the sound channel-based audio signal includes, but is not limited to, a mono signal, a stereo signal, a binaural signal, a 5.1, 7.1 surround sound signal, a 5.1.4, 7.1.4 surround sound signal, where ".4" stands for a height sound channel signal.

[0097] In some embodiments, the scene-based audio signal includes, but is not limited to, First Order Ambisonics (FOA), second Order High Ambisonics (HOA2), and third Order High Ambisonics (HOA3).

[0098] In some embodiments, the object-based audio signal includes audio data and metadata. In addition, IVAS also supports a metadata-assisted spatial audio (MASA) signal.

[0099] In some embodiments, a terminal that may support three-dimensional audio services may include, but is not limited to, a mobile phone, a computer, a tablet, conference system equipment, AR / VR equipment, a vehicle, or the like.

[0100] In some embodiments, in an application scenario of three-dimensional audio, the three-dimensional audio generally includes a signal in multiple audio formats, that is, a combined format audio signal. The encoding-end device receives the combined format audio signal, and an audio bitstream that is obtained by encoding the combined format audio signal is sent from a sending end to a receiving end. The decoder at the receiving end decodes the received audio bitstream and reconstructs the combined format audio signal.

[0101] However, the encoding and decoding method of the combined format audio signal in the related art is usually as follows. The combined format audio signal is subjected to corresponding unified encoding and decoding processing, the processing process on the encoding end is to performs bit allocation for an available bit budget based on energy parameters of the combined format audio signal, and each sound channel is encoded based on the allocated bits using an encoding core to output an encoding parameter, and the encoding parameters are placed in the bitstream. In other words, the encoder in the related art uses the bit allocation based on the energy parameters for the inputted combined format audio signal, and does not perform any sound field analysis on the combined format audio signal, resulting in the inability to reconstruct the most suitable sound field for a specific application scenario, or the bit allocation cannot be adjusted based on external control parameters, resulting in the failure to reconstruct a desired sound field.

[0102] In view of this, embodiments of the disclosure provide methods for processing a signal and devices thereof, which may solve the problem that the encoding and decoding method in the related art cannot reconstruct the desired and most suitable sound field to improve the reconstruction effect of the combined format audio signal.

[0103] FIG. 2A is a schematic diagram illustrating an interaction of a method for processing a signal according to an embodiment of the disclosure. As illustrated in FIG. 2A, the method for processing a signal according to an embodiment of the disclosure may be applied to a communication system 100, and the method includes but is not limited to the following.

[0104] At step S2101, an encoding-end device 101 obtains a combined format audio signal (or called mixed format audio, MFA).

[0105] In some embodiments, the encoding-end device 101 may be a terminal or a base station, and the terminal may be a device that provides voice and / or data connectivity to the user. The terminal may communicate with one or more core networks via Radio Access Network (RAN). The terminal may be an Internet of Things terminal, such as a sensor device, a mobile phone (or a "cellular" phone), or a computer with an Internet of Things terminal. For example, the terminal may be a fixed, portable, pocket-sized, handheld, computer-built-in or vehicle-mounted device. For example, it may be a station (STA), a subscriber unit, a subscriber station, a mobile station, a mobile, a remote station, an access point, a remote terminal, an access terminal, a user terminal or a user agent. Or the terminal may be a device of an unmanned aerial vehicle. Or the terminal may be a vehicle-mounted device. For example, it may be a driving computer with wireless communication function, or a wireless terminal connected to an external driving computer. Or the UE may also be a roadside device (such as, a street lamp, a signal lamp or other roadside device with a wireless communication function).

[0106] In some embodiments, the combined format audio signal includes audio signals in at least one of following formats: a sound channel-based audio signal, an object-based audio signal, a scene-based audio signal, or a metadata-based three-dimensional audio signal.

[0107] In embodiments of the disclosure, the combined format audio signal may include any one of the sound channel-based audio signal, the object-based audio signal, the scene-based audio siganl, or the metadata-based three-dimensional audio signal.

[0108] In embodiments of the disclosure, the combined format audio signal may include the object-based audio signal and the scene-based audio signal.

[0109] In embodiments of the disclosure, the combined format audio signal may include the object-based audio signal and the sound channel-based audio signal.

[0110] In embodiments of the disclosure, the combined format audio signal may include the sound channel-based audio signal and the scene-based audio signal.

[0111] In embodiments of the disclosure, the combined format audio signal may include the object-based audio signal and the metadata-based three-dimensional audio signal.

[0112] In embodiments of the disclosure, the combined format audio signal may include the sound channel-based audio signal and the metadata-based three-dimensional audio signal.

[0113] It should be noted that the above embodiments are not exhaustive but are only illustrations of some embodiments, and the above embodiments may be implemented individually or in combination. The above embodiments are only for illustration and are not intended to be specific limitations on the scope of protection of embodiments of the disclosure.

[0114] In some embodiments, the above audio signals in four audio formats are categorized based on their collection methods, and different formats of audio signals are also oriented towards distinct application scenarios.

[0115] In an embodiment of the disclosure, a main application scenario of the above-mentioned sound channel-based audio signal may be that a microphone collection / acquisition layout provided to a collection / acquisition end is the same as a speaker playback layout provided to a playback end, where the microphone of the collection / acquisition end may be used to collect or acquire the sound channel-based audio signal of the 5.0 format; the speaker of the playback end may play back the sound channel-based audio signal of the 5.0 format collected by the collection / acquisition end.

[0116] In another embodiment of the disclosure, the above-mentioned object-based audio signal is usually recorded by an independent microphone for a sound-emitting object, and its main application scenario is that independent control operations need to be performed on this audio signal at the playback end, such as sound switch, volume adjustment, sound image orientation adjustment, frequency band equalization processing and other control operations.

[0117] In another embodiment of the disclosure, a main application scenario of the above-mentioned scene-based audio signal may be that it needs to record a complete sound field where the collection / acquisition end is located, such as live recording of a concert, live recording of a football game, etc.

[0118] In some embodiments, the "collection / acquisition end" and the "encoding-end device 101" may be located on the same device or on different devices. As an example, in a case where the "collection / acquisition end" has an encoding function, the "collection / acquisition end" and the "encoding-end device 101" may be located on the same device, such as being interchangeable. As another example, in a case where the "collection / acquisition end" does not have an encoding function, the "collection / acquisition end" and the "encoding-end device 101" may be located on different devices, such as the "encoding-end device 101" being located on a network device.

[0119] In some embodiments, the "playback end" and the "decoding-end device 102" may be located on the same device or on different devices. As an example, in a case where the "playback end" has a decoding function, the "playback end" and the "decoding-end device 102" may be located on the same device, such as being interchangeable. As another example, in a case where the "playback end" does not have a decoding function, the "playback end" and the "decoding-end device 101" may be located on different devices, such as the "decoding-end device 102" being located on a network device.

[0120] In some embodiments, the combined format audio signal may be a multi- channel audio signal. In an implementation, the sound channel-based audio signal may include one or more sound channels, the object-based audio signal may include one or more sound channels, the scene-based audio signal may include one or more sound channels, and / or the metadata-based three-dimensional audio signal may include one or more sound channels.

[0121] At step S2102, the encoding-end device 101 performs filtering processing on the combined format audio signal.

[0122] In some embodiments, the encoding-end device 101 may perform high-pass filtering on the combined format audio signal. For example, a cutoff frequency of the filter may be set to 20 Hz. As an example, the filter formula may be as shown in a following formula (1): H 20 Z = b 0 + b 1 z − 1 + b 2 z − 2 1 + a 1 z − 1 + a 2 z − 2 , where H 20 represents a filter with the cutoff frequency of 20 Hz, Z represents the combined format audio signal, a 1 , a 2 , b 0 , b 1 , b 2 are all preset constants. For example, b 0 =0. 9981492 b 1 =-1. 9963008, b 2 =0. 9981498 a 1 =1. 9962990, a 2 =· -0. 9963056

[0123] At step S2103, the encoding-end device 101 determines one or more sound field analysis parameters of the combined format audio signal.

[0124] In some embodiments, the above-mentioned sound field analysis parameter includes at least one of: the total number of sound images; relative priority of sound images; or a background sound. In an implementation, the above-mentioned sound field analysis parameter includes any one of the total number of sound images, the relative priority of sound images, or the background sound. In another implementation, the above-mentioned sound field analysis parameter includes any two of the total number of sound images, the relative priority of sound images, or the background sound. In yet another implementation, the above-mentioned sound field analysis parameters includes the total number of sound images, the relative priority of sound images, and the background sound.

[0125] In some embodiments, the encoding-end device 101 may directly perform sound field analysis on the inputted combined format audio signal to determine the sound field analysis parameter of the combined format audio signal. Or in some embodiments, the encoding-end device 101 may perform the sound field analysis on the combined format audio signal that has been filtered (such as processed with the high-pass filtering) to determine the above-mentioned sound field analysis parameter. The sound field analysis refers to analyzing the total number of sound images, the relative priority of sound images, or the background sound of the combined format audio signal.

[0126] In some embodiments, the sound field analysis may be performed on the combined format audio signal based on Direction-Of-Arrival (DOA, also called wave arrival direction) estimation and / or Multichannel Cross-Correlation Coefficients (MCCC) method to obtain the sound field analysis parameter of the combined format audio signal. For example, sound source localization estimation (including azimuth and pitch angle) may be performed through the DOA estimation, and the MCCC method may be used for solution. Based on this method, the sound field analysis parameter of the combined format audio signal is identified, such as the total number of sound images, the relative priority of sound images, the background sound, or the like.

[0127] At step S2104, the encoding-end device 101 determines energy parameter(s) of one or more sound channels among the plurality of sound channels.

[0128] In some embodiments, the encoding-end device 101 may directly perform the energy calculation on the input combined format audio signal to determine the energy parameter(s) of one or more sound channels among the above multiple sound channels. For example, the encoding-end device 101 calculates cross-correlation coefficients between the multiple sound channels, performs sound channel combination operation based on the cross-correlation coefficients, and determines inter-sound channel parameters for sound channels in a same sound channel combination, aligns sound channels and downmixing aligned sound channels with the inter-channel parameters, and performs energy calculation on downmixed sound channels to obtain a respective energy parameter of each sound channel.

[0129] In some embodiments, the encoding-end device 101 may perform the energy calculation on the filtered combined format audio signal (that has been processed with, such as, high-pass filtering) to determine the energy parameter(s) of one or more sound channels among the above-mentioned multiple sound channels. For example, after filtering the multiple sound channels, the encoding-end device 101 calculates the cross-correlation coefficients between each two of the filtered sound channels, performs sound channel combination operation based on the cross-correlation coefficients, calculates the inter-sound channel parameters for sound channels in the same sound channel combination, uses the inter-sound channel parameters to align sound channels and perform downmixing, and performs energy calculation on downmixed sound channels to obtain the energy parameters of the sound channels respectively.

[0130] In some embodiments, the above-mentioned inter-sound channel parameter may be one or more of Interaural Level Difference (ILD), Interaural Time Difference (ITD), Interaural Phase Difference (IPD), or the like.

[0131] In some embodiments, the above energy parameter calculation method may include performing any one of following processing on sample points of a frame of a downmixed sound channel: calculating a sum of levels of the sample points; calculating a sum of squared levels of the sample points; or calculating a root mean square of levels of the sample points.

[0132] For example, the sum of levels of the sample points of a frame of a downmixed sound channel may be calculated to achieve the energy calculation. For example, a respective sum of levels of the sample points of a frame of each sound channel is calculated, and the obtained respective sum of levels of the sample points of a frame of each sound channel is determined as the energy parameter of the corresponding sound channel. For example, assuming that the combined format audio signal includes 3 sound channels, such as sound channel a, sound channel b, and sound channel c, a respective sum of levels of the sample points of a frame of each of the 3 sound channels is calculated, and the calculated sum of levels of the sample points of a frame of the sound channel a is determined as the energy parameter of the sound channel a, the calculated sum of levels of the sample points of a frame of the sound channel b is determined as the energy parameter of the sound channel b, and the calculated sum of levels of the sample points of a frame of the sound channel c is determined as the energy parameter of the sound channel c. As an example, a formula for calculating the above-mentioned sum of levels of the sample points is expressed as follows: S m = ∑ m = 0 N − 1 X m , where S m represents the sum of levels of the sample points of a frame of an m th< sound channel, a frame of an m th< sound channel includes N sample points, and X m represents an m th< sound channel.

[0133] For example, the sum of squared levels of the sample points of a frame of a downmixed sound channel may be calculated to achieve the energy calculation. For example, a respective sum of squared levels of the sample points of a frame of each sound channel is calculated, and the obtained respective sum of squared levels of the sample points of a frame of each sound channel is determined as the energy parameter of the corresponding sound channel. For example, assuming that the combined format audio signal includes three sound channels, such as sound channel a, sound channel b, and sound channel c, the respective sum of squared levels of the sample points of a frame of each sound channel is calculated respectively, the calculated sum of squared levels of the sample points of a frame of the sound channel a is determined as the energy parameter of the sound channel a, the calculated sum of squared levels of the sample points of a frame of the sound channel b is determined as the energy parameter of the sound channel b, and the calculated sum of squared levels of the sample points of a frame of the sound channel c is determined as the energy parameter of the sound channel c. As an example, the formula for calculating the sum of squared levels of the sample points is expressed as follows: S m = ∑ m = 0 N − 1 X m ∗ X m .

[0134] For example, the root mean square of levels of the sample points of a frame of a downmixed sound channel may be calculated to achieve the energy calculation. For example, a respective root mean square of levels of the sample points of a frame of each sound channel is calculated, and the respective root mean square of levels of the sample points of a frame of each sound channel is determined as the energy parameter of the corresponding sound channel. For example, assuming that the combined format audio signal includes three sound channels, such as sound channel a, sound channel b, and sound channel c, respective root mean squares of levels of the sample points of the three sound channels are calculated respectively, the calculated root mean square of levels of the sample points of a frame of the sound channel a is determined as the energy parameter of the sound channel a, the calculated root mean square of levels of the sample points of a frame of the sound channel b is determined as the energy parameter of the sound channel b, and the calculated root mean square of levels of the sample points of a frame of the sound channel c is determined as the energy parameter of the sound channel c. As an example, the formula for calculating the root mean square of levels of the sample points is expressed as follows: S m = 1 N ∑ m = 0 N − 1 X m ∗ X m .

[0135] For example, the above energy parameters may also be calculated and obtained by other methods, which are not limited in the disclosure and will not be described in detail.

[0136] At step S2105, the encoding-end device 101 performs bit allocation for the multiple sound channels based on the sound field analysis parameter of the combined format audio signal and the energy parameter(s) of one or more sound channels among the multiple sound channels to obtain bit allocation parameters.

[0137] In some embodiments, a bit allocation parameter is used for describing the number of bits allocated to one of the multiple sound channels.

[0138] In some embodiments, the encoding-end device 101 may perform the bit allocation for the multiple sound channels using the available bit budget based on the sound field analysis parameter of the combined format audio signal and the energy parameter(s) of one or more sound channels among the multiple sound channels, to obtain the bit allocation parameters.

[0139] In some embodiments, taking the above-mentioned combined format audio signal including the object-based audio signal and the scene-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and the scene-based audio signal based on the sound field analysis parameter of the combined format audio signal and the energy parameter(s) of one or more sound channels among the multiple sound channels to obtain the bit allocation parameters. For example, the object-based audio signal is a main sound element in the sound field, and the scene-based audio signal is an ambient background sound element in the sound field. To reconstruct an expected sound field, the object-based audio signal is required to have a smaller distortion, while the scene-based audio signal used as the background sound is allowed to have a certain degree of distortion. Therefore, in allocating bits, more bits may be allocated to the object-based audio signal compared to the scene-based audio signal. For example, the number of bits allocated to any sound channel of the object-based audio signal is more than the number of bits allocated to any sound channel of the scene-based audio signal. The encoding-end device 101 may also adjust the number of bits allocated to each sound channel by considering the energy parameter(s) of one or more sound channels of the combined format audio signal. For example, the sound channel with a larger energy parameter is allocated with relatively more bits.

[0140] In some embodiments, taking the combined format audio signal including the object-based audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and of the sound channel-based audio signal based on the sound field analysis parameter of the combined format audio signal and the energy parameter(s) of one or more sound channels among the multiple sound channels to obtain the bit allocation parameters. For example, the object-based audio signal is a main sound element in the sound field, while the sound channel-based audio signal is an ambient background sound element in the sound field. To reconstruct an expected sound field, the object-based audio signal is required to have a smaller distortion, while the sound channel-based audio signal used as a background sound is allowed to have a certain degree of distortion. Therefore, in allocating bits, more bits may be allocated to the object-based audio signal than to the sound channel-based audio signal. For example, the number of bits allocated to any sound channel of the object-based audio signal is more than the number of bits allocated to any sound channel of the sound channel-based audio signal. The encoding-end device 101 may also adjust the number of bits allocated to each sound channel by considering the energy parameter of each sound channel in the combined format audio signal. For example, a sound channel with a larger energy parameter is allocated with relatively more bits.

[0141] In some embodiments, taking the combined format audio signal including the scene-based audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for multiple sound channels in the combined format audio signal composed of the scene-based audio signal and the sound channel-based audio signal based on the sound field analysis parameter of the combined format audio signal and the energy parameter(s) of one or more sound channels among the multiple sound channels to obtain the bit allocation parameters. For example, the sound channel-based audio signal is a main sound element in the sound field, while the scene-based audio signal is an ambient background sound element in the sound field. To reconstruct an expected sound field, the sound channel-based audio signal is required to have a smaller distortion, while the scene-based audio signal used as a background sound is allowed to have a certain degree of distortion. Therefore, in allocating bits, more bits may be allocated to the sound channel-based audio signal than to the scene-based audio signal. For example, the number of bits allocated to any sound channel of the sound channel-based audio signal is more than the number of bits allocated to any sound channel of the scene-based audio signal. The encoding-end device 101 may also adjust the number of bits allocated to each sound channel by considering the energy parameter(s) of one or more sound channels of the combined format audio signal. For example, a sound channel with a larger energy parameter is allocated with relatively more bits.

[0142] In some embodiments, taking the above-mentioned combined format audio signal including the object-based audio signal and the metadata-based three-dimensional audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and of the metadata-based three-dimensional audio signal based on the sound field analysis parameter of the combined format audio signal and the energy parameter(s) of one or more sound channels among the multiple sound channels to obtain the bit allocation parameters. For example, the object-based audio signal is a main sound element in the sound field, while the metadata-based three-dimensional audio signal is an ambient background sound element in the sound field. To reconstruct an expected sound field, the object-based audio signal is required to have a smaller distortion, while the metadata-based three-dimensional audio signal used as a background sound is allowed to have a certain degree of distortion. Therefore, in allocating bits, more bits may be allocated to the object-based audio signal than to the metadata-based three-dimensional audio signal. For example, the number of bits allocated to any sound channel of the object-based audio signal is more than the number of bits allocated to any sound channel of the metadata-based three-dimensional audio signal. The encoding-end device 101 may also adjust the number of bits allocated to each sound channels by considering the energy parameter(s) of one or more sound channels of the combined format audio signal. For example, a sound channel with a larger energy parameter is allocated with relatively more bits.

[0143] In some embodiments, taking the above-mentioned combined format audio signal including the metadata-based three-dimensional audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for multiple sound channels of the combined format audio signal composed of the metadata-based three-dimensional audio signal and of the sound channel-based audio signal based on the sound field analysis parameter of the combined format audio signal and the energy parameter(s) of one or more sound channels among the multiple sound channels to obtain the bit allocation parameters. For example, the sound channel-based audio signal is a main sound element in the sound field, while the metadata-based three-dimensional audio signal is an ambient background sound element in the sound field. To reconstruct an expected sound field, the sound channel-based audio signal is required to have a smaller distortion, while the metadata-based three-dimensional audio signal used as a background sound is allowed to have a certain degree of distortion. Therefore, in allocating bits, more bits may be allocated to the sound channel-based audio signal than to the metadata-based three-dimensional audio signal. For example, the number of bits allocated to any sound channel of the sound channel-based audio signal is more than the number of bits allocated to any sound channel of the metadata-based three-dimensional audio signal. The encoding-end device 101 may also adjust the number of bits allocated to each sound channels by considering the energy parameter(s) of one or more sound channels of the combined format audio signal. For example, a sound channel with a larger energy parameter is allocated with relatively more bits.

[0144] In some embodiments, the number of bits allocated to a sound channel corresponding to a primary sound image of the object-based audio signal is greater than the number of bits allocated to a sound channel corresponding to a secondary sound image of the object-based audio signal. In other words, in the object-based audio signal, each object signal has a different importance in terms of its part in the combined format audio signal. For example, at a concert, the importance of a lead singer's voice signal is higher than that of a backing singer's voice signal, or a lead singer's voice signal is more important than that of a backing singer's voice signal, and the lead singer's object audio signal are allocated more bits than the backing singer's voice signal.

[0145] It should be noted that the composition of the combined format audio signal shown in the above embodiments is for the convenience of those skilled in the field to understand how to allocate bits to each sound channel to be encoded based on the sound field analysis parameter and the energy parameter(s). In other words, the bit allocation may be carried out for the combined format audio signal composed of the audio signals in other formats based on the above-mentioned sound field analysis parameter and the energy parameters, which are not repeated here.

[0146] In some embodiments, the number of bits of each sound channel of the audio signals in different formats of the above-mentioned combined format audio signal may be the same (the bits are evenly allocated), or the bits may be allocated based on other factors (such as a factor of whether it is primary or secondary). The above is not limited in the disclosure and is not repeated here.

[0147] At step S2106, the encoding-end device 101 encodes the multiple sound channels based on the bit allocation parameters to obtain sound channel signal encoding parameters, and writes the bit allocation parameters and the sound channel signal encoding parameters in a bitstream.

[0148] In some embodiments, the encoding-end device 101 may encode the multiple sound channels using corresponding encoding cores based on the bit allocation parameters to obtain the sound channel signal encoding parameters, and write the bit allocation parameters and the sound channel signal encoding parameters in a bitstream.

[0149] In some embodiments, the encoding cores corresponding to audio signals in different formats may be the same or different. For example, all sound channels of the combined format audio signal may be encoded using one encoding core based on the bit allocation parameters. Or the combined format audio signal includes audio signals in two different formats, such as the object-based audio signal and the scene-based audio signal, assuming that the object-based audio signal includes a sound channel a and a sound channel b, and that the scene-based audio signal includes a sound channel c, the bit allocation parameters of the sound channel a, the sound channel b, and the sound channel c are 3, 2, and 1 respectively. That is, the number of bits allocated to the sound channel a is 3, the number of bits allocated to the sound channel b is 2, and the number of bits allocated to the sound channel c is 1. The encoding-end device 101 may encode the sound channel a based on 3 bits using a encoding core corresponding to the object-based audio signal, encode the sound channel b based on 2 bits using the encoding core corresponding to the object-based audio signal, and encode the sound channel c based on 1 bit using an encoding core corresponding to the scene-based audio signal.

[0150] In some embodiments, each sound channel of the combined format audio signal corresponds to a respective encoding core, and the encoding cores corresponding to the sound channels may be the same or different, and each sound channel of the combined format audio signal may be encoded using the respective encoding core based on the number of allocated bits.

[0151] At step S2107, the encoding-end device 101 sends the bitstream.

[0152] In some embodiments, the encoding-end device 101 may send the bitstream to the decoding-end device 102, In other words, the bitstream is sent by encoding-end device 101 to the decoding-end device 102.

[0153] In some embodiments, the encoding-end device 101 may send the bitstream to the decoding-end device 102 based on multiplexing (MUX). In some embodiments, the decoding-end device 102 receives the bitstream. For example, the decoding-end device 102 receives the bitstream sent by the encoding-end device 101. The above bitstream is used by the decoding-end device 102 to reconstruct the above combined format audio signal.

[0154] At step S2108, the decoding-end device 102 parses the received bitstream to obtain the bit allocation parameters and the sound channel signal encoding parameters for the multiple sound channels of the combined format audio signal respectively.

[0155] At step S2109, the decoding-end device 102 reconstructs the combined format audio signal based on the bit allocation parameters and the sound channel signal encoding parameters for the multiple sound channels.

[0156] In some embodiments, the decoding-end device 102 decodes the sound channel signal encoding parameter based on the bit allocation parameter of each sound channel. The decoding-end device 102 uses the decoded sound channels to reconstruct the combined format audio signal. In some embodiments, the decoding-end device 102 decodes the sound channel signal encoding parameter based on the bit allocation parameter of each sound channel, and the decoding-end device 102 uses the decoded sound channels and the bitstream to reconstruct the combined format audio signal.

[0157] In some embodiments, the names of information are not limited to those used in embodiments, and terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "code element", "codebook", "codeword", "codepoint", "bit", "data", "program", and "chip" may be used interchangeably.

[0158] In some embodiments, terms such as "obtain", "acquire", "get", "receive", "transmit", "bidirectional transmission", "send and / or receive" may be used interchangeably and may be interpreted as receiving from other entities, obtaining based on a protocol, obtaining from a higher layer, obtaining by self-processing, autonomous implementation, or the like.

[0159] In some embodiments, terms such as "send", "transfer", "report", "issue", "transmit", "bidirectional transmission", "send and / or receive" may be used interchangeably.

[0160] In some embodiments, terms such as "bit" and "the number of bits" may be used interchangeably.

[0161] In some embodiments, terms such as "encoding-end device", "encoder" and "encoding end" may be used interchangeably.

[0162] In some embodiments, terms such as "decoding-end device", "decoder" and "decoding end" may be used interchangeably.

[0163] In some embodiments, terms such as "combined format audio", "combined format audio signal" and "audio signal in combined formats" may be used interchangeably.

[0164] In some embodiments, terms such as "sound channel audio", "sound channel signal", "sound channel audio signal", and "sound channel" may be used interchangeably.

[0165] In some embodiments, terms such as "auxiliary metadata-based spatial audio signal" and "metadata-based three-dimensional audio signal" may be used interchangeably.

[0166] In some embodiments, terms such as "bitstream" and "encoded bitstream" may be used interchangeably. The encoded bitstream may refer to a bitstream including the above-mentioned bit allocation parameters and the above-mentioned sound channel signal encoding parameters.

[0167] The method according to embodiments of the disclosure may include at least one of steps S2101 to S2109. For example, step S2101+step S2102+step S2103+step S2104+step S2105 may be implemented as an independent embodiment, step S2101+step S2103+step S2104+step S2105 may be implemented as an independent embodiment, step S2101+step S2102+step S2103+step S2104+step S2105+step S2106+step S2107 may be implemented as an independent embodiment, step S2101+step S2103+step S2104+step S2105+step S2106+step S2107 may be implemented as an independent embodiment, step S2108+step S2109 may be implemented as an independent embodiment, step S2101+step S2102+step S2103+step S2104+step S2105+step S2106+step S2107+step S2108+step S2109 may be implemented as an independent embodiment, step S2101+step S2103+step S2104+step S2105+step S2106+step S2107+step S2108+step S2109 may be implemented as an independent embodiment, but the disclosure is not limited thereto.

[0168] In some embodiments, step S2103 and step S2104 may be executed in an interchangeable order or simultaneously.

[0169] In some embodiments, step S2106, step S2107, step S2108, and step S2109 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0170] In some embodiments, step S2102, step S2106, step S2107, step S2108, and step S2109 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0171] In some embodiments, step S2108 and step S2109 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0172] In some embodiments, step S2102, step S2108, and step S2109 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0173] In some embodiments, step S2101, step S2102, step S2103, step S2104, step S2105, step S2106, and step S2107 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0174] In some embodiments, step S2102 is optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0175] In some embodiments, reference may be made to other optional implementations described before or after the specification corresponding to FIG. 2A.

[0176] FIG. 2B is a schematic diagram illustrating an interaction of a method for processing a signal according to an embodiment of the disclosure. As illustrated in FIG. 2B, the method for processing a signal according to an embodiment of the disclosure may be performed by a communication system 100, and the method includes but is not limited to the following.

[0177] At step S2201, the encoding-end device 101 obtains a combined format audio signal.

[0178] For optional implementations of step S2201, reference may be made to the optional implementations of step S2101 in FIG. 2A and other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0179] At step S2202, the encoding-end device 101 performs filtering processing on the combined format audio signal.

[0180] For optional implementations of step S2202, reference may be made to optional implementations of step S2102 in FIG. 2A and other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0181] At step S2203, the encoding-end device 101 determines a control parameter of the combined format audio signal.

[0182] In some embodiments, the control parameter may include a second parameter and / or a third parameter. In some embodiments, the second parameter is used for describing an importance ranking of the multiple sound channels of the combined format audio signal. For example, the importance may be set manually or may be obtained by other processing other than encoding processing.

[0183] In some embodiments, the third parameter is used for describing whether each of the sound channels is a diegetic sound or a non-diegetic sound.

[0184] In some embodiments, the encoding-end device 101 may determine the control parameter of the combined format audio signal based on a functional category of each sound channel determined in collecting the combined format audio signal by an audio signal collection device. The audio signal collection device is configured to collect the combined format audio signal. In some embodiments, terms such as "audio signal collection device" and "collection end" may be used interchangeably.

[0185] In some embodiments, the encoding-end device 101 may determine the control parameter of the above-mentioned combined format audio signal based on a setting parameter required for reconstructing a desired sound field by the decoding-end device 102.

[0186] In some embodiments, the encoding-end device 101 may determine the control parameter of the combined format audio signal based on the functional category of each sound channel determined in collecting the combined format audio signal by the audio signal collection device and the setting parameter required for the reconstruction of the desired sound field performed by the decoding-end device 102.

[0187] For example, the above control parameter refers to an importance ranking of the multiple sound channels in the combined format audio signal that is set in advance based on external information. This ranking may be determined based on the respective functional category of each sound channel in collecting the audio signal or may be set based on a requirement for the reconstruction of a desired sound field performed by the playback end. As an example, voice of a lead singer in a concert may be set as a most important sound channel. As another example, in order to focus on the sound of a certain instrument (such as the sound of a violin) in the reconstructed sound field, the sound of the violin may be set as the most important sound channel signal.

[0188] At step S2204, the encoding-end device 101 determines energy parameter(s) of one or more sound channels among the multiple sound channels.

[0189] For optional implementations of step S2204, reference may be made to optional implementations of step S2104 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0190] At step S2205, the encoding-end device 101 performs bit allocation for the multiple sound channels based on the energy parameter(s) of the one or more sound channels and the control parameter of the combined format audio signal to obtain bit allocation parameters.

[0191] In some embodiments, each bit allocation parameter is used for describing the number of bits allocated to a respective one of the multiple sound channels.

[0192] In some embodiments, the encoding-end device 101 may perform the bit allocation for the multiple sound channels using an available bit budget based on the control parameter of the combined format audio signal and the respective energy parameter of each of the multiple sound channels, to obtain the bit allocation parameters.

[0193] In some embodiments, taking the above-mentioned combined format audio signal including the object-based audio signal and the scene-based audio signal as an example, the encoding-end device 101 may allocate bits to multiple sound channels of the combined format audio signal composed of the object-based audio signal and the scene-based audio signal based on the control parameter of the combined format audio signal and the respective energy parameter of each sound channel of the multiple sound channels to obtain the bit allocation parameters. For example, the voice of the lead singer in the concert may be set as the most important sound channel, and the bit allocation parameters may be obtained by allocating bits to multiple sound channels of the combined format audio signal composed of the object-based audio signal and the scene-based audio signal based on the importance ranking of the sound channels in the combined format audio signal and the respective energy parameter of each sound channel of the multiple sound channels. For example, the number of bits allocated to any sound channel of the object-based audio signal is greater than the number of bits allocated to any sound channel of the scene-based audio signal. The encoding-end device 101 may also adjust the number of bits allocated to each sound channel by considering the respective energy parameter of each sound channel of the combined format audio signal. For example, a sound channel with a larger energy parameter is allocated with relatively more bits.

[0194] In some embodiments, taking the above-mentioned combined format audio signal including the object-based audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for multiple sound channels of the combined format audio signal composed of the object-based audio signal and the sound channel-based audio signal based on the control parameter of the combined format audio signal and the respective energy parameter of each of the multiple sound channels to obtain the bit allocation parameters.

[0195] In some embodiments, taking the above-mentioned combined format audio signal including the scene-based audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the scene-based audio signal and the sound channel-based audio signal based on the control parameter of the combined format audio signal and the respective energy parameter of each of the multiple sound channels to obtain the bit allocation parameters.

[0196] In some embodiments, taking the above-mentioned combined format audio signal including the object-based audio signal and the metadata-based three-dimensional audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and the metadata-based three-dimensional audio signal based on the control parameter of the combined format audio signal and the respective energy parameter of each of the multiple sound channels to obtain the bit allocation parameters.

[0197] In some embodiments, taking the above-mentioned combined format audio signal including the metadata-based three-dimensional audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the metadata-based three-dimensional audio signal and the sound channel-based audio signal based on the control parameter(s) of the combined format audio signal and the respective energy parameter of each of the multiple sound channels to obtain the bit allocation parameters.

[0198] In some embodiments, the number of bits allocated to a sound channel corresponding to a primary sound image of the object-based audio signal is greater than the number of bits allocated to a sound channel corresponding to a secondary sound image of the object-based audio signal. In other words, in the object-based audio signal, each object signal has a different importance in terms of its part in the combined format audio signal. For example, at a concert, the importance of a lead singer's voice signal is higher than that of a backing singer's voice signal, or a lead singer's voice signal is more important than that of a backing singer's voice signal, and the lead singer's object audio signal are allocated more bits than the backing singer's voice signal.

[0199] It should be noted that the composition of the combined format audio signal shown in the above embodiments is for the convenience of those skilled in the field to understand how to allocate bits to each sound channel to be encoded based on the control parameter and the energy parameter(s). In other words, the bit allocation may also be carried out for the combined format audio signal composed of the audio signals in other formats based on the above-mentioned control parameter and the energy parameters, which are not repeated here.

[0200] In some embodiments, the number of bits of each sound channel of the audio signals in different formats of the above-mentioned combined format audio signal may be the same (the bits are evenly allocated), or the bits may be allocated based on other factors (such as a factor of whether it is primary or secondary). The above is not limited in the disclosure and is not repeated here.

[0201] At step S2206, the encoding-end device 101 encodes the multiple sound channels based on the bit allocation parameters to obtain sound channel signal encoding parameters, and writes the bit allocation parameters and the sound channel signal encoding parameters into a bitstream.

[0202] For optional implementations of step S2206, reference may be made to optional implementations of step S2106 in FIG. 2A and to other related parts of embodiments involved in FIG. 2A, which are not described in detail here.

[0203] At step S2207, the encoding-end device 101 sends the bitstream.

[0204] For optional implementations of step S2207, reference may be made to optional implementations of step S2107 in FIG. 2A and to other related parts of embodiments involved in FIG. 2A, which are not described in detail here.

[0205] At step S2208, the decoding-end device 102 parses the received bitstream to obtain bit allocation parameters and sound channel signal encoding parameters for multiple sound channels of a combined format audio signal.

[0206] For optional implementations of step S2208, reference may be made to optional implementations of step S2108 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0207] At step S2209, the decoding-end device 102 reconstructs the combined format audio signal based on the r bit allocation parameters and the sound channel signal encoding parameters for the multiple sound channels.

[0208] For optional implementation of step S2209, reference may be made to optional implementations of step S2109 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0209] The method involved in embodiments of the disclosure may include at least one of steps S2201 to S2209. For example, step S2201+step S2202+step S2203+step S2204+step S2205 may be implemented as an independent embodiment, step S2201+step S2203+step S2204+step S2205 may be implemented as an independent embodiment, step S2201+step S2202+step S2203+step S2204+step S2205+step S2206+step S2207 may be implemented as an independent embodiment, step S2201+step S2203+step S2204+step S2205+step S2206+step S2207 may be implemented as an independent embodiment, and step S2201+step S2203+step S2204+step S2205+step S2206+step S2207 may be implemented as an independent embodiment, step S2208+step S2209 may be implemented as an independent embodiment, step S2201+step S2202+step S2203+step S2204+step S2205+step S2206+step S2207+step S2208+step S2209 may be implemented as an independent embodiment, step S2201+step S2203+step S2204+step S2205+step S2206+step S2207+step S2208+step S2209 may be implemented as an independent embodiment, but the disclosure is not limited thereto.

[0210] In some embodiments, step S2203 and step S2204 may be executed in an interchangeable order or simultaneously.

[0211] In some embodiments, step S2206, step S2207, step S2208, and step S2209 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0212] In some embodiments, step S2202, step S2206, step S2207, step S2208, and step S2209 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0213] In some embodiments, step S2208 and step S2209 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0214] In some embodiments, step S2202, step S2208, and step S2209 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0215] In some embodiments, step S2201, step S2202, step S2203, step S2204, step S2205, step S2206, and step S2207 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0216] In some embodiments, step S2202 is optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0217] In some embodiments, reference may be made to other optional implementations before or after the description corresponding to FIG. 2B.

[0218] FIG. 2C is a schematic diagram illustrating an interaction of a method for processing a signal according to an embodiment of the disclosure. As illustrated in FIG. 2C, the method for processing a signal according to an embodiment of the disclosure may be performed by a communication system 100, and the method includes but is not limited to the following.

[0219] At step S2301, the encoding-end device 101 obtains a combined format audio signal.

[0220] For optional implementations of step S2301, reference may be made to optional implementations of step S2101 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0221] At step S2302, the encoding-end device 101 performs filtering processing on the combined format audio signal.

[0222] For optional implementations of step S2302, reference may be made to optional implementations of step S2102 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0223] At step S2303, the encoding-end device 101 determines a sound field analysis parameter and / or a control parameter of the combined format audio signal.

[0224] In some embodiments, the encoding-end device 101 determines the sound field analysis parameter of the combined format audio signal. For optional implementations thereof, reference may be made to optional implementations of step S2103 of FIG. 2A and to other related parts of embodiments involved in FIG. 2A, which are not repeated here.

[0225] In some embodiments, the encoding-end device 101 determines the control parameter of the combined format audio signal. For optional implementations thereof, reference may be made to optional implementations of step S2203 in FIG. 2B and to other related parts in the embodiment involved in FIG. 2B, which are not described in detail here.

[0226] In some embodiments, the encoding-end device 101 determines the sound field analysis parameter and the control parameter of the combined format audio signal. For optional implementations thereof, reference may be made to optional implementations of step S2103 of FIG. 2A, to optional implementations of step S2203 of FIG. 2B, to other related parts in embodiments involved in FIG. 2A, and to other related parts in embodiments involved in FIG. 2B, which are not repeated here.

[0227] At step S2304, the encoding-end device 101 performs bit allocation for multiple sound channels based on the sound field analysis parameter and / or the control parameter of the combined format audio signal to obtain the bit allocation parameters.

[0228] In some embodiments, the encoding-end device 101 performs the bit allocation for the above-mentioned multiple sound channels using the available bit budget based on the sound field analysis parameter and / or the control parameter of the combined format audio signal, to obtain the bit allocation parameters.

[0229] In some embodiments, the encoding-end device 101 performs the bit allocation for the multiple sound channels based on the sound field analysis parameter of the combined format audio signal to obtain the bit allocation parameters.

[0230] In an implementation, taking the combined format audio signal including the object-based audio signal and the scene-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and the scene-based audio signal based on the sound field analysis parameter of the combined format audio to obtain the bit allocation parameters. For example, the object-based audio signal is a main sound element in the sound field, and the scene-based audio signal is an ambient background sound element in the sound field. To reconstruct an expected sound field, the object-based audio signal is required to have a smaller distortion, while the scene-based audio signal used as a background sound is allowed to have a certain degree of distortion. Therefore, in allocating bits, more bits may be allocated to the object-based audio signal than to the scene-based audio signal. For example, the number of bits allocated to any sound channel in the object-based audio signal is more than the number of bits allocated to any sound channel in the scene-based audio signal.

[0231] In another implementation, taking the combined format audio signal including the object-based audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and of the sound channel-based audio signal based on the sound field analysis parameter of the combined format audio signal to obtain the bit allocation parameters. For example, the object-based audio signal is a main sound element in the sound field, and the sound channel-based audio signal is an ambient background sound element in the sound field. To reconstruct an expected sound field, the object-based audio signal is required to have a smaller distortion, while the sound channel-based audio signal used as a background sound is allowed to have a certain degree of distortion. Therefore, in allocating bits, more bits may be allocated to the object-based audio signal than to the sound channel-based audio signal. For example, the number of bits allocated to any sound channel of the object-based audio signal is more than the number of bits allocated to any sound channel of the sound channel-based audio signal.

[0232] In another implementation, taking the combined format audio signal including the scene-based audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may allocate bits to the multiple sound channels of the combined format audio signal composed of the scene-based audio signal and the sound channel-based audio signal based on the sound field analysis parameter of the combined format audio signal to obtain the bit allocation parameters. For example, the sound channel-based audio signal is a main sound element in the sound field, and the scene-based audio signal is an ambient background sound element in the sound field. To reconstruct an expected sound field, the sound channel-based audio signal is required to have a smaller distortion, while the scene-based audio signal used as a background sound is allowed to have a certain degree of distortion, and in allocating bits, more bits may be allocated to the sound channel-based audio signal than to the scene-based audio signal. For example, the number of bits allocated to any sound channel of the sound channel-based audio signal is more than the number of bits allocated to any sound channel of the scene-based audio signal.

[0233] In another implementation, taking the combined format audio signal including the object-based audio signal and the metadata-based three-dimensional audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and the metadata-based three-dimensional audio signal based on the sound field analysis parameter of the combined format audio signal to obtain the bit allocation parameters. For example, the object-based audio signal is a main sound element in the sound field, and the metadata-based three-dimensional audio signal is an ambient background sound element in the sound field. To reconstruct an expected sound field, the object-based audio signal is required to have a smaller distortion, while the metadata-based three-dimensional audio signal used as a background sound is allowed to have a certain degree of distortion. Therefore, in allocating bits, more bits may be allocated to the object-based audio signal than to the metadata-based three-dimensional audio signal. For example, the number of bits allocated to any sound channel of the object-based audio signal is more than the number of bits allocated to any sound channel of the metadata-based three-dimensional audio signal.

[0234] In another implementation, taking the combined format audio signal including the metadata-based three-dimensional audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the metadata-based three-dimensional audio signal and the sound channel-based audio signal based on the sound field analysis parameter of the combined format audio signal to obtain the bit allocation parameters. For example, the sound channel-based audio signal is a main sound element in the sound field, and the metadata-based three-dimensional audio signal is an ambient background sound element in the sound field. To reconstruct an expected sound field, the sound channel-based audio signal is required to have a smaller distortion, while the metadata-based three-dimensional audio signal used as a background sound is allowed to have a certain degree of distortion, and in allocating bits, more bits may be allocated to the sound channel-based audio signal than to the metadata-based three-dimensional audio signal. For example, the number of bits allocated to any sound channel of the sound channel-based audio signal is more than the number of bits allocated to any sound channel of the metadata-based three-dimensional audio signal.

[0235] In some embodiments, the number of bits allocated to a sound channel corresponding to a primary sound image of the object-based audio signal is greater than the number of bits allocated to a sound channel corresponding to a secondary sound image of the object-based audio signal. In other words, in the object-based audio signal, each object signal has a different importance in terms of its part in the combined format audio signal. For example, at a concert, the importance of a lead singer's voice signal is higher than that of a backing singer's voice signal, or a lead singer's voice signal is more important than that of a backing singer's voice signal, and the lead singer's object audio signal are allocated more bits than the backing singer's voice signal.

[0236] It should be noted that the composition of the combined format audio signal shown in the above embodiments is for the convenience of those skilled in the field to understand how to allocate bits to each sound channel to be encoded based on the sound field analysis parameter. In other words, the bit allocation may be carried out for the combined format audio signal composed of the audio signals in other formats based on the above-mentioned sound field analysis parameter, which are not repeated here.

[0237] In some embodiments, the encoding-end device 101 may allocate bits to multiple sound channels based on the control parameter of the combined format audio signal to obtain the bit allocation parameters.

[0238] In an implementation, taking the above-mentioned combined format audio signal including the object-based audio signal and the scene-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and the scene-based audio signal based on the control parameter of the combined format audio signal to obtain the bit allocation parameters.

[0239] In another implementation, taking the above-mentioned combined format audio signal including the object-based audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and the sound channel-based audio signal based on the control parameter of the combined format audio signal to obtain the bit allocation parameters.

[0240] In another implementation, taking the above-mentioned combined format audio signal including the scene-based audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the scene-based audio signal and the sound channel-based audio signal based on the control parameter of the combined format audio signal to obtain the bit allocation parameters.

[0241] In another implementation, taking the above-mentioned combined format audio signal including the object-based audio signal and the metadata-based three-dimensional audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and the metadata-based three-dimensional audio signal based on the control parameter of the combined format audio signal to obtain the bit allocation parameters.

[0242] In another implementation, taking the above-mentioned combined format audio signal including the metadata-based three-dimensional audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the metadata-based three-dimensional audio signal and the sound channel-based audio signal based on the control parameter of the combined format audio signal to obtain the bit allocation parameters.

[0243] It should be noted that the composition of the combined format audio signal shown in the above embodiments is for the convenience of those skilled in the field to understand how to allocate bits to each sound channel to be encoded based on the energy parameters. In other words, the bit allocation may be carried out for the combined format audio signal composed of the audio signals in other formats based on the above-mentioned energy parameters, which are not repeated here.

[0244] In some embodiments, the encoding-end device 101 performs the bit allocation for the multiple sound channels based on the sound field analysis parameter and the control parameter of the combined format audio signal to obtain the bit allocation parameters.

[0245] In an implementation, taking the above-mentioned combined format audio signal including the object-based audio signal and the scene-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and the scene-based audio signal based on the sound field analysis parameter and the control parameter of the combined format audio signal to obtain the bit allocation parameters.

[0246] In another implementation, taking the above-mentioned combined format audio signal including the object-based audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and the sound channel-based audio signal based on the sound field analysis parameter and the control parameter of the combined format audio signal to obtain the bit allocation parameters.

[0247] In another implementation, taking the above-mentioned combined format audio signal including the scene-based audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the scene-based audio signal and the sound channel-based audio signal based on the sound field analysis parameter and the control parameter of the combined format audio signal to obtain the bit allocation parameters.

[0248] In another implementation, taking the above-mentioned combined format audio signal including the object-based audio signal and the metadata-based three-dimensional audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and the metadata-based three-dimensional audio signal based on the sound field analysis parameter and the control parameter of the combined format audio signal to obtain the bit allocation parameters.

[0249] In another implementation, taking the above-mentioned combined format audio signal including the metadata-based three-dimensional audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the metadata-based three-dimensional audio signal and of the sound channel-based audio signal based on the sound field analysis parameter and the control parameter of the combined format audio signal to obtain the bit allocation parameters.

[0250] In some embodiments, the number of bits allocated to a sound channel corresponding to a primary sound image of the object-based audio signal is greater than the number of bits allocated to a sound channel corresponding to a secondary sound image of the object-based audio signal. In other words, in the object-based audio signal, each object signal has a different importance in terms of its part in the combined format audio signal. For example, at a concert, the importance of a lead singer's voice signal is higher than that of a backing singer's voice signal, or a lead singer's voice signal is more important than that of a backing singer's voice signal, and the lead singer's object audio signal are allocated more bits than the backing singer's voice signal.

[0251] It should be noted that the composition of the combined format audio signal shown in the above embodiments is for the convenience of those skilled in the field to understand how to allocate bits to each sound channel to be encoded based on the sound field analysis parameter and the energy parameter(s). In other words, the bit allocation may be carried out for the combined format audio signal composed of the audio signals in other formats based on the above-mentioned sound field analysis parameter and the energy parameter(s), which are not repeated here.

[0252] At step S2305, the encoding-end device 101 encodes the multiple sound channels based on the bit allocation parameters to obtain sound channel signal encoding parameters, and writes the bit allocation parameters and the sound channel signal encoding parameters into a bitstream.

[0253] For optional implementations of step S2305, reference may be made to optional implementations of step S2106 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0254] At step S2306, the encoding-end device 101 sends the bitstream.

[0255] For optional implementations of step S2306, reference may be made to optional implementations of step S2107 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0256] At step S2307, the decoding-end device 102 parses the received bitstream to obtain bit allocation parameters and sound channel signal encoding parameters for multiple sound channels of a combined format audio signal respectively.

[0257] For optional implementations of step S2307, reference may be made to optional implementations of step S2108 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0258] At step S2308, the decoding-end device 102 reconstructs the combined format audio signal based on the bit allocation parameters and the sound channel signal encoding parameters for the multiple sound channels.

[0259] For optional implementations of step S2308, reference may be made to optional implementations of step S2109 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0260] The method involved in embodiments of the disclosure may include at least one of step S2301 to step S2308. For example, step S2301+step S2302+step S2303+step S2304 may be implemented as an independent embodiment, step S2301+step S2303+step S2304 may be implemented as an independent embodiment, step S2301+step S2302+step S2303+step S2304+step S2305+step S2306+step S2306 may be implemented as an independent embodiment, step S2301+step S2303+step S2304+step S2305+step S2306+step S2306 may be implemented as an independent embodiment, step S2307+step S2308 may be implemented as an independent embodiment, step S2301+step S2302+step S2303+step S2304+step S2305+step S2306+step S2307+step S2308 may be implemented as an independent embodiment, step S2301+step S2303+step S2304+step S2305+step S2306+step S2307+step S2308 may be implemented as an independent embodiment, but the disclosure is not limited thereto.

[0261] In some embodiments, step S2305, step S2306, step S2307, and step S2308 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0262] In some embodiments, step S2302, step S2305, step S2306, step S2307, and step S2308 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0263] In some embodiments, step S2307 and step S2308 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0264] In some embodiments, step S2302, step S2307, and step S2308 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0265] In some embodiments, step S2301, step S2302, step S2303, step S2304, step S2305, and step S2306 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0266] In some embodiments, step S2302 is optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0267] In some embodiments, reference may be made to other optional implementations before or after the description corresponding to FIG. 2C.

[0268] FIG. 2D is a schematic diagram illustrating an interaction of a method for processing a signal according to an embodiment of the disclosure. As illustrated in FIG. 2D, the method for processing a signal according to an embodiment of the disclosure may be performed by a communication system 100, and the method includes but is not limited to the following.

[0269] At step S2401, the encoding-end device 101 obtains a combined format audio signal.

[0270] For optional implementations of step S2401, reference may be made to optional implementations of step S2101 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0271] At step S2402, the encoding-end device 101 performs filtering processing on the combined format audio signal.

[0272] For optional implementations of step S2402, reference may be made to optional implementations of step S2102 in FIG. 2A and to other related parts in the embodiment involved in FIG. 2A, which are not described in detail here.

[0273] At step S2403, the encoding-end device 101 determines a sound field analysis parameter of the combined format audio signal.

[0274] For optional implementations of step S2403, reference may be made to optional implementations of step S2103 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0275] At step S2404, the encoding-end device 101 determines energy parameter(s) of one or more sound channels among the multiple sound channels.

[0276] For optional implementations of step S2404, reference may be made to optional implementations of step S2104 in FIG. 2A and to other related parts of embodiments involved in FIG. 2A, which are not described in detail here.

[0277] At step S2405, the encoding-end device 101 determines a control parameter of the combined format audio signal.

[0278] For optional implementations of step S2405, reference may be made to optional implementations of step S2203 in FIG. 2B and to other related parts in embodiments involved in FIG. 2B, which are not described in detail here.

[0279] At step S2406, the encoding-end device 101 performs the bit allocation for the multiple sound channels based on the energy parameter(s) of the one or more sound channels, the control parameter of the combined format audio signal, and the sound field analysis parameter of the combined format audio signal, to obtain bit allocation parameters.

[0280] In some embodiments, the encoding-end device 101 may allocate bits to the multiple sound channels using the available bit budget based on the control parameter of the combined format audio signal, the sound field analysis parameter of the combined format audio signal, and the energy parameter of each of the multiple sound channels, to obtain the bit allocation parameters.

[0281] In some embodiments, taking the above-mentioned combined format audio signal including the object-based audio signal and the scene-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and the scene-based audio signal based on the sound field analysis parameter, the control parameter, and the energy parameter of each sound channel of the combined format audio signal to obtain the bit allocation parameters.

[0282] In some embodiments, taking the above-mentioned combined format audio signal including the object-based audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and the sound channel-based audio signal based on the sound field analysis parameter, the control parameter, and the energy parameter of each sound channel of the combined format audio signal to obtain the bit allocation parameters.

[0283] In some embodiments, taking the above-mentioned combined format audio signal including the scene-based audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the scene-based audio signal and the sound channel-based audio signal based on the sound field analysis parameter, the control parameter, and the energy parameter of each of the multiple sound channels of the combined format audio signal to obtain the bit allocation parameters.

[0284] In some embodiments, taking the above-mentioned combined format audio signal including the object-based audio signal and the metadata-based three-dimensional audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the object-based audio signal and the metadata-based three-dimensional audio signal based on the sound field analysis parameter, the control parameter, and the energy parameter of each sound channel of the combined format audio signal to obtain the bit allocation parameters.

[0285] In some embodiments, taking the above-mentioned combined format audio signal including the metadata-based three-dimensional audio signal and the sound channel-based audio signal as an example, the encoding-end device 101 may perform the bit allocation for the multiple sound channels of the combined format audio signal composed of the metadata-based three-dimensional audio signal and the sound channel-based audio signal based on the sound field analysis parameter, the control parameter, and the energy parameter of each of the multiple sound channels of the combined format audio signal to obtain the bit allocation parameters.

[0286] In some embodiments, the number of bits allocated to a sound channel corresponding to a primary sound image of the object-based audio signal is greater than the number of bits allocated to a sound channel corresponding to a secondary sound image of the object-based audio signal. In other words, in the object-based audio signal, each object signal has a different importance in terms of its part in the combined format audio signal. For example, at a concert, the importance of a lead singer's voice signal is higher than that of a backing singer's voice signal, or a lead singer's voice signal is more important than that of a backing singer's voice signal, and the lead singer's object audio signal are allocated more bits than the backing singer's voice signal.

[0287] It should be noted that the composition of the combined format audio signal shown in the above embodiments is for the convenience of those skilled in the field to understand how to allocate bits to each sound channel to be encoded based on the sound field analysis parameter, the control parameter, and the energy parameter(s). In other words, the bit allocation may be carried out for the combined format audio signal composed of the audio signals in other formats based on the above-mentioned sound field analysis parameter, the control parameter, and the energy parameters, which are not repeated here.

[0288] In some embodiments, the number of bits of each sound channel of the audio signals in different formats of the above-mentioned combined format audio signal may be the same (the bits are evenly allocated), or the bits may be allocated based on other factors (such as a factor of whether it is primary or secondary). The above is not limited in the disclosure and is not repeated here.

[0289] At step S2407, the encoding-end device 101 encodes the multiple sound channels based on the bit allocation parameters to obtain sound channel signal encoding parameters, and writes the bit allocation parameters and the sound channel signal encoding parameters into a bitstream.

[0290] For optional implementations of step S2407, reference may be made to optional implementations of step S2106 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0291] At step S2408, the encoding-end device 101 sends the bitstream.

[0292] For optional implementations of step S2408, reference may be made to optional implementations of step S2107 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0293] At step S2409, the decoding-end device 102 parses the received bitstream to obtain bit allocation parameters and sound channel signal encoding parameters for multiple sound channels of a combined format audio signal.

[0294] For optional implementations of step S2409, reference may be made to optional implementations of step S2108 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0295] At step S2410, the decoding-end device 102 reconstructs the combined format audio signal based on the bit allocation parameters and the sound channel signal encoding parameters for the multiple sound channels.

[0296] For optional implementations of step S2410, reference may be made to optional implementations of step S2109 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0297] The method involved in embodiments of the disclosure may include at least one of step S2401 to step S2410. For example, step S2401+step S2402+step S2403+step S2404+step S2405+step S2406 may be implemented as an independent embodiment, step S2401+step S2403+step S2404+step S2405+step S2406 may be implemented as an independent embodiment, step S2401+step S2402+step S2403+step S2404+step S2405+step S2406+step S2407+step S2408 may be implemented as an independent embodiment, step S2401+step S2403+step S2404+step S2405+step S2406+step S2407+step S2408 may be implemented as an independent embodiment, step S2401+step S2403+step S2404+step S2405+step S2406+step S2407+step S2408 may be implemented as an independent embodiment, step S2409+step S2410 may be implemented as an independent embodiment, step S2401+step S2402+step S2403+step S2404+step S2405+step S2406+step S2407+step S2408+step S2409+step S2410 may be implemented as an independent embodiment, step S2401+step S2403+step S2404+step S2405+step S2406+step S2407+step S2408+step S2409+step S2410 may be implemented as an independent embodiment, but the disclosure is not limited thereto.

[0298] In some embodiments, step S2403, step S2404, and step S2405 may be executed in an interchangeable order or simultaneously.

[0299] In some embodiments, step S2407, step S2408, step S2409, and step S2410 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0300] In some embodiments, step S2402, step S2407, step S2408, step S2409, and step S2410 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0301] In some embodiments, step S2409 and step S2410 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0302] In some embodiments, step S2402, step S2409, and step S2410 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0303] In some embodiments, step S2401, step S2402, step S2403, step S2404, step S2405, step S2406, step S2407, and step S2408 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0304] In some embodiments, step S2402 is optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0305] In some embodiments, reference may be made to other optional implementations before or after the description corresponding to FIG. 2D.

[0306] FIG. 3 is a flow chart illustrating a method for processing a signal according to an embodiment of the disclosure. As illustrated in FIG. 3, embodiments of the disclosure relate to a method for processing a signal, which may be performed by the encoding-end device 101, and the method may include but is not limited to the following.

[0307] At step S3101, a combined format audio signal is obtained.

[0308] In some embodiments, the combined format audio signal is a multi-channel audio signal including multiple sound channels.

[0309] For optional implementations of step S3101, reference may be made to optional implementations of step S2101 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0310] At step S3102, filtering processing is performed on the combined format audio signal.

[0311] For optional implementations of step S3102, reference may be made to optional implementations of step S2102 in FIG. 2A and to other related parts in implementation examples involved in FIG. 2A, which are not repeated here.

[0312] At step S3103, a first parameter of the combined format audio signal and / or energy parameter(s) of one or more sound channels among the multiple sound channels are determined.

[0313] In some embodiments, the first parameter may include at least one of: a sound field analysis parameter or a control parameter. In an implementation, the first parameter may be (or may include) a sound field analysis parameter. In another implementation, the first parameter may be (or may include) a control parameter. In yet another implementation, the first parameter may be (or may include) a sound field analysis parameter and a control parameter.

[0314] In some embodiments, the above-mentioned sound field analysis parameter include at least one of: the total number of sound images, relative priority of sound images, or a background sound. In an implementation, the above-mentioned sound field analysis parameter includes any one of the total number of sound images, the relative priority of sound images, or the background sound. In another implementation, the above-mentioned sound field analysis parameter include any two of the total number of sound images, the relative priority of sound images, or the background sound. In yet another implementation, the above-mentioned sound field analysis parameter include the total number of sound images, the relative priority of sound images, and the background sound.

[0315] In some embodiments, the control parameter may include a second parameter and / or a third parameter. In some embodiments, the second parameter is used for describing an importance ranking of the multiple sound channels in the combined format audio signal. For example, the importance may be set manually or may be obtained by other processing other than encoding processing.

[0316] In some embodiments, the third parameter is used for describing whether each of the sound channels is a diegetic sound or a non-diegetic sound.

[0317] In some embodiments, a first parameter of the combined format audio signal may be determined.

[0318] For example, the sound field analysis parameter of the combined format audio signal may be determined. For optional implementations thereof, reference may be made to optional implementations of step S2103 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0319] For example, the control parameter of the combined format audio signal may be determined. For optional implementations thereof, reference may be made to optional implementations of step S2203 in FIG. 2B and to other related parts in embodiments involved in FIG. 2B, which are not described in detail here.

[0320] For example, the sound field analysis parameter and the control parameter of the combined format audio signal may be determined. For optional implementations thereof, reference may be made to optional implementations of step S2103 of FIG. 2A, to optional implementations of step S2203 of FIG. 2B, to other related parts in embodiments involved in FIG. 2A, and to other related parts in embodiments involved in FIG. 2B, which are not repeated here.

[0321] In some embodiments, the energy parameter(s) of one or more of the multiple sound channels may be determined. For optional implementations thereof, reference may be made to optional implementations of step S2104 in FIG. 2A and to other related parts of embodiments involved in FIG. 2A, which are not described in detail here.

[0322] In some embodiments, a first parameter of the combined format audio signal and the energy parameter(s) of one or more sound channels of the plurality of sound channels may be determined.

[0323] For example, the sound field analysis parameter of the combined format audio signal and the energy parameter(s) of one or more of the multiple sound channels may be determined. For optional implementations thereof, reference may be made to optional implementations of step S2103 of FIG. 2A, to optional implementations of step S2104 of FIG. 2A, and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0324] For example, the control parameter of the combined format audio signal and the energy parameter(s) of one or more of the multiple sound channels may be determined. For optional implementations thereof, reference may be made to optional implementations of step S2203 of FIG. 2B, to optional implementations of step S2104 of FIG. 2A, to other related parts in embodiments involved in FIG. 2B, and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0325] For example, the sound field analysis parameter, the control parameter, and the energy parameter(s) of one or more of the above-mentioned multiple sound channels of the combined format audio signal may be determined. For optional implementations thereof, reference may be made to optional implementations of step S2103 of FIG. 2A, to optional implementations of step S2104, to optional implementations of step S2203 of FIG. 2B, to other related parts in embodiments involved in FIG. 2A, and to other related parts in embodiments involved in FIG. 2B, which are not repeated here.

[0326] At step S3104, bit allocation is performed for the multiple sound channels based on the first parameter of the combined format audio signal and / or the energy parameter(s) of one or more sound channels among the multiple sound channels to obtain bit allocation parameters.

[0327] In some embodiments, the bit allocation parameters may be obtained by performing the bit allocation for the multiple sound channels based on the first parameter of the combined format audio signal.

[0328] For example, the bit allocation parameters may be obtained by performing the bit allocation for the multiple sound channels based on the sound field analysis parameter of the combined format audio signal. For optional implementations thereof, reference may be made to optional implementations of step S2304 of FIG. 2C and to other related parts of embodiments involved in FIG. 2C, which are not described in detail here.

[0329] For example, the bit allocation parameters may be obtained by performing the bit allocation for the multiple sound channels based on the control parameter of the combined format audio signal. For optional implementations thereof, reference may be made to optional implementations of step S2304 of FIG. 2C and to other related parts of embodiments involved in FIG. 2C, which are not described in detail here.

[0330] For example, the bit allocation parameters may be obtained by performing the bit allocation for the multiple sound channels based on the sound field analysis parameter and the control parameter of the combined format audio signal. For optional implementations thereof, reference may be made to optional implementations of step S2304 of FIG. 2C and to other related parts of embodiments involved in FIG. 2C, which are not described in detail here.

[0331] In some embodiments, the bit allocation may be performed for the multiple sound channels based on the energy parameter(s) of one or more sound channels among the multiple sound channels of the combined format audio signal to obtain the bit allocation parameters.

[0332] In some embodiments, the bit allocation parameters may be obtained by performing the bit allocation for the multiple sound channels based on the first parameter of the combined format audio signal and the energy parameter(s) of one or more sound channels among the multiple sound channels.

[0333] For example, the bit allocation parameters may be obtained by performing the bit allocation for the multiple sound channels based on the sound field analysis parameter of the combined format audio signal and the energy parameter(s) of one or more sound channels among the multiple sound channels. For optional implementations thereof, reference may be made to optional implementations of step S2105 of FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0334] For example, the bit allocation parameters may be obtained by performing the bit allocation for the multiple sound channels based on the control parameter of the combined format audio signal and the energy parameter(s) of one or more sound channels among the multiple sound channels. For optional implementations thereof, reference may be made to optional implementations of step S2205 of FIG. 2B and to other related parts in embodiments involved in FIG. 2B, which are not described in detail here.

[0335] For example, the bit allocation parameters may be obtained by performing the bit allocation for the multiple sound channels based on the sound field analysis parameter, the control parameter, and the energy parameter(s) of one or more sound channels among the multiple sound channels of the combined format audio signal. For optional implementations thereof, reference may be made to optional implementations of step S2406 of FIG. 2D and to other related parts of embodiments involved in FIG. 2D, which are not described in detail here.

[0336] At step S3105, the multiple sound channels are encoded based on the bit allocation parameters to obtain the sound channel signal encoding parameters, and the bit allocation parameters and the sound channel signal encoding parameters are written into a bitstream.

[0337] For optional implementations of step S3105, reference may be made to optional implementations of step S2106 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0338] At step S3106, the bitstream is sent.

[0339] For optional implementations of step S3106, reference may be made to optional implementations of step S2107 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0340] The method involved in embodiments of the disclosure may include at least one of step S3101 to step S3106. For example, step S3101+step S3102+step S3103+step S3104 may be implemented as an independent embodiment, step S3101+step S3103+step S3104 may be implemented as an independent embodiment, step S3101+step S3102+step S3103+step S3104+step S3105+step S3106 may be implemented as an independent embodiment, and step S3101+step S3103+step S3104+step S3105+step S3106 may be implemented as an independent embodiment, but the disclosure is not limited thereto.

[0341] In some embodiments, step S3105 and step S3106 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0342] In some embodiments, step S3102, step S3105, and step S3106 are optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0343] In some embodiments, step S3102 is optional, and one or more of these steps may be omitted or replaced in different embodiments.

[0344] FIG. 4 is a flowchart illustrating a method for processing a signal according to an embodiment of the disclosure. As illustrated in FIG. 4, embodiments of the disclosure relate to a method for processing a signal, which may be executed by a decoding-end device 102, and the method may include but is not limited to the following.

[0345] At step S4101, a bitstream is received.

[0346] In some embodiments, the decoding-end device 102 receives the bitstream. For example, the decoding-end device 102 receives the bitstream sent by the encoding-end device 101. The above bitstream is used for reconstructing the combined format audio signal by the decoding-end device 102. For optional implementations of the above bitstream, reference may be made to optional implementations of any of the above embodiments, and to other related parts of any of the above embodiments, which are not repeated here.

[0347] At step S4102, the bitstream is parsed to obtain bit allocation parameters and sound channel signal encoding parameters for multiple sound channels of a combined format audio signal.

[0348] For optional implementations of step S4102, reference may be made to optional implementations of step S2108 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0349] At step S4103, the combined format audio signal is reconstructed based on the bit allocation parameters and the sound channel signal encoding parameters for the multiple sound channels.

[0350] For optional implementations of step S4103, reference may be made to optional implementations of step S2109 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0351] FIG. 5 is a schematic diagram illustrating a method for processing a signal according to an embodiment of the disclosure. As illustrated in FIG. 5, the method involved in embodiments of the disclosure may be performed by a communication system 100, and the method includes but is not limited to the following.

[0352] At step S5101, the encoding-end device 101 obtains a combined format audio signal, where the combined format audio signal is a multi-channel audio signal including multiple sound channels.

[0353] For optional implementations of step S5101, reference may be made to optional implementations of step S2101 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0354] At step S5102, the encoding-end device 101 determines a first parameter of the combined format audio signal and / or energy parameter(s) of one or more sound channels among the multiple sound channels.

[0355] For optional implementations of step S5102, reference may be made to optional implementations of step S2103 in FIG. 2A, of step S2203 in FIG. 2B, of step S2104 in FIG. 2A, and of step S3103 in FIG. 3, and to other related parts in embodiments involved in FIGS. 2A, 2B, and 3, which are not repeated here.

[0356] At step S5103, the encoding-end device 101 performs the bit allocation for the multiple sound channels based on the first parameter of the combined format audio signal and / or the energy parameter(s) of one or more sound channels among the multiple sound channels to obtain bit allocation parameters.

[0357] For optional implementations of step S5103, reference may be made to optional implementations of step S2105 in FIG. 2A, of step S2205 in FIG. 2B, of step S2304 in FIG. 2C, of step S2406 in FIG. 2D, of step S3104 in FIG. 3, and to other related parts in embodiments involved in FIGS. 2A, 2B, 2C, 2D, and 3, which are not repeated here.

[0358] At step S5104, the encoding-end device 101 encodes the multiple sound channels based on the bit allocation parameters to obtain sound channel signal encoding parameters, and writes the bit allocation parameters and the sound channel signal encoding parameters into a bitstream.

[0359] For optional implementations of step S5104, reference may be made to optional implementations of step S2106 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0360] At step S5105, the encoding-end device 101 sends the bitstream.

[0361] For optional implementations of step S5105, reference may be made to optional implementations of step S2107 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0362] At step S5106, the decoding-end device 102 parses the received bitstream to obtain bit allocation parameters and sound channel signal encoding parameters for multiple sound channels of a combined format audio signal.

[0363] For optional implementations of step S5106, reference may be made to optional implementations of step S2108 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0364] At step S5107, the decoding-end device 102 reconstructs the combined format audio signal based on the bit allocation parameters and the sound channel signal encoding parameters for the multiple sound channels.

[0365] For optional implementations of step S5107, reference may be made to optional implementations of step S2109 in FIG. 2A and to other related parts in embodiments involved in FIG. 2A, which are not described in detail here.

[0366] In some embodiments, the encoding and decoding processing of the combined format audio signal disclosed herein is as follows. When the combined format audio signal is inputted to the encoder, the encoder performs the energy calculation and the sound field analysis on the combined format audio signal, the encoder allocates bits to each encoded sound channel signal based on the energy calculation and sound field analysis results, the encoder encodes each inputted audio sound channel signal using the allocated bits to obtain the bitstream, and the decoding-end decodes the bitstream to reconstruct the combined format audio signal.

[0367] In some embodiments, the encoding and decoding processing of the combined format audio signal disclosed herein is as follows. When the combined format audio signal is inputted to the encoder, the encoder performs the energy calculation on the combined format audio signal, allocates bits to each encoded sound channel signal based on control parameter and the energy calculation result, encodes the audio signal using the allocated bits to obtain a bitstream, and the decoding end decodes the bitstream to reconstruct the combined format audio signal.

[0368] The purpose of the disclosure is to design and complete the encoding and decoding processing of the combined format audio signal, determine the encoding mode based on the external control parameter and the sound field analysis result, and perform encoding, to achieve the reconstruction of the desired combined format audio signal at the decoding end.

[0369] The combined format audio signal inputted to the encoder in the disclosure includes any combination of audio signals in the following four formats, namely, sound channel-based audio signal, object-based audio signal, scene-based audio signal, or metadata-based three-dimensional audio signal.

[0370] All combined format audio signals inputted are subjected to high-pass filtering, and the cutoff frequency of the filter may be set to 20 Hz. As an example, the filter formula used may be shown in the following formula (1): H 20 Z = b 0 + b 1 z − 1 + b 2 z − 2 1 + a 1 z − 1 + a 2 z − 2

[0371] As illustrated in FIG. 6A, the audio signal processed with the high-pass filtering is subjected to energy calculation and sound field analysis.

[0372] The energy calculation is to calculate the energy value of each sound channel. The formula of calculating the square sum of a sound channel is shown as follows: S m = 1 N ∑ m = 0 N − 1 X m ∗ X m where S m represents the sum of levels of the sample point of a frame of an m th< sound channel, a frame of the m th< sound channel includes N sample points, and X m represents the m th< sound channel.

[0373] The sound field analysis refers to analyzing the total number of sound images in combined format audio signal, the relative priority of sound images, or a background sound. The example is as follows.

[0374] Case 1, the combined format audio signal is composed of object-based audio signal and scene-based audio signal. The object-based audio signal is a main sound element in the sound field, and the scene-based audio signal is an ambient background sound element in the sound field. To reconstruct an expected sound field, the object-based audio signal is required to have a smaller distortion, while the scene-based audio signal used as a background sound is allowed to have a certain degree of distortion. Therefore, in allocating bits, more bits are allocated to the object-based audio signal than to the scene-based audio signal.

[0375] In the object-based audio, each object signal has a different importance in terms of its part in the combined format audio signal. For example, at a concert, the importance of a lead singer's voice signal is higher than that of a backing singer's voice signal, or a lead singer's voice signal is more important than that of a backing singer's voice signal, and the lead singer's object audio signal are allocated more bits than the backing singer's voice signal.

[0376] Case 2, the combined format audio signal is composed of an object-based audio signal and a sound channel-based audio signal. The object-based audio signal is a main sound element in the sound field, and the sound channel-based audio signal is an ambient background sound element in the sound field. To reconstruct an expected sound field, the object-based audio signal is required to have a smaller distortion, while the sound channel-based audio signal used as a background sound is allowed to have a certain degree of distortion. Therefore, in allocating bits, more bits are allocated to the object-based audio signal than to the sound channel-based audio signal.

[0377] Case 3, the combined format audio signal is composed of a sound channel-based audio signal and a scene-based audio signal. The sound channel-based audio signal is a main sound element in the sound field, and the scene-based audio signal is an ambient background sound element in the sound field. To reconstruct an expected sound field, the sound channel-based audio signal is required to have a smaller distortion, while the scene-based audio signal used as a background sound is allowed to have a certain degree of distortion. Therefore, in allocating bits, more bits are allocated to the sound channel-based audio signal than to the scene-based audio signal.

[0378] Case 4, in the above three cases, the bit allocation of each sound channel is performed based on the result of energy calculation of each sound channel.

[0379] After obtaining the sound field analysis parameter of the combined format audio signal and the energy parameter of each sound channel, the corresponding encoding core may be used to allocate bits to multiple sound channels based on the sound field analysis parameter of the combined format audio signal and the energy parameter of each sound channel to obtain the bit allocation parameters. Multiple sound channels are encoded based on the bit allocation parameters to obtain sound channel signal encoding parameters, and the bit allocation parameters and sound channel signal encoding parameters are written into the bitstream. The encoding-end sends the bitstream to the decoding end. The decoding end decodes the received bitstream to reconstruct the above-mentioned combined format audio signal.

[0380] In some embodiments, after high-pass filtering is performed on all input combined format audio signal, as illustrated in FIG. 6B, the energy calculation and the sound field analysis may be performed on the audio signal that has been processed with the high-pass filtering, and the external control parameter need to be considered when allocating bits to sound channels. An example case is as follows. Case 5, in the aforementioned three cases (such as the above-mentioned case 1, the above-mentioned case 2, and the above-mentioned case 3), the bit allocation of each sound channel is performed in combination with the external control parameter and the result of the energy calculation of each sound channel.

[0381] For example, assuming that the combined format audio signal contains a 5.1 format multi-channel signal and four object audio signals, the combined format audio signal is subjected to high-pass filtering, the cross-correlation coefficients between the sound channels are calculated, the sound channel combination operation is performed based on the cross-correlation coefficients, the inter-channel parameters are calculated for the sound channels in the same sound channel combination, the inter-channel parameters may be ILD, ITD, IPD, etc., the sound channels are aligned using the inter-channel parameters, the downmixing is performed, and the energy of the downmixed sound channels is calculated. The energy calculation may be the sum of levels of sample points, the sum of squared levels of sample points, the root mean square of levesl of sample points, or others.

[0382] The sound field analysis is performed on the combined format audio signal. The sound field analysis is to calculate the total number of sound images, the relative priority of sound images, a background sound, or the like. The DOA estimation may be performed through sound source localization estimation, including azimuth and pitch angles. The multi-channel cross-correlation coefficient (MCCC) method may be used to solve a sound channel as a more important sound channel.

[0383] The control parameter of the combined format audio signal is determined. The control parameter refer to the importance ranking of sound channels in the combined format audio signal, that is determined in advance based on external information. This ranking may be determined based on the functional category of each sound channel determined in acquiring the combined format audio signal, or based on the setting required for reconstructing a desired sound field performed by the playback end (for example, in case 1, the voice of the lead singer in the concert may be set as the most important sound channel; in case 2, in order to focus on the sound of a certain instrument (such as the sound of a violin) in the reconstructed sound field, the sound of the violin may be set as the most important sound channel signal).

[0384] After obtaining the control parameter, the sound field analysis parameter, and the energy parameter of each sound channel of the combined format audio signal, the corresponding encoding core may be used to perform the bit allocation for the multiple sound channels based on the sound field analysis parameter, the control parameter, and the energy parameter of each sound channel of the combined format audio signal to obtain the bit allocation parameters. Multiple sound channels are encoded based on the bit allocation parameters to obtain the sound channel signal encoding parameters, and the bit allocation parameters and the sound channel signal encoding parameters are written into the bitstream. The encoding end sends the bitstream to the decoding end. The decoding end decodes the received bitstream to reconstruct the above-mentioned combined format audio signal.

[0385] Embodiments of the disclosure also provide apparatuses for implementing any of the above methods. As an example, an apparatus is provided. The above apparatus includes a unit or module for implementing each step performed by the encoding-end device in any of above methods. As another example, another apparatus is also provided. The apparatus includes a unit or module for implementing each step performed by the decoding-end device in any of above methods.

[0386] It should be understood that the division of the units or modules in above apparatuses is only a division of logical functions, and in actual implementation, they may be fully or partially integrated into one physical entity, or they may be physically separated. In addition, the units or modules in apparatuses may be implemented in the form of a processor calling software. For example, the apparatus includes a processor. The processor is connected to a memory. The memory has instructions stored therein. The processor calls the instructions stored in the memory to implement any of above methods or implement functions of the units or modules of above apparatuses. The processor is, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory is a memory inside the apparatus or a memory outside the apparatus. Or the units or modules in the apparatus may be implemented in the form of hardware circuits, and the functions of some or all of the units or modules may be realized by designing the hardware circuits. The above hardware circuits may be understood as one or more processors. For example, in an implementation, the above hardware circuit is an application-specific integrated circuit (ASIC), and the functions of some or all of the above units or modules are realized by designing the logical relationship of the components in the circuit. For example, in another implementation, the above hardware circuit may be realized by a programmable logic device (PLD), taking a field programmable gate array (FPGA) as an example, which may include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by a configuration file, to realize the functions of some or all of the above units or modules. All units or modules of the above apparatuses may be realized in the form of a processor calling software, or in the form of a hardware circuit, or in part by a processor calling software and the rest by a hardware circuit.

[0387] In embodiments of the disclosure, the processor is a circuit with signal processing capability. In an implementation, the processor may be a circuit with instruction reading and running capability, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which may be understood as a microprocessor), a digital signal processor (DSP), or others. In another implementation, the processor may realize certain functions through the logical relationship of the hardware circuit, and the logical relationship of the above hardware circuit is fixed or reconfigurable, such as a hardware circuit implemented by a processor as an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the processor loads a configuration document to implement the process of hardware circuit configuration, which may be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. In addition, it may also be a hardware circuit designed for artificial intelligence, which may be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), etc.

[0388] FIG. 7A is a schematic diagram illustrating a structure of a decoding-end device according to an embodiment of the disclosure. As illustrated in FIG. 7A, the decoding-end device 7100 may include at least one of a transceiver module 7101 or a processing module 7102. In some embodiments, the processing module is configured to obtain a combined format audio signal. The combined format audio signal is a multi-channel audio signal including multiple sound channels. The processing module is also configured to determine a first parameter of the combined format audio signal and / or one or more energy parameters of one or more sound channels among the multiple sound channels. The processing module is also configured to allocate bits to the multiple sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of one or more sound channels among the multiple sound channels to obtain bit allocation parameters. The processing module is also configured to encode the multiple sound channels based on the bit allocation parameters to obtain sound channel signal encoding parameters and write the bit allocation parameters and the sound channel signal encoding parameters into a bitstream. The transceiver module is configured to send the bitstream. Or the above-mentioned transceiver module is configured to perform at least one of the communication steps such as sending and / or receiving performed by the encoding-end device 101 in any of above methods (for example, but is not limited to, step S2107, step S2207, step S2306, step S2408), which are not repeated here. Or the above-mentioned processing module is configured to execute at least one of other steps performed by the terminal 102 in any of above methods (such as, but is not limited to, step S2101, step S2102, step S2103, step S2104, step S2105, step S2106, step S2201, step S2202, step S2203, step S2204, step S2205, step S2206, step S2301, step S2302, step S2303, step S2304, step S2305, step S2401, step S2402, step S2403, step S2404, step S2405, step S2406, step S2407), which are not repeated here.

[0389] FIG. 7B is a schematic diagram illustrating a structure of a decoding-end device according to an embodiment of the disclosure. As illustrated in FIG. 7B, the decoding-end device 7200 may include at least one of a transceiver module 7201, a processing module 7202, or the like. In some embodiments, the above-mentioned transceiver module is configured to receive a bitstream. The processing module is configured to parse the bitstream to obtain bit allocation parameters and sound channel signal encoding parameters for multiple sound channels of a combined format audio signal. The processing module is also configured to reconstruct the combined format audio signal based on the bit allocation parameters and the sound channel signal encoding parameters for the multiple sound channels. Or the above-mentioned transceiver module is configured to perform at least one of the communication steps such as sending and / or receiving performed by the decoding-end device 7200 in any of above methods, which are not repeated here. Or the above-mentioned processing module is configured to perform at least one of other steps (such as, but is not limited to, step S2108, step S2109, step S2208, step S2209, step S2307, step S2308, step S2409, step S2410) performed by the decoding-end device 102 in any of above methods, which are not repeated here.

[0390] In some embodiments, the transceiver module may include a sending module and / or a receiving module, and the sending module and the receiving module may be separate or integrated. For example, the term "transceiver module" may be interchangeable with the term "transceiver".

[0391] In some embodiments, the processing module may be a module or include multiple submodules. For example, the multiple submodules respectively execute all or part of the steps required to be executed by the processing module. For example, the term "processing module" may be interchangeable with replaced with the term "processor".

[0392] FIG. 8A is a schematic diagram illustrating a structure of a communication device 8100 according to an embodiment of the disclosure. The communication device 8100 may be an encoding-end device, a decoding-end device, or a chip, chip system, or processor that supports the encoding-end device to implement any of above methods, or a chip, chip system, or processor that supports the decoding-end device to implement any of above methods. The communication device 8100 may be configured to implement the methods described in above method embodiments, and for the details, reference may be made to the description in above method embodiments.

[0393] As illustrated in FIG. 8A, the communication device 8100 includes one or more processors 8101. The one or more processors 8101 may be a general-purpose processor or a dedicated processor, for example, a baseband processor or a central processing unit. The baseband processor may be configured to process the communication protocol and the communication data, and the central processing unit may be configured to control the communication device (such as a base station, a baseband chip, a terminal device, a terminal device chip, a DU or a CU, etc.), execute a program, and process the data of the program. The communication device 8100 is configured to perform any of the above methods.

[0394] In some embodiments, the communication device 8100 further includes one or more memories 8102 for storing instructions. For example, all or some of the memories 8102 may also be outside the communication device 8100.

[0395] In some embodiments, the communication device 8100 further includes one or more transceivers 8103. When the communication device 8100 includes one or more transceivers 8103, the one or more transceivers 8103 perform at least one of the communication steps such as sending and / or receiving in above methods (for example, but is not limited to, step S2107, step S2207, step S2306, step S2408), and the processor 8101 performs other steps (for example, but is not limited to, step S2101, step S2102, step S2103, step S2104, step S2105, step S2106, step S2201, step S2202, step S2203, step S2204, step S2205, step S2206, step S2301, step S2302, step S2303, step S2304, step S2305, step S2401, step S2402, step S2403, step S2404, step S2405, step S2406, step S2407, step S2108, step S2109, step S2208, step S2209, step S2307, step S2308, step S2409, step S2410).

[0396] In some embodiments, the transceiver may include a receiver and / or a transmitter, and the receiver and the transmitter may be separate or integrated. For example, the terms such as transceiver, transceiver unit, transceiver, transceiver circuit may be replaced with each other, the terms such as transmitter, transmission unit, transmitter, transmission circuit may be replaced with each other, and the terms such as receiver, receiving unit, receiver, receiving circuit may be replaced with each other.

[0397] In some embodiments, the communication device 8100 may include one or more interface circuits 8104. For example, the one or more interface circuits 8104 are connected to the one or more memories 8102, and the one or more interface circuits 8104 may be configured to receive signals from the one or more memories 8102 or other devices, and may be configured to send signals to the one or more memories 8102 or other devices. For example, the one or more interface circuits 8104 may read instructions stored in the one or more memories 8102 and send the instructions to the one or more processors 8101.

[0398] The communication device 8100 described in above embodiments may be an encoding-end device or a decoding-end device, but the scope of the communication device 8100 described in the disclosure is not limited thereto, and the structure of the communication device 8100 may not be limited by FIG. 8A. The communication device may be an independent device or may be part of a larger device. For example, the communication device may be: 1) an independent integrated circuit IC, a chip, a chip system, or subsystem; (2) a collection of one or more ICs, for example, the above collection may also include a storage component for storing data and programs; (3) an ASIC, such as a modem; (4) a module that may be embedded in other devices; (5) a receiver, a terminal, an intelligent terminal, a cellular phone, a wireless device, a handheld device, a mobile unit, a vehicle-mounted device, a network device, a cloud device, an artificial intelligence device; (6) others.

[0399] FIG. 8B is a schematic diagram illustrating a structure of a chip 8200 according to an embodiment of the disclosure. In the case where the communication device 8100 may be a chip or a chip system, reference may be made to the schematic diagram of the structure of the chip 8200 shown in FIG. 8B, but the disclosure is not limited thereto.

[0400] The chip 8200 includes one or more processors 8201, and the chip 8200 is configured to perform any of above methods.

[0401] In some embodiments, the chip 8200 further includes one or more interface circuits 8202. For example, the one or more interface circuits 8202 are connected to a memory 8203. The one or more interface circuits 8202 may be configured to receive signals from the memory 8203 or other devices, and the one or more interface circuits 8202 may be configured to send signals to the memory 8203 or other devices. For example, the one or more interface circuits 8202 may read instructions stored in the memory 8203 and send the instructions to the processor 8201.

[0402] In some embodiments, the one or more interface circuits 8202 performs at least one of the communication steps such as sending and / or receiving in above method (for example, but is not limited to, step S2107, step S2207, step S2306, step S2408), and the processor 8201 performs other steps (for example, but is not limited to, step S2101, step S2102, step S2103, step S2104, step S2105, step S2106, step S2201, step S2202, step S2203, step S2204, step S2 205, step S2206, step S2301, step S2302, step S2303, step S2304, step S2305, step S2401, step S2402, step S2403, step S2404, step S2405, step S2406, step S2407, step S2108, step S2109, step S2208, step S2209, step S2307, step S2308, step S2409, step S2410).

[0403] In some embodiments, terms such as interface circuit, interface, transceiver pin, and transceiver may be used interchangeably.

[0404] In some embodiments, the chip 8200 further includes one or more memories 8203 for storing instructions. For example, all or some of the memories 8203 may be outside the chip 8200.

[0405] The disclosure also provides a storage medium, on which instructions are stored. When the instructions are executed on the communication device 8100, the communication device 8100 executes any of above methods. For example, the storage medium is an electronic storage medium. For example, the storage medium is a computer-readable storage medium, but is not limited thereto, and it may also be a storage medium readable by other devices. For example, the storage medium may be a non-transitory storage medium, but is not limited thereto, and it may also be a temporary storage medium.

[0406] The disclosure also provides a program product, which, when executed by the communication device 8100, enables the communication device 8100 to execute any of above methods. For example, the program product is a computer program product.

[0407] The disclosure also provides a computer program, which, when executed on a computer, causes the computer to execute any one of above methods.

[0408] In above embodiments, it may be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs. When the computer program is loaded and executed on a computer, the process or function described in embodiments of the disclosure is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer program may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program may be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer may access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a high-density digital video disc (DVD)), or a semiconductor medium (e.g., a solid state disk (SSD)).

[0409] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with embodiments disclosed herein may be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the disclosure.

[0410] Those skilled in the art may clearly understand that, for the convenience and brevity of description, for the specific working processes of the systems, devices and units described above, reference may be made to the corresponding processes in aforementioned method embodiments and are not repeated here.

[0411] The above is only a specific embodiment of the disclosure, but the protection scope of the disclosure is not limited thereto. Any person skilled in the art who is familiar with the technical field may easily think of changes or substitutions within the technical scope disclosed in the disclosure, which should be included in the protection scope of the disclosure. Therefore, the protection scope of the disclosure should be based on the protection scope of the claims.

Claims

1. A method for processing a signal, comprising: obtaining a combined format audio signal, wherein the combined format audio signal is a multi-channel audio signal comprising a plurality of sound channels; determining a first parameter of the combined format audio signal and / or one or more energy parameters of one or more sound channels among the plurality of sound channels; obtaining bit allocation parameters by performing bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels; obtaining sound channel signal encoding parameters by encoding the plurality of sound channels based on the bit allocation parameters, and writing the bit allocation parameters and the sound channel signal encoding parameters into a bitstream; and sending the bitstream.

2. The method of claim 1, wherein the first parameter comprises at least one of: a sound field analysis parameter; a control parameter, wherein the control parameter comprises a second parameter and / or a third parameter, the second parameter is used for describing an importance ranking of the plurality of sound channels in the combined format audio signal, and the third parameter is used for describing whether each of the plurality of sound channels is a diegetic sound or a non-diegetic sound.

3. The method of claim 1 or 2, wherein the combined format audio signal comprises at least one of: a sound channel-based audio signal; an object-based audio signal; a scene-based audio signal; or a metadata-based three-dimensional audio signal.

4. The method of claim 2 or 3, wherein determining the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels comprises: determining the one or more energy parameters of the one or more sound channels among the plurality of sound channels; and determining the sound field analysis parameter of the combined format audio signal.

5. The method of claim 4, wherein obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels comprises: obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the one or more energy parameters of the one or more sound channels and the sound field analysis parameter of the combined format audio signal.

6. The method of claim 2 or 3, wherein determining the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels comprises: determining the one or more energy parameters of the one or more sound channels among the plurality of sound channels; and determining the control parameter of the combined format audio signal.

7. The method of claim 6, wherein obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels comprises: obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the one or more energy parameters of the one or more sound channels and the control parameter of the combined format audio signal.

8. The method of claim 2 or 3, wherein determining the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels comprises: determining the sound field analysis parameter and / or the control parameter of the combined format audio signal.

9. The method of claim 8, wherein obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels comprises: obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the sound field analysis parameter and / or the control parameter of the combined format audio signal.

10. The method of claim 2 or 3, wherein determining the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels among the plurality of sound channels comprises: determining the one or more energy parameters of the one or more sound channels among the plurality of sound channels; and determining the sound field analysis parameter and the control parameter of the combined format audio signal.

11. The method of claim 10, wherein obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels comprises: obtaining the bit allocation parameters by performing the bit allocation for the plurality of sound channels based on the one or more energy parameters of the one or more sound channels, the sound field analysis parameter and the control parameter of the combined format audio signal.

12. The method of any one of claims 4 to 7 or 10 to 11, wherein determining the one or more energy parameters of the one or more sound channels among the plurality of sound channels comprises: filtering the plurality of sound channels, and determining cross-correlation coefficients between a plurality of filtered sound channels; performing sound channel combination operation based on the cross-correlation coefficients, and determining inter-channel parameters for sound channels in a same sound channel combination; and aligning sound channels and downmixing aligned sound channels with the inter-channel parameters, and performing energy calculation on downmixed sound channels to obtain energy parameters of the plurality of sound channels respectively.

13. The method of claim 12, wherein performing the energy calculation comprises performing any of following processing on sample points of a frame of a downmixed sound channel: calculating a sum of levels of the sample points; calculating a sum of squared levels of the sample points; or calculating a root mean square of levels of the sample points.

14. The method of any one of claims 4 to 5 or 8 to 11, wherein determining the sound field analysis parameter of the combined format audio signal comprises: obtaining the sound field analysis parameter of the combined format audio signal by performing sound field analysis on the combined format audio signal based on a direction of arrival (DOA) estimation and / or a multi-channel cross-correlation coefficient (MCCC) method.

15. The method of any one of claims 6 to 11, wherein determining the control parameter of the combined format audio signal comprises at least one of: determining the control parameter of the combined format audio signal based on a respective functional category of each sound channel determined in collecting the combined format audio signal by an audio signal collection device, wherein the audio signal collection device is configured to collect the combined format audio signal; or determining the control parameter of the combined format audio signal based on a setting parameter required for reconstructing a desired sound field by a decoding-end device.

16. The method of any one of claims 2, 4 to 5, 8 to 11, or 14, wherein the sound field analysis parameter comprises at least one of: total number of sound images; relative priority of sound images; or a background sound.

17. The method of claim 2 or 3, wherein performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels comprises: performing the bit allocation to the combined format audio signal consisting of an object-based audio signal and a scene-based audio signal based on the sound field analysis parameter of the combined format audio signal, wherein a number of bits allocated to any sound channel of the object-based audio signal is greater than a number of bits allocated to any sound channel of the scene-based audio signal.

18. The method of claim 2 or 3, wherein performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels comprises: performing the bit allocation to the combined format audio signal consisting of an object-based audio signal and a sound channel-based audio signal based on the sound field analysis parameter of the combined format audio signal, wherein a number of bits allocated to any sound channel of the object-based audio signal is greater than a number of bits allocated to any sound channel of the sound channel-based audio signal.

19. The method of claim 2 or 3, wherein performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels comprises: performing the bit allocation on the combined format audio signal consisting of a sound channel-based audio signal and a scene-based audio signal based on the sound field analysis parameter of the combined format audio signal, wherein a number of bits allocated to any sound channel of the sound channel-based audio signal is greater than a number of bits allocated to any sound channel of the scene-based audio signal.

20. The method of claim 2 or 3, wherein performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels comprises: performing the bit allocation on the combined format audio signal consisting of an object-based audio signal and a metadata-based three-dimensional audio signal based on the sound field analysis parameter of the combined format audio signal, wherein a number of bits allocated to any sound channel of the object-based audio signal is greater than a number of bits allocated to any sound channel of the metadata-based three-dimensional audio signal.

21. The method of claim 2 or 3, wherein performing the bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels comprises: performing the bit allocation on the combined format audio signal consisting of a sound channel-based audio signal and a metadata-based three-dimensional audio signal based on the sound field analysis parameter of the combined format audio signal, wherein a number of bits allocated to any sound channel of the sound channel-based audio signal is greater than a number of bits allocated to any sound channel of the metadata-based three-dimensional audio signal.

22. The method of any one of claims 17, 18 or 20, wherein a number of bits allocated to a sound channel corresponding to a primary sound image in the object-based audio signal is greater than a number of bits allocated to a sound channel corresponding to a secondary sound image in the object-based audio signal.

23. A method for processing a signal, comprising: receiving a bitstream; parsing the bitstream to obtain bit allocation parameters and sound channel signal encoding parameters for a plurality of sound channels of a combined format audio signal respectively; and reconstructing the combined format audio signal based on the bit allocation parameters and the sound channel signal encoding parameters for the plurality of sound channels.

24. The method of claim 23, wherein reconstructing the combined format audio signal based on the bit allocation parameters and the sound channel signal encoding parameters for the plurality of sound channels comprises: decoding the sound channel signal encoding parameters based on the bit allocation parameters for the plurality of sound channels, and reconstructing the combined format audio signal based on a decoded sound channel bitstream.

25. A first communication device, comprising: a processing module, configured to: obtain a combined format audio signal, wherein the combined format audio signal is multi-channel audio signal comprising a plurality of sound channels; determine a first parameter of the combined format audio signal and / or one or more energy parameters of one or more sound channels among the plurality of sound channels; obtain bit allocation parameters by performing bit allocation for the plurality of sound channels based on the first parameter of the combined format audio signal and / or the one or more energy parameters of the one or more sound channels; and encode the plurality of sound channels based on the bit allocation parameters to obtain sound channel signal encoding parameters, and write the bit allocation parameters and the sound channel signal encoding parameters into a bitstream; and a transceiver module, configured to send the bitstream.

26. A second communication device, comprising: a transceiver module, configured to receive a bitstream; and a processing module, configured to: parse the bitstream to obtain bit allocation parameters and sound channel signal encoding parameters for a plurality of sound channels of a combined format audio signal respectively; and reconstruct the combined format audio signal based on the bit allocation parameters and the sound channel signal encoding parameters for the plurality of sound channels.

27. A communication system, comprising: an encoding-end device, configured to perform the method for processing the signal of any one of claims 1 to 22; and a decoding-end device, configured to perform the method for processing the signal of claim 23 or 24.

28. A communication device, comprising: one or more processors, wherein the one or more processors are configured to call instructions to enable the communication device to perform the method for processing the signal of any one of claims 1 to 22 and 23 to 24.

29. A storage medium, having instructions stored thereon, wherein when the instructions are executed on a communication device, the communication device is caused to perform the method for processing the of any one of claims 1 to 22 and 23 to 24.