Signal processing method and apparatus therefor
By acquiring the parameters and channel energy parameters of the mixed-format audio signal for bit allocation and encoding, the problem of the inability to reconstruct the desired sound field in the prior art is solved, thus improving the reconstruction effect of the audio signal.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING XIAOMI MOBILE SOFTWARE CO LTD
- Filing Date
- 2023-07-14
- Publication Date
- 2026-06-09
AI Technical Summary
In existing technologies, encoding and decoding methods for mixed-format audio signals cannot reconstruct the most suitable and/or desired sound field for specific application scenarios, resulting in poor reconstruction results.
By acquiring the parameters and channel energy parameters of the mixed-format audio signal, bit allocation and encoding are performed to generate bit allocation parameters and channel signal encoding parameters, which are written into the bitstream, and the audio signal is reconstructed at the decoding end based on these parameters.
It improves the reconstruction effect of mixed-format audio signals, making them closer to the real sound field of specific application scenarios, and solves the problem of being unable to reconstruct the desired sound field.
Smart Images

Figure CN122177126A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of communication technology, and in particular to a signal processing method and apparatus. Background Technology
[0002] With the increase in transmission bandwidth and the upgrading of signal acquisition equipment, signal processor performance, and terminal playback equipment, three signal formats—channel-based audio signals, object-based audio signals, and scene-based audio signals—can provide 3D audio services.
[0003] In 3D audio applications, 3D audio typically contains signals of multiple audio signal formats, i.e., mixed-format audio signals. Related technologies typically employ unified encoding and decoding methods for these mixed-format audio signals. For example, the encoder allocates available bits based on the energy parameters of the mixed-format audio signal, and each channel uses its corresponding encoding kernel to encode the allocated bits, outputting encoding parameters which are then written into the bitstream.
[0004] However, the encoders in related technologies use energy-based bit allocation for the input mixed-format audio signal. This bit allocation method is relatively simple and one-sided, which makes it impossible to reconstruct the most suitable and / or desired sound field for specific application scenarios, thus resulting in a deterioration in the reconstruction effect of mixed-format audio signals. Summary of the Invention
[0005] This disclosure presents a signal processing method and apparatus.
[0006] According to a first aspect of the present disclosure, a signal processing method is provided, the method comprising: Acquire a mixed-format audio signal, wherein the mixed-format audio signal is an audio signal including multiple channels; Determine a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels; Bit allocation parameters are obtained by bit-allocating the multiple channels based on the first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the multiple channels; The multiple channels are encoded according to the bit allocation parameters to obtain channel signal encoding parameters, and the bit allocation parameters and the channel signal encoding parameters are written into the bit stream. Send the bitstream.
[0007] According to a second aspect of the embodiments of this disclosure, a signal processing method is provided, comprising: Receive bitstream; The bitstream is parsed to obtain the bit allocation parameters and channel signal encoding parameters of each channel in the multiple channels of the mixed-format audio signal. The mixed-format audio signal is reconstructed based on the bit allocation parameters and channel signal encoding parameters of each channel.
[0008] According to a third aspect of the present disclosure, a first communication device is provided, comprising: The processing module is used to acquire audio signals of mixed format, wherein the audio signals of mixed format are audio signals including multiple channels; The processing module is further configured to determine a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels; The processing module is further configured to perform bit allocation parameters on the multiple channels according to the first parameter of the mixed format audio signal and / or the energy parameter of one or more of the multiple channels; The processing module is further configured to encode the plurality of channels according to the bit allocation parameters to obtain channel signal encoding parameters, and write the bit allocation parameters and the channel signal encoding parameters into the bit stream; A transceiver module is used to send the bitstream.
[0009] According to a fourth aspect of the embodiments of this disclosure, a second communication device is provided, comprising: The transceiver module is used to receive the bitstream; The processing module is used to perform bitstream parsing on the bitstream to obtain the bit allocation parameters and channel signal encoding parameters of each channel in the multiple channels of the mixed-format audio signal. The processing module is also used to reconstruct the mixed-format audio signal based on the bit allocation parameters and channel signal encoding parameters of each of the channels.
[0010] According to a fifth aspect of the embodiments of this disclosure, a communication system is provided, comprising: The encoding end device is configured to perform the optional implementation of the first aspect described above; The decoding device is configured to perform an optional implementation of the second aspect described above.
[0011] According to a sixth aspect of the present disclosure, a communication device is provided, comprising: one or more processors; The processor is used to invoke instructions to cause the communication device to execute the optional implementations of the first and second aspects mentioned above.
[0012] According to a seventh aspect of the present disclosure, a storage medium is provided that stores instructions which, when executed on a communication device, cause the communication device to perform optional implementations of the first and second aspects described above.
[0013] According to the technical solution disclosed herein, the problem that encoding and decoding methods in related technologies cannot reconstruct the desired and most suitable sound field can be solved, thereby improving the reconstruction effect of mixed-format audio signals. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings required for the description of the embodiments are introduced below. The following drawings are only some embodiments of this disclosure and do not impose specific limitations on the protection scope of this disclosure.
[0015] Figure 1 This is a schematic diagram of the architecture of a communication system provided in an embodiment of this disclosure; Figure 2A This is an interactive schematic diagram of a signal processing method according to an embodiment of the present disclosure; Figure 2B This is an interactive schematic diagram of a signal processing method according to an embodiment of the present disclosure; Figure 2C This is an interactive schematic diagram of a signal processing method according to an embodiment of the present disclosure; Figure 2D This is an interactive schematic diagram of a signal processing method according to an embodiment of the present disclosure; Figure 3 This is a schematic flowchart illustrating a signal processing method according to an embodiment of the present disclosure; Figure 4 This is a schematic flowchart illustrating another signal processing method according to an embodiment of the present disclosure; Figure 5 This is an interactive schematic diagram of a signal processing method according to an embodiment of the present disclosure; Figure 6A This is an example of the signal processing method proposed in the embodiments of this disclosure. Figure 1 ; Figure 6B Figure 2 is an example of the signal processing method proposed in the embodiments of this disclosure; Figure 7A This is a schematic diagram of the structure of the encoding end device proposed in the embodiments of this disclosure; Figure 7B This is a schematic diagram of the structure of the decoding device proposed in the embodiments of this disclosure; Figure 8A This is a schematic diagram of the structure of the communication device 8100 proposed in the embodiments of this disclosure; Figure 8BThis is a schematic diagram of the structure of the chip 8200 proposed in the embodiments of this disclosure. Detailed Implementation
[0016] This disclosure presents a signal processing method and apparatus.
[0017] In a first aspect, embodiments of this disclosure provide a signal processing method, the method comprising: Acquire a mixed-format audio signal, wherein the mixed-format audio signal is an audio signal including multiple channels; Determine a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels; Bit allocation parameters are obtained by bit-allocating the multiple channels based on the first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the multiple channels; The multiple channels are encoded according to the bit allocation parameters to obtain channel signal encoding parameters, and the bit allocation parameters and the channel signal encoding parameters are written into the bit stream. Send the bitstream.
[0018] In the above embodiments, by using a first parameter based on energy and / or a mixed format audio signal to perform bit allocation on multiple channels in the mixed format audio signal, the bit allocation parameters can be made closer to the real sound field situation of the specific application scenario, thereby reconstructing the desired and most suitable sound field. This solves the problem that the encoding and decoding methods in related technologies cannot reconstruct the desired and most suitable sound field, thereby improving the reconstruction effect of the mixed format audio signal.
[0019] In conjunction with some embodiments of the first aspect, in some embodiments, the first parameter includes at least one of the following: Sound field analysis parameters; The control parameters include a second parameter and / or a third parameter, wherein the second parameter describes the priority order of the roles of each of the plurality of channels in the mixed-format audio signal, and the third parameter describes whether each of the channels is narrative or non-narrative.
[0020] In the above embodiments, when allocating bits to multiple channels in a mixed-format audio signal, considering sound field analysis parameters and / or control parameters can make the bit allocation parameters more closely resemble the real sound field situation in the specific application scenario, thereby reconstructing the desired and most suitable sound field and improving the reconstruction effect of the mixed-format audio signal.
[0021] In conjunction with some embodiments of the first aspect, in some embodiments, the mixed-format audio signal includes audio signals of at least one of the following formats: Audio signals based on vocal tracts; Object-based audio signals; Scene-based audio signals; 3D audio signals based on metadata.
[0022] In the above embodiments, multiple different formats of mixed audio signals can be encoded and decoded, enabling three-dimensional audio communication services between the audio signal acquisition end and the audio signal playback end.
[0023] In conjunction with some embodiments of the first aspect, in some embodiments, determining the first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels includes: Determine the energy parameters of one or more of the plurality of audio channels; Determine the sound field analysis parameters for the audio signal in the mixed format.
[0024] In conjunction with some embodiments of the first aspect, in some embodiments, the step of bit-allocating the plurality of channels according to a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels to obtain bit allocation parameters includes: The bit allocation parameters are obtained by bit-allocating the plurality of channels based on the energy parameters of one or more of the channels and the sound field analysis parameters of the mixed-format audio signal.
[0025] In the above embodiments, when allocating bits to multiple channels in a mixed-format audio signal, in addition to considering energy parameters, sound field analysis parameters are also considered. This makes the bit allocation parameters closer to the real sound field situation in the specific application scenario, thereby reconstructing the desired and most suitable sound field and improving the reconstruction effect of the mixed-format audio signal.
[0026] In conjunction with some embodiments of the first aspect, in some embodiments, determining the first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels includes: Determine the energy parameters of one or more of the plurality of audio channels; The control parameters of the mixed-format audio signal are determined.
[0027] In conjunction with some embodiments of the first aspect, in some embodiments, the step of bit-allocating the plurality of channels according to a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels to obtain bit allocation parameters includes: The bit allocation parameters are obtained by bit-allocating the plurality of channels based on the energy parameters of one or more of the channels and the control parameters of the mixed-format audio signal.
[0028] In the above embodiments, when allocating bits to multiple channels in a mixed-format audio signal, in addition to considering energy parameters, external control parameters are also considered. The bit allocation can be adjusted according to the external control parameters, making the bit allocation parameters closer to the real sound field situation of the specific application scenario. This allows the desired and most suitable sound field to be reconstructed, thereby improving the reconstruction effect of the mixed-format audio signal.
[0029] In conjunction with some embodiments of the first aspect, in some embodiments, determining the first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels includes: Determine the sound field analysis parameters and / or the control parameters of the mixed-format audio signal.
[0030] In conjunction with some embodiments of the first aspect, in some embodiments, the step of bit-allocating the plurality of channels according to a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels to obtain bit allocation parameters includes: The bit allocation parameters are obtained by bit-allocating the plurality of channels based on the sound field analysis parameters and / or the control parameters of the mixed-format audio signal.
[0031] In the above embodiments, when allocating bits to multiple channels in a mixed-format audio signal, considering sound field analysis parameters and / or external control parameters can make the bit allocation parameters more closely resemble the real sound field situation in the specific application scenario, thereby reconstructing the desired and most suitable sound field and improving the reconstruction effect of the mixed-format audio signal.
[0032] In conjunction with some embodiments of the first aspect, in some embodiments, determining the first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels includes: Determine the energy parameters of one or more of the plurality of audio channels; The sound field analysis parameters and control parameters of the mixed-format audio signal are determined.
[0033] In conjunction with some embodiments of the first aspect, in some embodiments, the step of bit-allocating the plurality of channels according to a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels to obtain bit allocation parameters includes: The bit allocation parameters are obtained by bit-allocating the plurality of channels based on the energy parameters of one or more of the channels, the sound field analysis parameters of the mixed-format audio signal, and the control parameters.
[0034] In the above embodiments, when allocating bits to multiple channels in a mixed-format audio signal, in addition to considering energy parameters, external control parameters and sound field analysis parameters are also considered. This allows the bit allocation parameters to be closer to the real sound field situation in the specific application scenario, thereby reconstructing the desired and most suitable sound field and further improving the reconstruction effect of the mixed-format audio signal.
[0035] In conjunction with some embodiments of the first aspect, in some embodiments, determining the energy parameters of one or more of the plurality of channels includes: After filtering the multiple channels, the cross-correlation coefficients between the channels are calculated. Based on the cross-correlation coefficient, perform channel combination operations and calculate inter-channel parameters for channels within the same channel combination; After aligning the channels using the inter-channel parameters, downmixing is performed, and energy calculation is performed on the downmixed channels to obtain the energy parameters of each channel.
[0036] In the above embodiments, filtering the audio channels and performing energy calculations on the filtered channels can improve the accuracy of the calculation results and avoid noise interference.
[0037] In conjunction with some embodiments of the first aspect, in some embodiments, the method for calculating the energy parameters includes performing any of the following processing on samples of a frame of the audio channel: Calculate the sum of the amplitude values of the sample points; Calculate the sum of squares of the amplitudes of the sample points; Calculate the root mean square of the sample points.
[0038] In the above embodiments, the accuracy of the calculation results can be guaranteed by using any one of the following methods for energy calculation: sample amplitude sum, sample amplitude square sum, sample amplitude root mean square.
[0039] In conjunction with some embodiments of the first aspect, in some embodiments, determining the sound field analysis parameters of the mixed-format audio signal includes: Based on Direction-of-Arrival (DOA) estimation and / or Multi-Channel Cross-Correlation Coefficient (MCCC) method, sound field analysis is performed on the mixed-format audio signal to obtain the sound field analysis parameters of the mixed-format audio signal.
[0040] In conjunction with some embodiments of the first aspect, in some embodiments, the control parameters for determining the mixed-format audio signal include at least one of the following: Based on the functional division of each channel during audio signal acquisition by the audio signal acquisition device, the control parameters of the mixed-format audio signal are determined, and the audio signal acquisition device is used to acquire the mixed-format audio signal. The control parameters of the mixed-format audio signal are determined based on the setting parameters required for the desired sound field reconstructed by the decoding device.
[0041] In the above embodiments, by determining the control parameters based on the functional division of each channel during audio signal acquisition and / or the setting parameters required for the desired sound field reconstructed by the decoding device, the control parameters can be made closer to the actual situation of the application scenario. By adjusting the bit allocation using the control parameters, the adjusted bit allocation parameters can be made to further approximate the real sound field situation of the specific application scenario, thereby achieving the reconstruction of the desired mixed format audio signal at the decoding end.
[0042] In conjunction with some embodiments of the first aspect, in some embodiments, the sound field analysis parameters include at least one of the following: The number of sound images; Primary and secondary audio-visual elements; Background ambient sound.
[0043] In conjunction with some embodiments of the first aspect, in some embodiments, the bit allocation of the plurality of channels based on a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels includes: Bit allocation is performed on the mixed-format audio signal, which consists of the object-based audio signal and the scene-based audio signal, based on the sound field analysis parameters of the mixed-format audio signal. The number of bits allocated to any channel in the object-based audio signal is greater than the number of bits allocated to any channel in the scene-based audio signal.
[0044] In conjunction with some embodiments of the first aspect, in some embodiments, the bit allocation of the plurality of channels based on a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels includes: Bit allocation is performed on the mixed-format audio signal, which consists of the object-based audio signal and the channel-based audio signal, based on the sound field analysis parameters of the mixed-format audio signal. The number of bits allocated to any channel in the object-based audio signal is greater than the number of bits allocated to any channel in the channel-based audio signal.
[0045] In conjunction with some embodiments of the first aspect, in some embodiments, the bit allocation of the plurality of channels based on a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels includes: Bit allocation is performed on the mixed-format audio signal, which consists of the channel-based audio signal and the scene-based audio signal, based on the sound field analysis parameters of the mixed-format audio signal. The number of bits allocated to any channel in the channel-based audio signal is greater than the number of bits allocated to any channel in the scene-based audio signal.
[0046] In conjunction with some embodiments of the first aspect, in some embodiments, the bit allocation of the plurality of channels based on a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels includes: Bit allocation is performed on the mixed-format audio signal, which consists of the object-based audio signal and the metadata-based three-dimensional audio signal, based on the sound field analysis parameters of the mixed-format audio signal. The number of bits allocated to any channel in the object-based audio signal is greater than the number of bits allocated to any channel in the metadata-based three-dimensional audio signal.
[0047] In conjunction with some embodiments of the first aspect, in some embodiments, the bit allocation of the plurality of channels based on a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels includes: Bit allocation is performed on the mixed-format audio signal, which consists of the channel-based audio signal and the metadata-based three-dimensional audio signal, based on the sound field analysis parameters of the mixed-format audio signal. The number of bits allocated to any channel in the channel-based audio signal is greater than the number of bits allocated to any channel in the metadata-based three-dimensional audio signal.
[0048] In conjunction with some embodiments of the first aspect, in some embodiments, the number of bits allocated to the channel corresponding to the primary image in the object-based audio signal is greater than the number of bits allocated to the channel corresponding to the secondary image in the object-based audio signal.
[0049] Secondly, embodiments of this disclosure propose a signal processing method, including... Receive bitstream; The bitstream is parsed to obtain the bit allocation parameters and channel signal encoding parameters of each channel in the multiple channels of the mixed-format audio signal. The mixed-format audio signal is reconstructed based on the bit allocation parameters and channel signal encoding parameters of each channel.
[0050] In conjunction with some embodiments of the second aspect, in some embodiments, reconstructing the mixed-format audio signal based on the bit allocation parameters of each of the channels, the channel signal encoding parameters, and the bitstream includes: The channel signal encoding parameters are decoded based on the bit allocation parameters of each channel; The audio signal in the mixed format is reconstructed using the decoded channels.
[0051] Thirdly, embodiments of this disclosure provide an encoding end device, including at least one of a transceiver module and a processing module; wherein the encoding end device is used to execute an optional implementation of the first aspect.
[0052] Fourthly, embodiments of this disclosure provide a decoding end device, including at least one of a transceiver module and a processing module; wherein the decoding end device is used to execute an optional implementation of the second aspect.
[0053] Fifthly, embodiments of this disclosure provide a communication system, including: The encoding end device is configured as an optional implementation of the first aspect mentioned above; The decoding device is configured to perform an optional implementation of the second aspect described above.
[0054] In a sixth aspect, embodiments of this disclosure provide a communication device, comprising: one or more processors; wherein the processors are configured to invoke instructions to cause the communication device to perform an optional implementation of the first aspect described above.
[0055] In a seventh aspect, embodiments of this disclosure provide a communication device, comprising: one or more processors; wherein the processors are configured to invoke instructions to cause the communication device to perform an optional implementation of the second aspect described above.
[0056] Eighthly, embodiments of this disclosure provide a storage medium storing instructions that, when executed on a communication device, cause the communication device to perform optional implementations of the first and second aspects described above.
[0057] Ninthly, embodiments of this disclosure provide a program product that, when executed by a communication device, causes the communication device to perform the method as described in the optional implementations of the first and second aspects.
[0058] In a tenth aspect, embodiments of this disclosure provide a computer program that, when run on a computer, causes the computer to perform the methods described in alternative implementations of the first and second aspects.
[0059] Eleventhly, embodiments of this disclosure provide a chip or chip system. The chip or chip system includes processing circuitry configured to perform the methods described according to optional implementations of the first and second aspects above.
[0060] It is understood that the aforementioned encoding device, decoding device, communication system, storage medium, program product, computer program, chip, or chip system are all used to execute the methods proposed in the embodiments of this disclosure. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.
[0061] This disclosure provides a signal processing method and apparatus. In some embodiments, terms such as signal processing method, communication method, signal encoding / decoding method, signal coding method, and signal decoding method can be used interchangeably; terms such as signal processing apparatus and communication apparatus can be used interchangeably; and terms such as signal processing system, information processing system, and communication system can be used interchangeably.
[0062] This disclosure is not exhaustive, but merely illustrative of some embodiments, and is not intended to limit the scope of protection of this disclosure. Unless otherwise specified, each step in a particular embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a particular embodiment can also be implemented as an independent embodiment, and the order of the steps in a particular embodiment can be arbitrarily interchanged. Furthermore, the optional implementation methods in a particular embodiment can be arbitrarily combined; moreover, the embodiments can be arbitrarily combined, for example, some or all steps of different embodiments can be arbitrarily combined, and a particular embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.
[0063] In each of the disclosed embodiments, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of the embodiments are consistent and can be referenced by each other. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.
[0064] The terminology used in the embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure.
[0065] In this disclosure, unless otherwise stated, elements expressed in the singular form, such as "a," "an," "the," "the," "the," "the," "the," "the," "this," etc., can mean "one and only one," or "one or more," "at least one," etc. For example, when using articles such as "a," "an," "the," etc. in translation, the noun following the article can be understood as either a singular or a plural expression.
[0066] In the embodiments disclosed herein, "multiple" refers to two or more.
[0067] In some embodiments, the terms “at least one of,” “one or more,” “a plurality of,” and “multiple” may be used interchangeably.
[0068] In some embodiments, the notation "at least one of A and B", "A and / or B", "A in one case, B in another", "in response to one case A, in response to another case B", etc., may include the following technical solutions depending on the situation: in some embodiments, A (executes A regardless of B); in some embodiments, B (executes B regardless of A); in some embodiments, execution is selected from A and B (A and B are selectively executed); in some embodiments, both A and B are executed. The same applies when there are more branches such as A, B, C, etc.
[0069] In some embodiments, the notation "A or B" may include the following technical solutions, depending on the situation: in some embodiments, A (execution of A regardless of B); in some embodiments, B (execution of B regardless of A); in some embodiments, selective execution from A and B (A and B are selectively executed). The same applies when there are more branches such as A, B, and C.
[0070] The prefixes "first," "second," etc., used in the embodiments of this disclosure are merely for distinguishing different descriptive objects and do not impose restrictions on the position, order, priority, quantity, or content of the descriptive objects. The description of the descriptive objects is found in the claims or the context of the embodiments, and the use of prefixes should not constitute unnecessary restrictions. For example, if the descriptive object is a "field," the ordinal numbers preceding "field" in "first field" and "second field" do not restrict the position or order of the "fields." "First" and "second" do not restrict whether the "fields" they modify are in the same message, nor do they restrict the order of "first field" and "second field." Similarly, if the descriptive object is a "level," the ordinal numbers preceding "level" in "first level" and "second level" do not restrict the priority between "levels." Furthermore, the number of descriptive objects is not limited by ordinal numbers and can be one or more. For example, in "first device," the number of "devices" can be one or more. Furthermore, the objects modified by different prefixes can be the same or different. For example, if the object being described is "device", then "first device" and "second device" can be the same device or different devices, and their types can be the same or different. Similarly, if the object being described is "information", then "first information" and "second information" can be the same information or different information, and their content can be the same or different.
[0071] In some embodiments, “including A,” “containing A,” “for indicating A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.
[0072] In some embodiments, the terms “in response to…”, “in response to determining…”, “in the case of…”, “when…”, “if…”, “if…”, etc., can be used interchangeably.
[0073] In some embodiments, the terms “greater than,” “greater than or equal to,” “not less than,” “more than,” “more than or equal to,” “not less than,” “higher than,” “higher than or equal to,” “not lower than,” and “above” can be used interchangeably, as can the terms “less than,” “less than or equal to,” “not greater than,” “less than,” “less than or equal to,” “not more than,” “lower than,” “lower than or equal to,” “not higher than,” and “below”.
[0074] In some embodiments, the apparatus and device may be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. In some cases, they may also be understood as "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "body", etc.
[0075] In some embodiments, "network" can be interpreted as devices included in the network, such as access network devices, core network devices, etc.
[0076] In some embodiments, "access network device (AN device)" may also be referred to as "radio access network device (RAN device)," "base station (BS)," "radio base station," or "fixed station." In some embodiments, it may also be understood as "node," "access point," "transmission point (TP)," "reception point (RP)," "transmission and / or reception point (TRP)," "panel," "antenna panel," "antenna array," "cell," "macro cell," "small cell," "femto cell," "pico cell," "sector," "cellgroup," "serving cell," "carrier," "component carrier," or "bandwidth part (BWP)," etc.
[0077] In some embodiments, "terminal" or "terminal device" may be referred to as "user equipment (UE)," "user terminal," "mobile station (MS)," "mobile terminal (MT)," "subscriber station," "mobile unit," "subscriber unit," "wireless unit," "remote unit," "mobile device," "wireless device," "wireless communication device," "remote device," "mobile subscriber station," "access terminal," "mobile terminal," "wireless terminal," "remote terminal," "handset," "user agent," "mobile client," "client," etc.
[0078] In some embodiments, the acquisition of data, information, etc., may comply with the laws and regulations of the country where the location is situated.
[0079] In some embodiments, data, information, etc., may be obtained with the user's consent.
[0080] Figure 1 This is a schematic diagram of the architecture of a communication system according to an embodiment of this disclosure. The communication system may include, but is not limited to, an encoding end device and a decoding end device. Figure 1 The number and form of the devices shown are for illustrative purposes only and do not constitute a limitation on the embodiments of this disclosure. In actual applications, there may be two or more encoding end devices and two or more decoding end devices. Figure 1 The communication system 100 shown is exemplified by including an encoding end device 101 and a decoding end device 102.
[0081] In some embodiments, the encoding end device 101 may be a terminal with encoding capabilities, such as a terminal including an encoder. This terminal can be understood as a receiver of mixed-format audio signals. In some embodiments, the terminal may be a user-side entity used to receive or transmit signals, such as a mobile phone. It may also be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. The terminal may be at least one of the following: a car with communication capabilities, a smart car, a mobile phone, a wearable device, a tablet computer, a computer with wireless transceiver capabilities, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc. The embodiments disclosed herein do not limit the specific technology or device form used in the terminal.
[0082] In some embodiments, the encoding device 101 may be an encoder on a network device. For example, an encoder is deployed on the network device to encode the input mixed-format audio signal. For instance, the acquisition end sends the mixed-format audio signal to the network device. The encoder on the network device encodes the mixed-format audio signal and sends the resulting bitstream to the decoding device 102.
[0083] In some embodiments, the decoding device 102 may be a terminal with decoding capabilities, such as a terminal including an encoder. This terminal can be understood as a playback end for mixed-format audio signals. In some embodiments, the terminal may be a user-side entity used to receive or transmit signals, such as a mobile phone. It may also be referred to as a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. The terminal may be at least one of the following: a car with communication capabilities, a smart car, a mobile phone, a wearable device, a tablet computer, a computer with wireless transceiver capabilities, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc. The embodiments disclosed herein do not limit the specific technology or device form used in the terminal.
[0084] In some embodiments, the decoding device 102 may be a decoder on a network device. For example, a decoder is deployed on the network device, which decodes the received bitstream, reconstructs the mixed-format audio signal, and sends the reconstructed mixed-format audio signal to the playback end.
[0085] In some embodiments, the network device may be an access network device. In some embodiments, the access network device is, for example, a node or device that connects a terminal device to a wireless network. The access network device may include, but is not limited to, at least one of the following in a 5G communication system: evolved Node B (eNB), next-generation eNB (ng-eNB), next-generation Node B (gNB), node B (NB), home node B (HNB), home evolved node B (HeNB), radio backhaul device, radio network controller (RNC), base station controller (BSC), base transceiver station (BTS), base band unit (BBU), mobile switching center, base station in a 6G communication system, open RAN, cloud RAN, base station in other communication systems, and access node in a Wi-Fi system.
[0086] In some embodiments, the technical solutions of this disclosure can be applied to the Open RAN architecture. In this case, the interfaces between or within access network devices involved in the embodiments of this disclosure can be transformed into internal interfaces of Open RAN. The processes and information interactions between these internal interfaces can be implemented by software or programs.
[0087] In some embodiments, the access network device may be composed of a central unit (CU) and a distributed unit (DU). The CU may also be called a control unit. The CU-DU structure can separate the protocol layer of the access network device. Some of the protocol layer functions are centrally controlled by the CU, while the remaining part or all of the protocol layer functions are distributed in the DU and centrally controlled by the CU. However, this is not the only possibility.
[0088] It is understood that the communication system described in this disclosure is for the purpose of more clearly illustrating the technical solutions of this disclosure, and does not constitute a limitation on the technical solutions proposed in this disclosure. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions proposed in this disclosure are also applicable to similar technical problems.
[0089] The following embodiments of this disclosure can be applied to Figure 1The communication system 100 shown, or a part thereof, but not limited to it. Figure 1 The entities shown are illustrative; a communication system may include... Figure 1 All or part of the main body, or may include Figure 1 Other entities besides the main body, the number and form of each entity are arbitrary, each entity can be physical or virtual, the connection relationship between the entities is illustrative, the entities can be unconnected or connected, and the connection can be in any way, it can be a direct connection or an indirect connection, it can be a wired connection or a wireless connection.
[0090] The embodiments disclosed herein can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G new radio (NR), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New radio access (NX), Futuregeneration radio access (FX), Global System for Mobile communications (GSM), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi), IEEE 802.16 (WiMAX), and IEEE 802.20, Ultra-Wideband (UWB), Bluetooth (a registered trademark), Public Land Mobile Network (PLMN) networks, Device-to-Device (D2D) systems, Machine-to-Machine (M2M) systems, Internet of Things (IoT) systems, Vehicle-to-Everything (V2X) systems, systems utilizing other communication methods, and next-generation systems built upon them, etc. Furthermore, multiple systems can be combined (e.g., a combination of LTE or LTE-A with 5G).
[0091] It should be noted that the first generation of mobile communication technology (1G) was the first generation of wireless cellular technology, belonging to analog mobile communication networks. When 1G was upgraded to 2G, mobile phones transitioned from analog to digital communication, primarily using the GSM (Global System for Mobile Communications) network standard. Voice encoders employed AMR (Adaptive Multi-Rate), EFR (Enhanced Full Rate), FR (Full Rate), and HR (Half Rate) communication to provide single-channel narrowband voice service. The 3G mobile communication system was proposed by the ITU (International Telecommunication Union) for international mobile communications in 2000. China Mobile used TD-SCDMA, China Telecom used CDMA2000, and China Unicom used WCDMA. Their voice encoders used AMR-WB (Adaptive Multi-Rate Wideband) to provide single-channel broadband voice service. 4G is a further improvement on 3G technology. Both data and voice use the full IP (Internet protocol) approach, providing real-time HD / HD+ Voice services. The EVS codec used can achieve high-quality compression and reconstruction of both voice and audio.
[0092] The voice and audio communication services provided above have expanded from narrowband signals to ultra-wideband and even full-band services, but they are still all mono services. People's demand for high-quality audio is constantly increasing. Compared to mono audio, stereo audio provides a sense of orientation and distribution for each sound source and can improve clarity. This is further enhanced by the increase in transmission bandwidth, upgrades to terminal equipment signal acquisition devices, improvements in signal processor performance, and upgrades to terminal playback equipment.
[0093] In some embodiments, three signal formats—channel-based audio signals, object-based audio signals, and scene-based audio signals—can provide 3D audio services. The IVAS (Immersive Voice and Audio Services) codec, which is being standardized by 3GPP SA4, can support the encoding and decoding requirements of these three signal formats.
[0094] In some embodiments, the aforementioned channel-based audio signals include, but are not limited to: mono signal, stereo signal, binaural signal, 5.1, 7.1 surround sound signal, 5.1.4, 7.1.4 surround sound signal, where .4 represents the height channel signal.
[0095] In some embodiments, the above-mentioned scene-based audio signals include, but are not limited to: First Order Ambisonics (FOA), Second Order High Ambisonics (HOA2), and Third Order High Ambisonics (HOA3).
[0096] In some embodiments, the object-based audio signal described above includes audio data and metadata. In addition, IVAS also supports metadata-assisted spatial audio (MASA).
[0097] In some embodiments, terminals capable of supporting 3D audio services may include, but are not limited to, mobile phones, computers, tablets, conference system equipment, AR / VR devices, automobiles, etc.
[0098] In some embodiments, in the application scenarios of three-dimensional audio, three-dimensional audio typically contains signals of multiple audio signal formats, i.e., mixed format audio signals. The encoding end device receives the mixed format audio signals, and the encoded audio bitstream signal is sent from the sending end to the receiving end. The decoder at the receiving end decodes the received audio bitstream and reconstructs the mixed format audio signal.
[0099] However, the encoding and decoding methods for mixed-format audio signals in related technologies typically involve: performing unified encoding and decoding processing on the mixed-format audio signal; the encoding process involves allocating available bits based on the energy parameters of the mixed-format audio signal; each channel encodes the data using an encoding kernel based on the allocated bits, outputting encoding parameters, and then placing these parameters into the bitstream. In other words, the encoders in these technologies use energy-parameter-based bit allocation on the input mixed-format audio signal, without performing sound field analysis. This results in the inability to reconstruct the most suitable sound field for specific application scenarios, or the lack of adjustment of bit allocation based on external control parameters, leading to the failure to reconstruct the desired sound field.
[0100] Therefore, this disclosure provides a signal processing method and apparatus that can solve the problem that encoding and decoding methods in related technologies cannot reconstruct the desired and most suitable sound field, thereby improving the reconstruction effect of mixed-format audio signals.
[0101] Figure 2A This is an interactive schematic diagram illustrating a signal processing method according to an embodiment of the present disclosure. For example... Figure 2A As shown, the signal processing method disclosed herein can be applied to a communication system 100, and the method includes, but is not limited to, the following steps.
[0102] In step S2101, the encoding device 101 acquires the mixed format audio signal (MFA).
[0103] In some embodiments, the encoding end device 101 may be a terminal or a base station. The terminal may be a device that provides voice and / or data connectivity to a user. The terminal may communicate with one or more core networks via a RAN (Radio Access Network). The terminal may be an IoT terminal, such as a sensor device, a mobile phone (or "cellular" phone), or a computer with an IoT terminal. For example, it may be a fixed, portable, pocket-sized, handheld, computer-embedded, or vehicle-mounted device. Examples include a station (STA), subscriber unit, subscriber station, mobile station, mobile station, remote station, access point, remote terminal, access terminal, user terminal, or user agent. Alternatively, the terminal may be a device on an unmanned aerial vehicle (UAV). Alternatively, the terminal may be a vehicle-mounted device, such as a vehicle computer with wireless communication capabilities, or a wireless terminal connected to an external vehicle computer. Alternatively, the UE can also be a roadside device, such as a street light, traffic light, or other roadside device with wireless communication capabilities.
[0104] In some embodiments, the above-described mixed-format audio signal includes at least one of the following formats: channel-based audio signal; object-based audio signal; scene-based audio signal; and metadata-based 3D audio signal.
[0105] In this embodiment, the above-mentioned mixed-format audio signal may include any one of the following: channel-based audio signal, object-based audio signal, scene-based audio signal, and metadata-based three-dimensional audio signal.
[0106] In this embodiment, the above-mentioned mixed-format audio signal may include object-based audio signals and scene-based audio signals.
[0107] In this embodiment, the above-mentioned mixed-format audio signal may include object-based audio signals and channel-based audio signals.
[0108] In this embodiment, the above-mentioned mixed-format audio signal may include channel-based audio signals and scene-based audio signals.
[0109] In this embodiment, the above-mentioned mixed-format audio signal may include object-based audio signals and metadata-based three-dimensional audio signals.
[0110] In this embodiment, the above-mentioned mixed-format audio signal may include channel-based audio signals and metadata-based three-dimensional audio signals.
[0111] It should be noted that the above embodiments are not exhaustive, but only illustrative of some embodiments. Furthermore, the above embodiments can be implemented individually or in combination. The above embodiments are only illustrative and are not intended to limit the scope of protection of the embodiments disclosed herein.
[0112] In some embodiments, the four audio signal formats described above are specifically classified based on the signal acquisition format, and the application scenarios emphasized by different audio signal formats will also be different.
[0113] Specifically, in one embodiment of this disclosure, the main application scenario of the above-mentioned channel-based audio signal can be: the acquisition end and the playback end are respectively pre-set with the same microphone acquisition layout and speaker playback layout, wherein the microphone of the acquisition end can be used to acquire channel-based audio signals in 5.0 format; the speaker of the playback end can play back the channel-based audio signals in 5.0 format acquired by the acquisition end.
[0114] In another embodiment of this disclosure, the above-mentioned object-based audio signal is usually recorded by an independent microphone to record the sound of the sounding object. Its main application scenario is that the audio signal needs to be independently controlled at the playback end, such as sound on / off, volume adjustment, sound image orientation adjustment, frequency band equalization and other control operations.
[0115] In another embodiment of this disclosure, the main application scenario of the above-mentioned scene-based audio signal can be: the need to record the complete sound field where the acquisition end is located, such as live concert recording, live football match recording, etc.
[0116] In some embodiments, the "acquisition end" and the "encoding end device 101" may be located on the same device or on different devices. For example, if the "acquisition end" has encoding functionality, then the "acquisition end" and the "encoding end device 101" may be located on the same device, or they may be interchangeable. Alternatively, if the "acquisition end" does not have encoding functionality, then the "acquisition end" and the "encoding end device 101" may be located on different devices, such as the "encoding end device 101" being located on a network device.
[0117] In some embodiments, the "playback end" and the "decoding end device 102" may be located on the same device or on different devices. For example, if the "playback end" has decoding functionality, then the "playback end" and the "decoding end device 102" may be located on the same device, or they may be interchangeable. Alternatively, if the "playback end" does not have decoding functionality, then the "playback end" and the "decoding end device 101" may be located on different devices, such as the "decoding end device 102" being located on a network device.
[0118] In some embodiments, the above-described mixed-format audio signal may be a multi-channel audio signal. In one implementation, the above-described channel-based audio signal may include one or more channels; the above-described object-based audio signal may include one or more channels; the above-described scene-based audio signal may include one or more channels; and the above-described metadata-based 3D audio signal may include one or more channels.
[0119] In step S2102, the encoding device 101 performs filtering processing on the mixed-format audio signal.
[0120] In some embodiments, the encoding device 101 can perform high-pass filtering on the mixed-format audio signal. Optionally, the filter cutoff frequency can be set to 20Hz. As an example, the filter formula used can be shown in formula (1) below: (1) in, For a filter with a cutoff frequency of 20Hz, It is a mixed-format audio signal. , , , , These are all pre-set constants, for example. =0.9981492, =-1.9963008, =0.9981498, =1.9962990, = -0.9963056.
[0121] In step S2103, the encoding device 101 determines the sound field analysis parameters of the mixed-format audio signal.
[0122] In some embodiments, the sound field analysis parameters described above include at least one of the following: the number of sound images; the hierarchy of sound images; and ambient background sound. In one implementation, the sound field analysis parameters include any one of the number of sound images, the hierarchy of sound images, and ambient background sound. In another implementation, the sound field analysis parameters include any two of the number of sound images, the hierarchy of sound images, and ambient background sound. In yet another implementation, the sound field analysis parameters include the number of sound images, the hierarchy of sound images, and ambient background sound.
[0123] In some embodiments, the encoding device 101 can directly perform sound field analysis on the input mixed-format audio signal to determine the sound field analysis parameters of the mixed-format audio signal. Alternatively, in some embodiments, the encoding device 101 can perform sound field analysis on the mixed-format audio signal after filtering (such as high-pass filtering) to determine the aforementioned sound field analysis parameters. Here, sound field analysis refers to analyzing the number of sound images, the hierarchy of sound images, background ambient sounds, etc., in the mixed-format audio signal.
[0124] In some embodiments, sound field analysis of mixed-format audio signals can be performed based on DOA (Direction-Of-Arrival) estimation and / or MCCC (Multichannel Cross-Correlation Coefficients) to obtain sound field analysis parameters for the mixed-format audio signals. For example, sound source localization estimation, including azimuth and elevation angles, can be performed using DOA estimation, and the MCCC method can be used to solve for these parameters. Based on this method, sound field analysis parameters for the mixed-format audio signals can be identified, such as the number of sound images, the hierarchy of sound images, and background ambient noise.
[0125] In step S2104, the encoding device 101 determines the energy parameters of one or more of the multiple channels.
[0126] In some embodiments, the encoding device 101 can directly perform energy calculations on the input mixed-format audio signal to determine the energy parameters of one or more channels among the plurality of channels. For example, the encoding device 101 calculates the cross-correlation coefficients between the channels among the plurality of channels; performs channel combination operations based on the cross-correlation coefficients, calculates inter-channel parameters for channels within the same channel combination; aligns the channels using the inter-channel parameters, performs downmixing, and calculates the energy of the downmixed channels to obtain the energy parameters of each channel.
[0127] In some embodiments, the encoding device 101 can perform energy calculations on the mixed-format audio signal after filtering (such as high-pass filtering) to determine the energy parameters of one or more channels among the plurality of channels. For example, after filtering the plurality of channels, the encoding device 101 calculates the cross-correlation coefficients between the channels; performs channel combination operations based on the cross-correlation coefficients, calculates inter-channel parameters for channels within the same channel combination; aligns the channels using the inter-channel parameters, performs downmixing, and calculates the energy of the downmixed channels to obtain the energy parameters of each channel.
[0128] In some embodiments, the aforementioned inter-channel parameters may be one or more of ILD (Interaural Level Difference), ITD (Interaural Time Difference), and IPD (Interaural Phase Difference).
[0129] In some embodiments, the method for calculating the energy parameters described above may include performing any of the following processing on the samples of a frame of the above-mentioned channel: calculating the sum of amplitude values of the samples; calculating the sum of squared amplitudes of the samples; calculating the root mean square of the samples.
[0130] For example, the sum of amplitude values of the samples in a single frame after downmixing can be calculated to achieve energy calculation. For instance, the sum of amplitude values of the samples in a single frame for each channel can be calculated, and this sum is used as the energy parameter for the corresponding channel. For example, assuming the mixed audio signal includes three channels, such as channel a, channel b, and channel c, the sum of amplitude values of the samples in a single frame for each of these three channels can be calculated. The sum of amplitude values of the samples in a single frame for channel a can be used as the energy parameter for channel a, the sum of amplitude values of the samples in a single frame for channel b can be used as the energy parameter for channel b, and the sum of amplitude values of the samples in a single frame for channel c can be used as the energy parameter for channel c. As an example, the formula for calculating the sum of squared amplitudes of the samples is as follows: (2) in, The sum of amplitude values of the samples in one frame of the m-th channel, where one frame of the m-th channel includes... Sample points For the first One channel.
[0131] For example, the sum of squared amplitudes of the samples in a single frame of a channel after downmixing can be calculated to achieve energy calculation. For instance, the sum of squared amplitudes of the samples in a single frame of each channel can be calculated, and this sum of squared amplitudes is determined as the energy parameter of the corresponding channel. For example, assuming the mixed audio signal includes three channels, such as channel a, channel b, and channel c, the sum of squared amplitudes of the samples in a single frame of each of these three channels can be calculated. The sum of squared amplitudes of the samples in a single frame of channel a can be determined as the energy parameter of channel a, the sum of squared amplitudes of the samples in a single frame of channel b can be determined as the energy parameter of channel b, and the sum of squared amplitudes of the samples in a single frame of channel c can be determined as the energy parameter of channel c. As an example, the formula for calculating the sum of squared amplitudes of the samples is as follows: (3).
[0132] For example, the root mean square (RMS) of the samples in a single frame of a channel after downmixing can be calculated to achieve energy calculation. For instance, the RMS of the samples in a single frame of each channel can be calculated, and the resulting RMS of the samples in a single frame of each channel can be determined as the energy parameter of the corresponding channel. For example, assuming the mixed audio signal includes three channels, such as channel a, channel b, and channel c, the RMS of the samples in a single frame of each of these three channels can be calculated. The RMS of the samples in a single frame of channel a can be determined as the energy parameter of channel a, the RMS of the samples in a single frame of channel b can be determined as the energy parameter of channel b, and the RMS of the samples in a single frame of channel c can be determined as the energy parameter of channel c. As an example, the formula for calculating the RMS of the samples is as follows: (4).
[0133] Alternatively, the above-mentioned energy parameters can also be calculated in other ways, and this disclosure does not limit or elaborate on this.
[0134] In step S2105, the encoding device 101 performs bit allocation on multiple channels based on the sound field analysis parameters of the mixed-format audio signal and the energy parameters of one or more channels to obtain bit allocation parameters.
[0135] In some embodiments, the bit allocation parameters described above are used to describe the number of bits allocated to each of the plurality of channels.
[0136] In some embodiments, the encoding device 101 can allocate bits to the multiple channels using the available number of bits based on the sound field analysis parameters of the mixed-format audio signal and the energy parameters of one or more channels among the multiple channels, thereby obtaining bit allocation parameters.
[0137] In some embodiments, taking the above-mentioned mixed-format audio signal including object-based audio signals and scene-based audio signals as an example, the encoding device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of object-based audio signals and scene-based audio signals according to the sound field analysis parameters of the mixed-format audio signal and the energy parameters of one or more channels in the multiple channels. For example, the object-based audio signal is the main sound element in the sound field, and the scene-based audio signal is the ambient background sound element in the sound field. The expected reconstructed sound field requires the object-based audio signal to have lower distortion, while the scene-based audio signal, as ambient background sound, can tolerate a certain degree of distortion. Therefore, more bits can be allocated to the object-based audio signal than to the scene-based audio signal during bit allocation. For example, the number of bits allocated to any channel in the object-based audio signal is more than the number of bits allocated to any channel in the scene-based audio signal. The encoding device 101 can also adjust the number of bits allocated to the corresponding channel in combination with the energy parameters of one or more channels in the mixed-format audio signal. For example, the channel with a larger energy parameter is allocated a relatively larger number of bits.
[0138] In some embodiments, taking the above-mentioned mixed-format audio signal as including object-based audio signals and channel-based audio signals as an example, the encoding device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of object-based audio signals and channel-based audio signals according to the sound field analysis parameters of the mixed-format audio signal and the energy parameters of one or more channels in the multiple channels. For example, the object-based audio signal is the main sound element in the sound field, and the channel-based audio signal is the ambient background sound element in the sound field. The expected reconstructed sound field requires the object-based audio signal to have less distortion, while the channel-based audio signal, as ambient background sound, can allow a certain degree of distortion. Therefore, more bits can be allocated to the object-based audio signal than to the channel-based audio signal during bit allocation. For example, the number of bits allocated to any channel in the object-based audio signal is more than the number of bits allocated to any channel in the channel-based audio signal. The encoding device 101 can also adjust the number of bits allocated to the corresponding channel based on the energy parameters of each channel in the mixed audio signal. For example, the channel with the larger energy parameter is allocated more bits.
[0139] In some embodiments, taking the above-mentioned mixed-format audio signal as including scene-based audio signals and channel-based audio signals as an example, the encoding device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of scene-based audio signals and channel-based audio signals according to the sound field analysis parameters of the mixed-format audio signal and the energy parameters of one or more channels in the multiple channels. For example, the channel-based audio signal is the main sound element in the sound field, while the scene-based audio signal is the ambient background sound element in the sound field. The expected reconstructed sound field requires that the channel-based audio signal has less distortion, while the scene-based audio signal, as ambient background sound, can tolerate a certain degree of distortion. Therefore, more bits can be allocated to the channel-based audio signal than to the scene-based audio signal during bit allocation. For example, the number of bits allocated to any channel in the channel-based audio signal is more than the number of bits allocated to any channel in the scene-based audio signal. The encoding device 101 can also adjust the number of bits allocated to the corresponding channel by combining the energy parameters of one or more channels in the mixed format audio signal. For example, the channel with the larger energy parameter is allocated more bits.
[0140] In some embodiments, taking the above-mentioned mixed-format audio signal including an object-based audio signal and a metadata-based 3D audio signal as an example, the encoding device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of the object-based audio signal and the metadata-based 3D audio signal according to the sound field analysis parameters of the mixed-format audio signal and the energy parameters of one or more channels in the multiple channels. For example, the object-based audio signal is the main sound element in the sound field, and the metadata-based 3D audio signal is the ambient background sound element in the sound field. The expected reconstructed sound field requires the object-based audio signal to have lower distortion, while the metadata-based 3D audio signal, as ambient background sound, can allow a certain degree of distortion. Therefore, more bits can be allocated to the object-based audio signal than to the metadata-based 3D audio signal during bit allocation. For example, the number of bits allocated to any channel in the object-based audio signal is more than the number of bits allocated to any channel in the metadata-based 3D audio signal. The encoding device 101 can also adjust the number of bits allocated to the corresponding channel by combining the energy parameters of one or more channels in the mixed format audio signal. For example, the channel with the larger energy parameter is allocated more bits.
[0141] In some embodiments, taking the above-mentioned mixed-format audio signal including metadata-based 3D audio signals and channel-based audio signals as an example, the encoding device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of metadata-based 3D audio signals and channel-based audio signals according to the sound field analysis parameters of the mixed-format audio signal and the energy parameters of one or more channels among the multiple channels. For example, the channel-based audio signal is the main sound element in the sound field, and the metadata-based 3D audio signal is the ambient background sound element in the sound field. The expected reconstructed sound field requires that the channel-based audio signal has less distortion, while the metadata-based 3D audio signal, as ambient background sound, can allow a certain degree of distortion. Therefore, more bits can be allocated to the channel-based audio signal than to the metadata-based 3D audio signal during bit allocation. For example, the number of bits allocated to any channel in the channel-based audio signal is more than the number of bits allocated to any channel in the metadata-based 3D audio signal. The encoding device 101 can also adjust the number of bits allocated to the corresponding channel by combining the energy parameters of one or more channels in the mixed format audio signal. For example, the channel with the larger energy parameter is allocated more bits.
[0142] In some embodiments, the number of bits allocated to the channel corresponding to the main image in the object-based audio signal is greater than the number of bits allocated to the channel corresponding to the secondary image in the object-based audio signal. That is, in the object-based audio signal, each object signal has a different priority in the mixed format audio signal. For example, in a live concert, the lead singer's voice signal has a higher priority than the backing vocalist's voice signal, or in other words, the lead singer's voice signal is more important than the backing vocalist's voice signal, and therefore the lead singer's object audio signal will be allocated more bits than the backing vocalist's voice signal.
[0143] It should be noted that the hybrid audio signal composition method shown in the above embodiments is to facilitate those skilled in the art to understand how to allocate bits to each channel to be encoded based on sound field analysis parameters and energy parameters. In other words, bit allocation can also be performed on hybrid audio signals composed of signals of other formats based on the above sound field analysis parameters and energy parameters, which will not be elaborated here.
[0144] In some embodiments, the number of bits allocated to each channel in the different audio signal formats in the above-described mixed format audio signal may be evenly distributed, or the number of bits may be distributed based on other factors (such as primary and secondary factors). This disclosure does not limit this, nor will it elaborate further.
[0145] In step S2106, the encoding device 101 encodes multiple channels according to the bit allocation parameters to obtain channel signal encoding parameters, and writes the bit allocation parameters and channel signal encoding parameters into the bit stream.
[0146] In some embodiments, the encoding device 101 can encode multiple channels using the corresponding encoding kernel based on the bit allocation parameters to obtain channel signal encoding parameters, and write the bit allocation parameters and channel signal encoding parameters into the bit stream.
[0147] In some embodiments, the encoding kernels corresponding to different audio signal formats can be the same or different. For example, all channels in a mixed-format audio signal can be encoded using a single encoding kernel based on bit allocation parameters. Alternatively, for example, the mixed-format audio signal includes two different audio signal formats, such as an object-based audio signal and a scene-based audio signal. Assuming the object-based audio signal includes channels a and b, and the scene-based audio signal includes channel c, the bit allocation parameters for channels a, b, and c are 3, 2, and 1, respectively (i.e., channel a is allocated 3 bits, channel b is allocated 2 bits, and channel c is allocated 1 bit). The encoding device 101 can encode channel a using the encoding kernel corresponding to the object-based audio signal with 3 bits, encode channel b using the encoding kernel corresponding to the object-based audio signal with 2 bits, and encode channel c using the encoding kernel corresponding to the scene-based audio signal with 1 bit.
[0148] In some embodiments, each channel in the mixed-format audio signal corresponds to a coding core. The coding cores corresponding to each channel can be the same or different. In this way, each channel in the mixed-format audio signal can be encoded using the corresponding coding core based on the allocated number of bits.
[0149] Step S2107: Encoding device 101 sends the bit stream.
[0150] In some embodiments, the encoding device 101 may send the bit stream to the decoding device 102, or in other words, the encoding device 101 sends the bit stream to the decoding device 102.
[0151] In some embodiments, the encoding device 101 may send the bitstream to the decoding device 102 based on multiplex transmission (MUX). In some embodiments, the decoding device 102 receives the bitstream. Optionally, the decoding device 102 receives the bitstream sent by the encoding device 101. The bitstream is used by the decoding device 102 to reconstruct the mixed-format audio signal described above.
[0152] In step S2108, the decoding device 102 performs bitstream parsing on the received bitstream to obtain the bit allocation parameters and channel signal encoding parameters of each channel in the multiple channels of the mixed-format audio signal.
[0153] In step S2109, the decoding device 102 reconstructs the mixed-format audio signal based on the bit allocation parameters and channel signal encoding parameters of each channel.
[0154] In some embodiments, the decoding device 102 decodes the channel signal encoding parameters based on the bit allocation parameters of each channel; the decoding device 102 then reconstructs the mixed-format audio signal using the decoded channels.
[0155] In some embodiments, the names of information, etc., are not limited to the names described in the embodiments. Terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "symbol", "codebook", "codeword", "codepoint", "bit", "data", "program", and "chip" can be used interchangeably.
[0156] In some embodiments, “get,” “obtain,” “receive,” “transmit,” “bidirectional transmission,” and “send and / or receive” can be used interchangeably and can be interpreted as receiving from other entities, obtaining from protocols, obtaining from higher layers, obtaining through self-processing, or autonomous implementation, among other meanings.
[0157] In some embodiments, terms such as “send,” “transmit,” “report,” “distribute,” “transfer,” “bidirectional transmission,” “send and / or receive” can be used interchangeably.
[0158] In some embodiments, "bit" and "number of bits" can be used interchangeably.
[0159] In some embodiments, "encoding end device", "encoder" and "encoding end" can be used interchangeably.
[0160] In some embodiments, "decoding device", "decoder", and "decoding end" can be used interchangeably.
[0161] In some embodiments, "mixed format audio signal" and "mixed format audio signal" can be used interchangeably.
[0162] In some embodiments, “audio channel signal”, “audio channel signal”, and “audio channel” can be used interchangeably.
[0163] In some embodiments, "spatial audio signal based on auxiliary metadata" and "three-dimensional audio signal based on metadata" can be used interchangeably.
[0164] In some embodiments, "bitstream" and "encoded bitstream" can be used interchangeably. The encoded bitstream can refer to a bitstream that includes the bit allocation parameters and the channel signal encoding parameters described above.
[0165] The method involved in the embodiments of this disclosure may include at least one of steps S2101 to S2109. For example, step S2101 + step S2102 + step S2103 + step S2104 + step S2105 can be implemented as an independent embodiment, step S2101 + step S2103 + step S2104 + step S2105 can be implemented as an independent embodiment, step S2101 + step S2102 + step S2103 + step S2104 + step S2105 + step S2106 + step S2107 can be implemented as an independent embodiment, and step S2101 + step S2103 + step S2104 + step S2105 + step S2106 + step S2109 can be implemented as an independent embodiment. Step 2107 can be implemented as an independent embodiment, and steps S2108 and S2109 can be implemented as an independent embodiment. Steps S2101, S2102, S2103, S2104, S2105, S2106, S2107, S2108, and S2109 can be implemented as an independent embodiment, but are not limited thereto.
[0166] In some embodiments, steps S2103 and S2104 may be performed in an alternate order or simultaneously.
[0167] In some embodiments, steps S2106, S2107, S2108, and S2109 are optional and one or more of these steps may be omitted or substituted in different embodiments.
[0168] In some embodiments, steps S2102, S2106, S2107, S2108, and S2109 are optional and one or more of these steps may be omitted or substituted in different embodiments.
[0169] In some embodiments, steps S2108 and S2109 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0170] In some embodiments, steps S2102, S2108, and S2109 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0171] In some embodiments, steps S2101, S2102, S2103, S2104, S2105, S2106, and S2107 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0172] In some embodiments, step S2102 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0173] In some embodiments, see Figure 2A Other optional implementation methods described before or after the corresponding instruction manual.
[0174] Figure 2B This is an interactive schematic diagram illustrating a signal processing method according to an embodiment of the present disclosure. For example... Figure 2B As shown, the signal processing method disclosed herein can be applied to a communication system 100, and the method includes, but is not limited to, the following steps.
[0175] In step S2201, the encoding device 101 acquires the mixed-format audio signal.
[0176] For optional implementations of step S2201, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2101, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0177] In step S2202, the encoding device 101 performs filtering processing on the mixed-format audio signal.
[0178] For optional implementations of step S2202, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2102, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0179] In step S2203, the encoding device 101 determines the control parameters of the mixed-format audio signal.
[0180] In some embodiments, the control parameters described above may include a second parameter and / or a third parameter. In some embodiments, the second parameter is used to describe the priority order of the various channels in the mixed-format audio signal. Optionally, the priority may be manually set or obtained through processing other than encoding.
[0181] In some embodiments, the third parameter described above is used to describe whether each of the above audio channels is plot-driven or non-plot-driven.
[0182] In some embodiments, the encoding device 101 can determine the control parameters of the mixed-format audio signal based on the functional division of each channel during acquisition by the audio signal acquisition device, wherein the audio signal acquisition device is used to acquire the mixed-format audio signal. In some embodiments, "audio signal acquisition device" and "acquisition terminal" can be interchanged.
[0183] In some embodiments, the encoding device 101 can determine the control parameters of the above-mentioned mixed-format audio signal based on the setting parameters required for the desired sound field reconstructed by the decoding device 102.
[0184] In some embodiments, the encoding device 101 can determine the control parameters of the above-mentioned mixed-format audio signal based on the functional division of each channel during audio signal acquisition by the audio signal acquisition device and the setting parameters required for the desired sound field reconstructed by the decoding device 102.
[0185] For example, the aforementioned control parameters refer to the priority order of each channel in the mixed-format audio signal, which is pre-set based on external information. This order can be determined according to the functional division of each channel during acquisition, or it can be based on the settings required for the desired sound field reconstructed at the playback end. For instance, the lead singer's voice in a live concert can be set as the most important channel. Similarly, to emphasize the sound of a particular instrument (such as a violin) in the reconstructed sound field, the violin's sound can be set as the most important channel signal.
[0186] In step S2204, the encoding device 101 determines the energy parameters of one or more of the multiple channels.
[0187] For optional implementations of step S2204, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2104, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0188] In step S2205, the encoding device 101 performs bit allocation on multiple channels based on the energy parameters of one or more channels and the control parameters of the mixed-format audio signal to obtain bit allocation parameters.
[0189] In some embodiments, the bit allocation parameters described above are used to describe the number of bits allocated to each of the plurality of channels.
[0190] In some embodiments, the encoding device 101 can allocate bits to the multiple channels using the available number of bits based on the control parameters of the mixed-format audio signal and the energy parameters of each channel in the multiple channels, thereby obtaining the bit allocation parameters.
[0191] In some embodiments, taking the above-mentioned mixed-format audio signal as including object-based audio signals and scene-based audio signals as an example, the encoding device 101 can perform bit allocation parameters on multiple channels in the mixed-format audio signal composed of object-based audio signals and scene-based audio signals according to the control parameters of the mixed-format audio signal and the energy parameters of each channel in the multiple channels. For example, the lead singer's voice in a concert can be set as the most important channel. Bit allocation parameters can be obtained by performing bit allocation on multiple channels in the mixed-format audio signal composed of object-based audio signals and scene-based audio signals according to the priority order of each channel's role in the mixed-format audio signal and the energy parameters of each channel in the multiple channels. For example, the number of bits allocated to any channel in the object-based audio signal is more than the number of bits allocated to any channel in the scene-based audio signal. The encoding device 101 can also adjust the number of bits allocated to the corresponding channel in combination with the energy parameters of each channel in the mixed-format audio signal; for example, the channel with a larger energy parameter is allocated a relatively larger number of bits.
[0192] In some embodiments, taking the above-mentioned mixed-format audio signal as including object-based audio signal and channel-based audio signal as an example, the encoding device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of object-based audio signal and channel-based audio signal according to the control parameters of the mixed-format audio signal and the energy parameters of each channel in the multiple channels.
[0193] In some embodiments, taking the above-mentioned mixed-format audio signal as including scene-based audio signal and channel-based audio signal as an example, the encoding device 101 can perform bit allocation parameters on multiple channels in the mixed-format audio signal composed of scene-based audio signal and channel-based audio signal according to the control parameters of the mixed-format audio signal and the energy parameters of each channel in the multiple channels.
[0194] In some embodiments, taking the above-mentioned mixed-format audio signal as including object-based audio signal and metadata-based three-dimensional audio signal as an example, the encoding end device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of object-based audio signal and metadata-based three-dimensional audio signal according to the control parameters of the mixed-format audio signal and the energy parameters of each channel in multiple channels.
[0195] In some embodiments, taking the above-mentioned mixed-format audio signal as including metadata-based three-dimensional audio signal and channel-based audio signal as an example, the encoding device 101 can perform bit allocation parameters on multiple channels in the mixed-format audio signal composed of metadata-based three-dimensional audio signal and channel-based audio signal according to the control parameters of the mixed-format audio signal and the energy parameters of each channel in the multiple channels.
[0196] In some embodiments, the number of bits allocated to the channel corresponding to the main image in the object-based audio signal is greater than the number of bits allocated to the channel corresponding to the secondary image in the object-based audio signal. That is, in the object-based audio signal, each object signal has a different priority in the mixed format audio signal. For example, in a live concert, the lead singer's voice signal has a higher priority than the backing vocalist's voice signal, or in other words, the lead singer's voice signal is more important than the backing vocalist's voice signal, and therefore the lead singer's object audio signal will be allocated more bits than the backing vocalist's voice signal.
[0197] It should be noted that the mixed-format audio signal composition method shown in the above embodiments is to facilitate those skilled in the art to understand how to allocate bits to each channel to be encoded according to control parameters and energy parameters. In other words, bit allocation can also be performed on mixed-format audio signals composed of signals of other formats according to the above control parameters and energy parameters, which will not be elaborated here.
[0198] In some embodiments, the number of bits allocated to each channel in the different audio signal formats in the above-described mixed format audio signal may be evenly distributed, or the number of bits may be distributed based on other factors (such as primary and secondary factors). This disclosure does not limit this, nor will it elaborate further.
[0199] In step S2206, the encoding device 101 encodes multiple channels according to the bit allocation parameters to obtain channel signal encoding parameters, and writes the bit allocation parameters and channel signal encoding parameters into the bit stream.
[0200] For optional implementations of step S2206, please refer to [link / reference]. Figure 2AOptional implementation methods of step S2106, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0201] Step S2207: Encoding device 101 sends the bit stream.
[0202] For optional implementations of step S2207, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2107, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0203] In step S2208, the decoding device 102 performs bitstream parsing on the received bitstream to obtain the bit allocation parameters and channel signal encoding parameters of each channel in the multiple channels of the mixed-format audio signal.
[0204] For optional implementations of step S2208, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2108, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0205] In step S2209, the decoding device 102 reconstructs the mixed-format audio signal based on the bit allocation parameters and channel signal encoding parameters of each channel.
[0206] For optional implementations of step S2209, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2109, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0207] The method involved in the embodiments of this disclosure may include at least one of steps S2201 to S2209. For example, step S2201 + step S2202 + step S2203 + step S2204 + step S2205 can be implemented as an independent embodiment, step S2201 + step S2203 + step S2204 + step S2205 can be implemented as an independent embodiment, step S2201 + step S2202 + step S2203 + step S2204 + step S2205 + step S2206 + step S2207 can be implemented as an independent embodiment, and step S2201 + step S2203 + step S2204 + step S2205 + step S2206 + step S2209 can be implemented as an independent embodiment. Step 2207 can be implemented as an independent embodiment, and steps S2208 and S2209 can be implemented as an independent embodiment, as can steps S2201, S2202, S2203, S2204, S2205, S2206, S2207, S2208, and S2209, but are not limited thereto.
[0208] In some embodiments, steps S2203 and S2204 may be performed in an alternate order or simultaneously.
[0209] In some embodiments, steps S2206, S2207, S2208, and S2209 are optional and one or more of these steps may be omitted or substituted in different embodiments.
[0210] In some embodiments, steps S2202, S2206, S2207, S2208, and S2209 are optional and one or more of these steps may be omitted or substituted in different embodiments.
[0211] In some embodiments, steps S2208 and S2209 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0212] In some embodiments, steps S2202, S2208, and S2209 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0213] In some embodiments, steps S2201, S2202, S2203, S2204, S2205, S2206, and S2207 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0214] In some embodiments, step S2202 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0215] In some embodiments, see Figure 2B Other optional implementation methods described before or after the corresponding instruction manual.
[0216] Figure 2C This is an interactive schematic diagram illustrating a signal processing method according to an embodiment of the present disclosure. For example... Figure 2C As shown, the signal processing method disclosed herein can be applied to a communication system 100, and the method includes, but is not limited to, the following steps.
[0217] In step S2301, the encoding device 101 acquires the mixed-format audio signal.
[0218] For optional implementations of step S2301, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2101, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0219] In step S2302, the encoding device 101 performs filtering processing on the mixed-format audio signal.
[0220] For optional implementations of step S2302, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2102, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0221] In step S2303, the encoding device 101 determines the sound field analysis parameters and / or control parameters of the mixed-format audio signal.
[0222] In some embodiments, the encoding end device 101 determines sound field analysis parameters for the mixed-format audio signal. Optional implementations can be found in [reference needed]. Figure 2A Optional implementation methods of step S2103, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0223] In some embodiments, the encoding end device 101 determines control parameters for the mixed-format audio signal. Optional implementations can be found in [reference needed]. Figure 2BOptional implementation methods of step S2203, and Figure 2B Other related parts in the embodiments involved will not be described in detail here.
[0224] In some embodiments, the encoding device 101 determines sound field analysis parameters and control parameters for the mixed-format audio signal. Optional implementations can be found in [reference needed]. Figure 2A Optional implementation methods of step S2103 Figure 2B Optional implementation methods of step S2203, and Figure 2A Other related parts in the embodiments involved, and Figure 2B Other related parts in the embodiments involved will not be described in detail here.
[0225] In step S2304, the encoding device 101 performs bit allocation on multiple channels based on the sound field analysis parameters and / or control parameters of the mixed-format audio signal to obtain bit allocation parameters.
[0226] In some embodiments, the encoding device 101 allocates bits to the plurality of channels using the available number of bits based on the sound field analysis parameters and / or control parameters of the mixed-format audio signal, thereby obtaining bit allocation parameters.
[0227] In some embodiments, the encoding device 101 performs bit allocation on multiple channels based on the sound field analysis parameters of the mixed-format audio signal to obtain bit allocation parameters.
[0228] In one implementation, taking the aforementioned mixed-format audio signal comprising object-based audio signals and scene-based audio signals as an example, the encoding device 101 can obtain bit allocation parameters by bit-allocating multiple channels in the mixed-format audio signal composed of object-based audio signals and scene-based audio signals according to the sound field analysis parameters of the mixed-format audio signal. For example, the object-based audio signal is the main sound element in the sound field, while the scene-based audio signal is the ambient background sound element in the sound field. The expected reconstructed sound field requires the object-based audio signal to have lower distortion, while the scene-based audio signal, as ambient background sound, can tolerate a certain degree of distortion. Therefore, more bits can be allocated to the object-based audio signal compared to the scene-based audio signal during bit allocation. For example, the number of bits allocated to any channel in the object-based audio signal is greater than the number of bits allocated to any channel in the scene-based audio signal.
[0229] In another implementation, taking the aforementioned mixed-format audio signal, which includes both object-based and channel-based audio signals, as an example, the encoding device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of object-based and channel-based audio signals based on the sound field analysis parameters of the mixed-format audio signal. For example, the object-based audio signal is the main sound element in the sound field, while the channel-based audio signal is the ambient background sound element in the sound field. The expected reconstructed sound field requires the object-based audio signal to have lower distortion, while the channel-based audio signal, as ambient background sound, can tolerate a certain degree of distortion. Therefore, more bits can be allocated to the object-based audio signal compared to the channel-based audio signal during bit allocation. For example, the number of bits allocated to any channel in the object-based audio signal is greater than the number of bits allocated to any channel in the channel-based audio signal.
[0230] In another implementation, taking the aforementioned mixed-format audio signal, which includes scene-based audio signals and channel-based audio signals, as an example, the encoding device 101 can perform bit allocation parameters based on the sound field analysis parameters of the mixed-format audio signal to allocate bits to multiple channels within the mixed-format audio signal composed of scene-based and channel-based audio signals. For example, the channel-based audio signal is the main sound element in the sound field, while the scene-based audio signal is the ambient background sound element in the sound field. The expected reconstructed sound field requires lower distortion of the channel-based audio signal, while the scene-based audio signal, as ambient background sound, can tolerate a certain degree of distortion. Therefore, during bit allocation, more bits can be allocated to the channel-based audio signal compared to the scene-based audio signal. For example, the number of bits allocated to any channel in the channel-based audio signal is greater than the number of bits allocated to any channel in the scene-based audio signal.
[0231] In another implementation, taking the aforementioned mixed-format audio signal comprising object-based audio signals and metadata-based 3D audio signals as an example, the encoding device 101 can obtain bit allocation parameters by allocating bits to multiple channels in the mixed-format audio signal composed of object-based audio signals and metadata-based 3D audio signals according to the sound field analysis parameters of the mixed-format audio signal. For example, the object-based audio signal is the main sound element in the sound field, while the metadata-based 3D audio signal is the ambient background sound element in the sound field. The expected reconstructed sound field requires the object-based audio signal to have lower distortion, while the metadata-based 3D audio signal, as ambient background sound, can tolerate a certain degree of distortion. Therefore, more bits can be allocated to the object-based audio signal compared to the metadata-based 3D audio signal during bit allocation. For example, the number of bits allocated to any channel in the object-based audio signal is greater than the number of bits allocated to any channel in the metadata-based 3D audio signal.
[0232] In another implementation, taking the aforementioned mixed-format audio signal, which includes metadata-based 3D audio signals and channel-based audio signals, as an example, the encoding device 101 can perform bit allocation parameters based on the sound field analysis parameters of the mixed-format audio signal to allocate bits to multiple channels within the mixed-format audio signal composed of metadata-based 3D audio signals and channel-based audio signals. For example, the channel-based audio signal is the main sound element in the sound field, while the metadata-based 3D audio signal is the ambient background sound element in the sound field. The expected reconstructed sound field requires lower distortion of the channel-based audio signal, while the metadata-based 3D audio signal, as ambient background sound, can tolerate a certain degree of distortion. Therefore, during bit allocation, more bits can be allocated to the channel-based audio signal compared to the metadata-based 3D audio signal. For example, the number of bits allocated to any channel in the channel-based audio signal is greater than the number of bits allocated to any channel in the metadata-based 3D audio signal.
[0233] In some embodiments, the number of bits allocated to the channel corresponding to the main image in the object-based audio signal is greater than the number of bits allocated to the channel corresponding to the secondary image in the object-based audio signal. That is, in the object-based audio signal, each object signal has a different priority in the mixed format audio signal. For example, in a live concert, the lead singer's voice signal has a higher priority than the backing vocalist's voice signal, or in other words, the lead singer's voice signal is more important than the backing vocalist's voice signal, and therefore the lead singer's object audio signal will be allocated more bits than the backing vocalist's voice signal.
[0234] It should be noted that the mixed-format audio signal composition method shown in the above embodiments is to facilitate those skilled in the art to understand how to allocate bits to each channel to be encoded according to the sound field analysis parameters. In other words, the bit allocation method can also be used to allocate bits to a mixed-format audio signal composed of signals of other formats according to the above sound field analysis parameters, which will not be elaborated here.
[0235] In some embodiments, the encoding device 101 can perform bit allocation on multiple channels according to the control parameters of the mixed-format audio signal to obtain bit allocation parameters.
[0236] In one implementation, taking the above-mentioned mixed-format audio signal including object-based audio signal and scene-based audio signal as an example, the encoding end device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of object-based audio signal and scene-based audio signal according to the control parameters of the mixed-format audio signal.
[0237] In another implementation, taking the above-mentioned mixed-format audio signal as including object-based audio signals and channel-based audio signals as an example, the encoding device 101 can perform bit allocation parameters on multiple channels in the mixed-format audio signal composed of object-based audio signals and channel-based audio signals according to the control parameters of the mixed-format audio signal.
[0238] In another implementation, taking the above-mentioned mixed-format audio signal including scene-based audio signal and channel-based audio signal as an example, the encoding device 101 can perform bit allocation parameters on multiple channels in the mixed-format audio signal composed of scene-based audio signal and channel-based audio signal according to the control parameters of the mixed-format audio signal.
[0239] In another implementation, taking the above-mentioned mixed-format audio signal including object-based audio signal and metadata-based three-dimensional audio signal as an example, the encoding end device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of object-based audio signal and metadata-based three-dimensional audio signal according to the control parameters of the mixed-format audio signal.
[0240] In another implementation, taking the above-mentioned mixed-format audio signal including metadata-based three-dimensional audio signal and channel-based audio signal as an example, the encoding end device 101 can perform bit allocation parameters on multiple channels in the mixed-format audio signal composed of metadata-based three-dimensional audio signal and channel-based audio signal according to the control parameters of the mixed-format audio signal.
[0241] It should be noted that the mixed-format audio signal composition method shown in the above embodiments is to facilitate those skilled in the art to understand how to allocate bits to each channel to be encoded according to the control parameters. In other words, the mixed-format audio signal composed of signals of other formats can also be bit allocated according to the above control parameters, which will not be elaborated here.
[0242] In some embodiments, the encoding device 101 performs bit allocation on multiple channels based on the sound field analysis parameters and control parameters of the mixed-format audio signal to obtain bit allocation parameters.
[0243] In one implementation, taking the above-mentioned mixed-format audio signal including object-based audio signal and scene-based audio signal as an example, the encoding end device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of object-based audio signal and scene-based audio signal according to the sound field analysis parameters and control parameters of the mixed-format audio signal.
[0244] In another implementation, taking the above-mentioned mixed-format audio signal including object-based audio signals and channel-based audio signals as an example, the encoding device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of object-based audio signals and channel-based audio signals according to the sound field analysis parameters and control parameters of the mixed-format audio signal.
[0245] In another implementation, taking the above-mentioned mixed-format audio signal including scene-based audio signal and channel-based audio signal as an example, the encoding device 101 can perform bit allocation parameters on multiple channels in the mixed-format audio signal composed of scene-based audio signal and channel-based audio signal according to the sound field analysis parameters and control parameters of the mixed-format audio signal.
[0246] In another implementation, taking the above-mentioned mixed-format audio signal including object-based audio signal and metadata-based three-dimensional audio signal as an example, the encoding end device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of object-based audio signal and metadata-based three-dimensional audio signal according to the sound field analysis parameters and control parameters of the mixed-format audio signal.
[0247] In another implementation, taking the above-mentioned mixed-format audio signal including metadata-based three-dimensional audio signal and channel-based audio signal as an example, the encoding end device 101 can perform bit allocation parameters on multiple channels in the mixed-format audio signal composed of metadata-based three-dimensional audio signal and channel-based audio signal according to the sound field analysis parameters and control parameters of the mixed-format audio signal.
[0248] In some embodiments, the number of bits allocated to the channel corresponding to the main image in the object-based audio signal is greater than the number of bits allocated to the channel corresponding to the secondary image in the object-based audio signal. That is, in the object-based audio signal, each object signal has a different priority in the mixed format audio signal. For example, in a live concert, the lead singer's voice signal has a higher priority than the backing vocalist's voice signal, or in other words, the lead singer's voice signal is more important than the backing vocalist's voice signal, and therefore the lead singer's object audio signal will be allocated more bits than the backing vocalist's voice signal.
[0249] It should be noted that the mixed-format audio signal composition method shown in the above embodiments is to facilitate those skilled in the art to understand how to allocate bits to each channel to be encoded according to the sound field analysis parameters and control parameters. In other words, the mixed-format audio signal composed of signals of other formats can also be bit allocated according to the above sound field analysis parameters and control parameters, which will not be elaborated here.
[0250] In step S2305, the encoding device 101 encodes multiple channels according to the bit allocation parameters to obtain channel signal encoding parameters, and writes the bit allocation parameters and channel signal encoding parameters into the bit stream.
[0251] For optional implementations of step S2305, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2106, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0252] Step S2306: Encoding device 101 sends the bit stream.
[0253] For optional implementations of step S2306, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2107, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0254] In step S2307, the decoding device 102 performs bitstream parsing on the received bitstream to obtain the bit allocation parameters and channel signal encoding parameters of each channel in the multiple channels of the mixed-format audio signal.
[0255] For optional implementations of step S2307, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2108, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0256] In step S2308, the decoding device 102 reconstructs the mixed-format audio signal based on the bit allocation parameters and channel signal encoding parameters of each channel.
[0257] For optional implementations of step S2308, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2109, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0258] The method involved in the embodiments of this disclosure may include at least one of steps S2301 to S2308. For example, step S2301 + step S2302 + step S2303 + step S2304 can be implemented as an independent embodiment, step S2301 + step S2303 + step S2304 can be implemented as an independent embodiment, step S2301 + step S2302 + step S2303 + step S2304 + step S2305 + step S2306 + step S2306 can be implemented as an independent embodiment, and step S2301 + step S2303 + step S2304 + step S2305 + step S2306 + step S2308 can be implemented as an independent embodiment. Step 2306 can be implemented as an independent embodiment, and steps S2307 and S2308 can be implemented as an independent embodiment, as can steps S2301, S2302, S2303, S2304, S2305, S2306, S2307, and S2308, but are not limited thereto.
[0259] In some embodiments, steps S2305, S2306, S2307, and S2308 are optional and one or more of these steps may be omitted or substituted in different embodiments.
[0260] In some embodiments, steps S2302, S2305, S2306, S2307, and S2308 are optional and one or more of these steps may be omitted or substituted in different embodiments.
[0261] In some embodiments, steps S2307 and S2308 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0262] In some embodiments, steps S2302, S2307, and S2308 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0263] In some embodiments, steps S2301, S2302, S2303, S2304, S2305, and S2306 are optional and one or more of these steps may be omitted or substituted in different embodiments.
[0264] In some embodiments, step S2302 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0265] In some embodiments, see Figure 2C Other optional implementation methods described before or after the corresponding instruction manual.
[0266] Figure 2D This is an interactive schematic diagram illustrating a signal processing method according to an embodiment of the present disclosure. For example... Figure 2D As shown, the signal processing method disclosed herein can be applied to a communication system 100, and the method includes, but is not limited to, the following steps.
[0267] In step S2401, the encoding device 101 acquires the mixed-format audio signal.
[0268] For optional implementations of step S2401, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2101, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0269] In step S2402, the encoding device 101 performs filtering processing on the mixed-format audio signal.
[0270] For optional implementations of step S2402, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2102, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0271] In step S2403, the encoding device 101 determines the sound field analysis parameters of the mixed-format audio signal.
[0272] For optional implementations of step S2403, please refer to [link / reference]. Figure 2AOptional implementation methods of step S2103, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0273] In step S2404, the encoding device 101 determines the energy parameters of one or more of the multiple channels.
[0274] For optional implementations of step S2404, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2104, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0275] In step S2405, the encoding device 101 determines the control parameters of the mixed-format audio signal.
[0276] For optional implementations of step S2405, please refer to [link / reference]. Figure 2B Optional implementation methods of step S2203, and Figure 2B Other related parts in the embodiments involved will not be described in detail here.
[0277] In step S2406, the encoding device 101 performs bit allocation on multiple channels based on the energy parameters of one or more channels, the control parameters of the mixed-format audio signal, and the sound field analysis parameters to obtain bit allocation parameters.
[0278] In some embodiments, the encoding device 101 can allocate bits to the multiple channels using the available number of bits based on the control parameters of the mixed-format audio signal, the sound field analysis parameters, and the energy parameters of each channel in the multiple channels, thereby obtaining the bit allocation parameters.
[0279] In some embodiments, taking the above-mentioned mixed-format audio signal including object-based audio signal and scene-based audio signal as an example, the encoding end device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of object-based audio signal and scene-based audio signal according to the sound field analysis parameters, control parameters and energy parameters of each channel in the multiple channels of the mixed-format audio signal.
[0280] In some embodiments, taking the above-mentioned mixed-format audio signal as including object-based audio signal and channel-based audio signal as an example, the encoding device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of object-based audio signal and channel-based audio signal according to the sound field analysis parameters, control parameters and energy parameters of each channel in the mixed-format audio signal.
[0281] In some embodiments, taking the above-mentioned mixed-format audio signal as including scene-based audio signal and channel-based audio signal as an example, the encoding device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of scene-based audio signal and channel-based audio signal according to the sound field analysis parameters, control parameters and energy parameters of each channel in the mixed-format audio signal.
[0282] In some embodiments, taking the above-mentioned mixed-format audio signal as including object-based audio signal and metadata-based three-dimensional audio signal as an example, the encoding end device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of object-based audio signal and metadata-based three-dimensional audio signal according to the sound field analysis parameters, control parameters and energy parameters of each channel in the mixed-format audio signal.
[0283] In some embodiments, taking the above-mentioned mixed-format audio signal as including metadata-based three-dimensional audio signal and channel-based audio signal as an example, the encoding device 101 can perform bit allocation parameters on multiple channels of the mixed-format audio signal composed of metadata-based three-dimensional audio signal and channel-based audio signal according to the sound field analysis parameters, control parameters and energy parameters of each channel in the mixed-format audio signal.
[0284] In some embodiments, the number of bits allocated to the channel corresponding to the main image in the object-based audio signal is greater than the number of bits allocated to the channel corresponding to the secondary image in the object-based audio signal. That is, in the object-based audio signal, each object signal has a different priority in the mixed format audio signal. For example, in a live concert, the lead singer's voice signal has a higher priority than the backing vocalist's voice signal, or in other words, the lead singer's voice signal is more important than the backing vocalist's voice signal, and therefore the lead singer's object audio signal will be allocated more bits than the backing vocalist's voice signal.
[0285] It should be noted that the hybrid audio signal composition method shown in the above embodiments is to facilitate those skilled in the art to understand how to allocate bits to each channel to be encoded according to the sound field analysis parameters, control parameters and energy parameters. In other words, the bit allocation method can also be used to allocate bits to hybrid audio signals composed of signals of other formats according to the above sound field analysis parameters, control parameters and energy parameters, which will not be elaborated here.
[0286] In some embodiments, the number of bits allocated to each channel in the different audio signal formats in the above-described mixed format audio signal may be evenly distributed, or the number of bits may be distributed based on other factors (such as primary and secondary factors). This disclosure does not limit this, nor will it elaborate further.
[0287] In step S2407, the encoding device 101 encodes multiple channels according to the bit allocation parameters to obtain channel signal encoding parameters, and writes the bit allocation parameters and channel signal encoding parameters into the bit stream.
[0288] For optional implementations of step S2407, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2106, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0289] Step S2408: Encoding device 101 sends the bit stream.
[0290] For optional implementations of step S2408, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2107, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0291] In step S2409, the decoding device 102 performs bitstream parsing on the received bitstream to obtain the bit allocation parameters and channel signal encoding parameters of each channel in the multiple channels of the mixed-format audio signal.
[0292] For optional implementations of step S2409, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2108, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0293] In step S2410, the decoding device 102 reconstructs the mixed-format audio signal based on the bit allocation parameters and channel signal encoding parameters of each channel.
[0294] Optional implementations of step S2410 can be found in [reference]. Figure 2A Optional implementation methods of step S2109, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0295] The method involved in the embodiments of this disclosure may include at least one of steps S2401 to S2406. For example, step S2401+S2402+S2403+S2404+S2405+S2406 can be implemented as an independent embodiment, step S2401+S2403+S2404+S2405+S2406 can be implemented as an independent embodiment, step S2401+S2402+S2403+S2404+S2405+S2406+S2407+S2408 can be implemented as an independent embodiment, and step S2401+S2403+S2404+S2405+S2406+S2407+S2408 can be implemented as an independent embodiment. Steps S2407 and S2408 can be implemented as independent embodiments, as can steps S2409 and S2410, as can steps S2401, S2402, S2403, S2404, S2405, S2406, S2407, S2408, S2409, and S2410, but are not limited thereto.
[0296] In some embodiments, steps S2403, S2404, and S2405 may be performed in an alternate order or simultaneously.
[0297] In some embodiments, steps S2407, S2408, S2409, and S2410 are optional and one or more of these steps may be omitted or substituted in different embodiments.
[0298] In some embodiments, steps S2402, S2407, S2408, S2409, and S2410 are optional and one or more of these steps may be omitted or substituted in different embodiments.
[0299] In some embodiments, steps S2409 and S2410 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0300] In some embodiments, steps S2402, S2409, and S2410 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0301] In some embodiments, steps S2401, S2402, S2403, S2404, S2405, S2406, S2407, and S2408 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0302] In some embodiments, step S2402 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0303] In some embodiments, see Figure 2D Other optional implementation methods described before or after the corresponding instruction manual.
[0304] Figure 3 This is a schematic flowchart illustrating a signal processing method according to an embodiment of the present disclosure. Figure 3 As shown, the embodiments of this disclosure relate to a signal processing method, which can be executed by an encoding end device 101. The method may include, but is not limited to, the following steps.
[0305] Step S3101: Obtain the audio signal in the mixed format.
[0306] In some embodiments, the above-described mixed-format audio signal is an audio signal comprising multiple channels.
[0307] For optional implementations of step S3101, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2101, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0308] Step S3102: Filter the mixed-format audio signal.
[0309] Optional implementations of step S3102 can be found in [reference]. Figure 2A Optional implementation methods of step S2102, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0310] Step S3103: Determine the first parameter of the mixed-format audio signal and / or the energy parameter of one or more channels in the multiple channels.
[0311] In some embodiments, the first parameter described above may include at least one of the following: sound field analysis parameters; control parameters. In one implementation, the first parameter described above may include sound field analysis parameters. In another implementation, the first parameter described above may include control parameters. In yet another implementation, the first parameter described above may include both sound field analysis parameters and control parameters.
[0312] In some embodiments, the sound field analysis parameters described above include at least one of the following: the number of sound images; the hierarchy of sound images; and ambient background sound. In one implementation, the sound field analysis parameters include any one of the number of sound images, the hierarchy of sound images, and ambient background sound. In another implementation, the sound field analysis parameters include any two of the number of sound images, the hierarchy of sound images, and ambient background sound. In yet another implementation, the sound field analysis parameters include the number of sound images, the hierarchy of sound images, and ambient background sound.
[0313] In some embodiments, the control parameters described above may include a second parameter and / or a third parameter. In some embodiments, the second parameter is used to describe the priority order of the various channels in the mixed-format audio signal. Optionally, the priority may be manually set or obtained through processing other than encoding.
[0314] In some embodiments, the third parameter described above is used to describe whether each of the above audio channels is plot-driven or non-plot-driven.
[0315] In some embodiments, a first parameter of the mixed-format audio signal may be determined.
[0316] For example, sound field analysis parameters for mixed-format audio signals can be determined. Optional implementations can be found in [reference needed]. Figure 2A Optional implementation methods of step S2103, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0317] For example, control parameters for mixed-format audio signals can be determined. Optional implementations can be found in [reference needed]. Figure 2B Optional implementation methods of step S2203, and Figure 2B Other related parts in the embodiments involved will not be described in detail here.
[0318] For example, sound field analysis parameters and control parameters for mixed-format audio signals can be determined. Optional implementations can be found in [reference needed]. Figure 2A Optional implementation methods of step S2103 Figure 2B Optional implementation methods of step S2203, and Figure 2A Other related parts in the embodiments involved, and Figure 2B Other related parts in the embodiments involved will not be described in detail here.
[0319] In some embodiments, the energy parameters of one or more of the plurality of channels can be determined. Optional implementations can be found in [reference needed]. Figure 2AOptional implementation methods of step S2104, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0320] In some embodiments, a first parameter of the mixed-format audio signal and the energy parameter of one or more of the plurality of channels can be determined.
[0321] For example, sound field analysis parameters of a mixed-format audio signal and energy parameters of one or more of the aforementioned channels can be determined. Optional implementations can be found in [reference needed]. Figure 2A Optional implementation methods of step S2103 Figure 2A Optional implementation methods of step S2104, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0322] For example, control parameters for the mixed-format audio signal and energy parameters for one or more of the aforementioned multiple channels can be determined. Optional implementations can be found in [reference needed]. Figure 2B Optional implementation methods of step S2203 Figure 2A Optional implementation methods of step S2104, and Figure 2B Other related parts in the embodiments involved, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0323] For example, sound field analysis parameters, control parameters, and energy parameters of one or more channels among the multiple channels can be determined for a mixed-format audio signal. Optional implementations can be found in [reference needed]. Figure 2A Optional implementation methods for step S2103, optional implementation methods for step S2104, Figure 2B Optional implementation methods of step S2203, and Figure 2A Other related parts in the embodiments involved, and Figure 2B Other related parts in the embodiments involved will not be described in detail here.
[0324] Step S3104: Based on the first parameter of the mixed-format audio signal and / or the energy parameter of one or more channels, perform bit allocation on the multiple channels to obtain bit allocation parameters.
[0325] In some embodiments, bit allocation parameters can be obtained by bit-allocating multiple channels based on a first parameter of the mixed-format audio signal.
[0326] For example, bit allocation parameters can be obtained by bit-allocating multiple channels based on the sound field analysis parameters of the mixed-format audio signal. Optional implementations can be found in [reference needed]. Figure 2COptional implementation methods of step S2304, and Figure 2C Other related parts in the embodiments involved will not be described in detail here.
[0327] For example, bit allocation parameters can be obtained by bit-allocating multiple channels based on control parameters of the mixed-format audio signal. Optional implementations can be found in [reference needed]. Figure 2C Optional implementation methods of step S2304, and Figure 2C Other related parts in the embodiments involved will not be described in detail here.
[0328] For example, bit allocation parameters can be obtained by bit-allocating multiple channels based on the sound field analysis parameters and control parameters of the mixed-format audio signal. Optional implementation methods can be found in [reference needed]. Figure 2C Optional implementation methods of step S2304, and Figure 2C Other related parts in the embodiments involved will not be described in detail here.
[0329] In some embodiments, bit allocation parameters can be obtained by bit allocation of multiple channels based on the energy parameters of one or more channels in a mixed-format audio signal.
[0330] In some embodiments, bit allocation parameters can be obtained by bit allocation of multiple channels based on a first parameter of the mixed-format audio signal and the energy parameters of one or more channels among the multiple channels.
[0331] For example, bit allocation parameters can be obtained by bit-allocating multiple channels based on the sound field analysis parameters of the mixed-format audio signal and the energy parameters of one or more channels. Optional implementations can be found in [reference needed]. Figure 2A Optional implementation methods of step S2105, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0332] For example, bit allocation parameters can be obtained by bit-allocating multiple channels based on control parameters of the mixed-format audio signal and energy parameters of one or more channels. Optional implementations can be found in [reference needed]. Figure 2B Optional implementation methods of step S2205, and Figure 2B Other related parts in the embodiments involved will not be described in detail here.
[0333] For example, bit allocation parameters can be obtained by bit-allocating multiple channels based on sound field analysis parameters, control parameters, and energy parameters of one or more channels in a mixed-format audio signal. Optional implementations can be found in [reference needed]. Figure 2D Optional implementation methods of step S2406, and Figure 2D Other related parts in the embodiments involved will not be described in detail here.
[0334] Step S3105: Encode multiple channels according to the bit allocation parameters to obtain channel signal encoding parameters, and write the bit allocation parameters and channel signal encoding parameters into the bit stream.
[0335] For optional implementations of step S3105, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2106, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0336] Step S3106: Send the bitstream.
[0337] For optional implementations of step S3106, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2107, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0338] The method involved in the embodiments of this disclosure may include at least one of steps S3101 to S3106. For example, steps S3101+S3102+S3103+S3104 can be implemented as an independent embodiment, steps S3101+S3103+S3104 can be implemented as an independent embodiment, steps S3101+S3102+S3103+S3104+S3105+S3106 can be implemented as an independent embodiment, and steps S3101+S3103+S3104+S3105+S3106 can be implemented as an independent embodiment, but are not limited thereto.
[0339] In some embodiments, steps S3105 and S3106 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0340] In some embodiments, steps S3102, S3105, and S3106 are optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0341] In some embodiments, step S3102 is optional, and one or more of these steps may be omitted or substituted in different embodiments.
[0342] Figure 4 This is a schematic flowchart illustrating a signal processing method according to an embodiment of the present disclosure. Figure 4As shown, the embodiments of this disclosure relate to a signal processing method, which can be executed by a decoding terminal device 102. The above method may include, but is not limited to, the following steps.
[0343] Step S4101: Receive the bitstream.
[0344] In some embodiments, the decoding device 102 receives a bitstream. Optionally, the decoding device 102 receives a bitstream sent by the encoding device 101. The bitstream is used by the decoding device 102 to reconstruct the mixed-format audio signal described above. Optional implementations of the bitstream can be found in the optional implementations of any of the above embodiments and other related parts of any of the above embodiments, and will not be repeated here.
[0345] Step S4102: Perform bitstream parsing on the bitstream to obtain the bit allocation parameters and channel signal encoding parameters of each channel in the multiple channels of the mixed-format audio signal.
[0346] For optional implementations of step S4102, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2108, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0347] Step S4103: Based on the bit allocation parameters and channel signal encoding parameters of each channel, reconstruct the mixed-format audio signal.
[0348] For optional implementations of step S4103, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2109, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0349] Figure 5 This is an interactive schematic diagram illustrating a signal processing method according to an embodiment of the present disclosure. For example... Figure 5 As shown, the method disclosed in this embodiment can be applied to the communication system 100, and the method includes, but is not limited to, the following steps.
[0350] In step S5101, the encoding device 101 acquires a mixed-format audio signal, which is an audio signal including multiple channels.
[0351] For optional implementations of step S5101, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2101, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0352] In step S5102, the encoding device 101 determines the first parameter of the mixed-format audio signal and / or the energy parameter of one or more channels in the multiple channels.
[0353] Optional implementations of step S5102 can be found in [reference]. Figure 2A Step S2103 Figure 2B Step S2203 Figure 2A Step S2104 Figure 3 Optional implementation methods of step S3103, and Figure 2A , Figure 2B , Figure 3 Other related parts in the embodiments involved will not be described in detail here.
[0354] In step S5103, the encoding device 101 performs bit allocation on multiple channels according to the first parameter of the mixed-format audio signal and / or the energy parameter of one or more channels to obtain bit allocation parameters.
[0355] For optional implementations of step S5103, please refer to [link / reference]. Figure 2A Step S2105 Figure 2B Step S2205 Figure 2C Step S2304 Figure 2D Step S2406 Figure 3 Optional implementation methods of step S3104, and Figure 2A , Figure 2B , Figure 2C , Figure 2D , Figure 3 Other related parts in the embodiments involved will not be described in detail here.
[0356] In step S5104, the encoding device 101 encodes multiple channels according to the bit allocation parameters to obtain channel signal encoding parameters, and writes the bit allocation parameters and channel signal encoding parameters into the bit stream.
[0357] For optional implementations of step S5104, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2106, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0358] Step S5105: Encoding device 101 sends the bit stream.
[0359] For optional implementations of step S5105, please refer to [link / reference]. Figure 2A Optional implementation methods of step S2107, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0360] In step S5106, the decoding device 102 performs bitstream parsing on the received bitstream to obtain the bit allocation parameters and channel signal encoding parameters of each channel in the multiple channels of the mixed-format audio signal.
[0361] Optional implementations of step S5106 can be found in [reference needed]. Figure 2A Optional implementation methods of step S2108, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0362] In step S5107, the decoding device 102 reconstructs the mixed-format audio signal based on the bit allocation parameters and channel signal encoding parameters of each channel.
[0363] Optional implementations of step S5107 can be found in [reference needed]. Figure 2A Optional implementation methods of step S2109, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0364] In some embodiments, the encoding and decoding processing of mixed-format audio signals disclosed herein is as follows: when a mixed-format audio signal is input to an encoder, the encoder performs energy calculation and sound field analysis on the mixed-format audio signal, allocates bits to each encoded channel signal based on the results of energy calculation and sound field analysis, and encodes the audio signal using the allocated bits to obtain a bitstream. The decoder decodes the bitstream to reconstruct the obtained mixed-format audio signal.
[0365] In some embodiments, the encoding and decoding process for mixed-format audio signals disclosed herein is as follows: when a mixed-format audio signal is input to an encoder, the encoder performs energy calculation on the mixed-format audio signal, allocates bits to each encoded channel signal based on control parameters and energy calculation results, encodes the audio signal using the allocated bits to obtain a bitstream, and the decoder decodes the bitstream to reconstruct the obtained mixed-format audio signal.
[0366] The purpose of this disclosure is to design and complete the encoding and decoding process of mixed-format audio signals, determine the encoding mode based on external control parameters and sound field analysis results, and encode the signal to achieve the desired reconstruction of the mixed-format audio signal at the decoding end.
[0367] The mixed-format input audio signal of the encoder in this disclosure includes any combination of the following four audio formats: channel-based audio signal, object-based audio signal, scene-based audio signal, and metadata-based 3D audio signal.
[0368] All input mixed-format audio signals are subjected to high-pass filtering. The filter cutoff frequency can be set to 20Hz. As an example, the filter formula used is as shown in formula (1): (1).
[0369] like Figure 6A As shown, energy calculation and sound field analysis are performed on the audio signal after high-pass filtering: Energy calculation involves calculating the energy values for each channel. An example of calculating the sum of squares for one channel is shown below: (4).
[0370] in, For the first The sum of the amplitude values of the samples in one frame of each channel, the first Each channel per frame includes Sample points For the first One channel.
[0371] Sound field analysis refers to analyzing the number of sound images, their hierarchy, and background ambient noise in a mixed-format audio signal. An example is shown below: Scenario 1: The mixed-format audio signal consists of object-based audio signals and scene-based audio signals. The object-based audio signal is the main sound element in the sound field, while the scene-based audio signal is the ambient background sound element in the sound field. The expected reconstructed sound field requires the object-based audio signal to have less distortion, while the scene-based audio signal, as the ambient background sound, can tolerate a certain degree of distortion. Therefore, more bits are allocated to the object-based audio signal than to the scene-based audio signal during bit allocation.
[0372] In object-based audio signals, each object signal has a different priority in the mixed-format audio signal. For example, in a live concert, the lead singer's voice signal has a higher priority than the backing vocalist's voice signal, or in other words, the lead singer's voice signal is more important than the backing vocalist's voice signal. Therefore, the lead singer's object audio signal will be allocated more bits than the backing vocalist's voice signal.
[0373] Scenario 2: The mixed-format audio signal consists of object-based audio signals and channel-based audio signals. The object-based audio signal is the main sound element in the sound field, while the channel-based audio signal is the ambient background sound element in the sound field. The expected reconstructed sound field requires the object-based audio signal to have less distortion, while the channel-based audio signal, as the ambient background sound, can tolerate a certain degree of distortion. Therefore, more bits are allocated to the object-based audio signal than to the channel-based audio signal during bit allocation.
[0374] Scenario 3: The mixed-format audio signal consists of channel-based audio signals and scene-based audio signals. The channel-based audio signal is the main sound element in the sound field, while the scene-based audio signal is the ambient background sound element in the sound field. The expected reconstructed sound field requires the channel-based audio signal to have less distortion, while the scene-based audio signal, as the ambient background sound, can tolerate a certain degree of distortion. Therefore, more bits are allocated to the channel-based audio signal than to the scene-based audio signal during bit allocation.
[0375] Case 4: In the above three cases, the bit allocation for each channel is performed based on the results of the energy calculation for each channel.
[0376] After obtaining the sound field analysis parameters and energy parameters of each channel of the mixed-format audio signal, bit allocation parameters are obtained by using the corresponding encoding kernel to assign bits to multiple channels based on these parameters. Channel signal encoding parameters are then obtained by encoding multiple channels according to the bit allocation parameters, and these parameters are written into the bitstream. The encoding end sends the bitstream to the decoding end. The decoding end decodes the received bitstream to reconstruct the aforementioned mixed-format audio signal.
[0377] In some embodiments, after high-pass filtering of all input mixed-format audio signals, such as Figure 6B As shown, energy calculation and sound field analysis can be performed on the audio signal after high-pass filtering. When allocating bits to the channels, external control parameters also need to be considered. An example is as follows: Case 5: In the three cases mentioned above (such as Case 1, Case 2, and Case 3), the bit allocation of each channel is performed by combining the external control parameters and the results of the energy calculation for each channel.
[0378] For example, suppose a mixed-format audio signal contains a 5.1 format multi-channel signal and four object audio signals. After high-pass filtering the mixed-format audio signal, the cross-correlation coefficients between each channel are calculated. Based on the cross-correlation coefficients, channel combination operations are performed. Channel parameters are calculated for channels within the same channel combination. Channel parameters can be ILD, ITD, IPD, etc. After aligning the channels using the channel parameters, downmixing is performed. Energy calculation is performed on the downmixed channels. Energy calculation can be performed by sampling amplitude sum, sampling amplitude squared sum, sampling amplitude root mean square, etc.
[0379] Sound field analysis is performed on mixed-format audio signals. This analysis calculates the number of sound images, their relative importance, and ambient background noise. This can be achieved through source localization estimation (DOA), including azimuth and elevation angles, and can be solved using the Multichannel Cross-Correlation Coefficient (MCCC) method. Channels identified using this algorithm are then treated as more important channels.
[0380] Determine the control parameters for the mixed-format audio signal. These control parameters refer to the priority order of each channel in the mixed-format audio signal, pre-set based on external information. This order can be determined according to the functional division of each channel during acquisition, or according to the settings required for the desired sound field reconstructed at the playback end (e.g., in case 1, the lead singer's voice in a live concert can be set as the most important channel; in case 2, to emphasize the sound of a certain instrument (such as the violin) in the reconstructed sound field, the violin's sound can be set as the most important channel signal).
[0381] After obtaining the control parameters, sound field analysis parameters, and energy parameters of each channel of the mixed-format audio signal, bit allocation parameters are obtained by using the corresponding encoding kernel to assign bits to multiple channels based on these parameters. Channel signal encoding parameters are then obtained by encoding multiple channels according to the bit allocation parameters, and these parameters are written into the bitstream. The encoding end sends the bitstream to the decoding end. The decoding end decodes the received bitstream to reconstruct the aforementioned mixed-format audio signal.
[0382] This disclosure also provides embodiments of an apparatus for implementing any of the above methods. For example, an apparatus is provided that includes units or modules for implementing the steps performed by the encoding device in any of the above methods. Furthermore, another apparatus is provided that includes units or modules for implementing the steps performed by the decoding device in any of the above methods.
[0383] It should be understood that the division of units or modules in the above device is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the units or modules in the device can be implemented by a processor calling software: for example, the device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of the units or modules in the above device. The processor can be, for example, a general-purpose processor, such as a Central Processing Unit (CPU) or a microprocessor, and the memory can be internal or external to the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits. The functionality of some or all of the units or modules can be achieved through the design of these hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC), and the functionality of some or all of the units or modules is achieved through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a programmable logic device (PLD), such as a field-programmable gate array (FPGA), which can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files, thereby achieving the functionality of some or all of the units or modules. All units or modules of the above device can be implemented entirely through processor-called software, entirely through hardware circuits, or partially through processor-called software with the remaining parts implemented through hardware circuits.
[0384] In this embodiment, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction read and execute capabilities, such as a Central Processing Unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. The logical relationships of the aforementioned hardware circuits are fixed or reconfigurable. For example, the processor is a hardware circuit implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a Neural Network Processing Unit (NPU), a Tensor Processing Unit (TPU), or a Deep Learning Processing Unit (DPU).
[0385] Figure 7A This is a schematic diagram of the structure of the decoding device proposed in an embodiment of this disclosure. Figure 7AAs shown, the decoding end device 7100 may include at least one of a transceiver module 7101, a processing module 7102, etc. In some embodiments, the processing module is used to acquire a mixed-format audio signal, wherein the mixed-format audio signal is an audio signal including multiple channels; the processing module is further used to determine a first parameter of the mixed-format audio signal and / or the energy parameter of one or more channels among the multiple channels; the processing module is further used to perform bit allocation on the multiple channels according to the first parameter of the mixed-format audio signal and / or the energy parameter of one or more channels among the multiple channels to obtain bit allocation parameters; the processing module is further used to encode the multiple channels according to the bit allocation parameters to obtain channel signal encoding parameters, and write the bit allocation parameters and channel signal encoding parameters into the bitstream; the transceiver module is used to transmit the bitstream. Optionally, the transceiver module is used to perform at least one of the communication steps such as sending and / or receiving performed by the encoding end device 101 in any of the above methods (e.g., steps S2107, S2207, S2306, S2408, but not limited thereto), which will not be described in detail here. Optionally, the above processing module is used to execute at least one of the other steps executed by the terminal device 102 in any of the above methods (such as steps S2101, S2102, S2103, S2104, S2105, S2106, S2201, S2202, S2203, S2204, S2205, S2206, S2301, S2302, S2303, S2304, S2305, S2401, S2402, S2403, S2404, S2405, S2406, S2407, but not limited thereto), which will not be elaborated here.
[0386] Figure 7B This is a schematic diagram of the structure of the decoding device proposed in an embodiment of this disclosure. Figure 7BAs shown, the decoding device 7200 may include at least one of a transceiver module 7201 and a processing module 7202. In some embodiments, the transceiver module is used to receive a bitstream; the processing module is used to parse the bitstream to obtain bit allocation parameters and channel signal encoding parameters for each channel in a mixed-format audio signal; the processing module is also used to reconstruct the mixed-format audio signal based on the bit allocation parameters and channel signal encoding parameters for each channel. Optionally, the transceiver module is used to perform at least one of the communication steps such as sending and / or receiving performed by the decoding device 7200 in any of the above methods, which will not be elaborated here. Optionally, the processing module is used to perform at least one of the other steps performed by the decoding device 102 in any of the above methods (such as steps S2108, S2109, S2208, S2209, S2307, S2308, S2409, S2410, but not limited thereto), which will not be elaborated here.
[0387] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module, which may be separate or integrated. Optionally, the transceiver module may be interchangeable with a transceiver.
[0388] In some embodiments, the processing module may be a single module or may include multiple sub-modules. Optionally, the multiple sub-modules may each perform all or part of the steps required by the processing module. Optionally, the processing module may be interchangeable with a processor.
[0389] Figure 8A This is a schematic diagram of the structure of the communication device 8100 proposed in this embodiment. The communication device 8100 can be an encoding device, a decoding device, or a chip, chip system, or processor that supports the encoding device in implementing any of the above methods; similarly, it can be a chip, chip system, or processor that supports the decoding device in implementing any of the above methods. The communication device 8100 can be used to implement the methods described in the above method embodiments, as detailed in the descriptions within the above method embodiments.
[0390] like Figure 8A As shown, the communication device 8100 includes one or more processors 8101. The processor 8101 can be a general-purpose processor or a dedicated processor, such as a baseband processor or a central processing unit (CPU). The baseband processor can be used to process communication protocols and communication data, while the CPU can be used to control communication devices (e.g., base stations, baseband chips, terminal devices, terminal device chips, DUs or CUs, etc.), execute programs, and process program data. The communication device 8100 is used to execute any of the above methods.
[0391] In some embodiments, the communication device 8100 further includes one or more memories 8102 for storing instructions. Optionally, all or part of the memories 8102 may also be located outside the communication device 8100.
[0392] In some embodiments, the communication device 8100 further includes one or more transceivers 8103. When the communication device 8100 includes one or more transceivers 8103, the transceivers 8103 perform at least one of the communication steps such as sending and / or receiving in the above method (e.g., steps S2107, S2207, S2306, and S2408, but not limited thereto), and the processor 8101 performs other steps (e.g., steps S2101, S2102, S2103, S2104, S2105, S2106, S2201, S2202, S2203, and S2104). At least one of the following: step S204, step S2205, step S2206, step S2301, step S2302, step S2303, step S2304, step S2305, step S2401, step S2402, step S2403, step S2404, step S2405, step S2406, step S2407, step S2108, step S2109, step S2208, step S2209, step S2307, step S2308, step S2409, step S2410, but not limited to.
[0393] In some embodiments, a transceiver may include a receiver and / or a transmitter, which may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, transceiver circuit, etc., may be used interchangeably; the terms transmitter, transmitting unit, transmitter, transmitting circuit, etc., may be used interchangeably; and the terms receiver, receiving unit, receiver, receiving circuit, etc., may be used interchangeably.
[0394] In some embodiments, the communication device 8100 may include one or more interface circuits 8104. Optionally, the interface circuit 8104 is connected to the memory 8102, and the interface circuit 8104 can be used to receive signals from the memory 8102 or other devices, and can be used to send signals to the memory 8102 or other devices. For example, the interface circuit 8104 can read instructions stored in the memory 8102 and send the instructions to the processor 8101.
[0395] The communication device 8100 described in the above embodiments can be an encoding end device or a decoding end device, but the scope of the communication device 8100 described in this disclosure is not limited thereto, and the structure of the communication device 8100 is not limited thereto. Figure 8AThe limitations. The communication device can be a standalone device or part of a larger device. For example, the communication device can be: (1) a standalone integrated circuit IC, or chip, or chip system or subsystem; (2) a collection of one or more ICs, optionally including storage components for storing data and programs; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, terminal device, smart terminal device, cellular phone, wireless device, handheld device, mobile unit, vehicle device, network device, cloud device, artificial intelligence device, etc.; (6) others, etc.
[0396] Figure 8B This is a schematic diagram of the structure of chip 8200 according to an embodiment of this disclosure. For cases where the communication device 8100 can be a chip or a chip system, please refer to... Figure 8B The diagram shown is a schematic representation of the structure of chip 8200, but it is not limited to this.
[0397] Chip 8200 includes one or more processors 8201, which are used to perform any of the above methods.
[0398] In some embodiments, chip 8200 further includes one or more interface circuits 8202. Optionally, the interface circuit 8202 is connected to memory 8203, and the interface circuit 8202 can be used to receive signals from memory 8203 or other devices, and the interface circuit 8202 can be used to send signals to memory 8203 or other devices. For example, the interface circuit 8202 can read instructions stored in memory 8203 and send the instructions to processor 8201.
[0399] In some embodiments, the interface circuit 8202 performs at least one of the communication steps such as sending and / or receiving in the above method (e.g., steps S2107, S2207, S2306, and S2408, but not limited thereto), and the processor 8201 performs other steps (e.g., steps S2101, S2102, S2103, S2104, S2105, S2106, S2201, S2202, S2203, S2204, and S2105). At least one of the following: step S205, step S2206, step S2301, step S2302, step S2303, step S2304, step S2305, step S2401, step S2402, step S2403, step S2404, step S2405, step S2406, step S2407, step S2108, step S2109, step S2208, step S2209, step S2307, step S2308, step S2409, step S2410, but not limited to.
[0400] In some embodiments, the terms interface circuit, interface, transceiver pin, transceiver, etc., can be used interchangeably.
[0401] In some embodiments, chip 8200 further includes one or more memories 8203 for storing instructions. Optionally, all or part of the memories 8203 may be located outside of chip 8200.
[0402] This disclosure also proposes a storage medium storing instructions that, when executed on a communication device 8100, cause the communication device 8100 to perform any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but not limited thereto; it may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but not limited thereto; it may also be a temporary storage medium.
[0403] This disclosure also provides a program product that, when executed by the communication device 8100, causes the communication device 8100 to perform any of the above methods. Optionally, the program product is a computer program product.
[0404] This disclosure also proposes a computer program that, when run on a computer, causes the computer to perform any of the above methods.
[0405] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer programs. When the computer program is loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer program can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program can be transferred from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).
[0406] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0407] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0408] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A signal processing method, characterized in that, include: Acquire a mixed-format audio signal, wherein the mixed-format audio signal is an audio signal including multiple channels; Determine a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels; Bit allocation parameters are obtained by bit-allocating the multiple channels based on the first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the multiple channels; The multiple channels are encoded according to the bit allocation parameters to obtain channel signal encoding parameters, and the bit allocation parameters and the channel signal encoding parameters are written into the bit stream. Send the bitstream.
2. The method as described in claim 1, characterized in that, The first parameter includes at least one of the following: Sound field analysis parameters; The control parameters include a second parameter and / or a third parameter, wherein the second parameter describes the priority order of the roles of each of the plurality of channels in the mixed-format audio signal, and the third parameter describes whether each of the channels is narrative or non-narrative.
3. The method as described in claim 1 or 2, characterized in that, The mixed-format audio signal includes at least one of the following audio signal formats: Audio signals based on vocal tracts; Object-based audio signals; Scene-based audio signals; 3D audio signals based on metadata.
4. The method as described in claim 2 or 3, characterized in that, The determination of the first parameter of the audio signal in the mixed format and / or the energy parameter of one or more of the plurality of channels includes: Determine the energy parameters of one or more of the plurality of audio channels; Determine the sound field analysis parameters for the audio signal in the mixed format.
5. The method as described in claim 4, characterized in that, The step of bit-allocating the multiple channels according to a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the multiple channels to obtain bit allocation parameters includes: The bit allocation parameters are obtained by bit-allocating the plurality of channels based on the energy parameters of one or more of the channels and the sound field analysis parameters of the mixed-format audio signal.
6. The method as described in claim 2 or 3, characterized in that, The determination of the first parameter of the audio signal in the mixed format and / or the energy parameter of one or more of the plurality of channels includes: Determine the energy parameters of one or more of the plurality of audio channels; The control parameters of the mixed-format audio signal are determined.
7. The method as described in claim 6, characterized in that, The step of bit-allocating the multiple channels according to a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the multiple channels to obtain bit allocation parameters includes: The bit allocation parameters are obtained by bit-allocating the plurality of channels based on the energy parameters of one or more of the channels and the control parameters of the mixed-format audio signal.
8. The method as described in claim 2 or 3, characterized in that, The determination of the first parameter of the audio signal in the mixed format and / or the energy parameter of one or more of the plurality of channels includes: Determine the sound field analysis parameters and / or the control parameters of the mixed-format audio signal.
9. The method as described in claim 8, characterized in that, The step of bit-allocating the multiple channels according to a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the multiple channels to obtain bit allocation parameters includes: The bit allocation parameters are obtained by bit-allocating the plurality of channels based on the sound field analysis parameters and / or the control parameters of the mixed-format audio signal.
10. The method as described in claim 2 or 3, characterized in that, The determination of the first parameter of the audio signal in the mixed format and / or the energy parameter of one or more of the plurality of channels includes: Determine the energy parameters of one or more of the plurality of audio channels; The sound field analysis parameters and control parameters of the mixed-format audio signal are determined.
11. The method as described in claim 10, characterized in that, The step of bit-allocating the multiple channels according to a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the multiple channels to obtain bit allocation parameters includes: The bit allocation parameters are obtained by bit-allocating the plurality of channels based on the energy parameters of one or more of the channels, the sound field analysis parameters of the mixed-format audio signal, and the control parameters.
12. The method according to any one of claims 4-7 and 10-11, characterized in that, Determining the energy parameters of one or more of the plurality of channels includes: After filtering the multiple channels, the cross-correlation coefficients between the channels are calculated. Based on the cross-correlation coefficient, perform channel combination operations and calculate inter-channel parameters for channels within the same channel combination; After aligning the channels using the inter-channel parameters, downmixing is performed, and energy calculation is performed on the downmixed channels to obtain the energy parameters of each channel.
13. The method as described in claim 12, characterized in that, The method for calculating the energy parameters includes performing any of the following processing on the samples of one frame of the audio channel: Calculate the sum of the amplitude values of the sample points; Calculate the sum of squares of the amplitudes of the sample points; Calculate the root mean square of the sample points.
14. The method according to any one of claims 4-5 and 8-11, characterized in that, The determination of the sound field analysis parameters for the mixed-format audio signal includes: Based on Direction-of-Arrival (DOA) estimation and / or Multi-Channel Cross-Correlation Coefficient (MCCC) method, sound field analysis is performed on the mixed-format audio signal to obtain the sound field analysis parameters of the mixed-format audio signal.
15. The method according to any one of claims 6-11, characterized in that, The control parameters for determining the mixed-format audio signal include at least one of the following: Based on the functional division of each channel during audio signal acquisition by the audio signal acquisition device, the control parameters of the mixed-format audio signal are determined, and the audio signal acquisition device is used to acquire the mixed-format audio signal. The control parameters of the mixed-format audio signal are determined based on the setting parameters required for the desired sound field reconstructed by the decoding device.
16. The method according to any one of claims 2, 4-5, 8-11, and 14, characterized in that, The sound field analysis parameters include at least one of the following: The number of sound images; Primary and secondary audio-visual elements; Background ambient sound.
17. The method as described in claim 2 or 3, characterized in that, The step of bit allocation to the plurality of channels based on a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels includes: Bit allocation is performed on the mixed-format audio signal, which consists of the object-based audio signal and the scene-based audio signal, based on the sound field analysis parameters of the mixed-format audio signal. The number of bits allocated to any channel in the object-based audio signal is greater than the number of bits allocated to any channel in the scene-based audio signal.
18. The method as described in claim 2 or 3, characterized in that, The step of bit allocation to the plurality of channels based on a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels includes: Bit allocation is performed on the mixed-format audio signal, which consists of the object-based audio signal and the channel-based audio signal, based on the sound field analysis parameters of the mixed-format audio signal. The number of bits allocated to any channel in the object-based audio signal is greater than the number of bits allocated to any channel in the channel-based audio signal.
19. The method as described in claim 2 or 3, characterized in that, The step of bit allocation to the plurality of channels based on a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels includes: Bit allocation is performed on the mixed-format audio signal, which consists of the channel-based audio signal and the scene-based audio signal, based on the sound field analysis parameters of the mixed-format audio signal. The number of bits allocated to any channel in the channel-based audio signal is greater than the number of bits allocated to any channel in the scene-based audio signal.
20. The method as described in claim 2 or 3, characterized in that, The step of bit allocation to the plurality of channels based on a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels includes: Bit allocation is performed on the mixed-format audio signal, which consists of the object-based audio signal and the metadata-based three-dimensional audio signal, based on the sound field analysis parameters of the mixed-format audio signal. The number of bits allocated to any channel in the object-based audio signal is greater than the number of bits allocated to any channel in the metadata-based three-dimensional audio signal.
21. The method as described in claim 2 or 3, characterized in that, The step of bit allocation to the plurality of channels based on a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels includes: Bit allocation is performed on the mixed-format audio signal, which consists of the channel-based audio signal and the metadata-based three-dimensional audio signal, based on the sound field analysis parameters of the mixed-format audio signal. The number of bits allocated to any channel in the channel-based audio signal is greater than the number of bits allocated to any channel in the metadata-based three-dimensional audio signal.
22. The method according to any one of claims 17, 18, and 20, characterized in that, The number of bits allocated to the channel corresponding to the primary image in the object-based audio signal is greater than the number of bits allocated to the channel corresponding to the secondary image in the object-based audio signal.
23. A signal processing method, characterized in that, include: Receive bitstream; The bitstream is parsed to obtain the bit allocation parameters and channel signal encoding parameters of each channel in the multiple channels of the mixed-format audio signal. The mixed-format audio signal is reconstructed based on the bit allocation parameters and channel signal encoding parameter bitstream of each channel.
24. The method as described in claim 23, characterized in that, The process of reconstructing the mixed-format audio signal based on the bit allocation parameters and channel signal encoding parameters of each of the channels includes: The channel signal encoding parameters are decoded based on the bit allocation parameters of each channel; The audio signal in the mixed format is reconstructed using the decoded channel bitstream.
25. A first communication device, characterized in that, include: The processing module is used to acquire audio signals of mixed format, wherein the audio signals of mixed format are audio signals including multiple channels; The processing module is further configured to determine a first parameter of the mixed-format audio signal and / or the energy parameter of one or more of the plurality of channels; The processing module is further configured to perform bit allocation parameters on the multiple channels according to the first parameter of the mixed format audio signal and / or the energy parameter of one or more of the multiple channels; The processing module is further configured to encode the plurality of channels according to the bit allocation parameters to obtain channel signal encoding parameters, and write the bit allocation parameters and the channel signal encoding parameters into the bit stream; A transceiver module is used to send the bitstream.
26. A second communication device, characterized in that, include: The transceiver module is used to receive the bitstream; The processing module is used to perform bitstream parsing on the bitstream to obtain the bit allocation parameters and channel signal encoding parameters of each channel in the multiple channels of the mixed-format audio signal. The processing module is also used to reconstruct the mixed-format audio signal based on the bit allocation parameters and channel signal encoding parameter bitstream of each of the channels.
27. A communication system, characterized in that, include: An encoding end device is configured to perform the signal processing method as described in any one of claims 1-22; The decoding device is configured to perform the signal processing method as described in claim 23 or 24.
28. A communication device, characterized in that, include: One or more processors; The processor is used to invoke instructions to cause the communication device to execute the signal processing method according to any one of claims 1-22 and 23-24.
29. A storage medium storing instructions, characterized in that, When the instruction is executed on the communication device, the communication device performs the signal processing method according to any one of claims 1-22 and 23-24.