Audio data encoding method, device, system, and storage medium
By determining the similarity coefficient and coding pattern between channel data in multichannel coding, the problem of limited coding resources is solved and coding efficiency is improved.
Patent Information
- Application Number
- PCT/CN2024/089918
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-25
- Publication Date
- 2025-10-30
AI Technical Summary
In existing technologies, multi-channel coding algorithms cannot effectively encode audio data when coding resources are limited, resulting in low coding efficiency.
By determining the similarity coefficient between multi-channel input data, the channel data pair with the highest similarity is selected, and the corresponding encoding mode is determined according to the similarity coefficient to encode the channel data and generate bitstream information.
This reduces the amount of data carried by the bitstream information during the audio channel data encoding process, reduces the encoding pressure on the encoder, and improves encoding efficiency.
Smart Images

Figure CN2024089918_30102025_PF_FP_ABST
Abstract
Description
Audio data encoding methods, devices, systems and storage media Technical Field
[0001] This disclosure relates to the field of communication technology, and in particular to an audio data encoding method, device, system and storage medium. Background Technology
[0002] In related technologies, multi-channel coding algorithms are applied at full bitrate. The encoder pairs channel data with high cross-correlation coefficients, with each channel data pair containing two channel data. The selected channel data pairs are compressed and encoded to generate bit information, which is then transmitted to other terminals. Other terminals decode this bit information using their configured decoders to obtain the channel data pair.
[0003] Summary of the Invention
[0004] In order to overcome the technical problem that limited encoding resources in related technologies make it impossible to encode audio data, this disclosure provides an audio data encoding method, device, system and storage medium.
[0005] According to a first aspect of the present disclosure, an audio data encoding method is provided, executed by a first device, the method comprising:
[0006] Determine multiple similarity coefficients between the input data of each channel in the multi-channel input data;
[0007] Based on the multiple similarity coefficients, determine the multiple channel input data pairs with the highest similarity coefficients, and the multiple similarity coefficients corresponding to the multiple channel input data pairs;
[0008] Based on the multiple similarity coefficients, multiple encoding patterns corresponding one-to-one with the multiple vocal channel input data pairs are determined;
[0009] The multiple channel input data pairs are encoded according to the multiple encoding modes to generate multiple bitstream information corresponding to each of the multiple channel input data pairs.
[0010] According to a second aspect of the present disclosure, an audio data encoding method is provided, executed by a second device, the method comprising:
[0011] The device receives multiple bitstream information sent by a first device, wherein the multiple bitstream information is generated by the first device encoding multiple channel input data according to multiple encoding modes, and the multiple encoding modes are determined by the first device based on multiple similarity coefficients between the multiple channel input data;
[0012] The multiple bitstream information is decoded to generate the multiple audio channel input data.
[0013] According to a third aspect of the embodiments of this disclosure, a first device is provided, comprising:
[0014] The processing module is configured to determine multiple similarity coefficients between the input data of each channel in the multi-channel input data;
[0015] The processing module is further configured to determine, based on the plurality of similarity coefficients, a plurality of channel input data pairs with the highest similarity coefficients, and a plurality of similarity coefficients corresponding one-to-one to the plurality of channel input data pairs;
[0016] The processing module is further configured to determine, based on the plurality of similarity coefficients, a plurality of encoding modes corresponding one-to-one with the plurality of audio channel input data pairs;
[0017] The processing module is further configured to encode the multiple channel input data pairs according to the multiple encoding modes, and generate multiple bitstream information corresponding one-to-one with the multiple channel input data pairs.
[0018] According to a fourth aspect of the embodiments of this disclosure, a second device is provided, comprising:
[0019] The transceiver module is configured to receive multiple bitstream information sent by the first device. The multiple bitstream information is generated by the first device encoding multiple channel input data according to multiple encoding modes. The multiple encoding modes are determined by the first device based on multiple similarity coefficients between the multiple channel input data.
[0020] The processing module is configured to decode the multiple bitstream information to generate the multiple channel input data.
[0021] According to a fifth aspect of the embodiments of this disclosure, a first device is provided, comprising:
[0022] One or more processors;
[0023] The processor is configured to execute the audio data encoding method described in any one of the first aspects of this disclosure.
[0024] According to a sixth aspect of the embodiments of this disclosure, a second device is provided, comprising:
[0025] One or more processors;
[0026] The processor is configured to execute the audio data encoding method described in any of the second aspects of this disclosure.
[0027] According to a seventh aspect of the present disclosure, a communication system is provided, including a first device and a second device, wherein the first device is configured to implement the communication method described in any one of the first aspects of the present disclosure, and the first device is configured to implement the audio data encoding method described in any one of the second aspects of the present disclosure.
[0028] According to an eighth aspect of the present disclosure, a storage medium is provided that stores instructions that, when executed on a communication device, cause the communication device to perform an audio data encoding method as described in any one of the first or second aspects of the present disclosure.
[0029] According to a ninth aspect of the present disclosure, a computer program product is provided, comprising a computer program and / or instructions that, when executed by a communication device, implement the audio data encoding method as described in any one of the first aspects of the present disclosure, or when executed by a communication device, implement the audio data encoding method as described in any one of the second aspects of the present disclosure.
[0030] The above method determines multiple similarity coefficients between the input data of each channel in the multi-channel input data. Based on these similarity coefficients, the multiple input data pairs with the highest similarity coefficients are identified, along with multiple similarity coefficients corresponding to each pair. Based on these similarity coefficients, multiple encoding modes corresponding to each pair are determined. The multiple input data pairs are then encoded using these encoding modes, generating multiple bitstream information corresponding to each pair. Therefore, by using different encoding modes based on the similarity between the input data channels, the data load of the bitstream information during channel data encoding is reduced, the encoding pressure on the encoder is decreased, and the encoding efficiency is improved. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings required for the description of the embodiments are introduced below. The following drawings are only some embodiments of this disclosure and do not impose specific limitations on the protection scope of this disclosure.
[0032] Figure 1 is a schematic diagram of the architecture of a communication system according to an embodiment of the present disclosure.
[0033] Figure 2 is an interactive schematic diagram of an audio data encoding method according to an embodiment of the present disclosure.
[0034] Figure 3 is a flowchart illustrating an audio data encoding method according to an embodiment of the present disclosure.
[0035] Figure 4 is a flowchart illustrating an audio data encoding method according to an embodiment of the present disclosure.
[0036] Figure 5A is a flowchart illustrating an audio data encoding method according to an embodiment of the present disclosure.
[0037] Figure 5B is a schematic diagram of a sigmoid function curve according to an embodiment of the present disclosure.
[0038] Figure 5C is a schematic diagram illustrating the dynamic adjustment of the similarity threshold according to an embodiment of the present disclosure.
[0039] Figure 5D is a schematic diagram illustrating ODG scores according to an embodiment of the present disclosure.
[0040] Figure 5E is a schematic diagram illustrating encoder running time according to an embodiment of the present disclosure.
[0041] Figure 5F is a flowchart illustrating encoding mode 1 according to an embodiment of the present disclosure.
[0042] Figure 5G is a flowchart illustrating encoding mode 2 according to an embodiment of the present disclosure.
[0043] Figure 5H is a flowchart illustrating an audio data encoding method according to an embodiment of the present disclosure.
[0044] Figure 6 is a schematic diagram of the structure of a first device according to an embodiment of the present disclosure.
[0045] Figure 7 is a schematic diagram of the structure of the second device according to an embodiment of the present disclosure.
[0046] Figure 8 is a structural schematic diagram of a communication device 8100 according to an embodiment of the present disclosure.
[0047] Figure 9 is a schematic diagram of the structure of chip 8200 according to an embodiment of the present disclosure. Detailed Implementation
[0048] This disclosure provides an audio data encoding method, device, system, and storage medium.
[0049] In a first aspect, embodiments of this disclosure provide an audio data encoding method, executed by a first device, the method comprising:
[0050] Determine multiple similarity coefficients between the input data of each channel in the multi-channel input data;
[0051] Based on the multiple similarity coefficients, determine the multiple channel input data pairs with the highest similarity coefficients, and the multiple similarity coefficients corresponding to the multiple channel input data pairs;
[0052] Based on the multiple similarity coefficients, multiple encoding patterns corresponding one-to-one with the multiple vocal channel input data pairs are determined;
[0053] The multiple channel input data pairs are encoded according to the multiple encoding modes to generate multiple bitstream information corresponding to each of the multiple channel input data pairs.
[0054] In conjunction with some embodiments of the first aspect, determining multiple similarity coefficients between the input data of each channel in the multi-channel input data includes:
[0055] Determine the multiple modulation discrete cosine transform (MDCT) coefficients that correspond one-to-one with the multi-channel input data;
[0056] The plurality of similarity coefficients are determined based on the plurality of MDCT coefficients.
[0057] In the above embodiments, the similarity between the input data of each channel is determined by the MDCT coefficients of each channel input data. This is beneficial for identifying and comparing the audio features and similarity features between the channel input data, and can provide more accurate and effective audio feature analysis and processing methods to obtain the similarity coefficient results of pairwise comparison between multiple channel input data.
[0058] In conjunction with some embodiments of the first aspect, determining the plurality of modulation discrete cosine transform (MDCT) coefficients corresponding one-to-one with the multi-channel input data includes:
[0059] The first channel input data is subjected to frame-by-frame windowing processing to generate multi-frame channel input sub-data, wherein the first channel input data is any one of the channel input data in the multi-channel input data.
[0060] Based on the multi-frame channel input sub-data, determine the first MDCT spectral coefficients and the number of first MDCT coefficient points of the first channel input data;
[0061] The first MDCT coefficients of the first channel input data are generated based on the first MDCT spectral coefficients and the number of first MDCT coefficient points.
[0062] In the above embodiments, the input data for each channel is processed by frame segmentation and windowing, and the MDCT coefficients of the channel input data are determined based on the number of MDCT coefficient points and MDCT spectral coefficients of each subframe channel input data. This allows for frame segmentation processing and analysis of the channel input data, enabling various processing methods such as frequency domain analysis, compression encoding / decoding analysis, and audio transformation.
[0063] In conjunction with some embodiments of the first aspect, determining the plurality of similarity coefficients based on the plurality of MDCT coefficients includes:
[0064] The second similarity coefficient is determined by the following formula, wherein the second similarity coefficient is any one of the plurality of similarity coefficients:
[0065] Wherein, the sim cosθ The second similarity coefficient is defined as i, where i is the number of the vocal tract input sub-data, N is the number of the MDCT coefficient points, and X is the second similarity coefficient. i The second MDCT spectral coefficients of the second channel input data in the second channel input data pair, the Y i The third MDCT spectral coefficient of the third channel input data in the second channel input data pair is the third MDCT spectral coefficient of the third channel input data in the second channel input data pair. The second channel input data pair is any first channel input data pair and includes the second channel input data and the third channel input data.
[0066] In the above embodiments, the similarity between the input data of each channel is calculated based on cosine similarity, and specific similarity values are used to characterize the similarity of the input data of each channel, which improves the accuracy of comparison between multiple similarity coefficients and obtains multiple channel input data pairs with the highest similarity.
[0067] In conjunction with some embodiments of the first aspect, determining the multiple encoding patterns corresponding one-to-one with the multiple audio channel input data pairs based on the multiple similarity coefficients includes:
[0068] A first similarity threshold and a second similarity threshold are determined, wherein the second similarity threshold is less than 1, the second similarity threshold is greater than the first similarity threshold, and the first similarity threshold is greater than 0;
[0069] The plurality of encoding patterns are determined based on the relationship between the plurality of similarity coefficients and the first similarity threshold and the second similarity threshold.
[0070] In the above embodiments, the corresponding encoding mode is determined by the range of similarity coefficient values, thereby making full use of the encoding resources of the first device, reducing the encoding burden of the encoder in the first device, and improving encoding efficiency.
[0071] In conjunction with some embodiments of the first aspect, determining the plurality of encoding patterns based on the magnitude relationship between the plurality of similarity coefficients and the first similarity threshold and the second similarity threshold includes:
[0072] If the first similarity coefficient is greater than or equal to the second similarity threshold, the first encoding mode corresponding to the first similarity coefficient is determined to be the first set encoding mode, and the first similarity coefficient is any one of the plurality of similarity coefficients;
[0073] If the first similarity coefficient is less than the second similarity threshold, and the first similarity coefficient is greater than or equal to the first similarity threshold, then the first encoding model is determined to be the second set encoding mode.
[0074] If the first similarity coefficient is less than the first similarity threshold, the first encoding mode is determined to be the third set encoding mode.
[0075] In the above embodiments, the set encoding mode adopted under the current first similarity coefficient is determined according to the relationship between the similarity coefficient and the first similarity threshold and the second similarity threshold. Based on different set encoding modes, the audio channel input data pairs under different similarity coefficients are encoded, thereby flexibly adapting the encoding method based on different audio channel input data, reducing the encoding burden of the encoder, and improving encoding efficiency.
[0076] In conjunction with some embodiments of the first aspect, the first channel input data pair is any one of the plurality of channel input data pairs, the first channel input data pair including fourth channel input data and fifth channel input data, and determining the first similarity threshold and the second similarity threshold includes:
[0077] Based on the first similarity coefficient, determine the Holt predictor term for Holt exponential smoothing prediction;
[0078] Determine the channel energy ratio between the fourth channel input data and the fifth channel input data;
[0079] Obtain the first similarity benchmark threshold corresponding to the first similarity threshold, and obtain the second similarity benchmark threshold corresponding to the second similarity threshold;
[0080] The first similarity threshold is determined based on the Holt prediction term, the vocal tract energy ratio, and the first similarity benchmark threshold.
[0081] The second similarity threshold is determined based on the Holt prediction term, the vocal tract energy ratio, and the second similarity benchmark threshold.
[0082] In the above embodiments, the first similarity threshold and the second similarity threshold are dynamically adjusted based on different similarity coefficients to adapt to trend changes in the data and make the obtained encoding method more accurate.
[0083] In conjunction with some embodiments of the first aspect, determining the Holt predictor term for Holt exponent smoothing prediction based on the first similarity coefficient includes:
[0084] Based on the first similarity coefficient, determine the Holt baseline term and Holt trend term for Holt index smoothing prediction;
[0085] The Holt prediction term is determined based on the Holt baseline term and the Holt trend term.
[0086] In conjunction with some embodiments of the first aspect, the first encoding mode is the first preset encoding mode, the first channel input data pair includes sixth channel input data and seventh channel input data, and the step of encoding the first channel input data pair according to the first encoding mode to generate the first bitstream information of the first channel input pair includes:
[0087] Determine the second MDCT coefficients of the sixth channel input data, and determine the third MDCT coefficients of the seventh channel input data;
[0088] According to the spectral index of the MDCT coefficients, the second MDCT coefficients are divided into a first odd spectrum and a first even spectrum, and according to the spectral index, the third MDCT coefficients are divided into a second odd spectrum and a second even spectrum.
[0089] The first maximum correlation ratio (MCR) rotation angle is determined based on the first odd spectrum and the second odd spectrum.
[0090] The second MCR rotation angle is determined based on the first even spectrum and the second even spectrum;
[0091] Based on the first MCR rotation angle, the first odd spectrum and the second odd spectrum are rotated by MCR to generate a third odd spectrum;
[0092] Based on the second MCR rotation angle, the first even spectrum and the second even spectrum are rotated by MCR to generate a third even spectrum;
[0093] The first code stream information is generated based on the third odd spectrum and the third even spectrum.
[0094] In the above embodiments, by maximizing the similarity coefficient between the channel input data through MCR, the MCR is applied to the subframe input data corresponding to the odd / even spectrum of the channel input data, and the channel input data pairs are separated into odd and even spectra. MCR rotation and odd and even spectrum merging reduce the consumption of coding resources in the encoder and improve the coding efficiency of the encoder.
[0095] In conjunction with some embodiments of the first aspect, the first encoding mode is the second set encoding mode, the first channel input data pair includes eighth channel input data and ninth channel input data, and the step of encoding the first channel input data pair according to the first encoding mode to generate the first bitstream information of the first channel input pair includes:
[0096] Determine the fourth MDCT coefficient of the eighth channel input data, and determine the fifth MDCT coefficient of the ninth channel input data;
[0097] According to the spectral index of the MDCT coefficients, the fourth MDCT coefficient is divided into the fourth odd spectrum and the fourth even spectrum, and according to the spectral index, the fifth MDCT coefficient is divided into the fifth odd spectrum and the fifth even spectrum.
[0098] The third MCR rotation angle is determined based on the fourth odd spectrum and the fifth odd spectrum;
[0099] The fourth MCR rotation angle is determined based on the fourth even spectrum and the fifth even spectrum;
[0100] Based on the third MCR rotation angle, the fourth odd spectrum and the fifth odd spectrum are subjected to MCR rotation to generate the sixth odd spectrum and the seventh odd spectrum;
[0101] Based on the fourth MCR rotation angle, the fourth even spectrum and the fifth even spectrum are subjected to MCR rotation to generate the sixth even spectrum and the seventh even spectrum;
[0102] The sixth odd spectrum and the sixth even spectrum are combined to generate the sixth MDCT coefficients;
[0103] The seventh odd spectrum and the seventh even spectrum are combined to generate the seventh MDCT coefficients;
[0104] The first bitstream information is generated based on the sixth MDCT coefficient and the seventh MDCT coefficient.
[0105] In some embodiments, the similarity coefficient between the audio channel input data is maximized by MCR, and the MCR is applied to the subframe input data corresponding to the odd / even spectra of the audio channel input data. This separates the odd and even spectra of the audio channel input data pairs, performs MCR rotation and odd / even spectrum merging, and then performs encoding compression based on Mid / Side operations. This achieves the best compression effect, reduces the information carrying capacity of the bitstream, and improves the compression efficiency of the encoder.
[0106] In conjunction with some embodiments of the first aspect, the first encoding mode is the third preset encoding mode, the first channel input data pair includes tenth channel input data and eleventh channel input data, and the step of encoding the first channel input data pair according to the first encoding mode to generate the first bitstream information of the first channel input pair includes:
[0107] Determine the sixth MDCT coefficient of the tenth channel input data, and determine the seventh MDCT coefficient of the eleventh channel input data;
[0108] The first bitstream information is generated based on the sixth MDCT coefficient and the seventh MDCT coefficient.
[0109] In conjunction with some embodiments of the first aspect, the first bitstream information includes side information, the first bitstream information includes any of the plurality of bitstream information, and the side information includes at least one of the following:
[0110] The first encoding mode;
[0111] MCR rotation angle;
[0112] The first similarity coefficient.
[0113] In some embodiments, an encoding method with added edge information is adopted to reduce the encoding pressure on the encoder, make efficient use of the encoder's encoding resources, and improve the encoder's encoding efficiency.
[0114] In conjunction with some embodiments of the first aspect, the method further includes:
[0115] The plurality of bitstream information is sent to the second device, and the plurality of bitstream information is used to instruct the second device to decode the plurality of bitstream information to generate the plurality of audio channel input data pairs.
[0116] Secondly, embodiments of this disclosure provide an audio data encoding method, executed by a second device, the method comprising:
[0117] The device receives multiple bitstream information sent by a first device, wherein the multiple bitstream information is generated by the first device encoding multiple channel input data according to multiple encoding modes, and the multiple encoding modes are determined by the first device based on multiple similarity coefficients between the multiple channel input data;
[0118] The multiple bitstream information is decoded to generate the multiple audio channel input data.
[0119] By using the above method, the first device receives the bitstream information encoded and transmitted using different encoding modes based on the similarity between the audio channel input data, and decodes the bitstream information to obtain the audio channel input data, thereby reducing the decoding burden of the decoder and improving the output sound quality of the second device.
[0120] In conjunction with some embodiments of the second aspect, the first bitstream information includes side information, the first bitstream information includes any of the plurality of bitstream information, and the side information includes at least one of the following:
[0121] The first encoding mode;
[0122] MCR rotation angle;
[0123] The first similarity coefficient.
[0124] Thirdly, embodiments of this disclosure provide a first device, comprising:
[0125] The processing module is configured to determine multiple similarity coefficients between the input data of each channel in the multi-channel input data;
[0126] The processing module is further configured to determine, based on the plurality of similarity coefficients, a plurality of channel input data pairs with the highest similarity coefficients, and a plurality of similarity coefficients corresponding one-to-one to the plurality of channel input data pairs;
[0127] The processing module is further configured to determine, based on the plurality of similarity coefficients, a plurality of encoding modes corresponding one-to-one with the plurality of audio channel input data pairs;
[0128] The processing module is further configured to encode the multiple channel input data pairs according to the multiple encoding modes, and generate multiple bitstream information corresponding one-to-one with the multiple channel input data pairs.
[0129] Fourthly, embodiments of this disclosure provide a second device, comprising:
[0130] The transceiver module is configured to receive multiple bitstream information sent by the first device. The multiple bitstream information is generated by the first device encoding multiple channel input data according to multiple encoding modes. The multiple encoding modes are determined by the first device based on multiple similarity coefficients between the multiple channel input data.
[0131] The processing module is configured to decode the multiple bitstream information to generate the multiple channel input data.
[0132] Fifthly, embodiments of this disclosure provide a first device, comprising:
[0133] One or more processors;
[0134] The processor is configured to execute the audio data encoding method described in any one of the first aspects of this disclosure.
[0135] Sixthly, embodiments of this disclosure provide a second device, comprising:
[0136] One or more processors;
[0137] The processor is configured to execute the audio data encoding method described in any of the second aspects of this disclosure.
[0138] In a seventh aspect, embodiments of this disclosure provide a communication system including a first device and a second device, wherein the first device is configured to implement the audio data encoding method described in any one of the first aspects of this disclosure, and the second device is configured to implement the audio data encoding method described in any one of the second aspects of this disclosure.
[0139] Eighthly, embodiments of this disclosure provide a storage medium storing instructions that, when executed on a communication device, cause the communication device to perform an audio data encoding method as described in any one of the first or second aspects of this disclosure.
[0140] In a ninth aspect, embodiments of this disclosure provide a computer program product, including a computer program and / or instructions, wherein when the computer program and / or instructions are executed by a communication device, they implement the audio data encoding method as described in any one of the first aspects of this disclosure, or when the computer program and / or instructions are executed by a communication device, they implement the audio data encoding method as described in any one of the second aspects of this disclosure.
[0141] The above method determines multiple similarity coefficients between the input data of each channel in the multi-channel input data. Based on these similarity coefficients, the multiple input data pairs with the highest similarity coefficients are identified, along with multiple similarity coefficients corresponding to each pair. Based on these similarity coefficients, multiple encoding modes corresponding to each pair are determined. The multiple input data pairs are then encoded using these encoding modes, generating multiple bitstream information corresponding to each pair. Therefore, by using different encoding modes based on the similarity between the input data channels, the data load of the bitstream information during channel data encoding is reduced, the encoding pressure on the encoder is decreased, and the encoding efficiency is improved.
[0142] It is understood that the aforementioned terminal, access network equipment, first network element, second network element, core network equipment, communication system, storage medium, program product, computer program, chip or chip system are all used to execute the methods proposed in the embodiments of this disclosure. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects in the corresponding methods, and will not be repeated here.
[0143] This disclosure provides the invention title. In some embodiments, the terms audio data encoding method, information processing method, communication method, etc., can be used interchangeably; the terms audio data encoding device, information processing device, communication device, etc., can be used interchangeably; and the terms information processing system, communication system, etc., can be used interchangeably.
[0144] This disclosure is not exhaustive, but merely illustrative of some embodiments, and is not intended to limit the scope of protection of this disclosure. Unless otherwise specified, each step in a particular embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a particular embodiment can also be implemented as an independent embodiment, and the order of the steps in a particular embodiment can be arbitrarily interchanged. Furthermore, the optional implementation methods in a particular embodiment can be arbitrarily combined; moreover, the embodiments can be arbitrarily combined, for example, some or all steps of different embodiments can be arbitrarily combined, and a particular embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.
[0145] In each of the disclosed embodiments, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of the embodiments are consistent and can be referenced by each other. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.
[0146] The terminology used in the embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure.
[0147] In this embodiment of the disclosure, unless otherwise stated, elements expressed in the singular form, such as "a," "an," "the," "the," "the," "the," "the," "the," "this," etc., can mean "one and only one," or "one or more," "at least one," etc. For example, when using articles such as "a," "an," "the," etc. in translation, the noun following the article can be understood as either a singular expression or a plural expression.
[0148] In the embodiments disclosed herein, "multiple" refers to two or more.
[0149] In some embodiments, the terms “at least one (at least one item, at least one)”, “one or more”, “a plurality of”, “multiple”, etc., may be used interchangeably.
[0150] In some embodiments, the notation "at least one of A and B", "A and / or B", "A in one case, B in another", "in response to one case A, in response to another case B", etc., may include the following technical solutions depending on the situation: in some embodiments, A (execute A regardless of B); in some embodiments, B (execute B regardless of A); in some embodiments, execution is selected from A and B (A and B are selectively executed); in some embodiments, A and B (both A and B are executed). The same applies when there are more branches such as A, B, C, etc.
[0151] In some embodiments, the notation "A or B" may include the following technical solutions, depending on the situation: in some embodiments, A (execution of A regardless of B); in some embodiments, B (execution of B regardless of A); in some embodiments, execution is selected from A and B (A and B are selectively executed). The same applies when there are more branches such as A, B, C, etc.
[0152] The prefixes "first," "second," etc., used in the embodiments of this disclosure are merely for distinguishing different descriptive objects and do not impose restrictions on the position, order, priority, quantity, or content of the descriptive objects. The description of the descriptive objects is found in the claims or the context of the embodiments, and the use of prefixes should not constitute unnecessary restrictions. For example, if the descriptive object is a "field," the ordinal numbers preceding "field" in "first field" and "second field" do not restrict the position or order of the "fields." "First" and "second" do not restrict whether the "fields" they modify are in the same message, nor do they restrict the order of "first field" and "second field." Similarly, if the descriptive object is a "level," the ordinal numbers preceding "level" in "first level" and "second level" do not restrict the priority between "levels." Furthermore, the number of descriptive objects is not limited by ordinal numbers and can be one or more. For example, in "first device," the number of "devices" can be one or more. Furthermore, the objects modified by different prefixes can be the same or different. For example, if the object being described is "device", then "first device" and "second device" can be the same device or different devices, and their types can be the same or different. Similarly, if the object being described is "information", then "first information" and "second information" can be the same information or different information, and their content can be the same or different.
[0153] In some embodiments, “including A,” “containing A,” “for indicating A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.
[0154] In some embodiments, the terms “in response to…”, “in response to determining…”, “in the case of…”, “when…”, “if…”, “if…”, etc., can be used interchangeably.
[0155] In some embodiments, the terms “greater than,” “greater than or equal to,” “not less than,” “more than,” “more than or equal to,” “not less than,” “higher than,” “higher than or equal to,” “not lower than,” and “above” can be used interchangeably, as can the terms “less than,” “less than or equal to,” “not greater than,” “less than,” “less than or equal to,” “not more than,” “lower than,” “lower than or equal to,” “not higher than,” and “below”.
[0156] In some embodiments, the apparatus and device may be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. In some cases, they may also be understood as "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "body", etc.
[0157] In some embodiments, "network" can be interpreted as devices included in the network, such as access network devices, core network devices, etc.
[0158] In some embodiments, "access network device (AN device)" may also be referred to as "radio access network device (RAN device)," "base station (BS)," "radio base station," or "fixed station." In some embodiments, it may also be understood as "node," "access point," "transmission point (TP)," "reception point (RP)," "transmission / reception point (TRP)," "panel," "antenna panel," "antenna array," "cell," "macro cell," "small cell," "femto cell," "pico cell," "sector," "cell group," "serving cell," "carrier," "component carrier," or "bandwidth part (BWP)."
[0159] In some embodiments, "terminal" or "terminal device" may be referred to as "user equipment (UE)," "user terminal," "mobile station (MS)," "mobile terminal (MT)," "subscriber station," "mobile unit," "subscriber unit," "wireless unit," "remote unit," "mobile device," "wireless device," "wireless communication device," "remote device," "mobile subscriber station," "access terminal," "mobile terminal," "wireless terminal," "remote terminal," "handset," "user agent," "mobile client," "client," etc.
[0160] In some embodiments, the acquisition of data, information, etc., may comply with the laws and regulations of the country where the location is situated.
[0161] In some embodiments, data, information, etc., may be obtained with the user's consent.
[0162] Furthermore, each element, each row, or each column in the table of this disclosure can be implemented as an independent embodiment, and any combination of any element, any row, or any column can also be implemented as an independent embodiment.
[0163] Figure 1 is a schematic diagram of the architecture of a communication system according to an embodiment of the present disclosure. As shown in Figure 1, the communication system 100 includes a first device 101 and a second device 102.
[0164] In some embodiments, the first device 101 includes, for example, at least one of the following: a mobile phone, a wearable device, an Internet of Things device, a car with communication capabilities, a smart car, a tablet computer, a computer with wireless transceiver capabilities, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal device in industrial control, a wireless terminal device in self-driving, a wireless terminal device in remote medical surgery, a wireless terminal device in a smart grid, a wireless terminal device in transportation safety, a wireless terminal device in a smart city, and a wireless terminal device in a smart home, but is not limited thereto.
[0165] In some embodiments, the second device 102 may be a node or device for connecting a terminal to a wireless network. The access network device may include, but is not limited to, at least one of the following in a 5G communication system: evolved Node B (eNB), next-generation eNB (ng-eNB), next-generation Node B (gNB), node B (NB), home node B (HNB), home evolved node B (HeNB), wireless backhaul device, radio network controller (RNC), base station controller (BSC), base transceiver station (BTS), base band unit (BBU), mobile switching center, base station in a 6G communication system, open RAN, cloud RAN, base station in other communication systems, and access node in a Wi-Fi system.
[0166] In some embodiments, the first device 101 corresponds to the second device 102. The first device 101 is equipped with an encoder to encode the transmitted data and generate a bitstream information which is then sent to the second device 102. The second device 102 is equipped with a decoder to decode the bitstream information transmitted by the first device 101 and generate the transmitted data. For example, the first device 101 may also be a node or device that connects a terminal to a wireless network. The access network device may include, but is not limited to, at least one of the following in a 5G communication system: evolved Node B (eNB), next-generation eNB (ng-eNB), next-generation Node B (gNB), node B (NB), home node B (HNB), home evolved node B (HeNB), radio backhaul device, radio network controller (RNC), base station controller (BSC), base transceiver station (BTS), base band unit (BBU), mobile switching center, base station in a 6G communication system, open RAN, cloud RAN, base station in other communication systems, and access node in a Wi-Fi system. Correspondingly, the second device 102 can also be at least one of the following: mobile phone, wearable device, Internet of Things device, car with communication function, smart car, tablet computer, computer with wireless transceiver function, virtual reality (VR) terminal device, augmented reality (AR) terminal device, wireless terminal device in industrial control, wireless terminal device in self-driving, wireless terminal device in remote medical surgery, wireless terminal device in smart grid, wireless terminal device in transportation safety, wireless terminal device in smart city, and wireless terminal device in smart home, but is not limited thereto.
[0167] In some embodiments, the technical solutions of this disclosure can be applied to the Open RAN architecture. In this case, the interfaces between or within access network devices involved in the embodiments of this disclosure can be transformed into internal interfaces of Open RAN. The processes and information interactions between these internal interfaces can be implemented by software or programs.
[0168] In some embodiments, the access network device may be composed of a central unit (CU) and a distributed unit (DU). The CU may also be called a control unit. The CU-DU structure can separate the protocol layer of the access network device. Some of the protocol layer functions are centrally controlled by the CU, while the remaining part or all of the protocol layer functions are distributed in the DU and centrally controlled by the CU. However, this is not the only possibility.
[0169] It is understood that the communication system described in this disclosure is for the purpose of more clearly illustrating the technical solutions of this disclosure, and does not constitute a limitation on the technical solutions proposed in this disclosure. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions proposed in this disclosure are also applicable to similar technical problems.
[0170] The following embodiments of this disclosure can be applied to the communication system 100 shown in FIG1, or to some of the main bodies, but are not limited thereto. The main bodies shown in FIG1 are illustrative. The communication system may include all or some of the main bodies in FIG1, or may include other main bodies outside of FIG1. The number and form of each main body are arbitrary. Each main body may be physical or virtual. The connection relationship between the main bodies is illustrative. The main bodies may not be connected or may be connected. The connection can be in any way, it can be a direct connection or an indirect connection, it can be a wired connection or a wireless connection.
[0171] The embodiments disclosed herein can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G new radio (NR), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New radio access (NX), Future generation radio access (FX), Global System for Mobile communications (GSM), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi), IEEE 802.16 (WiMAX), and IEEE 802.20, Ultra-Wideband (UWB), Bluetooth (a registered trademark), Public Land Mobile Network (PLMN) networks, Device-to-Device (D2D) systems, Machine-to-Machine (M2M) systems, Internet of Things (IoT) systems, Vehicle-to-Everything (V2X) systems, systems utilizing other communication methods, and next-generation systems built upon them, etc. Furthermore, multiple systems can be combined (e.g., a combination of LTE or LTE-A with 5G).
[0172] In some embodiments, during multichannel data encoding, the encoder pairs channel data with high cross-correlation coefficients, that is, combines two channel data with high similarity to generate channel data pairs. A Mid / Side operation is then performed on these channel data pairs to convert them into mid-channel and side-channel data. The mid-channel data includes common information between the two channel data, while the side-channel data includes differences between them. For example, if the two channel data are left and right channel data, and the left and right channel data have high similarity, the resulting side-channel data after the Mid / Side operation will contain very little information. Therefore, efficient compression can be performed based on the similarity of the channel data pairs.
[0173] For example, the multi-channel input data includes left channel input data and right channel input data, which are compressed using the following formulas (1) and (2) respectively:
[0174] Wherein, Mid represents the center channel data, Side represents the side channel data, L represents the left channel input data, and R represents the right channel input data. Compression encoding is performed using the above method, and the pairing information of the channel data pairs is written into the side information, with simultaneous encoding and compression to generate the bitstream information. The decoder performs the inverse Mid / Side operation on the corresponding channel data based on the pairing information read from the side information to generate the channel data.
[0175] In some embodiments, the encoding operation is applied in a full-bitrate network environment, using a uniform encoding method to encode the channel data. This method lacks flexibility in adapting to the encoding environment, and encoding resources are limited in low-bitrate scenarios. More aggressive encoding strategies can be adopted to achieve efficient utilization of encoding resources. For example, an encoding method that transmits channel data plus descriptive information can be used to reduce the amount of channel data that needs to be encoded. In related technologies, channel data is not divided into frequency bands because users' sensitivity to sound varies significantly across different frequencies. This makes it impossible to handle the uniqueness of each frequency band, resulting in poor audio quality after decoding and conversion, and even audio distortion. Based on this technical problem, this embodiment proposes an audio data encoding method suitable for multi-channel data encoding in low-bitrate scenarios. It adopts an encoding method that transmits a small number of channels plus differential descriptive information in low-bitrate scenarios to encode and transmit channel data.
[0176] For example, by combining the similarity characteristics between the various channel data in the above method, different encoding modes are designed. Based on MCR (Maximum Correlation Rotation), the similarity features between channels are maximized. MCR is applied to the sub-band data of the odd / even spectrum corresponding to the channel data. At the same time, this embodiment combines historical input information to optimize the selection threshold of the encoding mode. In low bit rate scenarios, the encoder running time is significantly reduced, and the output sound quality of the channel data after decoding is improved.
[0177] In some embodiments, the above encoding method is applied to various communication systems, which may include terminals and network devices. These communication systems include Long Term Evolution (LTE), 5G mobile communication systems, 5G New Radio (NR) systems, or other next-generation mobile communication systems. Based on the above encoding method, the audio channel data encoding is applied to streaming media transmission systems or OTT (Over-the-Top) media transmission systems. For example, an encoder is configured in the terminal of the communication system, and a decoder is configured in the network device. The terminal encodes the audio channel data using the above method based on the decoder to generate a bitstream. The bitstream is then sent to the network device based on the communication system, and the network device decodes the bitstream based on the decoder to generate audio channel data. Optionally, the configuration location of the encoder and decoder is not limited in this embodiment. For example, the encoder can be configured in the network device, and the network device encodes the acquired audio channel data based on the encoder to generate a bitstream, which is then sent to the terminal. The decoder configured in the terminal decodes the bitstream to generate audio channel data.
[0178] Figure 2 is an interactive schematic diagram of an audio data encoding method according to an embodiment of the present disclosure. As shown in Figure 2, the embodiments of the present disclosure relate to an audio data encoding method, which includes:
[0179] Step S2101: The first device determines multiple MDCT coefficients that correspond one-to-one with the multi-channel input data.
[0180] In some embodiments, the multi-channel input data is audio input data transmitted from multiple channels received by the first device from other devices, or it may be audio input data from multiple channels generated by the first device based on current needs.
[0181] In some embodiments, the audio data encoding method is applied to a communication system, which consists of a first device and a second device constituting a mobile communication system. An encoder is configured in the first device to encode multi-channel input data to generate bitstream information. The bitstream information is sent to the second device through the mobile communication system. The decoder configured in the second device decodes the bitstream information to generate multi-channel input data.
[0182] For example, in this embodiment, the first device acquires multi-channel input data from multiple channels. This multi-channel input data consists of audio data from multiple channels, with the number of channels corresponding to the audio data being greater than or equal to 5. The multi-channel input data is analyzed according to different channels to determine the MDCT (Modulated Discrete Cosine Transform) coefficients of the input data corresponding to each channel. The MDCT coefficients are used to indicate the spectral distribution of the corresponding channel input data. By converting the audio data from the time domain to the frequency domain, the energy distribution of the audio data in different frequency bands and time periods is identified. In this embodiment, by determining the MDCT coefficients of the audio data input from each channel, the MDCT coefficients of each channel are compared pairwise to analyze the similarity between the multi-channel input data.
[0183] In some embodiments, the first device can be a terminal, configured with an encoder to encode the audio input data and generate a bitstream information to be sent to the second device. The second device can be another terminal or a network device, configured with a decoder to decode the bitstream information to obtain the audio input data. It should be noted that the encoding process occurs in the terminal, while the decoding process occurs in the other device or network device. Alternatively, the first device can also be a network device, configured with an encoder to encode the audio input data and generate a bitstream information to be sent to the second device. The second device can be a terminal, configured with a decoder to decode the received bitstream information to generate the audio input data.
[0184] Optionally, in some embodiments, step S2101 above includes:
[0185] The first device performs frame-by-frame windowing processing on the first channel input data to generate multiple frames of channel input sub-data.
[0186] The first device determines the first MDCT spectral coefficients and the number of first MDCT coefficient points of the first channel input data based on the multi-frame channel input sub-data.
[0187] The first device generates the first MDCT coefficients of the first channel input data based on the first MDCT spectral coefficients and the number of first MDCT coefficient points.
[0188] For example, the input data for each channel typically exists as a continuous time-domain signal. Therefore, it is necessary to perform frame segmentation processing on the channel input data, dividing it into many short time frames. In this embodiment, the first channel input data is taken as an example. This first channel input data can be any multi-channel input data. After performing frame segmentation processing on the first channel input data to generate multiple frames of initial channel input sub-data, windowing processing is then applied to each frame of initial channel input sub-data. A window function is applied to the initial channel input sub-data of each frame to generate multiple frames of channel input sub-data corresponding to the first channel input data. By performing frame segmentation and windowing processing on the channel input data, spectral leakage is reduced and frequency domain ripples caused by the time domain stage are avoided, thus preventing frame loss of channel input data during processing.
[0189] MDCT coefficients are frequency domain coefficients obtained through MDCT transform. The number of MDCT coefficient points indicates the number of MDCT coefficients output during the MDCT transform. In this embodiment, multi-frame channel input sub-data is analyzed to determine the number of coefficients existing in the corresponding frequency domain during the MDCT transform, as well as the MDCT spectral coefficients corresponding to the multi-frame channel input sub-data during the MDCT transform. Based on the number of MDCT coefficient points and the MDCT spectral coefficients, the MDCT coefficients of the first channel input data are determined.
[0190] In step S2102, the first device determines multiple similarity coefficients between the input data of each channel in the multiple channel input data based on multiple MDCT coefficients.
[0191] For example, after determining the MDCT coefficients of each channel input data through the above steps, pairwise comparisons are performed on the multi-channel input data based on the MDCT coefficients to obtain the similarity coefficients between each channel input data and other channel input data. It should be noted that in this embodiment, during the channel input data comparison process, it is necessary to determine the similarity coefficients between the current channel input data and all other channel input data. That is, it is necessary to compare the MDCT coefficients of the current channel input data with the MDCT coefficients of other channel input data one by one to obtain multiple similarity coefficients. For example, the multi-channel input data are: A, B, C, D, E, F, G; the channel input data pairs obtained through pairwise comparisons are: AB, AC, AD, AE, AF, AG, BC, BD, BE, BF, BG, CD, CE, CF, CG, DE, DF, DG, EF, EG, FG, and 21 corresponding similarity coefficients are generated.
[0192] Optionally, in some embodiments, step S2102 above includes:
[0193] The second similarity coefficient is determined using the following formula, where the second similarity coefficient can be any number of similarity coefficients:
[0194] Where, sim cosθ The second similarity coefficient is given by i, where i is the number of the vocal tract input sub-data, N is the number of MDCT coefficient points, and X is the second similarity coefficient. i Y represents the second MDCT spectral coefficient of the second channel input data in the second channel input data pair. i The third MDCT spectral coefficient is the third channel input data in the second channel input data pair. The second channel input data pair is any first channel input data pair, which includes the second channel input data and the third channel input data.
[0195] For example, sim cos(θ) This is the similarity coefficient between two vocal tract data points; the similarity coefficient is the cosine similarity, x. i and y i Let N be the MDCT spectral coefficients of the two channel data, N be the number of MDCT coefficient points in each frame of channel data, and i be the channel data number of the current subframe after framing the channel input data. The similarity between the two channel data is evaluated using the above formula (3), and the generated cosine similarity value ranges from [-1, 1]. Here, -1 indicates that the vectors corresponding to the two channel data are completely opposite, 0 indicates that the vectors corresponding to the two channel data are orthogonal, and 1 indicates that the vectors corresponding to the two channel data are completely identical in direction. By performing pairwise statistics on the above multi-channel data in this way, the similarity coefficient of each channel data pair is obtained.
[0196] In step S2103, the first device determines the multiple channel input data pairs with the highest similarity coefficients and the multiple similarity coefficients corresponding to the multiple channel input data pairs based on multiple similarity coefficients.
[0197] For example, the MDCT coefficients corresponding to each channel input data are compared, and based on the similarity of the MDCT coefficients, multiple similarity coefficients between the channel input data are determined. In this embodiment, based on the MDCT coefficients of each channel input data, pairwise comparisons are performed to determine the similarity coefficients between the channel input data. Then, the two channel input data with the highest similarity coefficients are selected to form a channel input data pair. For example, the multi-channel input data are A, B, C, D, E, and F; based on the MDCT coefficients of each channel input data, the multi-channel input data are paired and compared to generate initial channel input data pairs: AB, AC, AD, AE, AF, BC, BD, BE, BF, CD, CE, CF, DE, DF, and EF; after determining multiple initial similarity coefficients for each initial channel input data pair, the channel input data pair with the highest similarity coefficients is obtained: AB, CF, DE, and the corresponding similarity coefficients a, b, and c.
[0198] It should be noted that during the comparison of channel input data, the MDCT coefficients of each channel input data need to be compared with the MDCT coefficients of other channel input data to obtain multiple initial similarity coefficients for comparison among the multi-channel input data. Based on these initial similarity coefficients, they are then sorted in descending order of similarity, and the two channel input data with the highest similarity are selected to form a channel input data pair. This process continues until all channel input data forms channel input data pairs, thus generating multiple channel input data pairs and corresponding similarity coefficients.
[0199] Step S2104: The first device determines the first similarity threshold and the second similarity threshold.
[0200] In some embodiments, the second similarity threshold is greater than the first similarity threshold, and the first similarity threshold is greater than 0.
[0201] For example, in this embodiment, a first similarity threshold and a second similarity threshold can be set. These thresholds divide the range of similarity coefficient values into three intervals, each corresponding to a different encoding mode. For instance, if the similarity coefficient is greater than or equal to the second similarity threshold, the first encoding mode is used; if the similarity coefficient is less than the second similarity threshold but greater than or equal to the first similarity threshold, the second encoding mode is used; and if the similarity coefficient is less than the first similarity threshold, the third encoding mode is used.
[0202] Optionally, in some embodiments, the first channel input data pair is any plurality of channel input data pairs, the first channel input data pair including fourth channel input data and fifth channel input data, and the above step S2104 includes:
[0203] The first device determines the Holt prediction term for Holt exponential smoothing prediction based on the first similarity coefficient;
[0204] The first device determines the channel energy ratio between the fourth channel input data and the fifth channel input data;
[0205] The first device obtains a first similarity benchmark threshold corresponding to a first similarity threshold, and obtains a second similarity benchmark threshold corresponding to a second similarity threshold;
[0206] The first device determines the first similarity threshold based on the Holt prediction term, the channel energy ratio, and the first similarity benchmark threshold;
[0207] The first device determines the second similarity threshold based on the Holt prediction term, the channel energy ratio, and the second similarity benchmark threshold.
[0208] For example, the first similarity coefficient is the similarity coefficient of the first channel input data pair. Based on the first similarity coefficient, the Holt prediction term for Holt exponential smoothing prediction is calculated and determined. In this embodiment, the trend value and seasonal value are predicted separately to adapt to the trend changes in the channel input data through Holt exponential smoothing prediction. Holt exponential smoothing prediction is divided into a baseline term and a trend term. Based on the first similarity coefficient, the calculation method of the baseline term and the trend term is determined as follows: formula (4) and formula (5): L t =α*sim cos(θ) (t)+(1-α)*L t-1 +T t-1 (4) T t =β*[L t -L t-1 ]+(1-β)*T t-1 (5)
[0209] Among them, L t L is the baseline term at time t. t A weighted average can be calculated based on the difference between the historical baseline and the latest observation. The weight parameter 1 can be determined by the smoothing parameter α (0 < α < 1). The value of α can be adjusted according to the current encoding requirements. When the value of α approaches 1, it indicates that the latest observation has a greater weight; when the value of α approaches 0, it indicates that the historical baseline has a greater weight. cos(θ) (t) represents the cosine similarity coefficient generated at the current time using the above method, and L t-1 T is the reference term for the previous historical moment. t-1 This represents the trend at the previous historical moment. (T) t Let t be the trend term at time t. This trend term can be updated using a weighted average. When updating the trend term, the weight parameter 2 is a smoothing parameter β (0 < β < 1). β controls the sensitivity of the model to trend changes. For example, the higher the value of β, the stronger the sensitivity of the model to trend changes.
[0210] The encoding pattern of the vocal tract data pair is determined by the relationship between the cosine similarity of the vocal tract data pair and the similarity thresholds thr1 and thr2. thr1 and thr2 can be calculated using the following formulas (6) and (7): thr2=γ*sigmoid(ILD)+V+A (6) thr1=ε*sigmoid(ILD)+V+B (7)
[0211] Where thr1 and thr2 are similarity thresholds, B is the first threshold baseline value corresponding to thr1, and A is the second threshold baseline value corresponding to thr2. ILD is the energy ratio of the vocal tract data between pairs. γ > 0 is weight parameter 3, which adjusts the influence of ILD on the similarity threshold thr2. ε < 0 is weight parameter 4, which adjusts the influence of ILD on the similarity threshold thr1. V is the Holt prediction component.
[0212] For example, the function sigmoid() in formulas (6) and (7) above is a squeezing function, and the expression of the squeezing function is as follows: formula (10):
[0213] In the sigmoid() squeezing function mentioned above, the constant a controls the maximum value of the curve corresponding to the squeezing function, the coefficient k controls the steepness of the curve, and the constant h controls the horizontal shift of the curve.
[0214] For example, the sigmoid term does not have a significant suppressive effect when the ILD difference is small. However, when the difference is large and increases rapidly, the sigmoid term suppresses the transformation effect caused by the ILD difference. In the above formulas (4) and (5), γ > 0 and ε < 0. This weighting parameter is used to compress the probability interval of encoding mode 2 when the ILD difference is large. This helps to suppress the noise amplification problem caused by the Mid / side compression operation when noise is introduced during the channel data encoding process, and reduces the noise impact on the decoding end.
[0215] For example, in this embodiment, the Holt predictor is used to suppress mode jumps caused by slight perturbations near the reference value. The Holt predictor can be calculated using the following formula (8):
[0216] Where k is a constant greater than 0, u L <0 and u H >0 is the limiting constant, which is determined by u L and u H This limits the upper and lower bounds of the holt predictor V's ability to adjust the threshold.
[0217] For example, V is initialized in formula (8) above. t =0, V t The condition for the value to change is the previous two frames T. r-2 and T t-1When the overall trend is stable and the direction of the current similarity value changes compared to previous frames, V dynamically adjusts the similarity coefficient. As shown in Figure 6C, based on the baseline value A = 0.7, the dynamic adjustment effect of the Holt prediction component on the similarity threshold is illustrated. From frames 16 to 25, the similarity coefficient remains stable, but a slight perturbation below A occurs in frame 24. At this point, V's adjustment reduces two avoidable coding mode switches. The similarity coefficients in other frames are far from the baseline value, ensuring the selection of the applied coding mode. V considers the current level and trend changes of the data and adaptively fine-tunes according to the trend direction, thus resisting the influence of minor perturbations near the baseline value and avoiding overreaction to insignificant changes.
[0218] By using the above method, the first similarity threshold and the second similarity threshold are dynamically adjusted based on the similarity coefficient, thereby adapting to the trend changes in the vocal tract input data, avoiding the disorder caused by the disturbance of the vocal tract input data, and making the determined encoding pattern more accurate.
[0219] In step S2105, the first device determines multiple encoding modes that correspond one-to-one with the multiple channel input data pairs based on the relationship between multiple similarity coefficients and the first similarity threshold and the second similarity threshold.
[0220] For example, by using a first similarity threshold and a second similarity threshold, the range of similarity coefficient values is divided into three intervals. Based on the interval range of each similarity coefficient, the corresponding encoding mode is determined. In this embodiment, the first channel input data pair is used as an example, and the similarity coefficient of the first channel input data pair is the first similarity coefficient. If the first similarity coefficient is greater than or equal to the second similarity threshold, the first preset encoding mode is used to encode the first channel input data pair; if the first similarity coefficient is less than the second similarity threshold, and the first similarity coefficient is greater than or equal to the first similarity threshold, the second preset encoding mode is used; if the first similarity coefficient is less than the first similarity threshold, the third preset encoding mode is used.
[0221] In step S2106, the first device encodes multiple channel input data pairs according to multiple encoding modes to generate multiple bitstream information corresponding to each channel input data pair.
[0222] For example, after determining the encoding mode of each channel input data pair, the channel input data pairs are encoded based on the encoding mode to generate the bitstream information of each channel input data pair.
[0223] Optionally, in some embodiments, the encoding mode of the first channel input data pair is a first preset encoding mode, and the above step S2106 includes:
[0224] The first device determines the second MDCT coefficients of the sixth channel input data and the third MDCT coefficients of the seventh channel input data;
[0225] The first device divides the second MDCT coefficient into a first odd spectrum and a first even spectrum according to the spectral line index of the MDCT coefficient, and divides the third MDCT coefficient into a second odd spectrum and a second even spectrum according to the spectral line index.
[0226] The first device determines the first maximum correlation ratio (MCR) rotation angle based on the first odd spectrum and the second odd spectrum;
[0227] The first device determines the second MCR rotation angle based on the first even spectrum and the second even spectrum;
[0228] The first device performs MCR rotation on the first odd spectrum and the second odd spectrum based on the first MCR rotation angle to generate the third odd spectrum;
[0229] The first device performs MCR rotation on the first and second even spectra based on the second MCR rotation angle to generate a third even spectrum;
[0230] The first device generates the first code stream information based on the third odd spectrum and the third even spectrum.
[0231] For example, based on the spectral index of the MDCT coefficients, the MDCT coefficients are split. The second MDCT coefficient of the sixth channel input data pair is split into a first odd spectrum and a first even spectrum, and the third MDCT coefficient of the seventh channel input data is split into a second odd spectrum and a second even spectrum. Based on the first and second odd spectra, a first MCR rotation angle is determined. The first and second odd spectra are then subjected to MCR rotation according to the first MCR rotation angle to generate a third odd spectrum. Based on the first and second even spectra, a second MCR rotation angle is determined. The first and second even spectra are then subjected to MCR rotation according to the second MCR rotation angle to generate a third even spectrum. The third odd spectrum and the third even spectrum are then reassembled in an odd-even interleaving order to obtain the complete first bitstream information.
[0232] For example, the MDCT coefficients of the corresponding channel data are denoted as vectors L and R, respectively. Based on the spectral index, the MDCT coefficients are divided into odd and even spectra, that is, vector L is divided into Lodd and Leven spectra. odd and L even Divide vector R into R0 odd and R even Through L odd and R odd Calculate the MCR rotation angle θ1 based on L even and R evenCalculate the MCR rotation angle θ2, which can be calculated using the following formulas (11) and (12):
[0233] Among them, X l For input L odd or L even X r For input R odd Or R even θ MCR To calculate the generated MCR rotation angle.
[0234] By rotating the signal using MCR, the similarity between signals is enhanced, and then based on L... odd θ1 and R odd The output L' is obtained by performing MCR rotation. odd Based on L even R even The output L' is obtained by performing an MCR rotation with θ2. even The MCR rotation is calculated using the following formulas (13) and (14): Y l =cosθ*X l +sinθ*X r (13) Y r =-sinθ*X l +cosθ*X r (14)
[0235] Among them, Y l The output L' obtained after MCR rotation transformation odd Y r The output L' obtained after MCR rotation transformation even Finally, L' odd and L' even The data is reassembled according to the odd-even interleaving order to obtain the complete bitstream information L' for output.
[0236] Optionally, in some embodiments, the first bitstream information includes side information, which includes at least one of the following:
[0237] First encoding mode;
[0238] MCR rotation angle;
[0239] First similarity coefficient.
[0240] In some embodiments, the bitstream information may carry side information, and the encoding mode, MCR rotation angle θ1, and MCR rotation angle θ2 corresponding to the bitstream information are carried through the side information, thereby realizing the encoding method of channel transmission + side information description, realizing the encoded transmission of channel data, reducing the encoding pressure of the encoder and the encoding pressure of the decoder, and improving the encoding efficiency.
[0241] Optionally, in some embodiments, the first encoding mode is a second preset encoding mode, the first channel input data pair includes eighth channel input data and ninth channel input data, and the first channel input data pair is encoded according to the first encoding mode to generate the first bitstream information of the first channel input pair, including:
[0242] The first device determines the fourth MDCT coefficient of the eighth channel input data and the fifth MDCT coefficient of the ninth channel input data.
[0243] The first device divides the fourth MDCT coefficient into the fourth odd spectrum and the fourth even spectrum according to the spectral line index of the MDCT coefficient, and divides the fifth MDCT coefficient into the fifth odd spectrum and the fifth even spectrum according to the spectral line index.
[0244] The first device determines the third MCR rotation angle based on the fourth odd spectrum and the fifth odd spectrum;
[0245] The first device determines the fourth MCR rotation angle based on the fourth even spectrum and the fifth even spectrum;
[0246] The first device performs MCR rotation on the fourth and fifth odd spectra based on the third MCR rotation angle to generate the sixth and seventh odd spectra.
[0247] The first device performs MCR rotation on the fourth and fifth even spectra based on the fourth MCR rotation angle to generate the sixth and seventh even spectra.
[0248] The first device merges the sixth odd spectrum and the sixth even spectrum to generate the sixth MDCT coefficient;
[0249] The first device merges the seventh odd spectrum and the seventh even spectrum to generate the seventh MDCT coefficients;
[0250] The first device generates the first bitstream information based on the sixth and seventh MDCT coefficients.
[0251] For example, the fourth MDCT coefficient of the eighth channel input data and the fifth MDCT coefficient of the ninth channel input data are determined in the first channel data pair. Similarly, based on the spectral index of the MDCT coefficients, the MDCT coefficients are split, dividing the fourth MDCT coefficient into a fourth odd spectrum and a fourth even spectrum, and the fifth MDCT coefficient into a fifth odd spectrum and a fifth even spectrum. Based on the fourth and fifth odd spectra, the third MCR rotation angle is determined; based on the fourth and fifth even spectra, the fourth MCR rotation angle is determined. Based on the third MCR rotation angle, the fourth and fifth odd spectra are MCR rotated to generate the sixth and seventh odd spectra. Based on the fourth MCR rotation angle, the fourth and fifth even spectra are MCR rotated to generate the sixth and seventh even spectra. The sixth odd and sixth even spectra are interleaved and merged to generate the sixth MDCT coefficient; the seventh odd and seventh even spectra are interleaved and merged to generate the seventh MDCT coefficient. Finally, the sixth and seventh MDCT coefficients are Mid / Side compressed to generate the first bitstream information.
[0252] In some embodiments, compared to the first preset encoding mode described above, in the second preset encoding mode, after vectors L and R undergo spectral decomposition and MCR rotation transformation, the spectra of both vectors L and R are re-merged to obtain L' and R'. Then, Mid / Side compression is performed on L' and R' to obtain the compressed bitstream information.
[0253] In some embodiments, if the similarity coefficient of the channel data pair is less than thr1, the channel data pair is encoded using encoding mode 3, which is a transparent transmission mode. That is, when the similarity coefficient of the channel data pair is less than thr1, it means that there is a large difference between the two channel data of the channel data pair, and it is impossible to reduce the compression resource consumption based on MCR transformation. Therefore, the spectral coefficients of the channel data pair are not processed in any way, and the spectral coefficients of the channel data pair are transmitted directly.
[0254] Optionally, in some embodiments, the first encoding mode is a third preset encoding mode, the first channel input data pair includes tenth channel input data and eleventh channel input data, and the first channel input data pair is encoded according to the first encoding mode to generate the first bitstream information of the first channel input pair, including:
[0255] The first device determines the sixth MDCT coefficient of the tenth channel input data and the seventh MDCT coefficient of the eleventh channel input data.
[0256] The first device generates the first bitstream information based on the sixth and seventh MDCT coefficients.
[0257] For example, when the similarity coefficient of the first channel input data pair is less than the first similarity threshold, it indicates that the difference between the tenth and eleventh channel input data in the first channel input data pair is large, making it unsuitable for encoding and transmission using MCR transformation. Therefore, in this embodiment, when the first similarity coefficient is determined to be less than the first similarity threshold, the channel input data is first subjected to MDCT transformation to generate corresponding MDCT coefficients. Based on the MDCT coefficients, the corresponding first bitstream information is generated. That is, the MDCT spectral coefficients of the channel input data are not processed; instead, the MDCT coefficients are directly transmitted during transmission as the bitstream information of the first channel input data pair.
[0258] In step S2107, the first device sends multiple stream information to the second device.
[0259] For example, the first device encodes the channel input data pairs based on the encoding mode of each channel input data pair to generate bitstream information, and then sends the bitstream information to the second device, which decodes the bitstream information to generate multi-channel input data.
[0260] In some embodiments, the names of information, etc., are not limited to the names described in the embodiments. Terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "symbol", "codebook", "codeword", "codepoint", "bit", "data", "program", and "chip" can be used interchangeably.
[0261] In some embodiments, the terms "codebook," "codeword," and "precoding matrix" can be used interchangeably. For example, a codebook can be a collection of one or more codewords / precoding matrices.
[0262] In some embodiments, the terms "uplink", "uplink", and "physical uplink" can be used interchangeably, as can the terms "downlink", "downlink", and "physical downlink", as well as the terms "sidelink", "sidelink", "sidelink communication", "sidelink communication", "direct connection", "direct link", "direct communication", and "direct link communication".
[0263] In some embodiments, the terms “downlink control information (DCI),” “downlink (DL) assignment,” “DL DCI,” “uplink (UL) grant,” and “UL DCI” can be used interchangeably.
[0264] In some embodiments, terms such as "physical downlink shared channel (PDSCH)" and "DL data" can be used interchangeably, as can terms such as "physical uplink shared channel (PUSCH)" and "UL data".
[0265] In some embodiments, the terms “radio”, “wireless”, “radio access network (RAN)”, “access network (AN)”, and “RAN-based” can be used interchangeably.
[0266] In some embodiments, the terms "search space", "search space set", "search space configuration", "search space set configuration", "control resource set (CORESET)", and "CORESET configuration" can be used interchangeably.
[0267] In some embodiments, the terms "synchronization signal (SS)," "synchronization signal block (SSB)," "reference signal (RS)," "pilot," and "pilot signal" can be used interchangeably.
[0268] In some embodiments, terms such as “moment,” “point in time,” “time,” and “time location” can be used interchangeably, as can terms such as “duration,” “segment,” “time window,” “window,” and “time.”
[0269] In some embodiments, the terms "component carrier (CC)," "cell," "frequency carrier," and "carrier frequency" can be used interchangeably.
[0270] In some embodiments, the terms “resource block (RB)”, “physical resource block (PRB)”, “sub-carrier group (SCG)”, “resource element group (REG)”, “PRB pair”, “RB pair”, “resource element (RE)”, and “sub-carrier” can be used interchangeably.
[0271] In some embodiments, terms such as wireless access scheme and waveform can be used interchangeably.
[0272] In some embodiments, the terms "precoding", "precoder", "weight", "precoding weight", "quasi-co-location (QCL)", "transmission configuration indication (TCI) status", "spatial relation", "spatial domain filter", "transmission power", "phase rotation", "antenna port", "antenna port group", "layer", "the number of layers", "rank", "resource", "resource set", "resource group", "beam", "beam width", "beam angular degree", "antenna", "antenna element", and "panel" can be used interchangeably.
[0273] In some embodiments, the terms “frame”, “radio frame”, “subframe”, “slot”, “sub-slot”, “mini-slot”, “symbol”, “symbol”, and “transmission time interval (TTI)” can be used interchangeably.
[0274] In some embodiments, “get,” “obtain,” “receive,” “transmit,” “bidirectional transmission,” and “send and / or receive” can be used interchangeably and can be interpreted as receiving from other entities, obtaining from protocols, obtaining from higher layers, obtaining through self-processing, or autonomous implementation, among other meanings.
[0275] In some embodiments, terms such as “send,” “transmit,” “report,” “distribute,” “transfer,” “bidirectional transmission,” “send and / or receive” can be used interchangeably.
[0276] In some embodiments, terms such as "certain," "preset," "default," "set," "indicated," "a certain," "any," and "first" can be used interchangeably. "Certain A," "preset A," "default A," "set A," "indicated A," "a certain A," "any A," and "first A" can be interpreted as A pre-defined in a protocol or the like, or as A obtained through setting, configuration, or instruction, or as specific A, a certain A, any A, or first A, but are not limited thereto.
[0277] In some embodiments, the determination or judgment can be made by a value represented by 1 bit (0 or 1), or by a true or false value (boolean), or by a comparison of numerical values (e.g., a comparison with a predetermined value), but is not limited thereto.
[0278] In some embodiments, "not expecting to receive" can be interpreted as not receiving on time domain resources and / or frequency domain resources, or as not performing subsequent processing on the data after receiving it; "not expecting to send" can be interpreted as not sending, or as sending but not expecting the receiver to respond to the sent content.
[0279] The above method determines multiple similarity coefficients between the input data of each channel in the multi-channel input data. Based on these similarity coefficients, the multiple input data pairs with the highest similarity coefficients are identified, along with multiple similarity coefficients corresponding to each pair. Based on these similarity coefficients, multiple encoding modes corresponding to each pair are determined. The multiple input data pairs are then encoded using these encoding modes, generating multiple bitstream information corresponding to each pair. Therefore, by using different encoding modes based on the similarity between the input data channels, the data load of the bitstream information during channel data encoding is reduced, the encoding pressure on the encoder is decreased, and the encoding efficiency is improved.
[0280] Figure 3 is a flowchart illustrating an audio data encoding method according to an embodiment of the present disclosure. As shown in Figure 3, the embodiment of the present disclosure relates to an audio data encoding method, executed by a first device, the method comprising:
[0281] Step S3101: Determine multiple similarity coefficients between the input data of each channel in the multi-channel input data.
[0282] In some embodiments, step S3101 above includes:
[0283] Determine the multiple modulation discrete cosine transform (MDCT) coefficients that correspond one-to-one with the multi-channel input data;
[0284] Multiple similarity coefficients are determined based on multiple MDCT coefficients.
[0285] In some embodiments, the step of "determining the multiple modulation discrete cosine transform (MDCT) coefficients corresponding one-to-one with the multi-channel input data" includes:
[0286] The first channel input data is processed by frame segmentation and windowing to generate multi-frame channel input sub-data. The first channel input data is any channel input data in the multi-channel input data.
[0287] Based on the multi-frame channel input sub-data, determine the first MDCT spectral coefficients and the number of first MDCT coefficient points of the first channel input data;
[0288] The first MDCT coefficients of the first channel input data are generated based on the first MDCT spectral coefficients and the number of first MDCT coefficient points.
[0289] The optional implementation of step S3101 can be found in the optional implementation of steps S2101-S2102 in Figure 2, as well as other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0290] Step S3102: Based on multiple similarity coefficients, determine the multiple channel input data pairs with the highest similarity coefficients, and the multiple similarity coefficients corresponding to the multiple channel input data pairs.
[0291] In some embodiments, the above step "determining multiple similarity coefficients based on multiple MDCT coefficients" includes:
[0292] The second similarity coefficient is determined using the following formula, where the second similarity coefficient can be any number of similarity coefficients:
[0293] Where, sim cosθ The second similarity coefficient is given by i, where i is the number of the vocal tract input sub-data, N is the number of MDCT coefficient points, and X is the second similarity coefficient. i Y represents the second MDCT spectral coefficient of the second channel input data in the second channel input data pair. i The third MDCT spectral coefficient is the third channel input data in the second channel input data pair. The second channel input data pair is any first channel input data pair, which includes the second channel input data and the third channel input data.
[0294] The optional implementation of step S3102 can be found in the optional implementation of step S2103 in Figure 2 and other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0295] Step S3103: Based on multiple similarity coefficients, determine multiple encoding modes that correspond one-to-one with the input data pairs of multiple audio channels.
[0296] In some embodiments, step S3103 above includes:
[0297] Determine a first similarity threshold and a second similarity threshold. The second similarity threshold is less than 1, the second similarity threshold is greater than the first similarity threshold, and the first similarity threshold is greater than 0.
[0298] Multiple encoding patterns are determined based on the relationship between multiple similarity coefficients and the first and second similarity thresholds.
[0299] In some embodiments, step S3103 above includes:
[0300] If the first similarity coefficient is greater than or equal to the second similarity threshold, the first encoding mode corresponding to the first similarity coefficient is determined as the first set encoding mode, and the first similarity coefficient can be any multiple similarity coefficients;
[0301] If the first similarity coefficient is less than the second similarity threshold, and the first similarity coefficient is greater than or equal to the first similarity threshold, then the first coding model is determined to be the second set coding mode.
[0302] If the first similarity coefficient is less than the first similarity threshold, the first encoding mode is determined to be the third set encoding mode.
[0303] In some embodiments, the above step "determining the first similarity threshold and the second similarity threshold" includes:
[0304] Based on the first similarity coefficient, determine the Holt predictor term for the Holt exponential smoothing prediction;
[0305] Determine the channel energy ratio between the fourth channel input data and the fifth channel input data;
[0306] Obtain the first similarity benchmark threshold corresponding to the first similarity threshold, and obtain the second similarity benchmark threshold corresponding to the second similarity threshold;
[0307] The first similarity threshold is determined based on the Holt prediction term, the vocal tract energy ratio, and the first similarity benchmark threshold.
[0308] The second similarity threshold is determined based on the Holt prediction term, the vocal tract energy ratio, and the second similarity baseline threshold.
[0309] In some embodiments, the step "determining the Holt predictor term for Holt exponential smoothing prediction based on the first similarity coefficient" includes:
[0310] Based on the first similarity coefficient, determine the Holt baseline term and Holt trend term for Holt index smoothing prediction;
[0311] Determine the Holt forecast term based on the Holt baseline term and the Holt trend term.
[0312] The optional implementation of step S3103 can be found in the optional implementation of steps S2104-S2105 in Figure 2, as well as other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0313] Step S3104: Encode multiple channel input data pairs according to multiple encoding modes to generate multiple bitstream information corresponding to each channel input data pair.
[0314] In some embodiments, the first encoding mode is a first preset encoding mode, and the above step "the first channel input data pair includes sixth channel input data and seventh channel input data, and the first channel input data pair is encoded according to the first encoding mode to generate the first bitstream information of the first channel input pair" includes:
[0315] Determine the second MDCT coefficients for the sixth channel input data, and determine the third MDCT coefficients for the seventh channel input data;
[0316] Based on the spectral index of the MDCT coefficients, the second MDCT coefficients are divided into the first odd spectrum and the first even spectrum, and based on the spectral index, the third MDCT coefficients are divided into the second odd spectrum and the second even spectrum.
[0317] The first maximum correlation ratio (MCR) rotation angle is determined based on the first odd spectrum and the second odd spectrum.
[0318] The second MCR rotation angle is determined based on the first even spectrum and the second even spectrum;
[0319] Based on the first MCR rotation angle, the first odd spectrum and the second odd spectrum are rotated by MCR to generate the third odd spectrum;
[0320] Based on the second MCR rotation angle, the first even spectrum and the second even spectrum are rotated by MCR to generate the third even spectrum;
[0321] The first code stream information is generated based on the third odd spectrum and the third even spectrum.
[0322] In some embodiments, the first encoding mode is a second preset encoding mode, and the above step "the first channel input data pair includes eighth channel input data and ninth channel input data, and the first channel input data pair is encoded according to the first encoding mode to generate the first bitstream information of the first channel input pair" includes:
[0323] Determine the fourth MDCT coefficient of the eighth channel input data, and determine the fifth MDCT coefficient of the ninth channel input data;
[0324] Based on the spectral index of the MDCT coefficients, the fourth MDCT coefficient is divided into the fourth odd spectrum and the fourth even spectrum, and based on the spectral index, the fifth MDCT coefficient is divided into the fifth odd spectrum and the fifth even spectrum.
[0325] The third MCR rotation angle is determined based on the fourth and fifth odd spectra.
[0326] The fourth MCR rotation angle is determined based on the fourth even spectrum and the fifth even spectrum;
[0327] Based on the third MCR rotation angle, the fourth and fifth odd spectra are rotated by MCR to generate the sixth and seventh odd spectra.
[0328] Based on the fourth MCR rotation angle, the fourth even spectrum and the fifth even spectrum are rotated by MCR to generate the sixth even spectrum and the seventh even spectrum.
[0329] The sixth odd spectrum and the sixth even spectrum are combined to generate the sixth MDCT coefficients;
[0330] The seventh odd spectrum and the seventh even spectrum are combined to generate the seventh MDCT coefficients;
[0331] The first bitstream information is generated based on the sixth and seventh MDCT coefficients.
[0332] In some embodiments, the first encoding mode is a third preset encoding mode, and the above step "the first channel input data pair includes tenth channel input data and eleventh channel input data, and the first channel input data pair is encoded according to the first encoding mode to generate the first bitstream information of the first channel input pair" includes:
[0333] Determine the sixth MDCT coefficient of the tenth channel input data, and determine the seventh MDCT coefficient of the eleventh channel input data;
[0334] The first bitstream information is generated based on the sixth and seventh MDCT coefficients.
[0335] In some embodiments, the first bitstream information includes side information, and the first bitstream information includes any one or more bitstream information, the side information including at least one of the following:
[0336] First encoding mode;
[0337] MCR rotation angle;
[0338] First similarity coefficient.
[0339] The optional implementation of step S3104 can be found in the optional implementation of step S2106 in Figure 2, as well as other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0340] Step S3105: Send multiple bitstream information to the second device.
[0341] In some embodiments, multiple bitstream information is used to instruct a second device to decode multiple bitstream information to generate multiple channel input data pairs.
[0342] The optional implementation of step S3105 can be found in the optional implementation of step S2107 in Figure 2, as well as other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0343] The above method determines multiple similarity coefficients between the input data of each channel in the multi-channel input data. Based on these similarity coefficients, the multiple input data pairs with the highest similarity coefficients are identified, along with multiple similarity coefficients corresponding to each pair. Based on these similarity coefficients, multiple encoding modes corresponding to each pair are determined. The multiple input data pairs are then encoded using these encoding modes, generating multiple bitstream information corresponding to each pair. Therefore, by using different encoding modes based on the similarity between the input data channels, the data load of the bitstream information during channel data encoding is reduced, the encoding pressure on the encoder is decreased, and the encoding efficiency is improved.
[0344] Figure 4 is a flowchart illustrating an audio data encoding method according to an embodiment of the present disclosure. As shown in Figure 4, this embodiment relates to an audio data encoding method executed by a second device, the method comprising:
[0345] Step S4101: Receive multiple bitstream information sent by the first device.
[0346] In some embodiments, the multiple bitstream information is generated by the first device encoding multiple channel input data according to multiple encoding modes, and the multiple encoding modes are determined by the first device based on multiple similarity coefficients between the multiple channel input data.
[0347] In some embodiments, the first bitstream information includes side information, and the first bitstream information includes any one or more bitstream information, the side information including at least one of the following:
[0348] First encoding mode;
[0349] MCR rotation angle;
[0350] First similarity coefficient.
[0351] The optional implementation of step S4101 can be found in the optional implementation of step S2107 in Figure 2, as well as other related parts in the embodiments involved in Figure 2, which will not be repeated here.
[0352] Step S4102: Decode multiple bitstream information to generate multiple channel input data.
[0353] For example, the first device encodes the channel input data pairs based on the encoding mode of each channel input data pair to generate bitstream information, and then sends the bitstream information to the second device, which decodes the bitstream information to generate multi-channel input data.
[0354] In this manner, the first device encodes the multi-channel input data using different encoding modes based on the similarity coefficient, generating a bitstream. The second device then decodes this bitstream to generate the multi-channel input data. This reduces the decoding burden on the second device and improves decoding efficiency while maintaining the corresponding sound quality of the channel input data.
[0355] Figure 5A is a flowchart illustrating an audio data encoding method according to an embodiment of the present disclosure. As shown in Figure 5A, this disclosure relates to an audio data encoding method, which includes:
[0356] Step S5101: Preprocess and perform MDCT transformation on the multi-channel data to generate MDCT coefficients for each channel.
[0357] For example, this embodiment is applied to a first device, which is equipped with an encoder. The encoder performs frame-by-frame windowing processing on the input Z channel data in the front-end section, where Z is the number of channels in the currently input multi-channel data. Channel A data is any multi-channel data. The A channel data is subjected to frame-by-frame windowing processing to generate multi-frame sub-channel data. Based on the multi-frame sub-channel data, the MDCT coefficients of the A channel data are calculated and determined.
[0358] Step S5102: Based on the MDCT coefficients, perform similarity evaluation on the multi-channel data to generate multiple channel data pairs with the highest similarity coefficients and their corresponding multiple similarity coefficients.
[0359] For example, based on the MDCT coefficients of each channel data, pairwise comparisons are performed on each channel data pair to determine the similarity coefficient between the channel data pairs. Then, the two channel data pairs with the highest similarity coefficients are selected to form a channel data pair. For example, the multi-channel data are A, B, C, D, E, and F; based on the MDCT coefficients of each channel data, the multi-channel data are compared pairwise to generate initial channel data pairs: AB, AC, AD, AE, AF, BC, BD, BE, BF, CD, CE, CF, DE, DF, and EF; after determining multiple initial similarity coefficients for each initial channel data pair, the channel data pair with the highest similarity coefficient, AB, CF, and DE, and their corresponding similarity coefficients a, b, and c are obtained.
[0360] In some embodiments, the similarity of each channel data is evaluated based on cosine similarity. For example, the cosine similarity between two channel data can be determined based on the following formula (3):
[0361] Where, sim cos(θ) This is the similarity coefficient between two vocal tract data points; the similarity coefficient is the cosine similarity, x. i and y i Let N be the MDCT spectral coefficients of the two channel data, N be the number of MDCT coefficient points in each frame of channel data, and i be the channel data number of the current subframe after framing the channel input data. The similarity between the two channel data is evaluated using the above formula (3), and the generated cosine similarity value ranges from [-1, 1]. Here, -1 indicates that the vectors corresponding to the two channel data are completely opposite, 0 indicates that the vectors corresponding to the two channel data are orthogonal, and 1 indicates that the vectors corresponding to the two channel data are completely identical in direction. By performing pairwise statistics on the above multi-channel data in this way, the similarity coefficient of each channel data pair is obtained.
[0362] Step S5103: Determine the encoding mode of each channel data pair based on the relationship between multiple similarity coefficients and the first similarity threshold and the second similarity threshold.
[0363] Obtain from largest to smallest similarity After processing each channel data pair, the similarity coefficient of each channel data pair is compared with similarity thresholds thr1 and thr2, where 0... <thr1<thr2<1。
[0364] (1) Compare the similarity coefficient and similarity threshold of each channel data pair. If the similarity coefficient is ≥thr2, then select encoding mode 1 and encode the channel data pair.
[0365] (2) If thr1≤similarity coefficient<thr2, then select encoding mode 2 and encode the audio channel data pair.
[0366] (3) If the similarity coefficient is <thr1, then select encoding mode 3 and encode the audio channel data pair.
[0367] Based on the above steps (1)-(3), the encoding processing of each channel data pair is completed, and the code stream information corresponding to the multi-channel input data is generated.
[0368] In some embodiments, the similarity thresholds thr1 and thr2 are adjusted based on historical similarity information of the vocal tract data pairs. For example, Holt exponential smoothing prediction is used to predict the trend of similarity between vocal tract data pairs to adapt to trend changes in the vocal tract data.
[0369] Holt exponential smoothing forecasts consist of a baseline term and a trend term, which are calculated using the following formulas (4) and (5): L t =α*sim cos(θ) (t)+(1-α)*L t-1 +T t-1 (4) T t =β*[L t -L t-1 ]+(1-β)*T t-1 (5)
[0370] Among them, L t L is the baseline term at time t. t A weighted average can be calculated based on the difference between the historical baseline and the latest observation. The weight parameter 1 can be determined by the smoothing parameter α (0 < α < 1). The value of α can be adjusted according to the current encoding requirements. When the value of α approaches 1, it indicates that the latest observation has a greater weight; when the value of α approaches 0, it indicates that the historical baseline has a greater weight. cos(θ) (t) represents the cosine similarity coefficient generated at the current time using the above method, and L t-1 T is the reference term for the previous historical moment. t-1 This represents the trend at the previous historical moment. (T) t Let t be the trend term at time t. This trend term can be updated using a weighted average. When updating the trend term, the weight parameter 2 is a smoothing parameter β (0 < β < 1). β controls the sensitivity of the model to trend changes. For example, the higher the value of β, the stronger the sensitivity of the model to trend changes.
[0371] The encoding pattern of the vocal tract data pair is determined by the relationship between the cosine similarity of the vocal tract data pair and the similarity thresholds thr1 and thr2. thr1 and thr2 can be calculated using the following formulas (6) and (7): thr2=γ*sigmoid(ILD)+V+A (6) thr1=ε*sigmoid(ILD)+V+B (7)
[0372] Where thr1 and thr2 are similarity thresholds, B is the first threshold baseline value corresponding to thr1, and A is the second threshold baseline value corresponding to thr2. ILD is the energy ratio of the vocal tract data between pairs. γ > 0 is weight parameter 3, which adjusts the influence of ILD on the similarity threshold thr2. ε < 0 is weight parameter 4, which adjusts the influence of ILD on the similarity threshold thr1. V is the Holt prediction component.
[0373] For example, the function sigmoid() in formulas (6) and (7) above is a squeezing function, and the expression of the squeezing function is as follows: formula (10):
[0374] Figure 5B is a schematic diagram of a sigmoid function curve according to an embodiment of the present disclosure. As shown in Figure 5B, in the above-mentioned sigmoid() squeezing function, the maximum value of the curve corresponding to the squeezing function is controlled by a constant a, the steepness of the curve corresponding to the squeezing function is controlled by a coefficient k, and the horizontal shift value of the curve corresponding to the squeezing function is controlled by a constant h.
[0375] For example, the sigmoid term does not have a significant suppressive effect when the ILD difference is small. However, when the difference is large and increases rapidly, the sigmoid term suppresses the transformation effect caused by the ILD difference. In the above formulas (4) and (5), γ > 0 and ε < 0. This weighting parameter is used to compress the probability interval of encoding mode 2 when the ILD difference is large. This helps to suppress the noise amplification problem caused by the Mid / side compression operation when noise is introduced during the channel data encoding process, and reduces the noise impact on the decoding end.
[0376] For example, in this embodiment, the Holt predictor is used to suppress mode jumps caused by slight perturbations near the reference value. The Holt predictor can be calculated using the following formula (8):
[0377] Where k is a constant greater than 0, u L <0 and u H >0 is the limiting constant, which is determined by u L and y HThis limits the upper and lower bounds of the holt predictor V's ability to adjust the threshold.
[0378] For example, Figure 5C is a schematic diagram illustrating the dynamic adjustment of the similarity threshold according to an embodiment of the present disclosure. As shown in Figure 5C, V is initialized in the above formula (8). t =0, V t The condition for the value to change is the previous two frames T. t-2 and T t-1 When the overall trend is stable and the direction of the current similarity value changes compared to previous frames, V dynamically adjusts the similarity coefficient. As shown in Figure 5C, based on the baseline value A = 0.7, the dynamic adjustment effect of the Holt prediction component on the similarity threshold is illustrated. From frames 16 to 25, the similarity coefficient remains stable, but a slight perturbation below A occurs in frame 24. At this point, V's adjustment reduces two avoidable coding mode switches. The similarity coefficients in other frames are far from the baseline value, ensuring the selection of the applied coding mode. V considers the current level and trend changes of the data and adaptively fine-tunes according to the trend direction, thus resisting the influence of minor perturbations near the baseline value and avoiding overreaction to insignificant changes.
[0379] Figure 5D is a schematic diagram of the ODG score according to an embodiment of this disclosure. As shown in Figure 5D, the ODG (Objective Difference Grade) is used to evaluate the specific impact of the above encoding method on encoder performance. The overall runtime of the encoder as a single system is observed to test how changes to the modified modules affect the overall efficiency and processing speed of the encoder. The test content is the total time required for the encoder to complete every 300 frames. Simultaneously, the runtime of the encoder before and after modification is compared at the same bitrate and encoding the same audio, thereby determining the impact of the above audio data encoding method on encoder performance.
[0380] For example, Figure 5E is a schematic diagram of encoder runtime according to an embodiment of the present disclosure. As shown in Figure 5E, Ori represents the encoder before improvement, and Var represents the percentage of runtime of the encoder after improvement. For example, tests were conducted at different bitrates. Based on the above audio data encoding mode, it was determined that the time reduction rate in the encoder runtime percentage was relatively stable, decreasing by about 11% in all cases. This indicates that the above audio data encoding method reduces the time in the audio data encoding process and improves the encoder's encoding efficiency.
[0381] In some embodiments, after MCR transformation, the energy ratio ILD between the two channel data pairs is 1. However, MCR transformation is not performed across the entire frequency band during channel data processing; therefore, it is typically performed in the core frequency band at low frequencies. This means the energy level relationship between the channel data needs to be considered. In this embodiment, the energy ratio between the channel data is determined using the following formula (9):
[0382] The ILD is calculated using the sqrt() function, x i Let y be the energy value of the i-th subframe channel data in the x-channel data. i Let be the energy value of the i-th subframe channel data in the y-channel data.
[0383] Step S5104: Encode the channel data pairs according to the encoding mode to generate multi-channel data bitstream information.
[0384] For example, in this embodiment, there are three encoding modes. After determining the encoding mode of the channel data pair according to the above steps, the channel data pair is encoded according to the corresponding encoding mode to generate the first bitstream information. The above steps are repeated to encode each channel data pair to generate the bitstream information of multi-channel data.
[0385] Figure 5F is a flowchart illustrating encoding mode 1 according to an embodiment of the present disclosure. As shown in Figure 5F, when the similarity coefficient is determined to be ≥thr2, encoding mode 1 is used to encode the vocal channel data pairs.
[0386] For example, the MDCT coefficients of the corresponding channel data are denoted as vectors L and R, respectively. Based on the spectral index, the MDCT coefficients are divided into odd and even spectra, that is, vector L is divided into Lodd and Leven spectra. odd and L even Divide vector R into R0 odd and R even Through L odd and R odd Calculate the MCR rotation angle θ1 based on L even and R even Calculate the MCR rotation angle θ2, which can be calculated using the following formulas (11) and (12):
[0387] Among them, X l For input L odd or L even X r For input R odd or R even θ MCRTo calculate the generated MCR rotation angle.
[0388] By rotating the signal using MCR, the similarity between signals is enhanced, and then based on L... odd θ1 and R odd The output L' is obtained by performing MCR rotation. odd Based on L even R even The output L' is obtained by performing an MCR rotation with θ2. even The MCR rotation is calculated using the following formulas (13) and (14): Y l =cosθ*X l +sinθ*X r (13) Y r =-sinθ*X l +cosθ*X r (14)
[0389] Among them, Y l The output L' obtained after MCR rotation transformation odd Y r The output L' obtained after MCR rotation transformation even Finally, L' odd and L' even The data is reassembled according to the odd-even interleaving order to obtain the complete bitstream information L, which is then output.
[0390] In some embodiments, the bitstream information may carry side information, and the encoding mode, MCR rotation angle θ1, and MCR rotation angle θ2 corresponding to the bitstream information are carried through the side information, thereby realizing the encoding method of channel transmission + side information description, realizing the encoded transmission of channel data, reducing the encoding pressure of the encoder and the encoding pressure of the decoder, and improving the encoding efficiency.
[0391] Figure 5G is a flowchart illustrating encoding mode 2 according to an embodiment of the present disclosure. As shown in Figure 5G, when it is determined that thr1 ≤ similarity coefficient < thr2, encoding mode 2 is used to encode the vocal channel data pair.
[0392] For example, compared to encoding mode 1 above, in encoding mode 2, after spectral decomposition and MCR rotation transformation of vectors L and R, the spectra of both vectors L and R are re-merged to obtain L' and R'. Then, Mid / Side compression is performed on L' and R' to obtain the compressed bitstream information.
[0393] In some embodiments, if the similarity coefficient of the channel data pair is less than thr1, the channel data pair is encoded using encoding mode 3, which is a transparent transmission mode. That is, when the similarity coefficient of the channel data pair is less than thr1, it means that there is a large difference between the two channel data of the channel data pair, and it is impossible to reduce the compression resource consumption based on MCR transformation. Therefore, the spectral coefficients of the channel data pair are not processed in any way, and the spectral coefficients of the channel data pair are transmitted directly.
[0394] Figure 5H is a flowchart illustrating an audio data encoding method according to an embodiment of the present disclosure. As shown in Figure 5H, this disclosure relates to an audio data encoding method, which includes:
[0395] (1) Input multi-channel data;
[0396] (2) Preprocess and perform MDCT transformation on the input multichannel data to generate MDCT coefficients for each channel data.
[0397] (3) Based on the MDCT coefficients of each channel data, evaluate the similarity between the channel data; obtain the multiple channel data pairs with the highest similarity, and the similarity coefficient of each channel data pair;
[0398] (4) Based on the similarity coefficient of each channel data pair, determine the encoding mode of the channel data pair. If the similarity coefficient is ≥thr2, then select encoding mode 1 to encode the channel data pair; if the similarity coefficient is ≤thr1 and the similarity coefficient is <thr2, then select encoding mode 2 to encode the channel data pair; if the similarity coefficient is <thr1, then select encoding mode 3 to encode the channel data pair.
[0399] (5) Based on the encoding mode determined by the above method, the audio channel data pairs are encoded to generate code stream information.
[0400] The above approach employs a coding mode more suitable for low bitrate scenarios: "few channels transmitted + difference description information." By combining the characteristics of the channel data, different coding modes are used to encode the channel data, and MCR (Multi-Channel Recognition) is introduced to maximize the similarity between channels. This optimizes coding capabilities in low bitrate scenarios, significantly reduces encoder runtime, and improves the output audio quality.
[0401] This disclosure also provides an apparatus for implementing any of the above methods. For example, an apparatus is provided that includes units or modules for implementing the steps performed by the terminal in any of the above methods. Alternatively, another apparatus is provided that includes units or modules for implementing the steps performed by a network device (e.g., an access network device, a core network functional node, a core network device, etc.) in any of the above methods.
[0402] It should be understood that the division of units or modules in the above device is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the units or modules in the device can be implemented by a processor calling software: for example, the device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of the units or modules in the above device. The processor can be, for example, a general-purpose processor, such as a Central Processing Unit (CPU) or a microprocessor, and the memory can be internal or external to the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits. The functionality of some or all of the units or modules can be achieved through the design of these hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC). The functionality of some or all of the units or modules is achieved through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a programmable logic device (PLD), such as a field-programmable gate array (FPGA). This PLD can include a large number of logic gates, and the connection relationships between these logic gates are configured through configuration files, thereby achieving the functionality of some or all of the units or modules. All units or modules of the above device can be implemented entirely through processor-called software, entirely through hardware circuits, or partially through processor-called software with the remaining parts implemented through hardware circuits.
[0403] In this embodiment, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction read and execute capabilities, such as a Central Processing Unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. The logical relationships of the aforementioned hardware circuits are fixed or reconfigurable. For example, the processor is a hardware circuit implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a Neural Network Processing Unit (NPU), a Tensor Processing Unit (TPU), or a Deep Learning Processing Unit (DPU).
[0404] Figure 6 is a schematic diagram of the structure of a first device according to an embodiment of the present disclosure. As shown in Figure 6, the first device 6100 may include: a processing module 6101, a processing module 6102, a processing module 6103, and a processing module 6104. In some embodiments, the processing module 6101 is configured to determine multiple similarity coefficients between the channel input data in the multi-channel input data; the processing module 6102 is configured to determine multiple channel input data pairs with the highest similarity coefficients, and multiple similarity coefficients corresponding one-to-one to the multiple channel input data pairs, based on the multiple similarity coefficients; the processing module 6103 is configured to determine multiple encoding modes corresponding one-to-one to the multiple channel input data pairs, based on the multiple similarity coefficients; and the processing module 6104 is configured to encode the multiple channel input data pairs according to the multiple encoding modes to generate multiple bitstream information corresponding one-to-one to the multiple channel input data pairs. Optionally, the processing modules 6101, 6102, 6103 and 6104 described above are used to perform at least one of the communication steps such as determination and / or generation performed by the first device in any of the above methods, which will not be described in detail here.
[0405] In some embodiments, processing modules 6101, 6102, 6103, and 6104 may each include an execution module and a determination module. The execution module and the determination module may be separate or integrated together. Optionally, the execution module may be interchangeable with an executor, and the determination module may be interchangeable with a calculator.
[0406] Figure 7 is a schematic diagram of the structure of a second device according to an embodiment of the present disclosure. As shown in Figure 7, the second device 7100 may include a transceiver module 7101 and a processing module 7102. In some embodiments, the transceiver module 7101 is configured to receive multiple bitstream information sent by a first device, wherein the multiple bitstream information is generated by the first device encoding multiple channel input data according to multiple encoding modes, and the multiple encoding modes are determined by the first device based on multiple similarity coefficients between the multiple channel input data; the processing module 7102 is configured to decode the multiple bitstream information to generate multiple channel input data. Optionally, the transceiver module 7101 and the processing module 7102 are used to perform at least one of the communication steps such as determination and / or acquisition performed by the second device in any of the above methods, which will not be elaborated here.
[0407] In some embodiments, the transceiver module 7101 may include a receiving module and a transmitting module, which may be separate or integrated. Optionally, the transmitting module may be interchangeable with a transmitter. The receiving module may be interchangeable with a receiver.
[0408] In some embodiments, the processing module 7102 may include an execution module and a determination module. The execution module and the acquisition module may be separate or integrated together. Optionally, the execution module may be interchangeable with an executor.
[0409] Figure 8 is a schematic diagram of the structure of a communication device 8100 according to an embodiment of this disclosure. The communication device 8100 can be a network device (e.g., access network device, core network device, etc.), a terminal (e.g., user equipment, etc.), a chip, chip system, or processor that supports the network device in implementing any of the above methods, or a chip, chip system, or processor that supports the terminal in implementing any of the above methods. The communication device 8100 can be used to implement the methods described in the above method embodiments; for details, please refer to the descriptions in the above method embodiments.
[0410] As shown in Figure 8, the communication device 8100 includes one or more third processors 8101. The third processor 8101 can be a general-purpose processor or a dedicated processor, such as a baseband processor or a central processing unit (CPU). The baseband processor can be used to process communication protocols and communication data, while the CPU can be used to control communication devices (e.g., base stations, baseband chips, terminal devices, terminal device chips, DUs or CUs, etc.), execute programs, and process program data. Optionally, the communication device 8100 can be used to execute any of the above methods. Optionally, one or more third processors 8101 can be used to invoke instructions to cause the communication device 8100 to execute any of the above methods.
[0411] In some embodiments, the communication device 8100 further includes one or more third transceivers 8102. When the communication device 8100 includes one or more third transceivers 8102, the third transceiver 8102 performs at least one of the communication steps such as sending and / or receiving in the above method, and the third processor 8101 performs at least one of the other steps. In optional embodiments, the transceiver may include a receiver and / or a transmitter, which may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, transceiver circuit, interface circuit, interface, etc., can be used interchangeably; the terms transmitter, sending unit, transmitter, sending circuit, etc., can be used interchangeably; and the terms receiver, receiving unit, receiver, receiving circuit, etc., can be used interchangeably.
[0412] In some embodiments, the communication device 8100 further includes one or more third memories 8103 for storing data. Optionally, all or part of the third memories 8103 may be located outside the communication device 8100. In optional embodiments, the communication device 8100 may include one or more first interface circuits 8104. Optionally, the first interface circuit 8104 is connected to the third memory 8103, and the first interface circuit 8104 can be used to receive data from the third memory 8103 or other devices, and can be used to send data to the third processor 8101 or other devices. For example, the first interface circuit 8104 can read data stored in the third memory 8103 and send the data to the third processor 8101.
[0413] The communication device 8100 described in the above embodiments may be a network device or a terminal, but the scope of the communication device 8100 described in this disclosure is not limited thereto, and the structure of the communication device 8100 may not be limited by FIG8. The communication device may be a standalone device or may be part of a larger device. For example, the communication device may be: (1) a standalone integrated circuit IC, or chip, or chip system or subsystem; (2) a collection of one or more ICs, optionally, the IC collection may also include storage components for storing data and programs; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, terminal device, smart terminal device, cellular phone, wireless device, handheld device, mobile unit, vehicle device, network device, cloud device, artificial intelligence device, etc.; (6) others, etc.
[0414] Figure 9 is a schematic diagram of the structure of chip 8200 according to an embodiment of the present disclosure. For cases where the communication device 8100 can be a chip or a chip system, the schematic diagram of chip 8200 shown in Figure 9 can be referenced, but is not limited thereto.
[0415] Chip 8200 includes one or more fourth processors 8201. Chip 8200 is used to perform any of the above methods.
[0416] In some embodiments, chip 8200 further includes one or more second interface circuits 8202. Optionally, terms such as interface circuit, interface, and transceiver pin can be used interchangeably. In some embodiments, chip 8200 further includes one or more fourth memories 8203 for storing data. Optionally, all or part of the fourth memories 8203 may be located outside chip 8200. Optionally, the second interface circuit 8202 is connected to the fourth memories 8203, and the second interface circuit 8202 can be used to receive data from the fourth memories 8203 or other devices, and the second interface circuit 8202 can be used to send data to the fourth memories 8203 or other devices. For example, the second interface circuit 8202 can read data stored in the fourth memories 8203 and send the data to the fourth processor 8201.
[0417] In some embodiments, the second interface circuit 8202 performs at least one of the communication steps such as sending and / or receiving in the above-described method. For example, the second interface circuit 8202 performing the communication steps such as sending and / or receiving in the above-described method refers to the second interface circuit 8202 performing data interaction between the fourth processor 8201, the chip 8200, the fourth memory 8203, or the transceiver device. In some embodiments, the fourth processor 8201 performs at least one of the other steps.
[0418] The modules and / or devices described in the various embodiments, such as virtual devices, physical devices, and chips, can be combined or separated arbitrarily as needed. Optionally, some or all steps can also be performed collaboratively by multiple modules and / or devices, which is not limited here.
[0419] This disclosure also proposes a storage medium storing instructions that, when executed on a communication device 8100, cause the communication device 8100 to perform any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but not limited thereto; it may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but not limited thereto; it may also be a temporary storage medium.
[0420] This disclosure also provides a program product that, when executed by the communication device 8100, causes the communication device 8100 to perform any of the above methods. Optionally, the program product is a computer program product.
[0421] This disclosure also proposes a computer program that, when run on a computer, causes the computer to perform any of the above methods.
Claims
1. An audio data encoding method, characterized in that, Performed by a first device, the method includes: Determine multiple similarity coefficients between the input data of each channel in the multi-channel input data; Based on the multiple similarity coefficients, determine the multiple channel input data pairs with the highest similarity coefficients, and the multiple similarity coefficients corresponding to the multiple channel input data pairs; Based on the multiple similarity coefficients, multiple encoding patterns corresponding one-to-one with the multiple vocal channel input data pairs are determined; The multiple channel input data pairs are encoded according to the multiple encoding modes to generate multiple bitstream information corresponding to each of the multiple channel input data pairs.
2. The method according to claim 1, characterized in that, The determination of multiple similarity coefficients between the input data of each channel in the multi-channel input data includes: Determine the multiple modulation discrete cosine transform (MDCT) coefficients that correspond one-to-one with the multi-channel input data; The plurality of similarity coefficients are determined based on the plurality of MDCT coefficients.
3. The method according to claim 2, characterized in that, The step of determining the multiple modulation discrete cosine transform (MDCT) coefficients corresponding one-to-one with the multi-channel input data includes: The first channel input data is subjected to frame-by-frame windowing processing to generate multi-frame channel input sub-data, wherein the first channel input data is any one of the channel input data in the multi-channel input data. Based on the multi-frame channel input sub-data, determine the first MDCT spectral coefficients and the number of first MDCT coefficient points of the first channel input data; The first MDCT coefficients of the first channel input data are generated based on the first MDCT spectral coefficients and the number of first MDCT coefficient points.
4. The method according to claim 3, characterized in that The step of determining the plurality of similarity coefficients based on the plurality of MDCT coefficients includes: The second similarity coefficient is determined by the following formula, wherein the second similarity coefficient is any one of the plurality of similarity coefficients: Wherein, the sim cosθ The second similarity coefficient is defined as i, where i is the number of the vocal tract input sub-data, N is the number of the MDCT coefficient points, and X is the second similarity coefficient. i The second MDCT spectral coefficients of the second channel input data in the second channel input data pair, the Y i The third MDCT spectral coefficient of the third channel input data in the second channel input data pair is the third MDCT spectral coefficient of the third channel input data in the second channel input data pair. The second channel input data pair is any first channel input data pair and includes the second channel input data and the third channel input data.
5. The method according to claim 1, characterized in that, The step of determining multiple encoding patterns corresponding one-to-one with the multiple vocal channel input data pairs based on the multiple similarity coefficients includes: A first similarity threshold and a second similarity threshold are determined, wherein the second similarity threshold is less than 1, the second similarity threshold is greater than the first similarity threshold, and the first similarity threshold is greater than 0; The plurality of encoding patterns are determined based on the relationship between the plurality of similarity coefficients and the first similarity threshold and the second similarity threshold.
6. The method according to claim 5, characterized in that, The step of determining the multiple encoding patterns based on the relationship between the multiple similarity coefficients and the first similarity threshold and the second similarity threshold includes: If the first similarity coefficient is greater than or equal to the second similarity threshold, the first encoding mode corresponding to the first similarity coefficient is determined to be the first set encoding mode, and the first similarity coefficient is any one of the plurality of similarity coefficients; If the first similarity coefficient is less than the second similarity threshold, and the first similarity coefficient is greater than or equal to the first similarity threshold, then the first encoding model is determined to be the second set encoding mode. If the first similarity coefficient is less than the first similarity threshold, the first encoding mode is determined to be the third set encoding mode.
7. The method according to claim 5, characterized in that, The first channel input data pair is any one of the plurality of channel input data pairs, the first channel input data pair including fourth channel input data and fifth channel input data, and determining the first similarity threshold and the second similarity threshold includes: Based on the first similarity coefficient, determine the Holt predictor term for Holt exponential smoothing prediction; Determine the channel energy ratio between the fourth channel input data and the fifth channel input data; Obtain the first similarity benchmark threshold corresponding to the first similarity threshold, and obtain the second similarity benchmark threshold corresponding to the second similarity threshold; The first similarity threshold is determined based on the Holt prediction term, the vocal tract energy ratio, and the first similarity benchmark threshold. The second similarity threshold is determined based on the Holt prediction term, the vocal tract energy ratio, and the second similarity benchmark threshold.
8. The method according to claim 7, characterized in that, The step of determining the Holt prediction term for Holt exponent smoothing prediction based on the first similarity coefficient includes: Based on the first similarity coefficient, determine the Holt baseline term and Holt trend term for Holt index smoothing prediction; The Holt prediction term is determined based on the Holt baseline term and the Holt trend term.
9. The method according to claim 6, characterized in that, The first encoding mode is the first preset encoding mode, the first channel input data pair includes sixth channel input data and seventh channel input data, and the step of encoding the first channel input data pair according to the first encoding mode to generate the first bitstream information of the first channel input pair includes: Determine the second MDCT coefficients of the sixth channel input data, and determine the third MDCT coefficients of the seventh channel input data; According to the spectral index of the MDCT coefficients, the second MDCT coefficients are divided into a first odd spectrum and a first even spectrum, and according to the spectral index, the third MDCT coefficients are divided into a second odd spectrum and a second even spectrum. The first maximum correlation ratio (MCR) rotation angle is determined based on the first odd spectrum and the second odd spectrum. The second MCR rotation angle is determined based on the first even spectrum and the second even spectrum; Based on the first MCR rotation angle, the first odd spectrum and the second odd spectrum are rotated by MCR to generate a third odd spectrum; Based on the second MCR rotation angle, the first even spectrum and the second even spectrum are rotated by MCR to generate a third even spectrum; The first code stream information is generated based on the third odd spectrum and the third even spectrum.
10. The method according to claim 6, characterized in that, The first encoding mode is the second set encoding mode, the first channel input data pair includes eighth channel input data and ninth channel input data, and the step of encoding the first channel input data pair according to the first encoding mode to generate the first bitstream information of the first channel input pair includes: Determine the fourth MDCT coefficient of the eighth channel input data, and determine the fifth MDCT coefficient of the ninth channel input data; According to the spectral index of the MDCT coefficients, the fourth MDCT coefficient is divided into the fourth odd spectrum and the fourth even spectrum, and according to the spectral index, the fifth MDCT coefficient is divided into the fifth odd spectrum and the fifth even spectrum. The third MCR rotation angle is determined based on the fourth odd spectrum and the fifth odd spectrum; The fourth MCR rotation angle is determined based on the fourth even spectrum and the fifth even spectrum; Based on the third MCR rotation angle, the fourth odd spectrum and the fifth odd spectrum are subjected to MCR rotation to generate the sixth odd spectrum and the seventh odd spectrum; Based on the fourth MCR rotation angle, the fourth even spectrum and the fifth even spectrum are subjected to MCR rotation to generate the sixth even spectrum and the seventh even spectrum; The sixth odd spectrum and the sixth even spectrum are combined to generate the sixth MDCT coefficients; The seventh odd spectrum and the seventh even spectrum are combined to generate the seventh MDCT coefficients; The first bitstream information is generated based on the sixth MDCT coefficient and the seventh MDCT coefficient.
11. The method according to claim 6, characterized in that, The first encoding mode is the third preset encoding mode, the first channel input data pair includes tenth channel input data and eleventh channel input data, and the step of encoding the first channel input data pair according to the first encoding mode to generate the first bitstream information of the first channel input pair includes: Determine the sixth MDCT coefficient of the tenth channel input data, and determine the seventh MDCT coefficient of the eleventh channel input data; The first bitstream information is generated based on the sixth MDCT coefficient and the seventh MDCT coefficient.
12. The method according to any one of claims 1-11, characterized in that, The first bitstream information includes side information, and the first bitstream information includes any of the plurality of bitstream information, wherein the side information includes at least one of the following: The first encoding mode; MCR rotation angle; The first similarity coefficient.
13. The method according to any one of claims 1-11, characterized in that, The method further includes: The plurality of bitstream information is sent to the second device, and the plurality of bitstream information is used to instruct the second device to decode the plurality of bitstream information to generate the plurality of audio channel input data pairs.
14. An audio data encoding method, characterized in that, Performed by a second device, the method includes: The device receives multiple bitstream information sent by a first device, wherein the multiple bitstream information is generated by the first device encoding multiple channel input data according to multiple encoding modes, and the multiple encoding modes are determined by the first device based on multiple similarity coefficients between the multiple channel input data; The multiple bitstream information is decoded to generate the multiple audio channel input data.
15. The method according to claim 14, characterized in that, The first bitstream information includes side information, and the first bitstream information includes any of the plurality of bitstream information, wherein the side information includes at least one of the following: The first encoding mode; MCR rotation angle; The first similarity coefficient.
16. A first device, characterized in that, include: The processing module is configured to determine multiple similarity coefficients between the input data of each channel in the multi-channel input data; The processing module is further configured to determine, based on the plurality of similarity coefficients, a plurality of channel input data pairs with the highest similarity coefficients, and a plurality of similarity coefficients corresponding one-to-one to the plurality of channel input data pairs; The processing module is further configured to determine, based on the plurality of similarity coefficients, a plurality of encoding modes corresponding one-to-one with the plurality of audio channel input data pairs; The processing module is further configured to encode the multiple channel input data pairs according to the multiple encoding modes, and generate multiple bitstream information corresponding one-to-one with the multiple channel input data pairs.
17. A second device, characterized in that, include: The transceiver module is configured to receive multiple bitstream information sent by the first device. The multiple bitstream information is generated by the first device encoding multiple channel input data according to multiple encoding modes. The multiple encoding modes are determined by the first device based on multiple similarity coefficients between the multiple channel input data. The processing module is configured to decode the multiple bitstream information to generate the multiple channel input data.
18. A first device, characterized in that, include: One or more processors; The processor is used to execute the audio data encoding method according to any one of claims 1-13.
19. A second device, characterized in that, include: One or more processors; The processor is used to execute the audio data encoding method according to any one of claims 14-15.
20. A communication system, characterized in that, The device includes a first device and a second device, wherein the first device is configured to implement the audio data encoding method of any one of claims 1-13, and the second device is configured to implement the audio data encoding method of any one of claims 14-15.
21. A storage medium storing instructions, characterized in that, When the instruction is executed on the communication device, the communication device performs the audio data encoding method as described in any one of claims 1-13 or 14-15.
22. A computer program product, comprising a computer program and / or instructions, characterized in that, When the computer program and / or instructions are executed by the communication device, they implement the audio data encoding method as described in any one of claims 1-13, or when the computer program and / or instructions are executed by the communication device, they implement the audio data encoding method as described in any one of claims 14-15.
Citation Information
Patent Citations
Stereo audio signal encoder
CN104364842A
Coding of multi-channel audio signals
CN113035212A
Coding method and device of multi-channel audio signal
CN115410584A
Device and method for processing a multi-channel signal
CN1926608A
Device and method for coding audio signal
JP2006003580A