Grouping method, encoder, decoder, and storage medium
By dividing channel packets including three channels in audio encoding and using channel similarity to grouping, the problem of single channel cannot be grouped is solved, and the accuracy and quality of audio encoding is improved.
Patent Information
- Application Number
- PCT/CN2023/128800
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2025-05-08
AI Technical Summary
In the prior art, when grouping channels, a single channel cannot be grouped, resulting in inaccurate channel grouping, which in turn affects the accuracy of audio encoding.
A method of dividing channel packets including three channels is proposed to ensure that each channel can have a corresponding group, and grouping is made using the similarity between channels to reduce redundancy between channels.
Through accurate channel grouping, the redundancy between channels is reduced, the accuracy of audio encoding is improved, and the quality of audio encoding is ensured.
Smart Images

Figure CN2023128800_08052025_PF_FP_ABST
Abstract
Description
Grouping method, encoder, decoder and storage medium Technical Field
[0001] The present disclosure relates to the field of multimedia technology, and in particular to a grouping method, an encoder, a decoder, and a storage medium. Background Art
[0002] With the rapid development of multimedia technology, audio can be applied in various fields. In addition, in order to improve the spatial and directional sense of audio, audio can be 3D encoded and 3D decoded to ensure that the audio heard by the user is indistinguishable from the audio heard in the actual environment.
[0003] Summary of the Invention
[0004] The present disclosure solves the problem that a single channel cannot be grouped when grouping channels, and provides a method for dividing a channel group including three channels to ensure that each channel can have a corresponding group when grouping channels, thereby ensuring the accuracy of channel grouping, thereby ensuring the reduction of redundancy between channels, ensuring the accuracy of channel encoding, and thereby ensuring the accuracy of audio encoding.
[0005] The embodiments of the present disclosure provide a grouping method, device, and storage medium.
[0006] According to a first aspect of an embodiment of the present disclosure, a grouping method is proposed. The method is performed by an encoder, and the method includes:
[0007] Multiple channels are grouped to obtain at least one channel group, wherein the at least one channel group includes N first channel groups, each of which includes three channels, and the at least one channel group includes M second channel groups, each of which includes two channels, where N is 1 and M is a non-negative integer.
[0008] According to a second aspect of an embodiment of the present disclosure, a grouping method is proposed, where the method is performed by a decoder and includes:
[0009] The first information is decoded to obtain channel information of at least one channel group, where N first channel groups exist in the at least one channel group, each of the first channel group including three channels, and M second channel groups exist in the at least one channel group, each of the second channel group including two channels, where N is 1 and M is a non-negative integer.
[0010] According to a third aspect of an embodiment of the present disclosure, a grouping method is proposed, the method comprising:
[0011] The encoder groups the multiple channels to obtain at least one channel group, wherein the at least one channel group includes N first channel groups, each of which includes three channels, and the at least one channel group includes M second channel groups, each of which includes two channels, where N is 1 and M is a non-negative integer.
[0012] The decoder decodes the first information to obtain channel information of at least one channel group.
[0013] According to a fourth aspect of the embodiments of the present disclosure, a coding and decoding device is provided, including:
[0014] a processing module configured to group the plurality of channels to obtain at least one channel group, wherein the at least one channel group comprises N first channel groups, each of which includes three channels; and M second channel groups, each of which includes two channels, wherein N is 1 and M is a non-negative integer.
[0015] According to a fifth aspect of the embodiments of the present disclosure, a coding and decoding device is provided, including:
[0016] a processing module, configured to decode the first information to obtain channel information of at least one channel grouping, wherein the at least one channel grouping includes N first channel groups, each of which includes three channels; and M second channel groups, each of which includes two channels, wherein N is 1 and M is a non-negative integer.
[0017] According to a sixth aspect of the embodiments of the present disclosure, a coding and decoding device is provided, including:
[0018] one or more processors;
[0019] The encoding and decoding device is used to execute any one of the methods described in the first aspect.
[0020] According to a seventh aspect of the embodiments of the present disclosure, a coding and decoding device is provided, including:
[0021] one or more processors;
[0022] The encoding and decoding device is used to execute any method described in the second aspect.
[0023] According to an eighth aspect of the embodiments of the present disclosure, a coding and decoding system is proposed, including:
[0024] An encoder and a decoder, wherein the encoder is configured to implement the grouping method described in the first aspect, and the decoder is configured to implement the grouping method described in the second aspect.
[0025] According to a ninth aspect of an embodiment of the present disclosure, a storage medium is proposed, wherein the storage medium stores instructions. When the instructions are executed on a communication device, the communication device executes a method as described in any one of the first aspect or the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings described herein are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the present disclosure. The illustrative embodiments of the embodiments of the present disclosure and their descriptions are used to explain the embodiments of the present disclosure and do not constitute an improper limitation on the embodiments of the present disclosure. In the drawings:
[0027] FIG1 is a schematic diagram of the architecture of a coding and decoding system according to an embodiment of the present disclosure;
[0028] FIG2A is an interactive schematic diagram illustrating a grouping method according to an embodiment of the present disclosure;
[0029] FIG2B is an interactive schematic diagram illustrating a grouping method according to an embodiment of the present disclosure;
[0030] FIG2C is an interactive schematic diagram illustrating a grouping method according to an embodiment of the present disclosure;
[0031] FIG2D is an interactive schematic diagram illustrating a grouping method according to an embodiment of the present disclosure;
[0032] FIG2E is an interactive schematic diagram illustrating a grouping method according to an embodiment of the present disclosure;
[0033] FIG3A is a schematic diagram showing a flow chart of a grouping method according to an embodiment of the present disclosure;
[0034] FIG3B is a flow chart of a grouping method according to an embodiment of the present disclosure;
[0035] FIG4A is a schematic diagram showing a flow chart of a grouping method according to an embodiment of the present disclosure;
[0036] FIG4B is a schematic diagram of a flow chart of a grouping method according to an embodiment of the present disclosure;
[0037] FIG5 is a flow chart of a grouping method according to an embodiment of the present disclosure;
[0038] FIG6 is a flow chart of a grouping method according to an embodiment of the present disclosure;
[0039] FIG7A is a schematic structural diagram of a coding and decoding device proposed in an embodiment of the present disclosure;
[0040] FIG7B is a schematic structural diagram of a coding and decoding device proposed in an embodiment of the present disclosure;
[0041] FIG8A is a schematic structural diagram of a communication device proposed in an embodiment of the present disclosure;
[0042] FIG8B is a schematic diagram of the structure of the chip proposed in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0043] The present disclosure provides a grouping method, device, and storage medium.
[0044] According to a first aspect of an embodiment of the present disclosure, a grouping method is proposed. The method is performed by an encoder, and the method includes:
[0045] Multiple channels are grouped to obtain at least one channel group, wherein the at least one channel group includes N first channel groups, each of which includes three channels, and the at least one channel group includes M second channel groups, each of which includes two channels, where N is 1 and M is a non-negative integer.
[0046] In the above embodiment, the problem of a single channel being unable to be grouped when grouping channels is solved, and a method of dividing a channel group including three channels is provided to ensure that each channel has a corresponding group when grouping channels, thereby ensuring the accuracy of channel grouping, thereby ensuring the reduction of redundancy between channels, ensuring the accuracy of channel encoding, and thereby ensuring the accuracy of audio encoding.
[0047] In conjunction with some embodiments of the first aspect, in some embodiments,
[0048] The grouping of the multiple channels to obtain at least one channel group includes:
[0049] Obtaining a similarity between any two channels of the plurality of channels;
[0050] The multiple channels are grouped based on the similarity between any two channels to obtain the at least one channel group.
[0051] In the above embodiment, the channels are grouped according to the similarity between every two channels, ensuring that the similarity of the channels included in each channel group meets the requirements, ensuring the accuracy of the channel grouping, and thus ensuring the reduction of redundancy between channels, ensuring the accuracy of channel encoding, and thus ensuring the accuracy of audio encoding.
[0052] In combination with some embodiments of the first aspect, grouping the multiple channels based on the similarity between any two channels to obtain the at least one channel group includes:
[0053] For any two channels among the multiple channels, determining the two channels with the greatest similarity as a candidate channel group;
[0054] Excluding channels other than the two channels with the greatest similarity, and determining the two channels with the greatest similarity as a candidate channel group, until one independent channel remains or no channels remain;
[0055] The at least one channel grouping is determined based on the obtained candidate channel groupings and / or the independent channels.
[0056] In the above embodiment, the channels are divided into the same channel group in descending order of similarity, ensuring that the channels included in each channel group have highly correlated similarities, thereby ensuring the accuracy of the channel grouping, thereby ensuring the reduction of redundancy between channels, ensuring the accuracy of channel encoding, and thus ensuring the accuracy of audio encoding.
[0057] In combination with some embodiments of the first aspect, determining the at least one channel grouping based on the obtained candidate channel grouping and / or the independent channel includes:
[0058] Obtaining a similarity between the independent channel and each channel in each candidate channel group;
[0059] When the similarity between the independent channel and each channel in the first candidate channel group is greater than a similarity threshold, determining the independent channel and each channel in the first candidate channel group as a first channel group;
[0060] Each of the remaining candidate channel groups is determined as a second channel group.
[0061] In the above embodiment, the remaining single channels are also divided into a channel group to ensure that there is no situation where a single channel is not divided into a channel group, thereby ensuring that redundancy between channels is reduced, ensuring the accuracy of channel encoding, and thus ensuring the accuracy of audio encoding.
[0062] In conjunction with some embodiments of the first aspect, in some embodiments, grouping the multiple channels to obtain at least one channel group includes:
[0063] The multiple channels are grouped to directly obtain the N first channel groups and the M second channel groups.
[0064] In conjunction with some embodiments of the first aspect, in some embodiments, grouping the multiple channels to directly obtain the N first channel groups and the M second channel groups includes:
[0065] By performing a global search on the multiple channels, the multiple channels are grouped according to the similarity between any two channels to obtain N first channel groups and the M second channel groups.
[0066] In conjunction with some embodiments of the first aspect, in some embodiments, grouping the multiple channels to obtain at least one channel group includes:
[0067] Grouping the multiple channels to obtain multiple candidate channel groups including two channels and at least one independent channel;
[0068] An independent channel among the at least one independent channel is divided into one candidate channel group among the multiple candidate channel groups to obtain the N first channel groups and the M second channel groups.
[0069] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0070] downmixing the audio signal in each of the channel groups to obtain a downmixed audio signal;
[0071] The downmixed audio signal is encoded to obtain an audio stream.
[0072] In the above embodiment, downmix encoding is performed on the audio signals of the channel groups to obtain an audio stream, which ensures that the encoded audio stream can be transmitted normally and also ensures the accuracy of the audio stream.
[0073] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0074] First information is sent, where the first information is used to indicate channel information of the plurality of channel groups.
[0075] In the above embodiment, the channel information of each channel group is notified through the first information, thereby ensuring the reliability of transmitting the channel group.
[0076] In conjunction with some embodiments of the first aspect, in some embodiments, the channel information includes at least one of the following:
[0077] The number of channel groups including two channels;
[0078] Channel identifiers of channels included in a channel group including two channels;
[0079] an energy parameter of a channel grouping comprising two channels;
[0080] The number of channel groups including three channels;
[0081] Channel identifiers of channels included in a channel group including three channels;
[0082] An energy parameter of a channel grouping comprising three channels;
[0083] The energy parameter is used to adjust the energy of the channels in the channel group.
[0084] In the above embodiment, the channel information includes information of channel groups of two channels or three channels, thereby ensuring the comprehensiveness of the channel information.
[0085] In a second aspect, an embodiment of the present disclosure provides a grouping method, which is performed by a decoder and includes:
[0086] The first information is decoded to obtain channel information of at least one channel group, where N first channel groups exist in the at least one channel group, each of the first channel group including three channels, and M second channel groups exist in the at least one channel group, each of the second channel group including two channels, where N is 1 and M is a non-negative integer.
[0087] In conjunction with some embodiments of the second aspect, in some embodiments, the channel information includes at least one of the following:
[0088] The number of channel groups including two channels;
[0089] Channel identifiers of channels included in a channel group including two channels;
[0090] an energy parameter of a channel grouping comprising two channels;
[0091] The number of channel groups including three channels;
[0092] Channel identifiers of channels included in a channel group including three channels;
[0093] An energy parameter of a channel grouping comprising three channels;
[0094] The energy parameter is used to adjust the energy of the channels in the channel group.
[0095] In conjunction with some embodiments of the second aspect, in some embodiments, the method further includes:
[0096] When it is determined that the channel groups do not include a channel group with three channels, two-channel upmixing is performed on the audio stream.
[0097] In conjunction with some embodiments of the second aspect, in some embodiments, the method further includes:
[0098] When it is determined that the channel groups include a channel group with three channels, three-channel upmixing is performed on the audio stream.
[0099] In a third aspect, an embodiment of the present disclosure provides a grouping method, the method comprising:
[0100] The encoder groups the multiple channels to obtain at least one channel group, wherein the at least one channel group includes N first channel groups, each of which includes three channels, and the at least one channel group includes M second channel groups, each of which includes two channels, where N is 1 and M is a non-negative integer.
[0101] The decoder decodes the first information to obtain channel information of the at least one channel group, wherein at least one channel group among the multiple channel groups includes three channels.
[0102] In a fourth aspect, an embodiment of the present disclosure provides a coding and decoding device, which includes at least one of a transceiver module and a processing module; wherein the encoder is used to execute the optional implementation methods of the first and third aspects.
[0103] In a fifth aspect, an embodiment of the present disclosure provides a coding and decoding device, which includes at least one of a transceiver module and a processing module; wherein the access network device is used to execute the optional implementation methods of the second and third aspects.
[0104] In a sixth aspect, an embodiment of the present disclosure provides a coding and decoding device, including:
[0105] one or more processors;
[0106] The encoding and decoding device is used to execute the method described in any one of the first and third aspects.
[0107] In a seventh aspect, an embodiment of the present disclosure provides a coding and decoding device, including:
[0108] one or more processors;
[0109] The encoding and decoding device is used to execute the method described in any one of the second and third aspects.
[0110] In an eighth aspect, an embodiment of the present disclosure provides a storage medium storing first information. When the first information is run on a communication device, the communication device executes a method as described in any one of the first, second and third aspects.
[0111] In a ninth aspect, an embodiment of the present disclosure proposes a program product. When the program product is executed by a communication device, the communication device executes any one of the methods described in the first, second and third aspects.
[0112] In a tenth aspect, an embodiment of the present disclosure proposes a computer program, which, when executed on a communication device, enables the communication device to execute any one of the methods described in the first, second, and third aspects.
[0113] In an eleventh aspect, an embodiment of the present disclosure provides a chip or a chip system, wherein the chip or chip system includes a processing circuit configured to execute any one of the methods described in the first, second, and third aspects.
[0114] It is understandable that the above-mentioned encoder, storage medium, program product, computer program, chip or chip system are all used to execute the method proposed in the embodiment of the present disclosure. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding method and will not be repeated here.
[0115] The present disclosure provides a grouping method, apparatus, and storage medium. In some embodiments, the terms "grouping method," "information grouping method," and "grouping method" are interchangeable; the terms "coding and decoding apparatus," "information processing apparatus," and "indicating apparatus" are interchangeable; and the terms "information processing system," "coding and decoding system," and "information processing system" are interchangeable.
[0116] The embodiments of the present disclosure are not exhaustive and are merely illustrative of some embodiments, and are not intended to be a specific limitation on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be arbitrarily exchanged. In addition, the optional implementation methods in a certain embodiment can be arbitrarily combined; in addition, the embodiments can be arbitrarily combined. For example, some or all steps of different embodiments can be arbitrarily combined, and a certain embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.
[0117] In each embodiment of the present disclosure, unless otherwise specified or provided for by logic, the terms and / or descriptions between the embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form a new embodiment based on their inherent logical relationships.
[0118] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.
[0119] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular, such as "a", "an", "the", "above", "said", "the", "the", etc., may mean "one and only one", or "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English in translation, the noun following the article may be understood as a singular expression or a plural expression.
[0120] In the embodiments of the present disclosure, “plurality” refers to two or more.
[0121] In some embodiments, the terms "at least one," "one or more," "a plurality of," "multiple," etc. may be used interchangeably.
[0122] In some embodiments, descriptions such as "at least one of A and B," "A and / or B," "A in one case, B in another case," or "in response to one case A, in response to another case B" may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed); and in some embodiments, A and B (both A and B are executed). The above is also applicable when there are more branches such as A, B, and C.
[0123] In some embodiments, "A or B" and other descriptions may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed). The above is also applicable when there are more branches such as A, B, C, etc.
[0124] The prefixes such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different description objects and do not constitute any restriction on the position, order, priority, quantity or content of the description objects. For the statement of the description object, please refer to the description in the context of the claims or embodiments, and no unnecessary restriction should be constituted due to the use of prefixes. For example, if the description object is a "field", the ordinal number before the "field" in the "first field" and the "second field" does not limit the position or order between the "fields". "First" and "second" do not limit whether the "fields" they modify are in the same message, nor do they limit the order of the "first field" and the "second field". For another example, if the description object is a "level", the ordinal number before the "level" in the "first level" and the "second level" does not limit the priority between the "levels". For another example, the number of description objects is not limited by the ordinal number and can be one or more. Taking "first device" as an example, the number of "devices" can be one or more. In addition, the objects modified by different prefixes can be the same or different. For example, if the description object is "device", then the "first device" and the "second device" can be the same device or different devices, and their types can be the same or different; for another example, if the description object is "information", then the "first information" and the "second information" can be the same information or different information, and their contents can be the same or different.
[0125] In some embodiments, “including A,” “comprising A,” “used to indicate A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.
[0126] In some embodiments, terms such as "time / frequency" and "time / frequency domain" refer to the time domain and / or the frequency domain.
[0127] In some embodiments, terms such as "in response to...", "in response to determining...", "in the case of...", "at the time of...", "when...", "if...", "if...", etc. can be used interchangeably.
[0128] In some embodiments, terms such as "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not less than", and "above" can be replaced with each other, and terms such as "less than", "less than or equal to", "not greater than", "less than", "less than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", and "below" can be replaced with each other.
[0129] In some embodiments, devices and equipment can be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. In some cases, they can also be understood as "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "subject", etc.
[0130] In some embodiments, "network" can be interpreted as devices included in the network, such as access network equipment, core network equipment, etc.
[0131] In some embodiments, "access network device (AN device)" may also be referred to as "radio access network device (RAN device)", "base station (BS)", "radio base station", "fixed station", and in some embodiments may also be understood as "node", "access point", "transmission point (TP)", "reception point (RP)", "transmission and / or reception point (TRP)" "panel", "antenna panel", "antenna array", "cell", "macro cell", "small cell", "femto cell", "pico cell", "sector", "cell group", "serving cell", "carrier", "component carrier", "bandwidth part (BWP)", etc.
[0132] In some embodiments, "encoder (terminal)" or "encoder device (terminal device)" may be referred to as "user equipment (encoder)", "user terminal" "mobile station (MS)", "mobile terminal (MT)", subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, etc.
[0133] In some embodiments, obtaining data, information, etc. may comply with the laws and regulations of the country where the data is obtained.
[0134] In some embodiments, data, information, etc. may be obtained with the user's consent.
[0135] In addition, each element, each row, or each column in the table of the embodiment of the present disclosure can be implemented as an independent embodiment, and the combination of any elements, any rows, and any columns can also be implemented as an independent embodiment.
[0136] FIG1 is a schematic diagram of the architecture of a codec system according to an embodiment of the present disclosure. As shown in FIG1 , the method provided in the embodiment of the present disclosure can be applied to a codec system 100, which can include an encoder 101 and a decoder 102. It should be noted that the codec system 100 can also include other devices, and the present disclosure does not limit the devices included in the codec system 100.
[0137] In some embodiments, the encoder 101 and the decoder 102 are both provided in a terminal. In some embodiments, the terminal can be various devices. For example, the terminal can be a mobile phone, a wearable device, an Internet of Things device, a car with communication functions, a smart car, a tablet computer, a computer with wireless transceiver functions, a virtual reality (VR) encoder device, an augmented reality (AR) encoder device, a wireless encoder device in industrial control, a wireless encoder device in self-driving, a wireless encoder device in remote medical surgery, a wireless encoder device in a smart grid, a wireless encoder device in transportation safety, a wireless encoder device in a smart city, and a wireless encoder device in a smart home, but is not limited thereto.
[0138] It can be understood that the encoding and decoding system described in the embodiment of the present disclosure is for the purpose of more clearly illustrating the technical solution of the embodiment of the present disclosure, and does not constitute a limitation on the technical solution proposed in the embodiment of the present disclosure. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution proposed in the embodiment of the present disclosure is also applicable to similar technical problems.
[0139] The following embodiments of the present disclosure may be applied to the codec system 100 shown in FIG1 , or a portion thereof, but are not limited thereto. The entities shown in FIG1 are illustrative only. The codec system may include all or part of the entities shown in FIG1 , or may include other entities outside of FIG1 . The number and form of the entities are arbitrary, and the entities may be physical or virtual. The connection relationship between the entities is illustrative only. The entities may be connected or disconnected, and the connection may be in any manner, including direct or indirect, wired or wireless.
[0140] The embodiments of the present disclosure can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G New Radio (NR), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New Radio Access (NX), Future Generation Radio Access (FX), Global System for Mobile Communications (GSM (registered trademark)), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, Ultra-WideBand (UWB), Bluetooth (registered trademark), Public Land Mobile Network (PLMN) networks, Device-to-Device (D2D) systems, Machine-to-Machine (M2M) systems, Internet of Things (IoT) systems, Vehicle-to-Everything (V2X), systems using other packet methods, and next-generation systems based on and extending these systems. Furthermore, a combination of multiple systems (e.g., a combination of LTE or LTE-A with 5G) may also be employed.
[0141] In some embodiments, the present disclosure is used for 3D audio coding. This 3D audio coding is a key technology in immersive audio technology. Compared to traditional audio, 3D audio enhances spatial and directional perception, allowing listeners to reproduce sounds they hear in the real world. This satisfies people's demand for highly realistic and immersive sound experiences, while also offering personalized choices and interactive experiences.
[0142] In some embodiments, in order to reproduce the spatial and directional sense of sound, the technology can rely on a channel-based approach, an object-based approach, a sound field-based approach, and a combination of the above three forms. Among them: channel-based audio is a group of interrelated channels, common ones are 5.1 channels, 7.1 channels, 5.1.4 channels, 7.1.4 channels, etc. Each format corresponds to a speaker layout, and the best playback effect can be obtained under the corresponding speaker layout.
[0143] In some embodiments, object-based audio is a collection of monophonic audio elements and corresponding metadata. The metadata indicates information such as the object's position, intensity, and size. During playback, the metadata is used to map the object to one or more speakers or render it binaurally for headphone playback to achieve the desired spatial audio effect.
[0144] In some embodiments, sound field-based audio is a 3D sound field modeling format defined on the surface of a sphere. The principle is that sound is transmitted as a pressure wave. For a sound scene at a given time, each point needs to be represented by several pressure functions. If the pressure value of each point in the space is known, the sound in the space can be reconstructed. There is a certain relationship between the pressure of each point in the space and its neighboring points. In order to give full play to the advantages of scene-based audio production, it is necessary to accurately obtain the coefficients and improve the encoding quality of the sound field space coefficients. The collected sound field signal is called Higher Order Ambisonics (HOA). The performance of the HOA system increases with the increase of the HOA order, but the number of HOA signals also increases.
[0145] FIG2 is an interactive diagram of a grouping method according to an embodiment of the present disclosure. As shown in FIG2 , the embodiment of the present disclosure relates to a grouping method, which includes:
[0146] Step S2101: The encoder groups multiple channels to obtain at least one channel group.
[0147] In some embodiments, channels refer to independent audio signals captured or played back at different spatial locations during recording or playback. Channels are elements of audio, and different types of audio have different numbers of channels. For example, an HOA3 signal has 16 channels, while an MC22.2 signal has 24 channels.
[0148] In some embodiments, the audio signal includes audio data and side information.
[0149] In some embodiments, a channel group refers to a group including at least two channels. The encoder encodes the channels in units of channel groups to obtain an audio stream.
[0150] In some embodiments, the name of the channel group is not limited, and can be, for example, a channel group, a channel group category, etc.
[0151] In some embodiments, there are N first channel groups in at least one channel group, the first channel group includes three channels, and M is a non-negative integer.
[0152] In some embodiments, there are M second channel groups in the at least one channel group, and the second channel group includes two channels, where N is 1.
[0153] In some embodiments, N may also be 0, that is, when a plurality of channels are grouped, no channel group including three channels is obtained, that is, 0 first channel groups.
[0154] Optionally, for the value of M, M should not be greater than half the number of channels. Optionally, if the number of channels for grouping channels is P, then the value of M is greater than or equal to 0 and less than or equal to P / 2, where P is a positive integer.
[0155] In some embodiments, before the channels are grouped, it is necessary to pre-process the signals included in the channels, and then group the channels after the pre-processing is completed.
[0156] Optionally, the method of preprocessing the signal included in the channel includes at least one of the following: transient detection, window type judgment, time-frequency transformation, frequency domain noise shaping, time domain noise shaping, and frequency band extension coding.
[0157] In some embodiments, grouping the multiple channels to obtain at least one channel group includes: obtaining similarity between any two channels in the multiple channels, and grouping the multiple channels based on the similarity between the any two channels to obtain at least one channel group.
[0158] In the embodiment of the present disclosure, the encoder may obtain the similarity between every two channels, and then group the multiple channels according to the magnitude relationship of the similarity between every two channels to obtain at least one channel group.
[0159] In some embodiments, grouping multiple channels based on similarity between any two channels to obtain at least one channel group includes: for any two channels among the multiple channels, determining the two channels with the greatest similarity as a candidate channel group; excluding channels other than the two channels with the greatest similarity, and determining the two channels with the greatest similarity as a candidate channel group until one independent channel or no channels remain; and determining at least one channel group based on the obtained candidate channel group and / or independent channels.
[0160] For example, the multiple channels are channel 1, channel 2, channel 3, channel 4, and channel 5. The similarity between each two channels is obtained respectively. For example, if the similarity between channel 1 and channel 2 is the highest, channel 1 and channel 2 are grouped as one channel. Then, channel 1 and channel 2 are excluded, and the channels with the highest similarity among channel 3, channel 4, and channel 5 are grouped as one channel. At this time, only one channel 5 remains, and channel 5 is the independent channel.
[0161] In some embodiments, determining at least one channel grouping based on the obtained candidate channel groups and / or independent channels includes: obtaining a similarity between the independent channel and each channel in each candidate channel group; when the similarity between the independent channel and each channel in the first candidate channel group is greater than a similarity threshold, determining the independent channel and each channel in the first candidate channel group as a first channel group; and determining each remaining candidate channel group as a second channel group.
[0162] In an embodiment of the present disclosure, when the similarity between an independent channel and each channel in the first candidate channel group is greater than the similarity threshold, it means that the independent channel is similar to both channels in the first candidate channel group. In this case, the independent channel can be divided into the first candidate channel group. At this time, the first candidate channel group includes three channels, that is, the first channel group, and the remaining channel group still includes two channels, that is, the second channel group.
[0163] In some embodiments, when grouping multiple channels based on the similarity between any two channels, if it is determined that an odd number of channels meets the grouping condition and at least one channel group is obtained, the similarity between the independent channel and each channel in each candidate channel group is obtained. When the similarity between the independent channel and each channel in the first candidate channel group is greater than a similarity threshold, the independent channel and each channel in the first candidate channel group are determined to be a first channel group, and each remaining candidate channel group is determined to be a second channel group.
[0164] It should be noted that the disclosed embodiment uses the example of obtaining independent channels. In another embodiment, for any two channels among multiple channels, the two channels with the greatest similarity are determined as a candidate channel group. After excluding the two channels with the greatest similarity, the two channels with the greatest similarity are also determined as a candidate channel group until no channels remain. Based on the obtained candidate channel groups, at least one channel group is determined.
[0165] It should be noted that, for any two channels among the multiple channels, if the similarity between any two channels is not greater than the similarity threshold, the channels are not grouped.
[0166] It should be noted that in the embodiment of the present disclosure, the similarity between an independent channel and two channels in multiple candidate channel groups may be greater than the similarity threshold. In this case, it is necessary to additionally determine into which candidate channel group the independent channel should be divided.
[0167] In some embodiments, the sum of similarities between the independent channel and two channels in each candidate channel group is obtained, the largest sum is selected from the sums corresponding to each candidate channel group, and the independent channel is divided into the channel group corresponding to the largest sum.
[0168] For example, the independent channel is channel 5, the first candidate channel group includes channel 1 and channel 2, and the second candidate channel group includes channel 3 and channel 4. If the sum of the similarities between channel 5 and channels 1 and 2 is 1.8, and the sum of the similarities between channel 5 and channels 3 and 4 is 1.9, then channel 5 is assigned to the second candidate channel group.
[0169] In some embodiments, the similarities between the independent channel and the two channels in each candidate channel group are determined, and the candidate channel group corresponding to the maximum similarity is determined, and the independent channel is divided into the channel group corresponding to the maximum similarity.
[0170] For example, the independent channel is channel 5, the first candidate channel group includes channel 1 and channel 2, and the second candidate channel group includes channel 3 and channel 4. If the similarity between channel 5 and channel 1 is 0.7, the similarity between channel 5 and channel 2 is 0.9, the similarity between channel 5 and channel 3 is 0.8, and the similarity between channel 5 and channel 4 is 0.6, then channel 5 is classified into the first candidate channel group.
[0171] It should be noted that, in the above embodiment, if there is an independent channel that is the same as the sum or maximum value of the channels in at least two candidate channel groups, one of the candidate channel groups is randomly selected and the independent channel is assigned to the selected channel group.
[0172] It should be noted that the grouping in the embodiment of the present disclosure includes two grouping schemes. Each grouping scheme is described below.
[0173] In some embodiments, multiple channels are grouped to directly obtain N first channel groups and M second channel groups. Optionally, in the embodiments of the present disclosure, when multiple channels are grouped, the N first channel groups and M second channel groups obtained by grouping are directly output, and no other grouping results are output during the process.
[0174] Optionally, grouping the multiple channels to directly obtain N first channel groups and M second channel groups includes: performing a global search on the multiple channels and grouping the multiple channels according to the similarity between any two channels to obtain N first channel groups and M second channel groups.
[0175] In one possible implementation, similarity between any two of every three channels in the plurality of channels is obtained. If the similarity between any two of every three channels is greater than a similarity threshold, the three channels are grouped into a candidate channel group. If a candidate channel group is ultimately obtained, the candidate channel group is determined as the first channel group. If multiple candidate channel groups are ultimately obtained, the first channel group is determined from the multiple candidate channel groups based on the similarity relationship between each candidate channel group.
[0176] Optionally, if multiple candidate channel groups are ultimately obtained, determining a first channel group from the multiple candidate channel groups based on the similarity relationship between each candidate channel group includes: obtaining the sum of the similarities between every two channels in each candidate channel group, selecting the largest sum from the sums corresponding to each candidate channel group, and determining the channel group corresponding to the largest sum as the first channel group.
[0177] For example, the first candidate channel group includes channel 1, channel 2, and channel 3, and the second candidate channel group includes channel 4, channel 5, and channel 6. If the sum of the similarities between each two channels in channel 1, channel 2, and channel 3 is 2.7, and the sum of the similarities between channels 4, channel 5, and channel 6 is 2.6, then the first candidate channel group is determined to be the first channel group.
[0178] Optionally, if multiple candidate channel groups are ultimately obtained, a first channel group is determined from the multiple candidate channel groups based on the similarity relationship between each candidate channel group, including: determining the candidate channel group corresponding to the greatest similarity between the independent channel and each two channels in each candidate channel group, and determining the candidate channel group corresponding to the greatest similarity as the first channel group.
[0179] In some embodiments, the process of grouping multiple channels is similar to the process of grouping based on similarity in the above embodiment, and will not be described in detail here.
[0180] In some embodiments, multiple channels are grouped to obtain multiple candidate channel groups including two channels and at least one independent channel, and one independent channel of the at least one independent channel is divided into one candidate channel group in the multiple candidate channel groups to obtain N first channel groups and M second channel groups.
[0181] It should be noted that the embodiment of the present disclosure is described by taking grouping of multiple channels as an example. In another embodiment, after the multiple channels are grouped, channel information of each channel group is also generated.
[0182] In some embodiments, the channel information includes at least one of the following:
[0183] (1) The number of channel groups including two channels;
[0184] In some embodiments, the number of channel groups including two channels may also be referred to as the number of second channel groups, or the number of second channel groups, or the number of channel group pairs, or the number of groups, etc., which is not limited in the embodiments of the present disclosure.
[0185] In some embodiments, the number of channel groups including two channels is used to represent the number of second channel group pairs in the current frame.
[0186] (2) channel identifiers of the channels included in the channel group including two channels;
[0187] In some embodiments, the channel identifiers of the channels included in the channel group including two channels are used to represent the index of the channel pair, and the index values of the two channels in the current channel pair can be obtained by parsing.
[0188] (3) Energy parameters of a channel grouping including two channels;
[0189] (4) The number of channel groups including three channels;
[0190] In some embodiments, the number of channel groups including three channels may also be referred to as the number of first channel groups, or the number of first channel groups, or the number of channel group pairs, or the number of groups, etc., which is not limited in the embodiments of the present disclosure.
[0191] In some embodiments, the number of channel groups including three channels is used to represent the number of first channel group pairs of the current frame.
[0192] (5) Channel identifiers of the channels included in the channel group including three channels;
[0193] In some embodiments, the channel identifiers of the channels included in the channel group including three channels are used to represent the index of the channel pair, and the index values of the three channels in the current channel pair can be obtained by parsing.
[0194] (6) Energy parameters of the channel grouping including three channels;
[0195] The energy parameter is used to adjust the energy of the channels in the channel group.
[0196] In some embodiments, the energy parameter is used to quantize an index of an inter-channel amplitude difference (ILD) parameter between a first channel and a second channel in a current channel pair for inter-channel energy / amplitude adjustment.
[0197] In some embodiments, multiple channels are grouped to directly obtain N first channel groups and M second channel groups, and channel information of the first channel groups and the second channel groups is generated.
[0198] In some embodiments, multiple channels are grouped to obtain multiple candidate channel groups including two channels and at least one independent channel, channel information of the multiple candidate channel groups and the at least one independent channel is generated, one independent channel of the at least one independent channel is divided into one candidate channel group of the multiple candidate channel groups to obtain N first channel groups and M second channel groups, and the channel information of the multiple candidate channel groups and the at least one independent channel is rewritten to obtain channel information of the first channel group and the second channel group.
[0199] Optionally, multiple channels are grouped to obtain M+1 candidate channel groups each including two channels and at least one independent channel, channel information of the multiple candidate channel groups and the at least one independent channel is generated, and one independent channel of the at least one independent channel is divided into one candidate channel group of the multiple candidate channel groups to obtain N first channel groups and M second channel groups.
[0200] It should be noted that the method of grouping multiple channels to obtain the first channel grouping can be called 3-channel sum and difference coding. The method of grouping multiple channels to obtain the second channel grouping can be called 2-channel sum and difference coding.
[0201] It should be noted that the execution method in the embodiment of the present disclosure is executed by the judgment module, or by other modules, and the embodiment of the present disclosure is not limited.
[0202] Step S2102: The encoder downmixes the audio signal in each channel group to obtain a downmixed audio signal.
[0203] In some embodiments, downmixing refers to mixing the grouped channels using an orthogonal normalized matrix to obtain a mixed channel for each channel.
[0204] Optionally, the orthogonal normalized matrix is pre-set, and the embodiment of the present disclosure does not limit the orthogonal normalized matrix.
[0205] For example, for a channel group consisting of two channels, the orthogonal normalization matrix used is Among them, the first row is the sum vector and the second row is the difference vector.
[0206] For a channel group consisting of three channels, the orthogonal normalization matrix used is Among them, the first row is the sum vector, and the second and third rows are the difference vectors.
[0207] In some embodiments, 3ch ms I 3ch ms =O 3ch ms , M 2ch ms I 2ch ms =O 2ch ms Among them, I 3ch ms It is a 1*3 column vector. In addition, this column vector refers to the audio data in the channel group including three channels. 2ch ms It is a 1*2 column vector. In addition, this column vector refers to the audio data in the channel group including two channels.
[0208] It should be noted that the steps performed in the embodiment of the present disclosure are performed by the downmix module, or are performed in other ways, which is not limited in the embodiment of the present disclosure.
[0209] Step S2103: The encoder encodes the downmixed audio signal to obtain an audio stream.
[0210] In some embodiments, encoding includes bit allocation quantization entropy encoding and code stream multiplexing.
[0211] The following is an example of how to execute the above steps.
[0212] For example, referring to FIG2B , the decision module (step S2101) is executed to determine which sum and difference coding method or a combination of the two is used. The channels include L channel, R channel, C channel, LS channel and RS channel, and the similarity between any two channels is shown in Table 1.
[0213] Table 1
[0214] The downmix threshold is 0.5. The decision module obtains the sum-difference coding method to be used through an algorithm.
[0215] 3CH M / S refers to a module in which three channels are divided into a first channel group, and 2CH M / S refers to a module in which two channels are divided into a second channel group.
[0216] 1. Use maximum similarity iteration to select the first channel pair L and R channels (maximum similarity cox_L_R=0.85), and the second channel pair LS and RS channels (maximum similarity cox_LS_RS=0.66 among the remaining un-downmixed channels).
[0217] 2. Calculate the similarity between all channels that are not downmixed (C channel) and the two channels of the first channel pair and the second channel pair.
[0218] The similarity between the C channel and the L channel cox_C_L=0.74 is greater than the downmix threshold, and the similarity between the C channel and the R channel cox_C_R=0.78 is greater than the downmix threshold.
[0219] 3. Output the judgment results of 3CH M / S and 2CH M / S. In this embodiment, the judgment result of 3CH M / S is output as L channel, R channel, and C channel. The judgment result of 2CH M / S is output as LS channel and RS channel.
[0220] Subsequently, the downmix module is executed. 2ch ms I 2ch ms =O 2ch ms M 3ch ms I 3ch ms =O 3ch ms
[0221] Among them, the two-channel downmix matrix M 2ch ms It is a 2x2 orthogonal normalized matrix, where the first row is the sum vector and the second row is the difference vector. 3ch ms is a 3x3 orthogonal normalized matrix, where the first row is the sum vector and the second and third rows are the difference vectors.
[0222] I 2ch ms is a 1x2 column vector, I 3ch ms It is a 1x3 column vector. The data contained in the vector is the pre-processed audio data, in units of sampling points or frequency points.
[0223] Generate channel information for each channel group.
[0224] For another example, referring to FIG. 2C , two-channel pairing decision is performed to obtain two-channel pairing information, including the number of pairs and the channel pair index.
[0225] The channels include an L channel, an R channel, a C channel, an LS channel, and an RS channel. The similarity between any two channels is shown in Table 1, and the downmix threshold is 0.5.
[0226] The two-channel pair decision uses maximum similarity iteration to screen out the first channel pair L and R channels (maximum similarity cox_L_R=0.85) and the second channel pair LS and RS channels (maximum similarity cox_LS_RS among the remaining un-downmixed channels=0.66).
[0227] The number of the second channel groups finally obtained is 2, including the first channel group L channel and R channel, and the second channel group LS channel and RS channel.
[0228] Secondly, a three-channel pairing decision is executed. If a three-channel pairing is generated, the above channel grouping result will be changed.
[0229] First, the similarity between all channels not downmixed (C channel) and the two channels of the first channel pair and the second channel pair is calculated.
[0230] The similarity between the C channel and the L channel cox_C_L=0.74 is greater than the downmix threshold, and the similarity between the C channel and the R channel cox_C_R=0.78 is greater than the downmix threshold.
[0231] Output the judgment results of 3CH M / S and 2CH M / S. In this embodiment, the judgment result of 3CH M / S outputs L channel, R channel, and C channel. The judgment result of 2CH M / S is changed to output one channel group, that is, output LS channel and RS channel.
[0232] Generates two-channel and three-channel channel information.
[0233] Executes two-channel and three-channel downmixing modules. 2ch ms I 2ch ms =O 2ch ms M 3ch ms I 3ch ms =O 3ch ms
[0234] Among them, the two-channel downmix matrix M 2ch ms It is a 2x2 orthogonal normalized matrix, where the first row is the sum vector and the second row is the difference vector. 3ch msis a 3x3 orthogonal normalized matrix, where the first row is the sum vector and the second and third rows are the difference vectors.
[0235] The above-mentioned channel information is rewritten to obtain new channel information.
[0236] In some embodiments, the encoder obtains an audio stream by executing the above steps. Referring to FIG2D , after obtaining the channel signals, the encoder preprocesses the channel signals, performs downmixing of the preprocessed channel signals into inter-channel pairs, performs bit allocation quantization entropy coding, and finally multiplexes the bitstreams to obtain the audio stream.
[0237] Preprocessing includes transient detection window type determination, time-frequency transformation, frequency-domain noise shaping, time-domain noise shaping, and band extension coding. Inter-channel downmixing involves downmixing the L, R, and C channels into three-channel pairs to produce channels M1, S11, and S12; and downmixing the LS and RS channels into two-channel pairs to produce channels M2 and S2. The LFE channel is not processed.
[0238] Step S2104: The encoder sends the first information and the audio stream.
[0239] In some embodiments, a decoder receives the first information and an audio stream.
[0240] In some embodiments, the encoder may send the first information and the audio stream separately. For example, the encoder may first send the first information and then send the audio stream. Alternatively, the encoder may first send the audio stream and then send the first information. In some embodiments, the encoder may send the first information and the audio stream simultaneously.
[0241] In some embodiments, the channel information includes at least one of the following:
[0242] The number of channel groups including two channels;
[0243] Channel identifiers of channels included in a channel group including two channels;
[0244] an energy parameter of a channel grouping comprising two channels;
[0245] The number of channel groups including three channels;
[0246] Channel identifiers of channels included in a channel group including three channels;
[0247] An energy parameter of a channel grouping comprising three channels;
[0248] The energy parameter is used to adjust the energy of the channels in the channel group.
[0249] Step S2105: The decoder decodes the first information to obtain channel information of at least one channel group.
[0250] Step S2106: When the decoder determines that the channel grouping does not include a channel grouping with three channels, it performs two-channel upmixing on the audio stream; when the decoder determines that the channel grouping includes a channel grouping with three channels, it performs three-channel upmixing on the audio stream.
[0251] In some embodiments, the decoder determines whether the encoder has performed a two-channel downmix and a three-channel downmix. If the encoder has performed a two-channel downmix, the decoder performs a two-channel upmix. If the encoder has performed a three-channel downmix, the decoder performs a three-channel upmix.
[0252] Optionally, audio post-processing includes but is not limited to general decoding processes, such as time-frequency inverse transform, time-domain noise shaping inverse transform, frequency-domain noise shaping inverse transform, frequency band extension inverse transform and other modules, and also includes decoding processing for certain types of signal characteristics, such as multi-channel decoding processing, HOA channel decoding processing, object metadata decoding processing, etc.
[0253] In some embodiments, the decoder decodes and obtains the channel signals by performing the above steps. Referring to FIG2E , after obtaining the audio stream, the audio stream is demultiplexed, and then subjected to bit allocation, inverse quantization, entropy coding, inter-channel upmixing, and post-processing to obtain decoded channel signals.
[0254] Post-processing includes band extension decoding, inverse time-domain noise shaping, inverse frequency-domain noise shaping, inverse time-frequency transform, etc. Inter-channel upmixing includes upmixing the M1, S11, and S12 channels into three-channel groups, namely, the L, R, and C channels, to obtain the channel; and upmixing the M2 and S2 channels into two-channel groups, namely, the LS and RS channels, to obtain the channel. The LFE channel is not processed.
[0255] In some embodiments, the names of information, etc. are not limited to the names described in the embodiments, and terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "symbol", "codeword", "codebook", "codeword", "codepoint", "bit", "data", "program", and "chip" can be used interchangeably.
[0256] In some embodiments, terms such as "uplink", "uplink", "physical uplink" can be interchangeable with each other, and terms such as "downlink", "downlink", "physical downlink" can be interchangeable with each other, and terms such as "side", "sidelink", "side communication", "sidelink communication", "direct connection", "direct link", "direct communication", "direct link communication" can be interchangeable with each other.
[0257] In some embodiments, "obtain", "get", "get", "receive", "transmit", "bidirectional transmission", "send and / or receive" can be interchangeable, and can be interpreted as receiving from other entities, obtaining from protocols, obtaining from higher layers, obtaining by self-processing, autonomous implementation, etc.
[0258] In some embodiments, terms such as "send", "transmit", "report", "download", "transmit", "bidirectional transmission", "send and / or receive" can be used interchangeably.
[0259] In some embodiments, terms such as "moment", "time point", "time", and "time position" can be replaced with each other, and terms such as "duration", "period", "time window", "window", and "time" can be replaced with each other.
[0260] In some embodiments, terms such as "certain", "preset", "preset", "setting", "indicated", "a certain", "any", and "first" can be interchangeable. "Specific A", "preset A", "preset A", "setting A", "indicated A", "a certain A", "any A", and "first A" can be interpreted as A pre-specified in a protocol, etc., or as A obtained through setting, configuration, or indication, etc., or as specific A, a certain A, any A, or first A, etc., but not limited to this.
[0261] The grouping method involved in the embodiment of the present disclosure may include at least one of steps S2101 to S2106. For example, step S2101 can be implemented as an independent embodiment, step S2102 can be implemented as an independent embodiment, step S2103 can be implemented as an independent embodiment, step S2104 can be implemented as an independent embodiment, step S2105 can be implemented as an independent embodiment, step S2106 can be implemented as an independent embodiment, step S2101 and step S2102 can be implemented as independent embodiments, step S2103 and step S2104 can be implemented as independent embodiments, step S2105 and step S2106 can be implemented as independent embodiments, step S2101, step S2102, step S2103, and step S2104 can be implemented as independent embodiments, step S2101, step S2102, step S2105, and step S2106 can be implemented as independent embodiments, but is not limited to this.
[0262] In some embodiments, step S2101 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0263] In some embodiments, step S2102 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0264] In some embodiments, step S2103 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0265] In some embodiments, step S2104 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0266] In some embodiments, step S2105 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0267] In some embodiments, step S2106 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0268] In some embodiments, step S2101 and step S2102 are optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0269] In some embodiments, step S2103 and step S2104 are optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0270] In some embodiments, step S215 and step S2106 are optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0271] In some embodiments, reference may be made to other optional implementations described before or after the description corresponding to FIG. 2 .
[0272] FIG3A is a flow chart of a grouping method according to an embodiment of the present disclosure, which is applied to an encoder. As shown in FIG3A , the embodiment of the present disclosure relates to a grouping method, which includes:
[0273] Step S3101: The encoder groups multiple channels to obtain at least one channel group.
[0274] The optional implementation of step S3101 can refer to the optional implementation of step S2101 in Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.
[0275] Step S3102: The encoder downmixes the audio signal in each channel group to obtain a downmixed audio signal.
[0276] The optional implementation of step S3102 can refer to the optional implementation of step S2102 in Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.
[0277] Step S3103: The encoder encodes the downmixed audio signal to obtain an audio stream.
[0278] The optional implementation of step S3103 can refer to the optional implementation of step S2103 in Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.
[0279] Step S3104: The encoder sends the first information and the audio stream.
[0280] The optional implementation of step S3103 can refer to the optional implementation of step S2104 in Figure 2 and other related parts in the embodiment involved in Figure 2, which will not be repeated here.
[0281] The grouping method involved in the embodiments of the present disclosure may include at least one of steps S3101 to S3104. For example, step S3101 may be implemented as an independent embodiment, step S3102 may be implemented as an independent embodiment, step S3103 may be implemented as an independent embodiment, step S3104 may be implemented as an independent embodiment, or at least two steps may be combined, but the present invention is not limited thereto.
[0282] In some embodiments, step S3101 is optional, step S3102 is optional, step S3103 is optional, and step S3104 is optional. In different embodiments, one or more of these steps may be omitted or replaced, but the present invention is not limited thereto.
[0283] FIG3B is a flow chart of a grouping method according to an embodiment of the present disclosure, which is applied to an encoder. As shown in FIG3B , the embodiment of the present disclosure relates to a grouping method, which includes:
[0284] Step S3201: The encoder groups multiple channels to obtain at least one channel group.
[0285] Optional implementations of step S3201 may refer to step S2101 in FIG. 2 , step S3101 in FIG. 3A , and other related parts in the embodiments involved in FIG. 2 and FIG. 3A , which will not be described in detail here.
[0286] FIG4A is a flow chart of a grouping method according to an embodiment of the present disclosure, which is applied to a decoder. As shown in FIG4A , the embodiment of the present disclosure relates to a grouping method, which includes:
[0287] Step S4101: The decoder decodes the first information to obtain channel information of at least one channel group.
[0288] The optional implementation of step S4101 can be found in step S2105 of FIG. 2 and other related parts of the embodiment involved in FIG. 2 , which will not be described in detail here.
[0289] Step S4102: When the decoder determines that the channel grouping does not include a channel grouping with three channels, it performs two-channel upmixing on the audio stream; when the decoder determines that the channel grouping includes a channel grouping with three channels, it performs three-channel upmixing on the audio stream.
[0290] Optional implementations of step S4102 may refer to step S2106 in FIG2 and other related parts of the embodiment involved in FIG2 , which will not be described in detail here.
[0291] The grouping method involved in the embodiment of the present disclosure may include at least one of steps S4101 and S4102. For example, step S4101 may be implemented as an independent embodiment, step S4102 may be implemented as an independent embodiment, or at least two steps may be combined, but the present invention is not limited thereto.
[0292] In some embodiments, step S4101 is optional, step S4102 is optional, and one or more of these steps may be omitted or replaced in different embodiments, but the present invention is not limited thereto.
[0293] FIG4B is a flow chart of a grouping method according to an embodiment of the present disclosure, which is applied to a decoder. As shown in FIG4B , the embodiment of the present disclosure relates to a grouping method, which includes:
[0294] Step S4201: The decoder decodes the first information to obtain channel information of at least one channel group.
[0295] The optional implementation of step S4201 can be found in step S2105 of FIG. 2 and other related parts of the embodiment involved in FIG. 2 , which will not be described in detail here.
[0296] In some embodiments, there are N first channel groups in the at least one channel group, each of which includes three channels, and there are M second channel groups in the at least one channel group, each of which includes two channels, where N is 1 and M is a non-negative integer.
[0297] In some embodiments, the channel information includes at least one of the following:
[0298] The number of channel groups including two channels;
[0299] Channel identifiers of channels included in a channel group including two channels;
[0300] an energy parameter of a channel grouping comprising two channels;
[0301] The number of channel groups including three channels;
[0302] Channel identifiers of channels included in a channel group including three channels;
[0303] An energy parameter of a channel grouping comprising three channels;
[0304] The energy parameter is used to adjust the energy of the channels in the channel group.
[0305] In some embodiments, the method further comprises:
[0306] When it is determined that the channel groups do not include a channel group with three channels, two-channel upmixing is performed on the audio stream.
[0307] In some embodiments, the method further comprises:
[0308] When it is determined that the channel groups include a channel group with three channels, three-channel upmixing is performed on the audio stream.
[0309] FIG5 is a flow chart of a grouping method according to an embodiment of the present disclosure. As shown in FIG5 , the embodiment of the present disclosure relates to a grouping method, which includes:
[0310] Step S5101: The encoder groups multiple channels to obtain at least one channel group.
[0311] In some embodiments, there are N first channel groups in the at least one channel group, each of which includes three channels, and there are M second channel groups in the at least one channel group, each of which includes two channels, where N is 1 and M is a non-negative integer.
[0312] Step S5102: The decoder decodes the first information to obtain channel information of at least one channel group.
[0313] Optional implementations of step S5101 may refer to step S2101 in FIG. 2 , step S3101 in FIG. 3A , and other related parts in the embodiments involved in FIG. 2 and FIG. 4A , which will not be described in detail here.
[0314] Optional implementations of step S5102 may refer to step S2105 of FIG. 2 , step S4101 of FIG. 4A , and other related parts of the embodiments involved in FIG. 2 and FIG. 3A , which will not be described in detail here.
[0315] In some embodiments, the above method may include the methods of the above embodiments of the coding and decoding system side, encoder side, decoder side, etc., which will not be repeated here.
[0316] FIG6 is a flow chart of a grouping method according to an embodiment of the present disclosure. As shown in FIG6 , the embodiment of the present disclosure relates to a grouping method, which includes:
[0317] In step S6101, the encoding end encodes multiple channels by using a combination of two-channel sum and difference encoding and three-channel sum and difference encoding.
[0318] The encoder completes audio preprocessing and enters the inter-channel downmixing module. Audio preprocessing includes, but is not limited to, common encoding processes such as transient analysis, time-frequency transformation, time-domain noise shaping, frequency-domain noise shaping, and band extension. It also includes processing tailored to specific signal characteristics, such as multi-channel encoding, HOA channel encoding, and object metadata encoding.
[0319] The inter-channel downmix module introduces a three-channel sum-difference coding method and combines it with a two-channel sum-difference coding framework. This module includes a decision module and a downmix module.
[0320] The decision module determines which sum / difference encoding method or combination to use. The decision criterion is the correlation between channels, which is compared with the pair threshold. The decision results in three-channel sum / difference encoding for the L, R, and C channels, two-channel sum / difference encoding for the LS and LS channels, and no processing for the LFE channel.
[0321] The downmix module performs 3-channel sum and difference encoding on the L, R, and C channels, and 2-channel sum and difference encoding on the LS and LS.
[0322] After passing through the inter-channel group downmix module, the three-channel sum and difference coded downmixed channels (M1, S11, and S12), the two-channel sum and difference coded downmixed channels (M2 and S2), and the undownmixed channel (LFE channel) all pass through the bit allocation quantization entropy coding module and are multiplexed to form the coded bitstream E.
[0323] In the embodiments of the present disclosure, some or all of the steps and their optional implementations may be arbitrarily combined with some or all of the steps in other embodiments, or may be arbitrarily combined with the optional implementations of other embodiments.
[0324] The present disclosure also provides an apparatus for implementing any of the above methods. For example, a device is provided that includes units or modules for implementing each step performed by an encoder in any of the above methods. For another example, another device is provided that includes units or modules for implementing each step performed by a decoder in any of the above methods.
[0325] It should be understood that the division of the various units or modules in the above device is merely a division of logical functions. In actual implementation, they may be fully or partially integrated into a physical entity, or they may be physically separated. In addition, the units or modules in the device may be implemented in the form of a processor calling software: for example, the device includes a processor, the processor is connected to a memory, and the memory stores instructions. The processor calls the instructions stored in the memory to implement any of the above methods or implement the functions of the various units or modules of the above device, wherein the processor is, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory is a memory within the device or a memory outside the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits, and the functions of some or all of the units or modules can be realized by designing the hardware circuits. The above-mentioned hardware circuits can be understood as one or more processors; for example, in one implementation, the above-mentioned hardware circuit is an application-specific integrated circuit (ASIC), which realizes the functions of some or all of the above units or modules by designing the logical relationship of the components in the circuit; for example, in another implementation, the above-mentioned hardware circuit can be realized by a programmable logic device (PLD). Taking a field programmable gate array (FPGA) as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by configuring the configuration file, thereby realizing the functions of some or all of the above units or modules. All units or modules of the above devices can be realized in the form of software called by the processor, or in the form of hardware circuits, or in part by the form of software called by the processor, and the rest by hardware circuits.
[0326] In the embodiments of the present disclosure, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationship of the hardware circuit. The logical relationship of the above-mentioned hardware circuit is fixed or reconfigurable. For example, the processor is a hardware circuit implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and implementing the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), etc.
[0327] Figure 7A is a structural diagram of the encoding and decoding device proposed in an embodiment of the present disclosure. As shown in Figure 7A, the encoding and decoding device 7100 may include: at least one of a transceiver module 7101, a processing module 7102, etc. In some embodiments, the processing module 7102 is used to group multiple channels to obtain at least one channel group, wherein there are N first channel groups in the at least one channel group, and the first channel group includes three channels, and there are M second channel groups in the at least one channel group, and the second channel group includes two channels, wherein N is 1 and M is a non-negative integer. Optionally, the above-mentioned transceiver module 7101 is used to perform at least one of the communication steps such as sending and / or receiving performed by the encoder in any of the above methods (for example, step S2101 but not limited thereto), which will not be repeated here. Optionally, the above-mentioned processing module is used to perform at least one of the other steps performed by the encoder in any of the above methods, which will not be repeated here.
[0328] Optionally, the processing module 7102 is used to execute at least one of the communication steps such as processing performed by the encoder in any of the above methods, which will not be repeated here.
[0329] Figure 7B is a structural diagram of the encoding and decoding device proposed in an embodiment of the present disclosure. As shown in Figure 7B, the encoding and decoding device 7200 may include: at least one of a transceiver module 7201, a processing module 7202, etc. In some embodiments, the processing module 7202 is used to decode the first information to obtain channel information of at least one channel grouping, wherein there are N first channel groups in the at least one channel grouping, and the first channel grouping includes three channels, and there are M second channel groups in the at least one channel grouping, and the second channel grouping includes two channels, wherein N is 1 and M is a non-negative integer. Optionally, the above-mentioned transceiver module is used to perform at least one of the communication steps such as sending and / or receiving performed by the decoder in any of the above methods (such as step S2102 but not limited thereto), which will not be repeated here.
[0330] Optionally, the processing module 7202 is used to execute at least one of the communication steps such as processing performed by the decoder in any of the above methods, which will not be repeated here.
[0331] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module, and the transmitting module and the receiving module may be separate or integrated. Optionally, the transceiver module may be interchangeable with the transceiver.
[0332] In some embodiments, the processing module can be a single module or can include multiple submodules. Optionally, the multiple submodules respectively execute all or part of the steps required to be executed by the processing module. Optionally, the processing module can be interchangeable with the processor.
[0333] Figure 8A is a schematic diagram of the structure of a communication device 8100 proposed in an embodiment of the present disclosure. Communication device 8100 can be a decoder (e.g., an access network device, a core network device, etc.), an encoder (e.g., a user equipment, etc.), a chip, a chip system, or a processor that supports a decoder to implement any of the above methods, or a chip, a chip system, or a processor that supports an encoder to implement any of the above methods. Communication device 8100 can be used to implement the methods described in the above method embodiments. For details, please refer to the description of the above method embodiments.
[0334] As shown in Figure 8A, the communication device 8100 includes one or more processors 8101. The processor 8101 can be a general-purpose processor or a dedicated processor, for example, a baseband processor or a central processing unit. The baseband processor can be used to process the communication protocol and communication data, and the central processing unit can be used to control the encoding and decoding device (such as a base station, a baseband chip, an encoder device, an encoder device chip, a DU or CU, etc.), execute programs, and process program data. The communication device 8100 is used to perform any of the above methods.
[0335] In some embodiments, the communication device 8100 further includes one or more memories 8102 for storing instructions. Optionally, all or part of the memories 8102 may be located outside the communication device 8100.
[0336] In some embodiments, the communication device 8100 further includes one or more transceivers 8103. When the communication device 8100 includes one or more transceivers 8103, the transceiver 8103 performs at least one of the communication steps such as sending and / or receiving in the above method (for example, step S2101, step S2102, step S2103, step S2104, but not limited thereto).
[0337] In some embodiments, a transceiver may include a receiver and / or a transmitter. The receiver and transmitter may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, and transceiver circuit may be used interchangeably; the terms transmitter, transmitting unit, transmitter, and transmitting circuit may be used interchangeably; and the terms receiver, receiving unit, receiver, and receiving circuit may be used interchangeably.
[0338] In some embodiments, the communication device 8100 may include one or more interface circuits 8104. Optionally, the interface circuit 8104 is connected to the memory 8102. The interface circuit 8104 may be configured to receive signals from the memory 8102 or other devices, and may be configured to send signals to the memory 8102 or other devices. For example, the interface circuit 8104 may read instructions stored in the memory 8102 and send the instructions to the processor 8101.
[0339] The communication device 8100 described in the above embodiments may be a decoder or an encoder, but the scope of the communication device 8100 described in the present disclosure is not limited thereto, and the structure of the communication device 8100 may not be limited by FIG. 8A. The communication device may be an independent device or may be part of a larger device. For example, the communication device may be: 1) an independent integrated circuit IC, or a chip, or a chip system or subsystem; (2) a collection of one or more ICs, optionally, the above IC collection may also include a storage component for storing data or programs; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, an encoder device, an intelligent encoder device, a cellular phone, a wireless device, a handheld device, a mobile unit, an in-vehicle device, a decoder, a cloud device, an artificial intelligence device, etc.; (6) others, etc.
[0340] FIG8B is a schematic diagram of the structure of a chip 8200 according to an embodiment of the present disclosure. If the communication device 8100 can be a chip or a chip system, please refer to the schematic diagram of the structure of the chip 8200 shown in FIG8B , but the present disclosure is not limited thereto.
[0341] The chip 8200 includes one or more processors 8201 , and the chip 8200 is configured to execute any of the above methods.
[0342] In some embodiments, the chip 8200 further includes one or more interface circuits 8202. Optionally, the interface circuit 8202 is connected to the memory 8203. The interface circuit 8202 can be used to receive signals from the memory 8203 or other devices, and can be used to send signals to the memory 8203 or other devices. For example, the interface circuit 8202 can read instructions stored in the memory 8203 and send the instructions to the processor 8201.
[0343] In some embodiments, the interface circuit 8202 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processor 8201 performs at least one of the other steps.
[0344] In some embodiments, terms such as interface circuit, interface, transceiver pin, and transceiver may be used interchangeably.
[0345] In some embodiments, the chip 8200 further includes one or more memories 8203 for storing instructions. Alternatively, all or part of the memories 8203 may be outside the chip 8200.
[0346] The present disclosure also proposes a storage medium having instructions stored thereon, which, when executed on the communication device 8100, causes the communication device 8100 to execute any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but is not limited thereto, and may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but is not limited thereto, and may also be a temporary storage medium.
[0347] The present disclosure also provides a program product, which, when executed by the communication device 8100, enables the communication device 8100 to perform any of the above methods. Optionally, the program product is a computer program product.
[0348] The present disclosure also proposes a computer program, which, when executed on a computer, causes the computer to perform any one of the above methods.
Claims
1. A grouping method, characterized in that: The method is performed by an encoder, and comprises: A plurality of channels are grouped to obtain at least one channel group, wherein the at least one channel group includes N first channel groups, the first channel group includes three channels, and the at least one channel group includes M second channel groups, the second channel group includes two channels, wherein N is 1 and M is a non-negative integer.
2. The method according to claim 1, characterized in that The step of grouping the plurality of channels to obtain at least one channel group includes: Obtaining a similarity between any two channels of the multiple channels; The plurality of channels are grouped based on the similarity between any two channels to obtain the at least one channel group.
3. The method according to claim 2, characterized in that The grouping the plurality of channels based on the similarity between any two channels to obtain the at least one channel group includes: For any two channels among the multiple channels, determine the two channels with the greatest similarity as a candidate channel group; Excluding the other channels except the two channels with the greatest similarity, the two channels with the greatest similarity are determined as a candidate channel group until one independent channel remains or no channel remains; The at least one channel grouping is determined based on the obtained candidate channel groupings and / or the independent channels.
4. The method according to claim 3, characterized in that: The determining the at least one channel grouping based on the obtained candidate channel groupings and / or the independent channels comprises: Obtaining a similarity between the independent channel and each channel in each candidate channel group; When the similarity between the independent channel and each channel in the first candidate channel group is greater than a similarity threshold, determining the independent channel and each channel in the first candidate channel group as a first channel group; Each of the remaining candidate channel groups is determined as a second channel group.
5. The method according to claim 1, characterized in that The step of grouping the plurality of channels to obtain at least one channel group includes: The multiple channels are grouped to directly obtain the N first channel groups and the M second channel groups.
6. The method according to claim 5, characterized in that The grouping of the plurality of channels to directly obtain the N first channel groups and the M second channel groups includes: By performing a global search on the multiple channels, the multiple channels are grouped according to the similarity between any two channels to obtain N first channel groups and the M second channel groups.
7. The method according to claim 1, characterized in that The step of grouping the plurality of channels to obtain at least one channel group includes: Grouping the multiple channels to obtain multiple candidate channel groups including two channels and at least one independent channel; An independent channel among the at least one independent channel is divided into one candidate channel group among the plurality of candidate channel groups to obtain the N first channel groups and the M second channel groups.
8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: Downmixing the audio signal in each of the channel groups to obtain a downmixed audio signal; The downmixed audio signal is encoded to obtain an audio stream.
9. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: First information is sent, where the first information is used to indicate channel information of the plurality of channel groups.
10. The method according to claim 9, characterized in that The channel information includes at least one of the following: The number of channel groups including two channels; channel identifiers of channels included in a channel group including two channels; An energy parameter of a channel grouping comprising two channels; The number of channel groups including three channels; channel identifiers of channels included in a channel group including three channels; An energy parameter of a channel grouping comprising three channels; The energy parameter is used to adjust the energy of the channels in the channel group.
11. A grouping method, characterized in that: The method is performed by a decoder, and comprises: The first information is decoded to obtain channel information of at least one channel grouping, where there are N first channel groups in the at least one channel grouping, the first channel grouping includes three channels, and there are M second channel groups in the at least one channel grouping, the second channel grouping includes two channels, where N is 1 and M is a non-negative integer.
12. The method according to claim 11, characterized in that The channel information includes at least one of the following: The number of channel groups including two channels; channel identifiers of channels included in a channel group including two channels; An energy parameter of a channel grouping comprising two channels; The number of channel groups including three channels; channel identifiers of channels included in a channel group including three channels; An energy parameter of a channel grouping comprising three channels; The energy parameter is used to adjust the energy of the channels in the channel group.
13. The method according to claim 11 or 12, characterized in that: The method further comprises: When it is determined that the channel groups do not include a channel grouping for three channels, two-channel upmixing is performed on the audio stream.
14. The method according to claim 11 or 12, characterized in that: The method further comprises: When it is determined that the channel groups include a channel grouping with three channels, three-channel upmixing is performed on the audio stream.
15. A grouping method, characterized in that: The method comprises: The encoder groups the multiple channels to obtain at least one channel group, wherein the at least one channel group includes N first channel groups, the first channel group includes three channels, and the at least one channel group includes M second channel groups, the second channel group includes two channels, wherein N is 1 and M is a non-negative integer; The decoder decodes the first information to obtain channel information of the at least one channel grouping, wherein at least one channel grouping among the multiple channel groups includes three channels.
16. A coding and decoding device, characterized in that: The encoding and decoding device comprises: A processing module is used to group multiple channels to obtain at least one channel group, wherein there are N first channel groups in the at least one channel group, the first channel group includes three channels, and there are M second channel groups in the at least one channel group, the second channel group includes two channels, wherein N is 1 and M is a non-negative integer.
17. A coding and decoding device, characterized in that: The encoding and decoding device comprises: A processing module is used to decode the first information to obtain channel information of at least one channel grouping, where there are N first channel groups in the at least one channel grouping, the first channel grouping includes three channels, and there are M second channel groups in the at least one channel grouping, the second channel grouping includes two channels, wherein N is 1 and M is a non-negative integer.
18. A coding and decoding device, characterized in that: The encoding and decoding device comprises: one or more processors; The processor is used to execute the grouping method according to any one of claims 1 to 10.
19. A coding and decoding device, characterized in that: The encoding and decoding device comprises: one or more processors; The processor is used to execute the grouping method described in any one of claims 11 to 14.
20. A coding and decoding system, characterized in that: The invention comprises an encoder and a decoder, wherein the encoder is configured to implement the grouping method according to any one of claims 1 to 10, and the decoder is configured to implement the grouping method according to any one of claims 11 to 14.
21. A storage medium storing instructions, characterized in that: When the instruction is executed on a communication device, the communication device is enabled to execute the grouping method according to any one of claims 1 to 10, or execute the grouping method according to any one of claims 11 to 14.
Citation Information
Patent Citations
Stereo audio signal encoder
CN104364842A
Parametric mixing of audio signals
CN107112020A
Apparatus and method for encoding or decoding a multi-channel signal
CN107592937A
Multi-channel audio signal coding and decoding methods and devices
CN113948095A
Parametric encoding and decoding of multichannel audio signals
US20170339505A1