Methods and apparatuses for encoding and decoding, and storage medium
By performing inter-channel decorrelation processing on multi-channel audio data, a second data that removes redundant information is generated, which solves the problem of low encoding accuracy caused by redundancy of multi-channel audio data, and improves the accuracy of audio data processing and machine analysis.
Patent Information
- Application Number
- PCT/CN2024/075055
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-31
- Publication Date
- 2025-08-07
AI Technical Summary
Multi-channel audio data redundancy leads to low encoding accuracy, affecting the accuracy of subsequent analysis.
By performing inter-channel decorrelation processing on multi-channel audio data, a second data that removes redundant information is generated, and processed using an encoder and a decoder to improve encoding accuracy.
Improve the accuracy of audio data processing and ensure the accuracy of subsequent machine analysis.
Smart Images

Figure CN2024075055_07082025_PF_FP_ABST
Abstract
Description
Coding and decoding method, device and storage medium Technical Field
[0001] The present disclosure relates to the field of communication technologies, and in particular to a coding and decoding method, device, and storage medium. Background Art
[0002] With the rapid development of multimedia technology, audio data can be processed and transmitted, and then the corresponding audio data can be obtained by decoding the data, ensuring that the audio data can be transmitted efficiently.
[0003] Summary of the Invention
[0004] The present disclosure solves the problem of excessive redundancy in multi-channel audio data. By decorrelating the multi-channel data, redundant information in the channel data is removed, the accuracy of encoding is improved, and the accuracy of subsequent analysis based on the encoded data is improved.
[0005] The embodiments of the present disclosure provide a coding and decoding method, apparatus, and storage medium.
[0006] According to a first aspect of an embodiment of the present disclosure, a coding method is proposed. The method is performed by an encoder, and the method includes:
[0007] Based on first data included in each of a plurality of channels, audio data included in the plurality of channels is processed to obtain second data, wherein the processing includes decorrelation between channels, the first data is used to describe the corresponding channel, the first data is generated by a first device, each of the channels includes one audio data, and the audio data is collected by the first device.
[0008] According to a second aspect of an embodiment of the present disclosure, a decoding method is proposed, where the method is performed by a decoder and includes:
[0009] Based on the third data included in each of the multiple channels, the second data is decoded to obtain the audio data included in the multiple channels, the third data is used to indicate the corresponding channel, the second data is obtained by performing inter-channel decorrelation processing on the audio data included in the multiple channels, the first data is generated by the first device, each of the channels includes one audio data, and the audio data is collected by the first device.
[0010] According to a third aspect of the embodiments of the present disclosure, a coding and decoding method is proposed, the method comprising:
[0011] The encoder processes audio data included in the multiple channels based on first data included in each channel to obtain second data, wherein the processing includes decorrelation between channels, the first data is used to describe the corresponding channel, the first data is generated by a first device, each channel includes one piece of audio data, and the audio data is collected by the first device;
[0012] The decoder decodes the second data based on the first data included in each of the multiple channels to obtain the audio data included in the multiple channels.
[0013] According to a fourth aspect of the embodiments of the present disclosure, an encoding device is provided, including:
[0014] A processing module is configured to process the audio data included in the multiple channels based on first data included in each channel to obtain second data, wherein the processing includes decorrelation between channels, the first data is used to describe the corresponding channel, the first data is generated by a first device, and each channel includes one audio data, and the audio data is collected by the first device.
[0015] According to a fifth aspect of the embodiments of the present disclosure, a decoding device is provided, including:
[0016] A processing module is configured to decode the second data based on third data included in each of the multiple channels to obtain audio data included in the multiple channels, where the third data is used to indicate the corresponding channel, the second data is obtained by performing inter-channel decorrelation processing on the audio data included in the multiple channels, the first data is generated by a first device, each channel includes one piece of audio data, and the audio data is collected by the first device.
[0017] According to a sixth aspect of the embodiments of the present disclosure, an encoding device is provided, including:
[0018] one or more processors;
[0019] The encoding and decoding device is used to execute any one of the methods described in the first aspect.
[0020] According to a seventh aspect of the embodiments of the present disclosure, a decoding device is provided, including:
[0021] one or more processors;
[0022] The encoding and decoding device is used to execute any method described in the second aspect.
[0023] According to an eighth aspect of the embodiments of the present disclosure, a coding and decoding system is proposed, including:
[0024] An encoder and a decoder, wherein the encoder is configured to implement the encoding and decoding method described in the first aspect, and the decoder is configured to implement the encoding and decoding method described in the second aspect.
[0025] According to a ninth aspect of an embodiment of the present disclosure, a storage medium is proposed, wherein the storage medium stores instructions. When the instructions are executed on a communication device, the communication device executes a method as described in any one of the first aspect or the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings described herein are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the present disclosure. The illustrative embodiments of the embodiments of the present disclosure and their descriptions are used to explain the embodiments of the present disclosure and do not constitute an improper limitation on the embodiments of the present disclosure. In the drawings:
[0027] FIG1 is a schematic diagram of the architecture of a coding and decoding system according to an embodiment of the present disclosure;
[0028] FIG2A is an interactive schematic diagram illustrating a coding and decoding method according to an embodiment of the present disclosure;
[0029] FIG2B is an interactive schematic diagram illustrating a coding and decoding method according to an embodiment of the present disclosure;
[0030] FIG2C is an interactive schematic diagram illustrating a coding and decoding method according to an embodiment of the present disclosure;
[0031] FIG2D is an interactive schematic diagram illustrating a coding and decoding method according to an embodiment of the present disclosure;
[0032] FIG2E is an interactive schematic diagram illustrating a coding and decoding method according to an embodiment of the present disclosure;
[0033] FIG2F is an interactive schematic diagram illustrating a coding and decoding method according to an embodiment of the present disclosure;
[0034] FIG3A is a schematic diagram showing a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0035] FIG3B is a schematic diagram showing a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0036] FIG4 is a schematic diagram of a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0037] FIG5 is a schematic diagram of a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0038] FIG6 is a schematic diagram of a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0039] FIG7A is a schematic structural diagram of a coding and decoding device proposed in an embodiment of the present disclosure;
[0040] FIG7B is a schematic structural diagram of a coding and decoding device proposed in an embodiment of the present disclosure;
[0041] FIG8A is a schematic structural diagram of a communication device proposed in an embodiment of the present disclosure;
[0042] FIG8B is a schematic diagram of the structure of the chip proposed in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0043] The present disclosure provides a coding and decoding method, device, and storage medium.
[0044] According to a first aspect of an embodiment of the present disclosure, a coding method is proposed. The method is performed by an encoder, and the method includes:
[0045] Based on first data included in each of a plurality of channels, audio data included in the plurality of channels is processed to obtain second data, wherein the processing includes decorrelation between channels, the first data is used to describe the corresponding channel, the first data is generated by a first device, each of the channels includes one audio data, and the audio data is collected by the first device.
[0046] In the above embodiment, the first signal indicates the coding information involved in processing the audio data, which solves the problem of low accuracy in processing the audio data. The audio data is processed by the coding information included in the first signal, thereby ensuring the accuracy of the audio data processing and further ensuring the accuracy of subsequent machine analysis of the audio data.
[0047] In combination with some embodiments of the first aspect, in some embodiments, the first data is used to indicate the type of audio data included in the channel.
[0048] In the above embodiment, the first data indicates the function of each channel, ensuring the accuracy of the indicated channel, thereby ensuring the accuracy of the processing of the audio data.
[0049] In combination with some embodiments of the first aspect, in some embodiments, the first data includes reference information, and the reference information is used to indicate that reference audio data included in the channel corresponding to the first data is used to match the audio data to be tested.
[0050] In combination with some embodiments of the first aspect, in some embodiments, the first data includes information to be tested, and the information to be tested is used to indicate that the audio data to be tested included in the channel corresponding to the first data is used to match with reference audio data.
[0051] In conjunction with some embodiments of the first aspect, in some embodiments, one of the multiple channels includes reference audio data, and the other channels include audio data to be measured, and processing the audio data included in the multiple channels based on the first data included in each channel of the multiple channels to obtain the second data includes:
[0052] Obtaining a degree of difference between the audio data to be tested and the reference audio data included in each channel;
[0053] The audio data is processed based on the difference between the reference audio data and each of the channels to obtain the second data.
[0054] In the above embodiment, the audio data is processed by obtaining the degree of difference, thereby ensuring the accuracy of encoding and eliminating the differences between the audio data.
[0055] In conjunction with some embodiments of the first aspect, in some embodiments, obtaining the degree of difference between the audio data to be tested and the reference audio data included in each channel includes:
[0056] A difference between the audio data to be tested and the reference audio data included in each channel is obtained and determined as the difference degree.
[0057] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0058] For the audio data to be tested included in each channel, performing phase rotation on the audio data to be tested of each channel so that the phase of the audio data to be tested of each channel is the same as that of the reference audio data;
[0059] The obtaining of the difference between the audio data to be tested and the reference audio data included in each channel includes:
[0060] The degree of difference between the audio data to be measured and the reference audio data after phase rotation is obtained in each channel.
[0061] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0062] For the audio data to be tested included in each channel, performing energy balancing on the audio data to be tested in each channel so that the amplitude of the audio data to be tested in each channel is the same as that of the reference audio data;
[0063] The obtaining of the difference between the audio data to be tested and the reference audio data included in each channel includes:
[0064] Obtain the degree of difference between the audio data to be measured and the reference audio data after energy balance is performed in each channel.
[0065] In conjunction with some embodiments of the first aspect, in some embodiments, processing the audio data included in the multiple channels to obtain the second data includes:
[0066] Perform kernel selection and kernel encoding on the audio data included in each of the channels to obtain the second data.
[0067] In conjunction with some embodiments of the first aspect, in some embodiments, the channel corresponds to first data, the first data is used to indicate a first core module of the channel, and performing core selection and core encoding on the audio data included in each of the channels to obtain the second data includes:
[0068] Determining a third core module based on the second core module selected from the audio data and the first core module, wherein the third core module is used to process the audio data;
[0069] The third core module is used to perform the core encoding on the audio data to obtain the second data.
[0070] In conjunction with some embodiments of the first aspect, in some embodiments, determining the third core module based on the second core module selected by the audio data and the first core module includes:
[0071] The first core module is the same as the second core module, and the third core module is determined to be either the first core module or the second core module.
[0072] In conjunction with some embodiments of the first aspect, in some embodiments, determining the third core module based on the second core module selected by the audio data and the first core module includes:
[0073] The first core module is different from the second core module, and the third core module is determined to be the first core module.
[0074] In the above embodiment, the accuracy of the encoding is ensured by selecting an accurate core module to perform core encoding.
[0075] In a second aspect, an embodiment of the present disclosure provides a decoding method, which is performed by a decoder and includes:
[0076] Based on the third data included in each of the multiple channels, the second data is decoded to obtain the audio data included in the multiple channels, the third data is used to indicate the corresponding channel, the second data is obtained by performing inter-channel decorrelation processing on the audio data included in the multiple channels, the first data is generated by the first device, each of the channels includes an audio data, and the audio data is collected by the first device.
[0077] In combination with some embodiments of the second aspect, in some embodiments, the third data is used to indicate the type of audio data included in the channel.
[0078] In combination with some embodiments of the second aspect, in some embodiments, the third data includes reference information, and the reference information is used to indicate that reference audio data included in the channel corresponding to the third data is used to match the audio data to be tested.
[0079] In combination with some embodiments of the second aspect, in some embodiments, the third data includes information to be tested, and the information to be tested is used to indicate that the audio data to be tested included in the channel corresponding to the third data is used to match with reference audio data.
[0080] In conjunction with some embodiments of the second aspect, in some embodiments, one of the multiple channels includes reference audio data, and the other channels include audio data to be tested, and decoding the second data based on the third data included in each of the multiple channels to obtain the audio data included in the multiple channels includes:
[0081] decoding the second data based on the third data included in each of the multiple channels to obtain a degree of difference between the audio data to be tested and the reference audio data included in each channel;
[0082] The audio data included in the multiple channels is determined based on the degree of difference between the audio data to be measured included in each channel and the reference audio data.
[0083] In combination with some embodiments of the second aspect, in some embodiments, the degree of difference is a difference between the audio data to be measured and the reference audio data included in each channel.
[0084] In a third aspect, an embodiment of the present disclosure provides a coding and decoding method, the method comprising:
[0085] The encoder processes audio data included in the multiple channels based on first data included in each channel to obtain second data, wherein the processing includes decorrelation between channels, the first data is used to describe the corresponding channel, the first data is generated by a first device, each channel includes one piece of audio data, and the audio data is collected by the first device;
[0086] The decoder decodes the second data based on third data included in each of the multiple channels to obtain audio data included in the multiple channels, where the third data is used to describe the corresponding channel.
[0087] In a fourth aspect, an embodiment of the present disclosure provides an encoding device, wherein the encoding and decoding device includes at least one of a transceiver module and a processing module; wherein the first device is used to execute the optional implementation methods of the first and third aspects.
[0088] In a fifth aspect, an embodiment of the present disclosure provides a decoding device, wherein the encoding and decoding device includes at least one of a transceiver module and a processing module; wherein the encoder is used to execute the optional implementation methods of the second and third aspects.
[0089] In a sixth aspect, an embodiment of the present disclosure provides an encoding device, including:
[0090] one or more processors;
[0091] The encoding and decoding device is used to execute the method described in any one of the first and third aspects.
[0092] In a seventh aspect, an embodiment of the present disclosure provides a decoding device, including:
[0093] one or more processors;
[0094] The encoding and decoding device is used to execute the method described in any one of the second and third aspects.
[0095] In an eighth aspect, an embodiment of the present disclosure provides a storage medium storing first information. When the first information is run on a communication device, the communication device executes a method as described in any one of the first, second and third aspects.
[0096] In a ninth aspect, an embodiment of the present disclosure proposes a program product. When the program product is executed by a communication device, the communication device executes any one of the methods described in the first, second and third aspects.
[0097] In a tenth aspect, an embodiment of the present disclosure proposes a computer program, which, when executed on a communication device, enables the communication device to execute any one of the methods described in the first, second, and third aspects.
[0098] In an eleventh aspect, an embodiment of the present disclosure provides a chip or a chip system, wherein the chip or chip system includes a processing circuit configured to execute any one of the methods described in the first, second, and third aspects.
[0099] It is understandable that the above-mentioned encoder, storage medium, program product, computer program, chip or chip system are all used to execute the method proposed in the embodiment of the present disclosure. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding method and will not be repeated here.
[0100] The present disclosure provides a coding method, apparatus, and storage medium. In some embodiments, the terms coding method, information coding method, and coding method are interchangeable; the terms coding apparatus, information coding apparatus, and indicating apparatus are interchangeable; and the terms information processing system, coding system, and so on are interchangeable.
[0101] The embodiments of the present disclosure are not exhaustive and are merely illustrative of some embodiments, and are not intended to be a specific limitation on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be arbitrarily exchanged. In addition, the optional implementation methods in a certain embodiment can be arbitrarily combined; in addition, the embodiments can be arbitrarily combined. For example, some or all steps of different embodiments can be arbitrarily combined, and a certain embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.
[0102] In each embodiment of the present disclosure, unless otherwise specified or provided for by logic, the terms and / or descriptions between the embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form a new embodiment based on their inherent logical relationships.
[0103] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.
[0104] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular, such as "a", "an", "the", "above", "said", "the", "the", etc., may mean "one and only one", or "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English in translation, the noun following the article may be understood as a singular expression or a plural expression.
[0105] In the embodiments of the present disclosure, “plurality” refers to two or more.
[0106] In some embodiments, the terms "at least one," "one or more," "a plurality of," "multiple," etc. may be used interchangeably.
[0107] In some embodiments, descriptions such as "at least one of A and B," "A and / or B," "A in one case, B in another case," or "in response to one case A, in response to another case B" may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed); and in some embodiments, A and B (both A and B are executed). The above is also applicable when there are more branches such as A, B, and C.
[0108] In some embodiments, "A or B" and other descriptions may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed). The above is also applicable when there are more branches such as A, B, C, etc.
[0109] The prefixes such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different description objects and do not constitute any restriction on the position, order, priority, quantity or content of the description objects. For the statement of the description object, please refer to the description in the context of the claims or embodiments, and no unnecessary restriction should be constituted due to the use of prefixes. For example, if the description object is a "field", the ordinal number before the "field" in the "first field" and the "second field" does not limit the position or order between the "fields". "First" and "second" do not limit whether the "fields" they modify are in the same message, nor do they limit the order of the "first field" and the "second field". For another example, if the description object is a "level", the ordinal number before the "level" in the "first level" and the "second level" does not limit the priority between the "levels". For another example, the number of description objects is not limited by the ordinal number and can be one or more. Taking "first device" as an example, the number of "devices" can be one or more. In addition, the objects modified by different prefixes can be the same or different. For example, if the description object is "device", then the "first device" and the "second device" can be the same device or different devices, and their types can be the same or different; for another example, if the description object is "information", then the "first information" and the "second information" can be the same information or different information, and their contents can be the same or different.
[0110] In some embodiments, “including A,” “comprising A,” “used to indicate A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.
[0111] In some embodiments, terms such as "time / frequency" and "time / frequency domain" refer to the time domain and / or the frequency domain.
[0112] In some embodiments, terms such as "in response to...", "in response to determining...", "in the case of...", "at the time of...", "when...", "if...", "if...", etc. can be used interchangeably.
[0113] In some embodiments, terms such as "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not less than", and "above" can be replaced with each other, and terms such as "less than", "less than or equal to", "not greater than", "less than", "less than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", and "below" can be replaced with each other.
[0114] In some embodiments, devices and equipment can be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. In some cases, they can also be understood as "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "subject", etc.
[0115] In some embodiments, "network" can be interpreted as devices included in the network, such as access network equipment, core network equipment, etc.
[0116] In some embodiments, "access network device (AN device)" may also be referred to as "radio access network device (RAN device)", "base station (BS)", "radio base station", "fixed station", and in some embodiments may also be understood as "node", "access point", "transmission point (TP)", "reception point (RP)", "transmission and / or reception point (TRP)" "panel", "antenna panel", "antenna array", "cell", "macro cell", "small cell", "femto cell", "pico cell", "sector", "cell group", "serving cell", "carrier", "component carrier", "bandwidth part (BWP)", etc.
[0117] In some embodiments, "encoder (terminal)" or "encoder device (terminal device)" may be referred to as "user equipment (encoder)", "user terminal" "mobile station (MS)", "mobile terminal (MT)", subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, etc.
[0118] In some embodiments, obtaining data, information, etc. may comply with the laws and regulations of the country where the data is obtained.
[0119] In some embodiments, data, information, etc. may be obtained with the user's consent.
[0120] In addition, each element, each row, or each column in the table of the embodiment of the present disclosure can be implemented as an independent embodiment, and the combination of any elements, any rows, and any columns can also be implemented as an independent embodiment.
[0121] FIG1 is a schematic diagram of the architecture of a codec system according to an embodiment of the present disclosure. As shown in FIG1 , the method provided in the embodiment of the present disclosure can be applied to a codec system 100, which can include an encoder 101 and a decoder 102. It should be noted that the codec system 100 can also include other devices, and the present disclosure does not limit the devices included in the codec system 100.
[0122] In some embodiments, the encoder 101 and the decoder 102 are both provided in a terminal. In some embodiments, the terminal can be various devices. For example, the terminal can be a mobile phone, a wearable device, an Internet of Things device, a car with communication functions, a smart car, a tablet computer, a computer with wireless transceiver functions, a virtual reality (VR) encoder device, an augmented reality (AR) encoder device, a wireless encoder device in industrial control, a wireless encoder device in self-driving, a wireless encoder device in remote medical surgery, a wireless encoder device in a smart grid, a wireless encoder device in transportation safety, a wireless encoder device in a smart city, and a wireless encoder device in a smart home, but is not limited thereto.
[0123] It can be understood that the encoding and decoding system described in the embodiment of the present disclosure is for the purpose of more clearly illustrating the technical solution of the embodiment of the present disclosure, and does not constitute a limitation on the technical solution proposed in the embodiment of the present disclosure. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution proposed in the embodiment of the present disclosure is also applicable to similar technical problems.
[0124] The following embodiments of the present disclosure may be applied to the codec system 100 shown in FIG1 , or a portion thereof, but are not limited thereto. The entities shown in FIG1 are illustrative only. The codec system may include all or part of the entities shown in FIG1 , or may include other entities outside of FIG1 . The number and form of the entities are arbitrary, and the entities may be physical or virtual. The connection relationship between the entities is illustrative only. The entities may be connected or disconnected, and the connection may be in any manner, including direct or indirect, wired or wireless.
[0125] The embodiments of the present disclosure can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G New Radio (NR), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New Radio Access (NX), Future Generation Radio Access (FX), Global System for Mobile Communications (GSM (registered trademark)), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, Ultra-WideBand (UWB), Bluetooth (registered trademark), Public Land Mobile Network (PLMN) networks, Device-to-Device (D2D) systems, Machine-to-Machine (M2M) systems, Internet of Things (IoT) systems, Vehicle-to-Everything (V2X), systems using other coding and decoding methods, and next-generation systems based on these. In addition, multiple systems can also be combined (for example, a combination of LTE or LTE-A with 5G, etc.) for application.
[0126] In some embodiments, the present disclosure further includes a first device, or may also be referred to as an acquisition device. Optionally, the first device is used to acquire audio data and generate first data of the audio data, and the audio data is described by the first data. Optionally, the first device is an audio sensor device.
[0127] FIG2A is an interactive diagram of a coding and decoding method according to an embodiment of the present disclosure. As shown in FIG2A , the embodiment of the present disclosure relates to a coding and decoding method, which includes:
[0128] Step S2101: The first device collects audio data of multiple channels.
[0129] In some embodiments, the first device refers to a device with a collection function. For example, the first device may also be referred to as a collection device, a collection terminal, etc., which is not limited in the embodiments of the present disclosure. In some embodiments, the audio data refers to audio, or it can also be understood as sound collected by the first device. In some embodiments, the audio data is used to indicate the audio generated by the device during operation.
[0130] In some embodiments, the audio data is used for machine analysis or machine training. Alternatively, it can be understood that machine analysis or machine training of the audio data can determine whether the operating status of the device indicated by the audio data is normal. Alternatively, it can be understood that the audio data indicates the operating status of the device, and whether the device is operating normally can be determined based on the audio data.
[0131] In some embodiments, the audio data is machine-analyzed by a second device to determine whether the operating status of the device indicated by the audio data is normal. Alternatively, determining the audio data can determine whether the device is operating normally. In some embodiments, the second device is a terminal, an analysis device, a detection device, etc., which is not limited in the present embodiment.
[0132] In some embodiments, the channel refers to the channel used for data transmission between the first device and the encoder. In some embodiments, each channel can transmit audio data, and the audio data is used to indicate the audio of a device, or it can be understood that one channel corresponds to one device. In some embodiments, the first device can transmit multiple channels simultaneously, and the first device can transmit audio data through multiple channels simultaneously.
[0133] In some embodiments, the embodiments of the present disclosure are applied to automated production and testing scenarios, where the first device collects audio data of the produced products (components), and then performs machine analysis on the collected audio data to determine whether the generated products (components) are qualified.
[0134] It should be noted that the embodiments of the present disclosure are described using the role of audio data as an example. Furthermore, after the decoder decodes and obtains the audio data in the following embodiments, the second device analyzes the audio data obtained by the decoder to determine whether the object monitored by the first device is operating normally. In some embodiments, the object can be understood as the first device, the product produced, or other information that the audio data refers to, which is not limited in the embodiments of the present disclosure. In some embodiments, the second device is used to control the first device, which can be understood as the second device being a controller of the first device.
[0135] In some embodiments, the embodiments of the present disclosure are applied in product testing, see Figure 2B, where audio data of the product is collected by a first device, the audio data is processed by an encoder, the encoded data is sent to a decoder, and the decoder decodes the encoded data to determine whether the product corresponding to each piece of audio data is qualified.
[0136] In some embodiments, after collecting audio data from each channel, the first device also generates first data for the audio data, where the first data is used to describe the corresponding channel. In some embodiments, the first data is used to describe the function of the corresponding channel. In some embodiments, the first data is used to describe the function of the audio data included in the corresponding channel. In some embodiments, the disclosed embodiments do not limit the name of the first data. For example, the first data may be metadata, descriptive data, or descriptive information.
[0137] Step S2102: The first device sends audio data of multiple channels to the encoder.
[0138] In some embodiments, the encoder receives audio data of multiple channels sent by the first device. In some embodiments, the first device sends audio data of multiple channels. In some embodiments, the encoder receives audio data of multiple channels.
[0139] In an embodiment of the present disclosure, after the first device collects audio data of multiple channels, it can send the audio data of multiple channels to the encoder. The subsequent encoder can process the audio data of multiple channels to facilitate the transmission of the encoded data, thereby ensuring the saving of transmission resources, and other devices can decode the encoded data and verify the obtained audio data.
[0140] In some embodiments, the first device not only sends the audio data of multiple channels to the encoder, but also sends the first data of multiple channels to the encoder, so that the encoder performs subsequent processes after receiving the audio data of multiple channels and the first data.
[0141] Step S2103: The encoder processes the audio data included in the multiple channels based on the first data included in each of the multiple channels to obtain second data.
[0142] In some embodiments, the first data is used to describe the corresponding channel. In some embodiments, the first data is used to describe the function of the corresponding channel. In some embodiments, the first data is used to describe the function of the audio data included in the corresponding channel. In some embodiments, the disclosed embodiments do not limit the name of the first data. For example, the first data may be metadata, descriptive data, or descriptive information.
[0143] It should be noted that while the disclosed embodiments illustrate the use of first data to describe a channel, in another embodiment, the first data may also be used to describe audio data, meaning that the information contained in the audio data is also included in the first data. Alternatively, it can be understood that the information contained in the first data is not limited to describing the corresponding channel, but is also used to describe the audio data in the channel.
[0144] In the embodiment of the present disclosure, the first data is used to describe the corresponding channel. Therefore, the audio data included in each channel can be processed based on the first data included in each channel to achieve compression of the audio data, reduce the amount of data, and ensure that the compressed data can be transmitted efficiently.
[0145] In some embodiments, the encoder also receives first data sent by the first device, and the encoder needs to first determine the parameters of the audio data indicated by the first data, and then process the audio data included in the multiple channels according to the parameters of the indicated audio data to obtain second data.
[0146] In some embodiments, the encoder determines encoding parameters corresponding to the first data based on the scene indicated by the first data. The encoding parameters are used to process the audio data. The encoding parameters are processed to obtain the encoded first data. In the disclosed embodiments, the encoder stores encoding parameters corresponding to each scene. After receiving the first data, the encoder can determine the encoding parameters corresponding to the first data, and then the encoder processes the encoding parameters of the first data before subsequently transmitting it.
[0147] In some embodiments, after receiving the first data, the encoder needs to first analyze the first data, then determine the encoding parameters corresponding to the first data, and then perform subsequent operations. Alternatively, it can be understood that the encoder needs to first convert the first signal to obtain the encoding parameters corresponding to the first data.
[0148] In some embodiments, each channel includes one audio data. In the disclosed embodiment, each channel corresponds to one audio data, which can also be understood as one channel being used to transmit one audio data. In some embodiments, the number of the channels is the same as the number of the audio data.
[0149] In some embodiments, processing the audio data includes performing inter-channel decorrelation processing on the audio data. In some embodiments, the second data is obtained by performing decorrelation processing on the audio data included in multiple channels. In the embodiments of the present disclosure, decorrelation can also be understood as similarity between the audio data included in the multiple channels. Therefore, the audio data included in the multiple channels can be processed such as de-redundancy to eliminate the correlation between the audio data included in the multiple channels and ensure that the audio data included in the multiple channels has reduced dimensionality.
[0150] In some embodiments, referring to FIG2C , the present disclosure processes the audio data by performing preprocessing, multi-channel processing, core module processing, and other processes on the input first audio data, second audio data, and third audio data, and the multi-channel processing also refers to the first data. In addition, the first data is also processed, and then the audio data and the data obtained by encoding the first data are mixed to obtain a data stream and transmitted.
[0151] In some embodiments, the first data is used to indicate the type of audio data included in the channel. Optionally, different types of audio data refer to audio data with different functions. The following describes different types of audio data.
[0152] In some embodiments, the first data includes reference information, which is used to indicate that the channel corresponding to the first data includes reference audio data for matching with the audio data to be tested. Optionally, the reference audio data can also be understood as standard audio data, indicating that the device is normal, that is, audio data generated when the device is in a normal operating state. Optionally, by matching the audio data to be tested with the reference audio data, it can be determined whether the audio data to be tested matches the reference audio data, and thus whether the audio data to be tested is normal.
[0153] In some embodiments, the first data includes information to be tested, where the information to be tested is used to indicate that audio data to be tested included in a channel corresponding to the first data is used to be matched with reference audio data.
[0154] Optionally, the audio data in the channel may be audio data to be tested. Optionally, the audio data to be tested may also be understood as audio data that needs to be analyzed and tested. Optionally, by matching the audio data to be tested with reference audio data, it is possible to determine whether the audio data to be tested matches the reference audio data, and further determine whether the audio data to be tested is normal.
[0155] In some embodiments, one channel among multiple channels includes reference audio data, and other channels include audio data to be tested. Based on the first data included in each of the multiple channels, the audio data included in the multiple channels are processed to obtain second data, including: obtaining the degree of difference between the audio data to be tested included in each channel and the reference audio data, processing the audio data according to the degree of difference between the reference audio data and each channel, and obtaining the second data.
[0156] In an embodiment of the present disclosure, if the audio data of one channel among multiple channels is reference audio data and the audio data of other channels is audio data to be tested, decorrelation processing can be performed on the audio data by obtaining the difference between the reference audio data and the audio data to be tested.
[0157] In some embodiments, the degree of difference between the audio data to be tested and the reference audio data refers to the distinction between the audio data to be tested and the reference audio data, or can be understood as the difference between the audio data to be tested and the reference audio data.
[0158] In some embodiments, obtaining the degree of difference between the audio data to be tested and the reference audio data included in each channel includes: obtaining the difference between the audio data to be tested and the reference audio data included in each channel, and determining the difference as the degree of difference. In the disclosed embodiment, the degree of difference between the audio data to be tested and the reference audio data is determined using the difference between the audio data to be tested and the reference audio data. In other words, the difference between the audio data to be tested and the reference audio data can express the difference between the two.
[0159] In some embodiments, referring to FIG. 2D , the first data indicates that the first audio data is reference audio data, and the first data indicates that the second audio data and the third audio data are audio data to be tested. Then, the first audio data is not processed, and the difference between the second audio data and the first audio data, as well as the difference between the third audio data and the first audio data, is obtained.
[0160] It should be noted that the above embodiment uses the example of directly obtaining the degree of difference between the audio data to be tested and the reference audio data. In another embodiment, before obtaining the degree of difference, the audio data of each channel may be processed, and the degree of difference between the audio data may be obtained based on the processed audio data.
[0161] In some embodiments, for the audio data to be tested included in each channel, phase rotation is performed on the audio data to be tested in each channel so that the phase of the audio data to be tested in each channel is the same as that of the reference audio data, and the degree of difference between the phase-rotated audio data to be tested and the reference audio data in each channel is obtained. In the disclosed embodiments, by rotating the audio data to be tested, the phase of the audio data to be tested can be made the same as that of the reference audio data, thereby ensuring that the degree of difference between the audio data to be tested and the reference audio data based on the same phase is more accurate.
[0162] In some embodiments, for the audio data to be tested included in each channel, energy balancing is performed on the audio data to be tested in each channel so that the amplitude of the audio data to be tested in each channel is the same as that of the reference audio data, and the degree of difference between the audio data to be tested and the reference audio data after energy balancing is obtained in each channel. In the disclosed embodiments, by adjusting the energy of the audio data to be tested so that the energy of the audio data to be tested is the same as that of the reference audio data, the degree of difference between the audio data to be tested and the reference audio data based on the same phase can be ensured to be more accurate.
[0163] It should be noted that the above embodiments illustrate phase rotation and energy balancing as examples. In another embodiment, phase rotation may be performed on the audio data first, followed by energy balancing, and then the difference between the reference audio data and the audio data to be tested after energy balancing is determined.
[0164] In some embodiments, referring to Figure 2E, the first data indicates that the first audio data is reference audio data, and the first data indicates that the second audio data and the third audio data are audio data to be tested, then the first audio data is not processed, and the second audio data and the third audio data are phase rotated and energy balanced, respectively, and then the difference between the processed second audio data and the first audio data, as well as the difference between the processed third audio data and the first audio data, are obtained.
[0165] In some embodiments, processing audio data included in multiple channels to obtain second data includes: performing kernel selection and kernel encoding on the audio data included in each channel to obtain the second data. In the disclosed embodiments, kernel selection and kernel encoding refer to selecting a kernel module for kernel encoding and then kernel encoding the audio data using the selected kernel module to obtain the second data.
[0166] In some embodiments, after the audio data is processed according to the above embodiment, the processed audio data is processed again based on the core module selected by the core to obtain second data.
[0167] In some embodiments, a channel corresponds to first data, and the first data is used to indicate the first core module of the channel, and the audio data included in each channel is core-selected and core-encoded to obtain second data, including: determining a third core module based on the second core module selected by the audio data and the first core module, the third core module is used to process the audio data, and the audio data is core-encoded using the third core module to obtain the second data. Optionally, the second core module selected based on the audio data refers to the core module required to be used after analyzing the audio data. In some embodiments, the type of audio data can be determined after signal analysis, and then the corresponding core module is determined according to the type of audio data. Optionally, the type of audio data and the core module are in a one-to-one correspondence.
[0168] In some embodiments, when performing core selection in the embodiments of the present disclosure, the encoder processes the identifier of the core module indicated by the first data so as to transmit the encoded identifier of the core module indicated by the first data.
[0169] Optionally, the encoding parameters corresponding to the first data include an identifier of a core module corresponding to the core selection, and the encoder may process the identifier of the core module to obtain the encoded second data.
[0170] In some embodiments, determining a third core module based on the second core module selected by the audio data and the first core module includes: if the first core module is identical to the second core module, and determining the third core module to be either the first core module or the second core module. In the disclosed embodiment, if the first core module is identical to the second core module, it means that the same core module is used regardless of which one is used, and therefore the third core module is determined to be either the first core module or the second core module.
[0171] In some embodiments, determining a third core module based on the second core module selected based on the audio data and the first core module includes: determining the third core module to be the first core module because the first core module is different from the second core module. In some embodiments, the priority of the core module indicated by the first data is higher than the priority of the core module selected based on the audio data. Alternatively, it can be understood that the priority of the core module indicated by the first data is higher than the priority of the core module determined based on signal analysis.
[0172] In some embodiments, determining a third core module based on the second core module selected based on the audio data and the first core module includes: determining the third core module to be the second core module because the first core module is different from the second core module. In some embodiments, the priority of the core module indicated by the first data is lower than the priority of the core module selected based on the audio data. Alternatively, it can be understood that the priority of the core module indicated by the first data is higher than the priority of the core module determined based on signal analysis.
[0173] In some embodiments, determining a third core module based on the second core module selected by the audio data and the first core module includes: if the first core module is different from the second core module, determining the third core module to be either the first core module or the second core module. In some embodiments, the priority of the core module indicated by the first data is not differentiated from the priority of the core module selected based on the audio data, and one of the core modules determined by the two methods is randomly selected.
[0174] In some embodiments, the encoder in the present disclosure pre-processes the audio data, performs at least one of multi-channel data encoding based on the first data, kernel selection, kernel encoding, and multi-channel data mixing.
[0175] In some embodiments, pre-processing further comprises at least one of the following:
[0176] 1. Divide into multiple frames;
[0177] Optionally, framing refers to inputting continuous PCM (Pulse Code Modulation) data and a time-varying signal, and then splitting the audio data into multiple audio data of fixed durations through a window function signal. Optionally, the fixed duration can be a value such as 5ms (milliseconds), 10ms, 20ms, or other lengths, which are not limited in the embodiments of the present disclosure. In the embodiments of the present disclosure, framing can ensure that the split audio data is stable within a fixed duration, thereby ensuring the stability of the audio data.
[0178] 2. Time-frequency transformation;
[0179] Optionally, time-frequency transform refers to converting audio data from the time domain to the frequency domain using MDCT (Modified Discrete Cosine Transform) or FFT (Fast Fourier Transform). The disclosed embodiments can process signals in different frequency ranges separately, thereby increasing processing efficiency.
[0180] 3. Time window adjustment;
[0181] Optionally, the time window adjustment refers to using a short window for transient audio data and a long window for non-transient audio data.
[0182] 4. Time domain noise shaping;
[0183] Optionally, time-domain noise shaping refers to inputting audio data with a shorter time domain or transient audio data, and then filtering the data using a time-domain noise shaping filter to obtain filtered audio data.
[0184] 5. Frequency domain noise shaping.
[0185] Optionally, frequency domain noise shaping refers to filtering frequency domain audio data using frequency domain noise shaping to obtain filtered audio data.
[0186] 6. Bandwidth filtering.
[0187] Optionally, bandwidth filtering refers to reducing the bandwidth of the audio data.
[0188] Step S2104: The encoder processes the first data to obtain encoded first data.
[0189] In the embodiment of the present disclosure, the encoder not only needs to process the audio data, but also needs to process the first data, so that the encoder can send the encoded first data to the decoder.
[0190] In some embodiments, the first data also includes parameter information used when processing the audio data. In some embodiments, if the encoder obtains the degree of difference between the audio data, the first data may indicate the degree of difference between the reference audio data and the audio data to be measured. In some embodiments, if the encoder performs phase rotation on the audio data, the first data may indicate the rotated phase of the audio data to be measured. In some embodiments, if the encoder performs energy balance on the audio data, the first data may indicate the adjusted energy of the audio data to be measured.
[0191] It should be noted that step S2104 in the embodiment of the present disclosure is an optional step, and in another embodiment, step S2104 may not be performed. The embodiment of the present disclosure does not limit this.
[0192] It should be noted that the encoder in the embodiments of the present disclosure may also be located within the first device, meaning that the first device has both acquisition and encoding functions. In some embodiments, the first device includes an acquisition module and an encoding module. The acquisition module is configured to perform steps S2101-S2102, and the encoding module is configured to perform steps S2103-S2105. Alternatively, the first device and the encoder are two separate devices.
[0193] Step S2105: The encoder sends the second data and the encoded first data to the decoder.
[0194] In some embodiments, the decoder receives the second data sent by the encoder and the encoded first data. In some embodiments, the encoder sends the second data and the encoded first data. In some embodiments, the decoder receives the second data and the encoded first data.
[0195] In some embodiments, if step S2104 is not performed in the embodiments of the present disclosure, step S2105 may be replaced by: the encoder sends the second data and the first data to the decoder.
[0196] Step S2106: The decoder decodes the second data based on the third data included in each of the multiple channels to obtain the audio data included in the multiple channels.
[0197] In the embodiment of the present disclosure, after receiving the second data, the decoder may decode the second data according to the third data included in each channel to obtain the audio data included in the multiple channels.
[0198] In some embodiments, the third data is obtained by decoding the encoded first data. Optionally, compared to the first data, the third data includes parameters used by the encoder to process the audio data. Optionally, compared to the first data, the third data suffers encoding and decoding losses. Alternatively, the third data may be understood to indicate the same content as the first data. Alternatively, the third data may be understood to be identical to the first data.
[0199] In some embodiments, after receiving the encoded first data and second data, the decoder first decodes the encoded first data to obtain third data, and then decodes the second data based on the third data to obtain audio data corresponding to the second data.
[0200] In some embodiments, the third data is used to indicate the corresponding channel. The third data is used to indicate the type of audio data included in the channel. In the embodiment of the present disclosure, the third data is similar to the first data in the above embodiment and will not be repeated here.
[0201] In some embodiments, the third data includes reference information, where the reference information is used to indicate that the channel corresponding to the first data includes reference audio data for matching with the audio data to be tested.
[0202] In some embodiments, the third data includes information to be tested, where the information to be tested is used to indicate that audio data to be tested included in a channel corresponding to the first data is used to be matched with reference audio data.
[0203] In the embodiment of the present disclosure, after the second data of each channel is determined by the third data, the second data of each channel can be decoded to obtain corresponding audio data.
[0204] In some embodiments, one of the multiple channels includes reference audio data, and the other channels include audio data to be tested. Decoding the second data based on the third data included in each of the multiple channels to obtain the audio data included in the multiple channels includes: decoding the second data based on the third data included in each of the multiple channels to obtain a degree of difference between the audio data to be tested included in each channel and the reference audio data; and determining the audio data included in the multiple channels based on the degree of difference between the audio data to be tested included in each channel and the reference audio data. The decoding process in the disclosed embodiments is the opposite of the encoding process described above and is not further described herein.
[0205] In some embodiments, the degree of difference is a difference between the audio data to be tested and the reference audio data included in each channel.
[0206] In some embodiments, referring to FIG2F , after receiving the data stream, the decoder demultiplexes the data stream to obtain second data and third data after encoding the audio data of multiple channels, and then performs core decoding, multi-channel decoding, and pre-processing decoding on the second data based on the third data to obtain audio data for each channel.
[0207] In some embodiments, the second device includes a decoder, meaning the decoder can also be located within the second device, meaning the second device has both decoding and detection capabilities. In some embodiments, the second device includes a decoding module and a detection module, wherein the decoding module is configured to perform step S2106 above, and the detection module is configured to perform machine analysis or machine training on the audio data. In some embodiments, the second device and the decoder are separate devices.
[0208] In some embodiments, the names of information, etc. are not limited to the names described in the embodiments, and terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "symbol", "codeword", "codebook", "codeword", "codepoint", "bit", "data", "program", and "chip" can be used interchangeably.
[0209] In some embodiments, terms such as "uplink", "uplink", "physical uplink" can be interchangeable with each other, and terms such as "downlink", "downlink", "physical downlink" can be interchangeable with each other, and terms such as "side", "sidelink", "side communication", "sidelink communication", "direct connection", "direct link", "direct communication", "direct link communication" can be interchangeable with each other.
[0210] In some embodiments, "obtain", "get", "get", "receive", "transmit", "bidirectional transmission", "send and / or receive" can be interchangeable, and can be interpreted as receiving from other entities, obtaining from protocols, obtaining from higher layers, obtaining by self-processing, autonomous implementation, etc.
[0211] In some embodiments, terms such as "send", "transmit", "report", "download", "transmit", "bidirectional transmission", "send and / or receive" can be used interchangeably.
[0212] In some embodiments, terms such as "moment", "time point", "time", and "time position" can be replaced with each other, and terms such as "duration", "period", "time window", "window", and "time" can be replaced with each other.
[0213] In some embodiments, terms such as "certain", "preset", "preset", "setting", "indicated", "a certain", "any", and "first" can be interchangeable. "Specific A", "preset A", "preset A", "setting A", "indicated A", "a certain A", "any A", and "first A" can be interpreted as A pre-specified in a protocol, etc., or as A obtained through setting, configuration, or indication, etc., or as specific A, a certain A, any A, or first A, etc., but not limited to this.
[0214] The encoding and decoding method involved in the embodiment of the present disclosure may include at least one of steps S2101 to S2106. For example, step S2101 can be implemented as an independent embodiment, step S2102 can be implemented as an independent embodiment, step S2103 can be implemented as an independent embodiment, step S2104 can be implemented as an independent embodiment, step S2105 can be implemented as an independent embodiment, step S2106 can be implemented as an independent embodiment, step S2101 and step S2102 can be implemented as independent embodiments, step S2103 and step S2104 can be implemented as independent embodiments, step S2105 and step S2106 can be implemented as independent embodiments, step S2101, step S2102, step S2103, and step S2104 can be implemented as independent embodiments, step S2101, step S2102, step S2105, and step S2106 can be implemented as independent embodiments, but is not limited to this.
[0215] In some embodiments, step S2101 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0216] In some embodiments, step S2102 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0217] In some embodiments, step S2103 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0218] In some embodiments, step S2104 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0219] In some embodiments, step S2105 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0220] In some embodiments, step S2106 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0221] In some embodiments, step S2101 and step S2102 are optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0222] In some embodiments, step S2103 and step S2104 are optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0223] In some embodiments, step S215 and step S2106 are optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0224] In some embodiments, reference may be made to other optional implementations described before or after the description corresponding to FIG. 2A .
[0225] FIG3A is a flow chart of a coding and decoding method according to an embodiment of the present disclosure, which is applied to an encoder. As shown in FIG3A , the embodiment of the present disclosure relates to a coding and decoding method, which includes:
[0226] Step S3101: The encoder processes the audio data included in the multiple channels based on the first data included in each of the multiple channels to obtain second data.
[0227] The optional implementation of step S3101 can be found in step S2103 of FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.
[0228] Step S3102: The encoder processes the first data to obtain encoded first data.
[0229] Optional implementations of step S3102 may refer to step S2104 in FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.
[0230] Step S3103: The encoder sends the second data and the encoded first data to the decoder.
[0231] The optional implementation of step S3103 can be found in step S2105 of FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.
[0232] The encoding and decoding method involved in the embodiments of the present disclosure may include at least one of steps S3101 to S3103. For example, step S3101 may be implemented as an independent embodiment, step S3102 may be implemented as an independent embodiment, and step S3103 may be implemented as an independent embodiment, or at least two steps may be combined, but the present invention is not limited thereto.
[0233] In some embodiments, step S3101 is optional, step S3102 is optional, and step S3103 is optional, and one or more of these steps may be omitted or replaced in different embodiments, but the present invention is not limited thereto.
[0234] FIG3B is a flow chart of a coding and decoding method according to an embodiment of the present disclosure, which is applied to an encoder. As shown in FIG3B , the embodiment of the present disclosure relates to a coding and decoding method, which includes:
[0235] Step S3201: The encoder processes the audio data included in the multiple channels based on the first data included in each of the multiple channels to obtain second data.
[0236] Optional implementations of step S3201 may refer to step S2103 in FIG. 2A , step S3101 in FIG. 3A , and other related parts in the embodiments involved in FIG. 2A and FIG. 3A , which will not be described in detail here.
[0237] FIG4 is a flow chart of a coding and decoding method according to an embodiment of the present disclosure, which is applied to a decoder. As shown in FIG4 , the embodiment of the present disclosure relates to a coding and decoding method, which includes:
[0238] Step S4101: The decoder decodes the second data based on the third data included in each of the multiple channels to obtain the audio data included in the multiple channels.
[0239] In some embodiments, the audio data is used for machine analysis, and the first signal includes encoded information for processing the audio data.
[0240] Optional implementations of step S4101 may refer to step S2106 in FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.
[0241] FIG5 is a flow chart of a coding and decoding method according to an embodiment of the present disclosure. As shown in FIG6 , an embodiment of the present disclosure relates to a coding and decoding method, which includes:
[0242] Step S5101: The encoder processes the audio data included in the multiple channels based on the first data included in each of the multiple channels to obtain second data.
[0243] In some embodiments, the audio data is used for machine analysis, and the first signal includes encoded information for processing the audio data.
[0244] Step S5102: The decoder decodes the second data based on the third data included in each of the multiple channels to obtain the audio data included in the multiple channels.
[0245] In some embodiments, the above method may include the methods of the above embodiments of the communication system side, the first device side, the encoder side, the decoder side, etc., which will not be repeated here.
[0246] FIG6 is a flow chart of a coding and decoding method according to an embodiment of the present disclosure. As shown in FIG6 , the embodiment of the present disclosure relates to a coding and decoding method, and the method includes:
[0247] Step S6101: The encoder processes the data to obtain encoded data.
[0248] In some embodiments, the input to the encoder is N audio channels and corresponding metadata.
[0249] In some embodiments, the metadata analysis module analyzes the metadata to obtain metadata that can guide the multi-channel encoding module. Furthermore, the metadata analysis module analyzes correlations between metadata, prepares metadata to be encoded, and prepares prediction parameters for metadata that can be predicted by other metadata. The metadata second data is written into the bitstream along with the audio second data.
[0250] In some embodiments, the N audio channels are first independently pre-processed. The pre-processing module performs one or more of the following operations on the audio signal: framing, time-frequency conversion, window length determination, long frame and short frame determination, time-domain noise shaping, frequency-domain noise shaping, etc.
[0251] In some embodiments, redundancy between the N pre-processed audio channels is then removed in a multi-channel encoding module. The output of the multi-channel encoding module is a multi-channel downmix and downmix channels, along with corresponding side information. The multi-channel downmix and downmix channels are processed together by the core encoding module to produce N encoded channel data. The N encoded channel data is mixed with encoding metadata and written to the bitstream.
[0252] In some embodiments, the bitstream is decomposed into N multi-channel encoded channel data, encoding metadata, and side information.
[0253] In some embodiments, the metadata decoding module decodes the compressed metadata to obtain final decoded metadata.
[0254] In some embodiments, the multi-channel decoding module obtains N independent channel data using N multi-channel data and multi-channel side information.
[0255] In some embodiments, the final N decoded audio channels are obtained by a post-processing module that performs one or more operations on the audio signal, including inverse time-frequency transform, inverse time-domain noise shaping transform, and inverse frequency-domain noise shaping transform.
[0256] In some embodiments, channel 1 has corresponding metadata, wherein one of the metadata information is "ground truth", which indicates the purpose of channel 1.
[0257] In some embodiments, channel 2 has corresponding metadata 2. One piece of information in the metadata 2 is "product 1 to be testing", which indicates the purpose of channel 2.
[0258] In some embodiments, channel 3 has corresponding metadata, wherein one of the metadata is "product 2 to be tested", which indicates the purpose of channel 3.
[0259] In some embodiments, before differential encoding, correlation analysis is performed to obtain phase correlation and amplitude difference. If there is a phase difference, it is rotated to make the phase consistent. If there is an amplitude difference, energy balance is performed to make the amplitude consistent. The rotation side information and energy balance side information are written into the bit stream together with the audio data.
[0260] In some embodiments, the core encoding module includes a core selection module and a core set, where task1 core, task2 core, and taskN core are used for machine audio encoding, and speech core / music core are used for human audio encoding. The core selection module includes a signal analysis module and a signal classification module for selecting the most appropriate encoding core.
[0261] In some embodiments, metadata captures a wealth of useful information about the recorded environment and also aids in selecting an appropriate encoding core.
[0262] In the embodiments of the present disclosure, some or all of the steps and their optional implementations may be arbitrarily combined with some or all of the steps in other embodiments, or may be arbitrarily combined with the optional implementations of other embodiments.
[0263] The present disclosure also provides an apparatus for implementing any of the above methods. For example, a device is provided that includes units or modules for implementing each step performed by an encoder in any of the above methods. For another example, another device is provided that includes units or modules for implementing each step performed by a decoder in any of the above methods.
[0264] It should be understood that the division of the various units or modules in the above device is merely a division of logical functions. In actual implementation, they may be fully or partially integrated into a physical entity, or they may be physically separated. In addition, the units or modules in the device may be implemented in the form of a processor calling software: for example, the device includes a processor, the processor is connected to a memory, and the memory stores instructions. The processor calls the instructions stored in the memory to implement any of the above methods or implement the functions of the various units or modules of the above device, wherein the processor is, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory is a memory within the device or a memory outside the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits, and the functions of some or all of the units or modules can be realized by designing the hardware circuits. The above-mentioned hardware circuits can be understood as one or more processors; for example, in one implementation, the above-mentioned hardware circuit is an application-specific integrated circuit (ASIC), which realizes the functions of some or all of the above units or modules by designing the logical relationship of the components in the circuit; for example, in another implementation, the above-mentioned hardware circuit can be realized by a programmable logic device (PLD). Taking a field programmable gate array (FPGA) as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by configuring the configuration file, thereby realizing the functions of some or all of the above units or modules. All units or modules of the above devices can be realized in the form of software called by the processor, or in the form of hardware circuits, or in part by the form of software called by the processor, and the rest by hardware circuits.
[0265] In the embodiments of the present disclosure, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. The logical relationships of the above-mentioned hardware circuits are fixed or reconfigurable. For example, the processor is a hardware circuit implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration file to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), etc.
[0266] Figure 7A is a structural diagram of the encoding and decoding device proposed in an embodiment of the present disclosure. As shown in Figure 7A, the encoding and decoding device 7100 may include: at least one of a transceiver module 7101, a processing module 7102, etc. In some embodiments, the processing module 7102 is used to process the audio data included in the multiple channels based on the first data included in each channel of the multiple channels to obtain second data, and the processing includes inter-channel decorrelation. The first data is used to describe the corresponding channel, and each channel includes one audio data. The first data is generated by the first device, and the audio data is collected by the first device. Optionally, the above-mentioned transceiver module 7101 is used to perform at least one of the communication steps such as sending and / or receiving performed by the encoder in any of the above methods (for example, step S2101 but not limited to this), which will not be repeated here. Optionally, the above-mentioned processing module is used to perform at least one of the other steps performed by the encoder in any of the above methods, which will not be repeated here.
[0267] Optionally, the processing module 7102 is used to execute at least one of the communication steps such as processing performed by the encoder in any of the above methods, which will not be repeated here.
[0268] Figure 7B is a structural diagram of the encoding and decoding device proposed in an embodiment of the present disclosure. As shown in Figure 7B, the encoding and decoding device 7200 may include: at least one of a transceiver module 7201, a processing module 7202, etc. In some embodiments, the processing module 7202 is used to decode the second data based on the third data included in each channel of the multiple channels to obtain the audio data included in the multiple channels, the third data is used to indicate the role of the corresponding channel, the second data is obtained by performing channel decorrelation processing on the audio data included in the multiple channels, each of the channels includes one audio data, the first data is generated by the first device, and the audio data is collected by the first device. Optionally, the above-mentioned transceiver module is used to perform at least one of the communication steps such as sending and / or receiving performed by the decoder in any of the above methods, which will not be repeated here.
[0269] Optionally, the processing module 7202 is used to execute at least one of the communication steps such as processing performed by the decoder in any of the above methods, which will not be repeated here.
[0270] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module, and the transmitting module and the receiving module may be separate or integrated. Optionally, the transceiver module may be interchangeable with the transceiver.
[0271] In some embodiments, the processing module can be a single module or can include multiple submodules. Optionally, the multiple submodules each execute all or part of the steps required to be executed by the processing module. Optionally, the processing module can be interchangeable with the processor.
[0272] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module, and the transmitting module and the receiving module may be separate or integrated. Optionally, the transceiver module may be interchangeable with the transceiver.
[0273] In some embodiments, the processing module can be a single module or can include multiple submodules. Optionally, the multiple submodules each execute all or part of the steps required to be executed by the processing module. Optionally, the processing module can be interchangeable with the processor.
[0274] Figure 8A is a schematic diagram of the structure of a communication device 8100 proposed in an embodiment of the present disclosure. Communication device 8100 can be a first device, an encoder, a decoder, or a chip, a chip system, or a processor that supports the first device, encoder, or decoder in implementing any of the above methods. Communication device 8100 can be used to implement the methods described in the above method embodiments. For details, please refer to the description of the above method embodiments.
[0275] As shown in Figure 8A, communication device 8100 includes one or more processors 8101. Processor 8101 can be a general-purpose processor or a dedicated processor, such as a baseband processor or a central processing unit. The baseband processor can be used to process communication protocols and communication data, while the central processing unit can be used to control the codec device, execute programs, and process program data. Communication device 8100 is configured to perform any of the above methods.
[0276] In some embodiments, the communication device 8100 further includes one or more memories 8102 for storing instructions. Optionally, all or part of the memories 8102 may be located outside the communication device 8100.
[0277] In some embodiments, the communication device 8100 further includes one or more transceivers 8103. When the communication device 8100 includes one or more transceivers 8103, the transceiver 8103 performs at least one of the communication steps of sending and / or receiving in the above method.
[0278] In some embodiments, a transceiver may include a receiver and / or a transmitter. The receiver and transmitter may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, and transceiver circuit may be used interchangeably; the terms transmitter, transmitting unit, transmitter, and transmitting circuit may be used interchangeably; and the terms receiver, receiving unit, receiver, and receiving circuit may be used interchangeably.
[0279] In some embodiments, the communication device 8100 may include one or more interface circuits 8104. Optionally, the interface circuit 8104 is connected to the memory 8102. The interface circuit 8104 may be configured to receive signals from the memory 8102 or other devices, and may be configured to send signals to the memory 8102 or other devices. For example, the interface circuit 8104 may read instructions stored in the memory 8102 and send the instructions to the processor 8101.
[0280] The communication device 8100 described in the above embodiments may be a first device, a decoder, or an encoder, but the scope of the communication device 8100 described in the present disclosure is not limited thereto, and the structure of the communication device 8100 may not be limited by FIG. 8A. The communication device may be an independent device or may be part of a larger device. For example, the communication device may be: 1) an independent integrated circuit IC, or a chip, or a chip system or subsystem; (2) a collection of one or more ICs, optionally, the above IC collection may also include a storage component for storing data or programs; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, an encoder device, an intelligent encoder device, a cellular phone, a wireless device, a handheld device, a mobile unit, an in-vehicle device, a decoder, a cloud device, an artificial intelligence device, etc.; (6) others, etc.
[0281] FIG8B is a schematic diagram of the structure of a chip 8200 according to an embodiment of the present disclosure. If the communication device 8100 can be a chip or a chip system, please refer to the schematic diagram of the structure of the chip 8200 shown in FIG8B , but the present disclosure is not limited thereto.
[0282] The chip 8200 includes one or more processors 8201 , and the chip 8200 is configured to execute any of the above methods.
[0283] In some embodiments, the chip 8200 further includes one or more interface circuits 8202. Optionally, the interface circuit 8202 is connected to the memory 8203. The interface circuit 8202 can be used to receive signals from the memory 8203 or other devices, and can be used to send signals to the memory 8203 or other devices. For example, the interface circuit 8202 can read instructions stored in the memory 8203 and send the instructions to the processor 8201.
[0284] In some embodiments, the interface circuit 8202 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processor 8201 performs at least one of the other steps.
[0285] In some embodiments, terms such as interface circuit, interface, transceiver pin, and transceiver may be used interchangeably.
[0286] In some embodiments, the chip 8200 further includes one or more memories 8203 for storing instructions. Alternatively, all or part of the memories 8203 may be outside the chip 8200.
[0287] The present disclosure also proposes a storage medium having instructions stored thereon, which, when executed on the communication device 8100, causes the communication device 8100 to execute any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but is not limited thereto, and may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but is not limited thereto, and may also be a temporary storage medium.
[0288] The present disclosure also provides a program product, which, when executed by the communication device 8100, enables the communication device 8100 to perform any of the above methods. Optionally, the program product is a computer program product.
[0289] The present disclosure also proposes a computer program, which, when executed on a computer, causes the computer to perform any one of the above methods.
Claims
1. A coding method, characterized in that: The method is performed by an encoder, and includes: Based on first data included in each of a plurality of channels, audio data included in the plurality of channels is processed to obtain second data, wherein the processing includes decorrelation between channels, the first data is used to describe the corresponding channel, the first data is generated by a first device, each of the channels includes one audio data, and the audio data is collected by the first device.
2. The method according to claim 1, characterized in that The first data is used to indicate a type of audio data included in the channel.
3. The method according to claim 2, characterized in that The first data includes reference information, where the reference information is used to indicate that reference audio data included in the channel corresponding to the first data is used to match the audio data to be tested.
4. The method according to claim 2, characterized in that The first data includes information to be tested, where the information to be tested is used to indicate that audio data to be tested included in the channel corresponding to the first data is used to be matched with reference audio data.
5. The method according to any one of claims 1 to 4, characterized in that: One of the multiple channels includes reference audio data, and the other channels include audio data to be measured, and the processing of the audio data included in the multiple channels based on the first data included in each of the multiple channels to obtain second data includes: Obtaining a degree of difference between the audio data to be tested and the reference audio data included in each channel; The audio data is processed based on the difference between the reference audio data and each of the channels to obtain the second data.
6. The method according to claim 5, characterized in that The obtaining of the difference between the audio data to be tested and the reference audio data included in each channel includes: A difference between the audio data to be tested and the reference audio data included in each channel is obtained and determined as the difference degree.
7. The method according to claim 5, characterized in that The method further comprises: For the audio data to be tested included in each channel, performing phase rotation on the audio data to be tested of each channel so that the phase of the audio data to be tested of each channel is the same as that of the reference audio data; The obtaining of the difference between the audio data to be tested and the reference audio data included in each channel includes: The degree of difference between the audio data to be measured and the reference audio data after phase rotation is obtained in each channel.
8. The method according to claim 5, characterized in that The method further comprises: For the audio data to be tested included in each channel, performing energy balancing on the audio data to be tested in each channel so that the amplitude of the audio data to be tested in each channel is the same as that of the reference audio data; The obtaining of the difference between the audio data to be tested and the reference audio data included in each channel includes: Obtain the degree of difference between the audio data to be measured and the reference audio data after energy balance is performed in each channel.
9. The method according to any one of claims 1 to 4, characterized in that: The processing of the audio data included in the plurality of channels to obtain the second data includes: Perform kernel selection and kernel encoding on the audio data included in each of the channels to obtain the second data.
10. The method according to claim 9, characterized in that The channel corresponds to first data, the first data being used to indicate a first core module of the channel, and performing core selection and core encoding on the audio data included in each of the channels to obtain the second data, including: Determining a third core module based on the second core module selected from the audio data and the first core module, wherein the third core module is used to process the audio data; The third core module is used to perform the core encoding on the audio data to obtain the second data.
11. The method according to claim 10, characterized in that The determining of the third core module based on the second core module selected based on the audio data and the first core module includes: The first core module is the same as the second core module, and the third core module is determined to be the first core module or the second core module. Any of the modules.
12. The method according to claim 10, characterized in that The determining of the third core module based on the second core module selected based on the audio data and the first core module includes: The first core module is different from the second core module, and the third core module is determined to be the first core module.
13. A decoding method, characterized in that: The method is performed by a decoder, and comprises: Based on the third data included in each of the multiple channels, the second data is decoded to obtain the audio data included in the multiple channels, the third data is used to indicate the function of the corresponding channel, the second data is obtained by performing channel decorrelation processing on the audio data included in the multiple channels, the first data is generated by the first device, each channel includes one audio data, and the audio data is collected by the first device.
14. The method according to claim 13, characterized in that The third data is used to indicate the type of audio data included in the channel.
15. The method according to claim 14, characterized in that The third data includes reference information, where the reference information is used to indicate that reference audio data included in the channel corresponding to the first data is used to match the audio data to be tested.
16. The method according to claim 14, characterized in that The third data includes information to be tested, where the information to be tested is used to indicate that the audio data to be tested included in the channel corresponding to the first data is used to match with reference audio data.
17. The method according to any one of claims 13 to 16, characterized in that: One of the multiple channels includes reference audio data, and the other channels include audio data to be measured, and decoding the second data based on the third data included in each of the multiple channels to obtain the audio data included in the multiple channels includes: decoding the second data based on the third data included in each of the multiple channels to obtain a degree of difference between the audio data to be tested and the reference audio data included in each channel; The audio data included in the multiple channels is determined based on the degree of difference between the audio data to be measured included in each channel and the reference audio data.
18. The method according to claim 17, characterized in that The difference degree is the difference between the audio data to be tested and the reference audio data included in each channel.
19. A coding and decoding method, characterized in that: The method comprises: The encoder processes audio data included in the multiple channels based on first data included in each channel to obtain second data, wherein the processing includes decorrelation between channels, the first data is used to describe the corresponding channel, the first data is generated by a first device, each channel includes one piece of audio data, and the audio data is collected by the first device; The decoder decodes the second data based on third data included in each of the multiple channels to obtain audio data included in the multiple channels, where the third data is used to indicate the corresponding channel.
20. An encoding device, characterized in that: The encoding device comprises: A processing module is configured to process the audio data included in the multiple channels based on first data included in each of the multiple channels to obtain second data, wherein the processing includes decorrelation between the channels, the first data is used to describe the corresponding channel, the first data is generated by a first device, and the second data is obtained by decorrelation processing of the audio data included in the multiple channels, each of the channels includes one audio data, and the audio data is collected by the first device.
21. A decoding device, characterized in that: The decoding device comprises: A processing module is configured to decode the second data based on third data included in each of the multiple channels to obtain audio data included in the multiple channels, where the third data is used to indicate the corresponding channel, the second data is obtained by performing inter-channel decorrelation processing on the audio data included in the multiple channels, the first data is generated by a first device, each channel includes one piece of audio data, and the audio data is collected by the first device.
22. A coding device, characterized in that: The encoding device comprises: one or more processors; The processor is configured to execute the encoding method according to any one of claims 1 to 12.
23. A decoding device, characterized in that: The decoding device comprises: one or more processors; The processor is configured to execute the decoding method according to any one of claims 13 to 18.
24. A coding and decoding system, characterized in that: The present invention comprises an encoder and a decoder, wherein the first device is configured to implement the encoding method according to any one of claims 1 to 12, and the encoder is configured to implement the decoding method according to any one of claims 13 to 18.
25. A storage medium storing instructions, characterized in that: When the instruction is executed on a communication device, the communication device is caused to execute the encoding method according to any one of claims 1 to 12, or execute the decoding method according to any one of claims 13 to 18.
Citation Information
Patent Citations
Methods for controlling inter-channel coherence of upmixed audio signals
CN104981867A
Signal decorrelation in an audio processing system
CN104995676A
Multi-track audio signal decorrelation coding method and device
CN106710600A
Audio signal processing with acoustic echo cancellation
CN111128210A
Coding and decoding method and device and storage medium
CN118202406A