Decoding method, encoding method, communication device, and storage medium
By extracting and encoding features from audio input data at the encoding end, decoding data for machine hearing is generated, solving the problem that existing technologies cannot meet the needs of AI tasks. This achieves efficient compression and decoding of audio data, meeting the analysis needs of AI tasks.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-23
- Publication Date
- 2026-03-26
AI Technical Summary
Existing audio encoding and decoding methods cannot meet the machine hearing requirements of different artificial intelligence tasks.
By extracting and encoding features from the audio input data at the encoding end, encoded data is generated to represent the feature information in the audio input data that is relevant to a specific task. Then, the data is decoded at the decoding end to generate decoded data for machine hearing, so as to meet the needs of specific AI tasks.
It achieves efficient compression and decoding of audio data, which can meet the analysis needs of different AI tasks and improve the accuracy and efficiency of machine hearing analysis.
Smart Images

Figure CN2024120495_26032026_PF_FP_ABST
Abstract
Description
Decoding, encoding method, communication device and storage medium TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of communication, and in particular, to a decoding, encoding method, communication device and storage medium. BACKGROUND
[0002] The audio encoding and decoding method defined by the related standard cannot meet the needs of machine hearing for different artificial intelligent (AI) tasks.
[0003] SUMMARY
[0004] Embodiments of the present disclosure provide a decoding, encoding method, communication device and storage medium to solve the technical problem that the audio encoding and decoding method defined by the related standard cannot meet the needs of machine hearing for different AI tasks.
[0005] According to a first aspect of embodiments of the present disclosure, a decoding method is provided, executed by a first communication device, and the method comprises: receiving first encoded data, the first encoded data comprising feature information related to a first task in audio input data; and decoding the first encoded data to obtain first decoded data, wherein the first decoded data is used to perform the first task based on machine hearing.
[0006] According to a second aspect of embodiments of the present disclosure, an encoding method is provided, executed by a second communication device, and the method comprises: encoding audio input data to obtain first encoded data; wherein the first encoded data comprises feature information related to a first task in the audio input data; and transmitting the first encoded data.
[0007] According to a third aspect of embodiments of the present disclosure, a decoding apparatus is provided, comprising: a transceiver module configured to receive first encoded data, the first encoded data comprising feature information related to a first task in audio input data; and a processing module configured to decode the first encoded data to obtain first decoded data, wherein the first decoded data is used to perform the first task based on machine hearing.
[0008] According to a fourth aspect of embodiments of the present disclosure, an encoding apparatus is provided, comprising: a processing module configured to encode audio input data to obtain first encoded data; wherein the first encoded data comprises feature information related to a first task in the audio input data; and a transceiver module configured to transmit the first encoded data.
[0009] According to a fifth aspect of the embodiments of the present disclosure, a communication device is provided, comprising: one or more processors; a memory coupled to the processors, the memory having stored therein executable instructions that, when executed by the processors, cause the communication device to perform the decoding method of the first aspect or the encoding method of the second aspect.
[0010] According to a sixth aspect of the embodiments of the present disclosure, a communication system is provided, comprising a first communication device and a second communication device, wherein the first communication device is configured to implement the decoding method of the first aspect, and the second communication device is configured to implement the encoding method of the second aspect.
[0011] According to a seventh aspect of the embodiments of the present disclosure, a storage medium is provided, the storage medium storing instructions that, when executed on a communication device, cause the communication device to perform the decoding or encoding method of the first aspect or the second aspect.
[0012] According to the embodiments of the present disclosure, by encoding the audio input data at the encoding end, first encoding information for representing feature information related to the first task in the audio input data is obtained and sent to the decoding end; the first encoding information is decoded by the decoding end to obtain first decoding data related to the first task, and input to the machine auditory module for the first task to analyze the first decoding data, so that the data obtained by encoding and decoding the audio input data can meet the needs of the first task, to analyze the analysis result required by the first task. BRIEF DESCRIPTION OF DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor.
[0014] FIG. 1 is an architecture schematic diagram of a communication system according to an embodiment of the present disclosure.
[0015] FIG. 2A is an interaction schematic diagram of a decoding and encoding method according to an embodiment of the present disclosure.
[0016] FIG. 2B is an interaction schematic diagram of a decoding and encoding method according to an embodiment of the present disclosure.
[0017] FIG. 2C is an interaction schematic diagram of a decoding and encoding method according to an embodiment of the present disclosure.
[0018] FIG. 2D is an interaction schematic diagram of a decoding and encoding method according to an embodiment of the present disclosure.
[0019] FIG. 2E is an interaction schematic diagram of a decoding and encoding method according to an embodiment of the present disclosure.
[0020] FIG. 2F is an interaction schematic diagram of a decoding and encoding method according to an embodiment of the present disclosure.
[0021] FIG. 3 is a schematic flowchart of a decoding method according to an embodiment of the present disclosure.
[0022] FIG. 4 is a schematic flowchart of an encoding method according to an embodiment of the present disclosure.
[0023] FIG. 5 is a schematic block diagram of an apparatus structure of a first communication device according to an embodiment of the present disclosure.
[0024] FIG. 6 is a schematic block diagram of an apparatus structure of a second communication device according to an embodiment of the present disclosure.
[0025] FIG. 7 is a schematic block diagram of a communication device according to an embodiment of the present disclosure.
[0026] FIG. 8 is a schematic block diagram of a chip according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] Embodiments of the present disclosure provide a decoding and encoding method, a communication device and a storage medium.
[0028] In a first aspect, embodiments of the present disclosure provide a decoding method, executed by a first communication device, comprising: receiving first encoded data, the first encoded data comprising feature information related to a first task in audio input data; decoding the first encoded data to obtain first decoded data, wherein the first decoded data is used to perform the first task based on machine hearing.
[0029] In the above embodiments, the audio input data is encoded at the encoding end to obtain first encoded information for representing feature information related to the first task in the audio input data, and is sent to the decoding end; the first encoded information is decoded by the decoding end to obtain first decoded data related to the first task, and is input to a machine hearing module for the first task to analyze the first decoded data, so that the data obtained by encoding and decoding the audio input data can meet the needs of the first task to analyze and obtain the analysis result required by the first task.
[0030] In combination with some embodiments of the first aspect. In some embodiments, the audio input data comprises audio data and label data; wherein the label data comprises description information of the audio data.
[0031] In some embodiments of the first aspect. In some embodiments, the first encoded data is audio feature data obtained by performing feature extraction on the audio input data.
[0032] In some embodiments of the first aspect. In some embodiments, the first encoded data is a bitstream obtained by encoding the audio feature data.
[0033] In some embodiments of the first aspect. In some embodiments, further comprising: receiving encoding information, the encoding information being used to indicate an encoding mode for encoding the audio feature data; and decoding the first encoded data based on a decoding mode corresponding to the encoding mode.
[0034] In some embodiments of the first aspect. In some embodiments, the encoding mode is determined by at least one of the following: a proportion of the audio feature data to the audio input data; a number of audio channels of the audio input data; a correlation between audio channels of the audio input data; and configuration information of the first task.
[0035] In some embodiments of the first aspect. In some embodiments, the configuration information of the first task comprises: the encoding information; and related information of the audio feature data required by the first task.
[0036] In some embodiments of the first aspect. In some embodiments, the encoding mode comprises at least one of the following: a first mode, the first mode being an encoding mode without compression of the audio feature data; a second mode, the second mode being an encoding mode with lossless compression of the audio feature data; and a third mode, the third mode being an encoding mode with lossy compression of the audio feature data.
[0037] In some embodiments of the first aspect. In some embodiments, the first mode is adopted when the proportion of the audio feature data to the audio input data is less than a first threshold; the second mode is adopted when the proportion of the audio feature data to the audio input data is greater than or equal to the first threshold and less than a second threshold; or the third mode is adopted when the proportion of the audio feature data to the audio input data is greater than or equal to the second threshold.
[0038] In some embodiments of the first aspect. In some embodiments, the encoding mode comprises: an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data.
[0039] In some embodiments of the first aspect. In some embodiments, the second mode comprises: an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data.
[0040] The third mode includes: an encoding mode suitable for single-channel audio data; and / or an encoding mode suitable for multi-channel audio data.
[0041] In a second aspect, embodiments of the present disclosure provide an encoding method, performed by a second communication device, the method comprising: encoding audio input data to obtain first encoding data; wherein the first encoding data comprises feature information of the audio input data related to a first task; and transmitting the first encoding data.
[0042] In some embodiments of the second aspect. In some embodiments, the audio input data comprises: audio data and label data; wherein the label data comprises description information of the audio data.
[0043] In some embodiments of the second aspect. In some embodiments, the first encoding data is audio feature data obtained by performing feature extraction on the audio input data.
[0044] In some embodiments of the second aspect. In some embodiments, the first encoding data is a bitstream obtained by encoding the audio feature data.
[0045] In some embodiments of the second aspect. In some embodiments, the method further comprises: transmitting encoding information, the encoding information being used to indicate an encoding mode of encoding the audio feature data.
[0046] In some embodiments of the second aspect. In some embodiments, the encoding mode is determined by at least one of the following information: a proportion of the audio feature data to the audio input data; a number of audio channels of the audio input data; a correlation between the audio channels of the audio input data; configuration information of the first task.
[0047] In some embodiments of the second aspect. In some embodiments, the configuration information of the first task comprises: encoding information; and related information of the audio feature data required by the first task.
[0048] In some embodiments of the second aspect. In some embodiments, the encoding mode can comprise at least one of the following: a first mode, the first mode being an encoding mode without compression of the audio feature data; a second mode, the second mode being an encoding mode with lossless compression of the audio feature data; and a third mode, the third mode being an encoding mode with lossy compression of the audio feature data.
[0049] Some embodiments combine the second aspect. In some embodiments, the first mode is adopted when the proportion of the audio feature data to the audio input data is less than a first threshold; the second mode is adopted when the proportion of the audio feature data to the audio input data is greater than or equal to the first threshold and less than a second threshold; or the third mode is adopted when the proportion of the audio feature data to the audio input data is greater than or equal to the second threshold.
[0050] Some embodiments combine the second aspect. In some embodiments, the encoding mode includes: an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data.
[0051] In some embodiments, the second mode includes: an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data;
[0052] In some embodiments, the third mode includes: an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data.
[0053] In a third aspect, a decoding apparatus is provided, the apparatus comprising: a transceiver module configured to receive first encoded data, the first encoded data comprising feature information of audio input data related to a first task; and a processing module configured to decode the first encoded data to obtain first decoded data, wherein the first decoded data is used to perform the first task based on machine hearing.
[0054] In a fourth aspect, an encoding apparatus is provided, the apparatus comprising: a processing module configured to encode audio input data to obtain first encoded data; wherein the first encoded data comprises feature information of the audio input data related to a first task; and a transceiver module configured to transmit the first encoded data.
[0055] In a fifth aspect, the embodiments of the present disclosure provide a communication device, which comprises: one or more processors; and a memory coupled to the processors, the memory having stored thereon executable instructions that, when executed by the processors, cause the processors to invoke the executable instructions to cause the communication device to perform the decoding and encoding methods as described in the first aspect and the second aspect, and the optional embodiments of the first aspect and the second aspect.
[0056] In a sixth aspect, the embodiments of the present disclosure provide a communication system, which comprises: a first communication device and a second communication device; wherein the first communication device is configured to perform the methods as described in the first aspect and the optional embodiments of the first aspect, and the second communication device is configured to perform the methods as described in the second aspect and the optional embodiments of the second aspect.
[0057] In a seventh aspect, the embodiments of the present disclosure provide a storage medium, which stores instructions. When the instructions are executed on a communication device, the communication device performs the method described in the first aspect and the second aspect, and the optional embodiments of the first aspect and the second aspect.
[0058] In an eighth aspect, the embodiments of the present disclosure provide a program product, which, when executed by a communication device, causes the communication device to perform the method described in the first aspect and the second aspect, and the optional embodiments of the first aspect and the second aspect.
[0059] In a ninth aspect, the embodiments of the present disclosure provide a computer program, which, when executed on a computer, causes the computer to perform the method described in the first aspect and the second aspect, and the optional embodiments of the first aspect and the second aspect.
[0060] It can be understood that the first communication device, the second communication device, the communication device, the communication system, the storage medium, the program product, and the computer program are all used to perform the method proposed in the embodiments of the present disclosure. Therefore, the beneficial effects they can achieve can refer to the beneficial effects in the corresponding method, which will not be described here.
[0061] The embodiments of the present disclosure propose a decoding and encoding method, a communication device, and a storage medium. In some embodiments, the terms of information sending method, information receiving method, information processing method, and communication method can be replaced with each other, the terms of terminal, network device, information processing device, and communication device can be replaced with each other, and the terms of information processing system and communication system can be replaced with each other.
[0062] The embodiments of the present disclosure are not exhaustive, but only illustrate some embodiments, and are not specific limitations on the protection scope of the present disclosure. In the case of no contradiction, each step in an embodiment can be implemented as an independent embodiment, and the steps can be combined arbitrarily, for example, the scheme after removing some steps in an embodiment can also be implemented as an independent embodiment, and the order of the steps in an embodiment can be exchanged arbitrarily, in addition, the optional embodiments in an embodiment can be combined arbitrarily; in addition, the embodiments can be combined arbitrarily, for example, some or all steps of different embodiments can be combined arbitrarily, and an embodiment can be combined with the optional embodiments of other embodiments.
[0063] In each embodiment of the present disclosure, the terms and / or descriptions between the embodiments are consistent if there is no special description and logical conflict, and can be referred to each other, and the technical features in different embodiments can be combined to form a new embodiment according to their inherent logical relationship.
[0064] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments, and not as a limitation on the present disclosure.
[0065] In the embodiments of the present disclosure, an element expressed in singular form, such as "a", "an", "the", "said", "the aforementioned", "the foregoing", "this", and the like, can represent "one and only one", or "one or more", "at least one", and the like, unless otherwise specified.
[0066] For example, in the case of using an article such as "a", "an", "the" in English in translation, the noun after the article can be understood as a singular expression, or can be understood as a plural expression.
[0067] In the embodiments of the present disclosure, "plurality" means two or more.
[0068] In some embodiments, the terms "at least one of", "one or more", "a plurality of", "multiple", and the like can be replaced with each other.
[0069] In some embodiments, the description manner such as "at least one of A, B", "A and / or B", "A in one case, B in another case", "in response to a case A, in response to another case B", and the like can include the following technical solutions according to the case: A in some embodiments (A is executed regardless of B); B in some embodiments (B is executed regardless of A); A and B are selectively executed in some embodiments (A and B are selected to be executed); A and B are executed in some embodiments (A and B are both executed). When there are more branches such as A, B, C, and the like, it is similar to the above.
[0070] In some embodiments, the description manner such as "A or B" and the like can include the following technical solutions according to the case: A in some embodiments (A is executed regardless of B); B in some embodiments (B is executed regardless of A); A and B are selectively executed in some embodiments (A and B are selected to be executed). When there are more branches such as A, B, C, and the like, it is similar to the above.
[0071] The prefix words "first", "second", and the like in the embodiments of the present disclosure are only used to distinguish different description objects, and do not constitute a limitation on the position, order, priority, quantity, or content of the description objects. The description of the description objects should be referred to the description in the context of the claims or embodiments, and should not constitute an unnecessary limitation because of the use of the prefix words.
[0072] For example, the ordinal numbers before the description object "field" in "the first field" and "the second field" do not limit the positions or orders between the "fields", and "the first" and "the second" do not limit whether the "fields" they modify are in the same message or not, nor the order of "the first field" and "the second field". For another example, the ordinal numbers before the description object "level" in "the first level" and "the second level" do not limit the priorities between the "levels". For another example, the quantity of the description object is not limited by the ordinal numbers, which can be one or more. For example, "the first device", the quantity of the "device" can be one or more. In addition, the description objects modified by different prefixes can be the same or different, for example, the description object is "device", "the first device" and "the second device" can be the same device or different devices, and their types can be the same or different; for another example, the description object is "information", "the first information" and "the second information" can be the same information or different information, and their contents can be the same or different.
[0073] In some embodiments, "comprising A", "including A", "for indicating A", "carrying A" can be interpreted as directly carrying A, or indirectly indicating A.
[0074] In some embodiments, the terms "in response to", "in response to determining", "in the case of", "when", "when", "if", "if" and the like can be replaced with each other.
[0075] In some embodiments, the terms "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not lower than", "above" and the like can be replaced with each other, and the terms "less than", "less than or equal to", "not greater than", "less than", "less than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", "below" and the like can be replaced with each other.
[0076] In some embodiments, the device and the like can be interpreted as physical or virtual, and its name is not limited to the name described in the embodiments, and the terms "device", "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "subject" and the like can be replaced with each other.
[0077] In some embodiments, "network" can be interpreted as a device (for example, access network device, core network device, etc.) contained in the network.
[0078] In some embodiments, "network" can be interpreted as a device (for example, access network device, core network device, etc.) contained in the network.
[0079] In some embodiments, the terms “access network device (AN device),” “radio access network device (RAN device),” “base station (BS),” “radio base station,” “fixed station,” “node,” “access point,” “transmission point (TP),” “reception point (RP),” “transmission / reception point (TRP),” “panel,” “antenna panel,” “antenna array,” “cell,” “macro cell,” “small cell,” “femto cell,” “pico cell,” “sector,” “cell group,” “serving cell,” “carrier,” “component carrier,” “bandwidth part (BWP),” and the like can be used interchangeably.
[0080] In some embodiments, the terms "terminal," "terminal device," "user equipment (UE)," "user terminal," "mobile station (MS)," "mobile terminal (MT)," "subscriber station," "mobile unit," "subscriber unit," "wireless unit," "remote unit," "mobile device," "wireless device," "wireless communication device," "remote device," "mobile subscriber station," "access terminal," "mobile terminal," "wireless terminal," "remote terminal," "handset," "user agent," "mobile client," "client," and so on can be replaced with each other.
[0081] In some embodiments, the access network device, the core network device, or the network device can be replaced with a terminal. For example, the embodiments of the present disclosure can also be applied to a structure in which communication between the access network device, the core network device, or the network device and the terminal is replaced with communication between a plurality of terminals (e.g., device-to-device (D2D), vehicle-to-everything (V2X), etc.). In this case, the terminal can also be configured to have all or part of the functions of the access network device. In addition, the terms "uplink," "downlink," and the like can also be replaced with terms corresponding to the inter-terminal communication (e.g., "side"). For example, the uplink channel, the downlink channel, and the like can be replaced with the side channel, and the uplink, the downlink, and the like can be replaced with the sidelink.
[0082] In some embodiments, the terminal can be replaced with the access network device, the core network device, or the network device. In this case, the access network device, the core network device, or the network device can also be configured to have all or part of the functions of the terminal.
[0083] In some embodiments, obtaining data, information, etc. can comply with laws and regulations of the country where the location is.
[0084] In some embodiments, data, information, etc. can be obtained after obtaining the consent of the user.
[0085] In addition, each element, each row, or each column in the table of the embodiments of the present disclosure can be implemented as an independent embodiment, and any combination of any element, any row, or any column can also be implemented as an independent embodiment.
[0086] FIG. 1 is an architecture schematic diagram of a communication system according to an embodiment of the present disclosure.
[0087] As shown in FIG. 1, the communication system 100 includes a first communication device 101 and a second communication device 102; wherein the first communication device 101 can be a terminal or a network device, and the second communication device 102 can also be a terminal or a network device, wherein the network device includes at least one of the following: an access network device, a core network device.
[0088] In some embodiments, the terminal includes at least one of the following, but is not limited thereto: a mobile phone, a wearable device, an Internet of Things device, a communication-capable automobile, a smart automobile, a Pad, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal device in industrial control, a wireless terminal device in self-driving, a wireless terminal device in remote medical surgery, a wireless terminal device in smart grid, a wireless terminal device in transportation safety, a wireless terminal device in smart city, a wireless terminal device in smart home, etc.
[0089] In some embodiments, the access network device is, for example, a node or device that accesses a terminal to a wireless network, and the access network device can include at least one of an evolved NodeB (eNB) in a 5G communication system, a next generation eNB (ng-eNB), a next generation NodeB (gNB), a node B (NB), a home node B (HNB), a home evolved node B (HeNB), a wireless backhaul device, a radio network controller (RNC), a base station controller (BSC), a base transceiver station (BTS), a base band unit (BBU), a mobile switching center, a base station in a 6G communication system, an Open RAN, a Cloud RAN, a base station in other communication systems, an access node in a Wi-Fi system, but is not limited thereto.
[0090] In some embodiments, the core network device can be one device including one or more network elements, or can be multiple devices or device groups including all or part of the one or more network elements described above. The network element can be virtual or physical. The core network includes, for example, at least one of an evolved packet core (EPC), a 5G core network (5GCN), and a next generation core (NGC).
[0091] In some embodiments, the technical solutions of the present disclosure can be applied to an Open RAN architecture, at which time the interfaces between or within the access network devices involved in the embodiments of the present disclosure can become internal interfaces of the Open RAN, and the processes and information interactions between these internal interfaces can be realized through software or programs.
[0092] In some embodiments, the access network device can be composed of a central unit (CU) and a distributed unit (DU), where the CU can also be referred to as a control unit. The CU-DU structure can split the protocol layers of the access network device, and some of the protocol layers are controlled by the CU, and the remaining or all of the protocol layers are distributed in the DU and controlled by the CU, but is not limited thereto.
[0093] It can be understood that the communication system described in the embodiments of the present disclosure is for more clearly illustrating the technical solutions of the embodiments of the present disclosure, and does not constitute a limitation on the technical solutions proposed by the embodiments of the present disclosure. Those skilled in the art can know that, with the evolution of system architecture and the appearance of new business scenarios, the technical solutions proposed by the embodiments of the present disclosure are also applicable to similar technical problems.
[0094] The following embodiments of the present disclosure can be applied to the communication system 100 shown in FIG. 1 or part of the subjects, but are not limited thereto. The subjects shown in FIG. 1 are exemplary, and the communication system can include all or part of the subjects in FIG. 1, or other subjects other than FIG. 1. The number and form of each subject is arbitrary, each subject can be physical or virtual, the connection relationship between each subject is exemplary, each subject can not be connected or can be connected, the connection can be in any way, can be direct connection or indirect connection, can be wired connection or wireless connection.
[0095] Embodiments of the present disclosure can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G new radio (NR), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New radio access (NX), Future generation radio access (FX), Global System for Mobile communications (GSM (registered trademark)), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, Ultra-WideBand (UWB), Bluetooth (Bluetooth (registered trademark)), Public Land Mobile Network (PLMN) network, Device-to-Device (D2D) system, Machine to Machine (M2M) system, Internet of Things (IoT) system, Vehicle-to-Everything (V2X), system using other communication methods, next-generation system expanded based thereon, and the like. Further, a plurality of systems can be applied in combination (for example, combination of LTE or LTE-A and 5G, and the like).
[0096] Related audio coding standards can include Motion Picture Experts Group (MPEG) standards, Advanced Audio Coding (AAC) standards, Free Lossless Audio Codec (FLAC) standards, and the like. However, none of the related standards can meet the needs of machine listening for different tasks.
[0097] For example, the fourth generation Motion Picture Experts Group Scalable Lossless Coding (MPEG-4 SLS) standard defines a method for lossless or near-lossless encoding of a small number of channels. MPEG-4 SLS cannot represent necessary metadata and does not provide tools for exploiting spatial redundancy between multiple channels.
[0098] For another example, the third generation Motion Picture Experts Group High-efficiency 3D Audio (MPEG-H 3D Audio) standard specifies a method for encoding spatial audio into separate channels, objects, and Higher Order Ambisonics (HOA) formats. MPEG-H 3D Audio cannot represent necessary metadata, and since MPEG-H 3D Audio is based on perceptual audio coding, it is neither lossless nor near-lossless.
[0099] Therefore, a new coding method for machine listening is needed to efficiently compress audio data for AI tasks.
[0100] FIG. 2A is an interaction diagram illustrating a decoding and encoding method according to an embodiment of the present disclosure.
[0101] As shown in FIG. 2A, the decoding and encoding method includes:
[0102] In step S201, the second communication device encodes the audio input data to obtain first encoded data.
[0103] In some embodiments, the second communication device can serve as an encoding end or an encoding end device of the audio data. The second communication device can obtain audio input data for the first task and encode the audio input data to obtain first encoded data. Through the encoding of the audio input data, the first encoded data can be used to represent audio features in the audio input data, where the audio features in the audio input data can include feature information related to the first task, or can also be considered as feature information required by the first task.
[0104] The first task can be an AI task that requires machine hearing, such as a voice assistant, autonomous driving, medical monitoring, a security system, etc.
[0105] In some embodiments, the audio input data can include audio data and label data. The audio data can be data in different audio formats, such as various uncompressed audio data, such as pulse code modulation (PCM) audio data, waveform audio file format (WAV) audio data, direct stream digital (DSD) audio data, etc. The label data corresponding to the audio data includes description information of the audio data, which can also be referred to as metadata of the audio data, such as collection information of the audio data, task information of the first task to which the audio data is directed, etc.
[0106] Correspondingly, the encoding of the audio input data by the second communication device can include encoding the audio data and the label data contained in the audio input data respectively, or jointly encoding the audio data and the label data, or encoding the audio data based on the label data.
[0107] In some embodiments, the audio input data can include data used in the process of training, verifying, and testing machine hearing.
[0108] In step S202, the second communication device sends the first encoded data to the first communication device.
[0109] In some embodiments, after the second communication device encodes the audio input data to obtain the first encoded data, the second communication device can send the first encoded data to the first communication device. The first encoded data can be a bitstream, which can be referred to as an encoded bitstream. The first communication device can serve as a decoding end device of the audio data.
[0110] In some embodiments, the second communication device can send the first encoded data to the first communication device through a transmission module.
[0111] The transmission module can employ some channel transmission techniques to ensure the secure and complete transmission of the bit stream. For example, wireless transmission techniques or wired transmission techniques can be employed, etc.
[0112] The input and output of the transmission module can both be bit streams.
[0113] Step S203: The first communication device decodes the first encoded data.
[0114] In some embodiments, the first communication device can serve as a decoding end or decoding end device. The first communication device can receive the first encoded data from the second communication device, decode the first encoded data, and obtain first decoded data.
[0115] The decoding manner for decoding the first encoded data can be determined based on the encoding manner of the second communication device for encoding the audio input data.
[0116] In some embodiments, the first decoded data can be equivalent to the audio input data. In the case where the audio input data includes audio data for the first task and label data, the first decoded data obtained by the first communication device after decoding can also include audio data for the first task and label data.
[0117] In some embodiments, during the encoding of the audio input data by the second communication device, only the feature information of the audio input data related to the first task can be retained. In this case, the first decoded data can also only include the audio data of the audio input data for representing the feature information related to the first task, i.e., the audio data for representing the feature information required by the first task.
[0118] In some embodiments, the first decoded data obtained by the first communication device by decoding the first encoded data can be used to perform the first task based on machine hearing. The first decoded data can be input into a machine hearing module for the first task, and the machine hearing module can analyze the first decoded data, such as training, verification, testing, etc., based on the first decoded data, and output the analysis result required by the first task.
[0119] For example, if the first task is to monitor the user's heartbeat, the audio input data obtained by the second communication device can include the user's body monitoring data. The second communication device encodes the audio input data to obtain first encoded data, which can be used to represent the feature information of the audio data related to the heartbeat in the audio input data. The first encoded data is sent to the first communication device through the transmission module; the first communication device decodes the first encoded data to obtain first decoded data, which can include the audio data related to the heartbeat in the audio input data. The first communication device inputs the first decoded data into the machine listening module for the first task, and the machine listening module analyzes the audio data related to the heartbeat to obtain the monitoring result of the user's heartbeat.
[0120] FIG. 2B is a whole block diagram of a decoding and encoding method according to an embodiment of the present disclosure. As shown in FIG. 2B, the whole block diagram includes the following links, and the purpose of each link is described as follows:
[0121] Audio input data (audio input) includes data used for machine listening training, verification, and measurement. The data can include PCM data and labels data. The labels data is used to describe the pcm data, including acquisition information, description information, etc.
[0122] Encoding end 211 (second communication device): also referred to as machine listening encoding, the input of the module is audio input, and the module can be used to encode the pcm data and the labels data respectively or jointly encode to form a bitstream. The output of the module is a bitstream (first encoded data).
[0123] Transmission module 212: also referred to as transmission of encoded bitstream, the input and output of the module are both bitstream, and the module adopts some channel transmission technology to ensure the safe and complete transmission of the bitstream.
[0124] Decoding end 213 (first communication device): also referred to as machine listening decoding, the input of the module is a bitstream, and the module decodes the bitstream to obtain decoded data (first decoded data). The decoded data can be lossy.
[0125] Machine hearing 214: The input of this module is the decoded data, which can include pcm audio data and labels data. This module analyzes the decoded data and outputs one or more of the following: analysis results, generated data.
[0126] In the above embodiments, by encoding the audio input data at the encoding end, the first encoding information used to represent the feature information related to the first task in the audio input data is obtained, and is sent to the decoding end; the first encoding information is decoded by the decoding end to obtain the first decoding data related to the first task, and is input to the machine hearing module for the first task to analyze the first decoding data, so that the data obtained after encoding and decoding the audio input data can meet the needs of the first task, so as to analyze the analysis results required by the first task.
[0127] In some embodiments, the second communication device can perform feature extraction on the obtained audio input data to obtain audio feature data in step S201.
[0128] In some embodiments, the second communication device can perform feature extraction on the audio input data based on the needs of the first task to obtain audio feature data.
[0129] In some embodiments, the second communication device can include a feature extraction module, and the feature extraction module can perform feature extraction on the audio input data based on the needs of the first task to obtain audio feature data.
[0130] Among them, the audio feature data can include audio data related to the first task.
[0131] Among them, the requirements of the first task can include requirements for any one or more of the audio data, the audio feature data, and the feature extraction method; wherein the requirements of the audio data can include requirements for content, type, format, etc., and the requirements of the audio feature data can include requirements for type, format, etc., which are not limited here.
[0132] For example, for audio tasks such as speech recognition, audio classification, user recognition, emotion recognition, and audio fingerprinting, mel frequency cepstral coefficients (MFCC) can be required to be extracted as audio feature data.
[0133] In some embodiments, after the second communication device extracts the audio feature data from the audio input data, the second communication device can send the audio feature data as the first encoded data to the first communication device through the transmission module; the decoding operation performed by the first communication device on the first encoded data can be not decoding the audio feature data, but directly inputting the audio feature data into the machine auditory module for the first task for analysis to obtain the analysis result required by the first task.
[0134] FIG. 2C is a whole block diagram of a decoding and encoding method according to an embodiment of the present disclosure. As shown in FIG. 2C, at the encoding end 211, the feature extraction module 2111 can perform feature extraction on the audio input data, and the extracted audio feature data is directly transmitted through the transmission module 212. The decoding end 213 does not need any operation, and directly inputs the received audio feature data into the machine auditory module for analysis.
[0135] The feature extraction module 2111 is configured to compress the audio input data and only retain the audio feature data required by the machine auditory module, which is equivalent to retaining the audio feature data required by the first task, thereby playing a role in storage space and transmission bandwidth.
[0136] In the above embodiment, by performing feature extraction on the audio input data at the encoding end and sending the extracted audio feature data to the decoding end, the storage space and transmission bandwidth can be reduced.
[0137] In some embodiments, in order to further improve the compression rate of the audio data and further eliminate the temporal redundancy and spatial redundancy of the audio feature data, in step S201, after the second communication device extracts the audio feature data from the audio input data through feature extraction, the second communication device can further encode the audio feature data to further compress the audio feature data to obtain the first encoded data, and send the first encoded data to the first communication device.
[0138] The first communication device receives the first encoded data through the transmission module, decodes the first encoded data to obtain the first decoded data, and inputs the first decoded data into the machine auditory module for the first task, and the machine auditory module analyzes the first decoded data to output the analysis result required by the first task.
[0139] In some embodiments, the second communication device can include a feature extraction module and a feature encoding module, the feature extraction module extracts the audio feature data from the audio input data, and the feature encoding module encodes the audio feature data to obtain the first encoded data, and sends the first encoded data to the first communication device.
[0140] The first communication device can comprise a feature decoding module, the received first encoded data is decoded by the feature decoding module to obtain first decoded data, and the first decoded data is input to the machine hearing module for the first task.
[0141] The obtained first decoded data can comprise audio feature data.
[0142] FIG. 2D is a whole block diagram of a decoding and encoding method according to an embodiment of the present disclosure. As shown in FIG. 2D, at the encoding end 211, the audio input data is feature-extracted by the feature extraction module 2111, the extracted audio feature data is further compressed by the designated feature encoding module 2112, and the compressed audio feature data is transmitted by the transmission module 212. The decoding end 213 can be decoded by the feature decoding module 2131 according to the designated feature decoding method, and the decoded audio feature data is input to the machine hearing module 214 for use.
[0143] The advantage of the present embodiment is that the feature encoding module and the feature decoding module are added, and the temporal redundancy and spatial redundancy of the audio feature data are further eliminated. The disadvantage of the present embodiment is that the encoding end and the decoding end can only use a specific feature compression method and decompression method, and the flexibility is not high enough. In some scenarios, the compression rate requirement is not so high, and the compression rate can be reduced to improve the quality of the audio data.
[0144] In the above embodiment, the audio feature data is encoded at the encoding end, and the received first encoded data is decoded at the decoding end, so that the temporal redundancy and spatial redundancy of the audio feature data are eliminated to a certain extent.
[0145] In some embodiments, the second communication device can flexibly select a corresponding encoding mode as needed when performing the encoding operation, and correspondingly, the first communication device can flexibly select a corresponding decoding mode as needed when performing the decoding operation.
[0146] In some embodiments, the second communication device can comprise a feature extraction module, a mode selection module, and a feature encoding module corresponding to each encoding mode, the feature extraction module extracts audio feature data from the audio input data, the mode selection module selects a currently required encoding mode from a plurality of available encoding modes, and the audio feature data is sent to the feature encoding module corresponding to the selected encoding mode for encoding to obtain the first encoded data.
[0147] The factors for determining the current required encoding mode can be various, and can include information related to at least one of the audio input data, the audio feature data, and the first task, such as a proportion of the audio feature data to the audio input data, a number of audio channels of the audio input data, a correlation between the audio channels of the audio input data, configuration information of the first task, and the like.
[0148] The configuration information of the first task can include encoding information, and / or related information of the audio feature data required by the first task.
[0149] The encoding information can be used to indicate an encoding mode suitable for the first task, or can also be used to indicate a requirement of the first task on the used encoding mode, such as a requirement on one or more of a proportion of the audio feature data to the audio input data, a number of audio channels of the audio input data, a correlation between the audio channels of the audio input data.
[0150] In some embodiments, the encoding mode that can be used by the second communication device can include at least one of: a first mode, which is an encoding mode without compression of the audio feature data; a second mode, which is an encoding mode with lossless compression of the audio feature data, such as MPEG-SLS; and a third mode, which is an encoding mode with lossy compression of the audio feature data, such as MPEG-4 ACC, MPEG-H 3D, and the like.
[0151] In some embodiments, the second communication device can extract the audio feature data from the audio input data by the feature extraction module, and determine the encoding mode required by the first task based on a proportion of the audio feature data to the audio input data by the mode selection module: if the proportion is less than a first threshold, the first mode is adopted, i.e., without compression of the audio feature data; if the proportion is greater than or equal to the first threshold and less than a second threshold, the second mode is adopted, i.e., the lossless compression encoding mode such as MPEG-SLS is adopted for the audio feature data; if the proportion is greater than or equal to the second threshold, the third mode is adopted, i.e., the lossy compression encoding mode such as MPEG-4 ACC is adopted for the audio feature data. The audio feature data is input into the feature encoding module corresponding to the determined encoding mode for encoding to obtain the first encoded data. The first threshold and the second threshold can be set according to actual needs, for example, can be set based on the first task, or indicated by the configuration information of the first task.
[0152] In some embodiments, for the second mode and the third mode, further division can be made according to actual needs, and the basis for the division can be derived from the number of audio channels of the audio input data, and / or the correlation between the audio channels of the audio input data.
[0153] The second mode can include an encoding mode suitable for single-channel audio data and / or an encoding mode suitable for multi-channel audio data.
[0154] The third mode can include an encoding mode suitable for single-channel audio data, such as MPEG-HE ACC V2, and / or an encoding mode suitable for multi-channel audio data, such as MPEG-MCT.
[0155] For example, the mode selection module can determine the encoding mode required by the first task based on the proportion of the audio feature data to the audio input data, the number of audio channels of the audio input data, and the correlation between the audio channels of the audio input data; in the case where the third mode that can be used includes an encoding mode suitable for single-channel audio data and an encoding mode suitable for multi-channel audio data, if the proportion is greater than or equal to a second threshold, the third mode can be determined to be used; and the encoding mode required can be further determined from the plurality of third modes based on the number of audio channels of the audio input data and the correlation between the audio channels of the audio input data: if the audio input data has a plurality of inter-frequency channels with high inter-channel correlation, an encoding mode suitable for multi-channel audio data, such as MPEG-MCT, can be determined to be used; if the audio input data has audio channels with low inter-channel correlation, an encoding mode suitable for single-channel audio data, such as MPEG-HE ACC V2, can be determined to be used to compress each channel separately.
[0156] In some embodiments, in step S202, the second communication device can send the first communication device encoding information, the encoding information being used to indicate the encoding mode of the audio feature data.
[0157] The encoding information sent by the second communication device to the first communication device can include a mode indication bit (mode_flag), and the value of the mode indication bit is used to indicate the encoding mode.
[0158] In some embodiments, in step S203, the decoding mode of the first communication device for decoding the received first encoded data corresponds to the encoding mode used by the second communication device.
[0159] In some embodiments, the decoding mode can be determined based on the same factors used to determine the encoding mode, which can include information related to at least one of the audio input data, the audio feature data, and the first task, such as the proportion of the audio feature data to the audio input data, the number of audio channels of the audio input data, the correlation between the audio channels of the audio input data, configuration information of the first task, and the like.
[0160] In some embodiments, the decoding mode can include a decoding mode corresponding to at least one of the first mode, the second mode, and the third mode.
[0161] In some embodiments, the decoding mode corresponding to the second mode can include a decoding mode suitable for single-channel audio data; and / or, a decoding mode suitable for multi-channel audio data.
[0162] In some embodiments, the decoding mode corresponding to the third mode can include a decoding mode suitable for single-channel audio data; and / or, a decoding mode suitable for multi-channel audio data.
[0163] In some embodiments, the first communication device, when receiving the first encoded data from the second communication device, can also receive encoding information, which can be used to indicate an encoding mode for encoding the audio feature data; and based on a decoding mode corresponding to the encoding mode, decode the first encoded data to obtain first decoded data.
[0164] The encoding information received by the first communication device can be derived from the encoding information sent by the second device; or can also be derived from the configuration information of the first task.
[0165] FIG. 2E is a whole block diagram of a decoding and encoding method according to an embodiment of the present disclosure. As shown in FIG. 2E, at the encoding end 211:
[0166] The feature extraction module 2111 can be used to extract audio features from the audio input data;
[0167] The adaptive mode selection module 2113 can take the audio feature data output by the feature extraction module 2111 as input, i.e., the audio features extracted from the audio input data. The module can calculate the proportion of the audio feature data relative to the audio input data, and when the proportion is less than a first threshold, the audio feature data is not compressed; when the proportion is between the first threshold and a second threshold, the audio feature data can be losslessly compressed; and when the proportion is greater than the second threshold, the audio feature data can be lossily compressed. The module can also adaptively select a lossy compression method suitable for the audio feature data by analyzing the characteristics of the audio feature data, such as using the MPEG-MCT algorithm for compression when the input is multiple audio channels with high inter-channel correlation; or using MPEG-HEAACV2 to compress each channel separately when the input is audio channels with low inter-channel correlation. The encoding mode that can be selected can include:
[0168] Mode 0: no compression on audio feature data;
[0169] Mode 1: lossless compression on audio feature data, such as MPEG-SLS;
[0170] Mode 2: audio feature uses lossy compression method 1, such as MPEG-HEAACV2;
[0171] Mode 3: audio feature uses lossy compression method 2, such as MPEG-MCT.
[0172] Different compression methods occupy different storage space or transmission bandwidth. Mode 0 occupies the highest storage space or transmission bandwidth. Mode 2 and mode 3 occupy lower storage space or transmission bandwidth than Mode 0 and Mode 1.
[0173] The mode selection module 2113 can also generate encoding information (mode_flag), which can be transmitted in the bit stream.
[0174] At the decoding end 213, the input of the feature decoding module 2131 includes the encoding information mode_flag and the compressed audio feature data. The module can select a decoding mode for decoding the compressed audio feature data according to the mode_flag, decode the compressed audio feature data according to the decoding mode corresponding to the mode_flag, and obtain the decoded audio feature data.
[0175] The machine hearing module 214 can use the decoded audio feature data for training, verification, and testing, and output the analysis results of the audio feature data or generated new data.
[0176] In some embodiments, the encoding mode for encoding the audio feature data can be determined based on the configuration information (user config) of the first task, and the encoding information can be included in the configuration information of the first task, which is used to indicate the encoding mode suitable for the first task. The mode selection module in the second communication device can determine the corresponding encoding mode to encode the audio feature data based on the configuration information of the first task, and obtain the first encoding data.
[0177] In some embodiments, if it is determined based on the configuration information of the first task that multiple encoding modes are applicable to the first task, the current required encoding mode for encoding the audio feature data can be selected from the multiple encoding modes based on at least one of the proportion of the audio feature data to the audio input data, the number of audio channels of the audio input data, and the correlation between the audio channels of the audio input data, to obtain the first encoded data.
[0178] In some embodiments, the decoding mode for decoding the audio feature data can also be determined based on the configuration information of the first task, and the configuration information of the first task can include encoding information indicating the decoding mode applicable to the first task. The first communication device can determine the corresponding decoding mode based on the configuration information of the first task to decode the first encoded data to obtain the first decoded data.
[0179] FIG. 2F is a general block diagram illustrating a decoding and encoding method according to an embodiment of the present disclosure. The difference between the implementation of FIG. 2F and the implementation of FIG. 2E is the source of mode selection. In FIG. 2F, an external mode selection module 2114 is used to determine the compression method of the audio feature data, i.e., the compression method of the audio feature data is determined by the configuration information (user config) received by the external mode selection module 2114, i.e., the compression method of the audio feature data can be determined by the user according to the use scenario.
[0180] In the above embodiments, the encoding end and the decoding end can flexibly select the encoding mode and the decoding mode according to actual needs.
[0181] The audio codec method of the embodiments of the present disclosure has the following differences compared with the decoding method defined in the related audio codec standard:
[0182] 1. Different receiving objects: the receiving object of the traditional technology is the human ear. The receiving object of the present application is a machine auditory module for specifying an AI task.
[0183] 2. Different encoding modes: the encoding mode defined in the related standard adopts different encoding modes according to the perception characteristics of the human ear, such as speech core, music core, and noise core. The encoding mode of the present application is to select a mode suitable for feature encoding according to the features required by the machine auditory module.
[0184] 3. Different compression methods are adopted: the compression method of the coding mode defined by the relevant standard is to compress the time domain expression or frequency domain expression of the audio signal. The compression method of the present application is to compress the audio features. The audio features can be extracted by AI methods or traditional methods for AI tasks. The audio features and the time domain expression or frequency domain expression of the traditional signal are different in data format and data content, and the compression methods adopted are also different.
[0185] The communication method related to the embodiments of the present disclosure can include at least one of steps S201 to S203. For example, step S201 can be implemented as an independent embodiment, step S202 can be implemented as an independent embodiment, step S203 can be implemented as an independent embodiment, any combination of steps S201 to S203 can be implemented as an independent embodiment, but is not limited thereto.
[0186] The communication method related to the embodiments of the present disclosure can include at least one of steps S201 to S203. For example, step S201 can be implemented as an independent embodiment, step S202 can be implemented as an independent embodiment, step S203 can be implemented as an independent embodiment, any combination of steps S201 to S203 can be implemented as an independent embodiment, but is not limited thereto.
[0187] In some embodiments, step S201 is optional, and one or more of the steps can be omitted or replaced in different embodiments.
[0188] In some embodiments, step S202 is optional, and one or more of the steps can be omitted or replaced in different embodiments.
[0189] In some embodiments, step S203 is optional, and one or more of the steps can be omitted or replaced in different embodiments.
[0190] In some embodiments, other optional embodiments described before or after the description corresponding to FIG. 2A can be referred to.
[0191] Embodiments of the present disclosure propose a decoding method. FIG. 3 is a schematic flowchart of a decoding method according to an embodiment of the present disclosure. The decoding method shown in the present embodiment can be performed by a first communication device.
[0192] As shown in FIG. 3, the decoding method can include the following steps:
[0193] In step S301, first encoded data is received, and the first encoded data includes feature information related to a first task in audio input data.
[0194] In some embodiments, the first communication device can receive the first encoded data from the second communication device; wherein the first encoded data can be obtained by the second communication device encoding the obtained audio input data, and the first encoded data can be used to represent the audio features in the audio input data. The audio features in the audio input data can include feature information related to the first task, or can also be considered as feature information required by the first task.
[0195] The first encoded data can be a bitstream, which can be referred to as an encoded bitstream.
[0196] The first task can be an AI task that requires machine hearing.
[0197] In some embodiments, the audio input data can include audio data and label data. The audio data can be data in different audio formats, such as various uncompressed audio data, such as PCM audio data. The label data corresponding to the audio data includes description information of the audio data, for example, can include collection information of the audio data, task information of the first task to which the audio data is directed, and the like.
[0198] Correspondingly, the second communication device encoding the audio input data can include encoding the audio data and the label data contained in the audio input data respectively, or jointly encoding the audio data and the label data, or encoding the audio data based on the label data.
[0199] In step S302, the first encoded data is decoded to obtain first decoded data, wherein the first decoded data is used to perform the first task based on machine hearing.
[0200] In some embodiments, the first communication device can be a decoding end or decoding end device, which decodes the first encoded data to obtain the first decoded data.
[0201] The decoding manner of the first encoded data can be determined based on the encoding manner of the second communication device encoding the audio input data.
[0202] In some embodiments, the first decoded data can be equivalent to the audio input data, and in the case that the audio input data includes audio data and label data for the first task, the first decoded data obtained by the first communication device after decoding can also include audio data and label data for the first task.
[0203] In some embodiments, during the encoding of the audio input data by the second communication device, only feature information related to the first task in the audio input data can be reserved, and in this case, the first decoded data can also only include audio data in the audio input data for representing the feature information related to the first task, that is, audio data for representing feature information required by the first task.
[0204] In some embodiments, after the first communication device obtains the first decoded data by decoding the first encoded data, the first decoded data can be input into the machine auditory module for the first task, and the machine auditory module can analyze the first decoded data, for example, training, verification, testing, etc. based on the first decoded data, and output analysis results required by the first task.
[0205] It should be noted that the embodiment shown in FIG. 3 can be independently implemented, or can be combined with at least one other embodiment of the present disclosure for implementation. The specific implementation can be selected as needed, and the present disclosure does not limit it.
[0206] In some embodiments, by encoding the audio input data at the encoding end, first encoded information for representing feature information related to the first task in the audio input data is obtained and sent to the decoding end; the decoding end decodes the first encoded information to obtain first decoded data related to the first task, and inputs the first decoded data to the machine auditory module for the first task to analyze the first decoded data, so that the data obtained after encoding and decoding the audio input data can meet the requirements of the first task, and the analysis results required by the first task are obtained by analysis.
[0207] In some embodiments, the first encoded data is audio feature data obtained after feature extraction of the audio input data. The audio feature data can be obtained by the second communication device based on the requirements of the first task.
[0208] The audio feature data can include audio data related to the first task.
[0209] The requirements of the first task can include requirements for any one or more of audio data, audio feature data, and feature extraction methods; the requirements for audio data can include requirements for content, type, format, etc., and the requirements for audio feature data can include requirements for type, format, etc., which are not limited here.
[0210] In some embodiments, the first encoded data received by the first communication device can be audio feature data, and the first communication device can not decode the audio feature data, but directly input the audio feature data into the machine auditory module for the first task for analysis to obtain the analysis results required by the first task.
[0211] In some embodiments, the first encoded data is a bitstream obtained by encoding the audio feature data. The audio feature data can be encoded by the second communication device after the audio feature data is extracted from the audio input data by feature extraction, so as to further compress the audio feature data.
[0212] The first communication device can receive the first encoded data by the transmission module, decode the first encoded data to obtain the first decoded data, input the first decoded data into the machine auditory module for the first task, and analyze the first decoded data by the machine auditory module to output the analysis result required by the first task.
[0213] In some embodiments, the first communication device can include a feature decoding module, which decodes the received first encoded data to obtain the first decoded data and inputs the first decoded data into the machine auditory module for the first task.
[0214] The obtained first decoded data can include the audio feature data.
[0215] In some embodiments, the second communication device can flexibly select a corresponding encoding mode according to needs when performing the encoding operation. Correspondingly, the first communication device can also flexibly select a corresponding decoding mode according to needs when performing the decoding operation.
[0216] In some embodiments, the encoding mode is determined by at least one of the following information: a proportion of the audio feature data to the audio input data; a number of audio channels of the audio input data; a correlation between the audio channels of the audio input data; and configuration information of the first task.
[0217] In some embodiments, the configuration information of the first task includes: encoding information; and related information of the audio feature data required by the first task.
[0218] The encoding information can be used to indicate the encoding mode suitable for the first task, or can also be used to indicate the requirement of the first task on the used encoding mode.
[0219] In some embodiments, the encoding mode can include at least one of the following: a first mode, which is an encoding mode without compressing the audio feature data; a second mode, which is an encoding mode for lossless compression of the audio feature data; and a third mode, which is an encoding mode for lossy compression of the audio feature data.
[0220] In some embodiments, the first mode is adopted when the proportion of the audio feature data to the audio input data is less than a first threshold; the second mode is adopted when the proportion of the audio feature data to the audio input data is greater than or equal to the first threshold and less than a second threshold; or the third mode is adopted when the proportion of the audio feature data to the audio input data is greater than or equal to the second threshold.
[0221] In some embodiments, the encoding mode includes an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data.
[0222] In some embodiments, the second mode includes an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data.
[0223] In some embodiments, the third mode includes an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data.
[0224] In some embodiments, the second communication device can send the first communication device encoding information, the encoding information being used to indicate the encoding mode for encoding the audio feature data. In the encoding information, a mode indication bit (mode_flag) can be included, and the value of the mode indication bit is used to indicate the encoding mode.
[0225] In some embodiments, the decoding mode of the first communication device for decoding the received first encoded data corresponds to the encoding mode used by the second communication device.
[0226] In some embodiments, the decoding mode can be determined based on the same factors used to determine the encoding mode, which can include information related to at least one of the audio input data, the audio feature data and the first task, such as the proportion of the audio feature data to the audio input data, the number of audio channels of the audio input data, the correlation between the audio channels of the audio input data; configuration information of the first task, etc.
[0227] In some embodiments, the decoding mode can include a decoding mode corresponding to at least one of the first mode, the second mode and the third mode.
[0228] In some embodiments, when the first communication device receives the first encoded data from the second communication device, it can also receive encoding information, the encoding information being used to indicate the encoding mode for encoding the audio feature data; and based on the decoding mode corresponding to the encoding mode, the first encoded data is decoded to obtain the first decoded data.
[0229] In some embodiments, the encoding information received by the first communication device can be derived from the encoding information sent by the second device; or it can also be derived from the configuration information of the first task.
[0230] In some embodiments, the encoding mode for encoding the audio feature data can be determined based on the configuration information of the first task (user config), in which the encoding information can be included, indicating the encoding mode applicable to the first task. The mode selection module in the second communication device can determine the corresponding encoding mode based on the configuration information of the first task to encode the audio feature data to obtain the first encoded data.
[0231] In some embodiments, if multiple encoding modes applicable to the first task are determined based on the configuration information of the first task, the current required encoding mode for encoding the audio feature data can be selected from the multiple encoding modes based on at least one of the proportion of the audio feature data to the audio input data, the number of audio channels of the audio input data, and the correlation between the audio channels of the audio input data, to obtain the first encoded data.
[0232] In some embodiments, the decoding mode for decoding the audio feature data can also be determined based on the configuration information of the first task, in which the encoding information can be included, indicating the decoding mode applicable to the first task. The first communication device can determine the corresponding decoding mode based on the configuration information of the first task to decode the first encoded data to obtain the first decoded data.
[0233] Embodiments of the present disclosure propose an encoding method. FIG. 4 is a schematic flowchart of an encoding method according to an embodiment of the present disclosure. The encoding method shown in the present embodiment can be performed by the second communication device.
[0234] As shown in FIG. 4, the encoding method can include the following steps:
[0235] In step S401, the audio input data is encoded to obtain the first encoded data; wherein the first encoded data includes the feature information related to the first task in the audio input data.
[0236] In some embodiments, the second communication device can be an encoding end or an encoding end device of the audio data. The second communication device can obtain the audio input data (audio input) for the first task, and encode the audio input data (Encoding) to obtain the first encoded data. Through the encoding of the audio input data, the first encoded data can be used to represent the audio features in the audio input data, wherein the audio features in the audio input data can include the feature information related to the first task, or can also be considered as the feature information required by the first task.
[0237] The first task can be an AI task requiring machine hearing.
[0238] In some embodiments, the audio input data can include audio data and labels data. The audio data can be in different audio formats, such as various uncompressed audio data, e.g., PCM audio data. The labels data corresponding to the audio data includes description information of the audio data, which can also be referred to as metadata of the audio data, e.g., can include collection information of the audio data, task information of a first task to which the audio data is directed, and the like.
[0239] Correspondingly, the second communication device encoding the audio input data can include encoding the audio data and the labels data contained in the audio input data respectively, or jointly encoding the audio data and the labels data, or encoding the audio data based on the labels data.
[0240] In some embodiments, the audio input data can include data used in the process of training, verifying and testing machine hearing.
[0241] In step S402, the first encoded data is sent.
[0242] In some embodiments, after the second communication device encodes the audio input data to obtain the first encoded data, the second communication device can send the first encoded data to the first communication device. The first encoded data can be a bitstream, which can be referred to as an encoded bitstream. The first communication device can be a decoding end device of the audio data.
[0243] In some embodiments, the second communication device can send the first encoded data to the first communication device through a transmission module.
[0244] The transmission module can use some channel transmission technology to ensure the safe and complete transmission of the bitstream, e.g., can use wireless transmission technology or wired transmission technology, and the like.
[0245] It should be noted that the embodiment shown in FIG. 4 can be independently implemented, or can be combined with at least one other embodiment of the present disclosure for implementation. The specific implementation can be selected as needed, and the present disclosure is not limited.
[0246] According to an embodiment of the present disclosure, the first encoding information is obtained by encoding the audio input data at an encoding end, and is sent to a decoding end; the first decoding data related to the first task is obtained by decoding the first encoding information at the decoding end, and is input into a machine auditory module for the first task to analyze the first decoding data, so that the data obtained by encoding and decoding the audio input data can meet the requirement of the first task, and the analysis result required by the first task is obtained by analysis.
[0247] In some embodiments, the first encoding data is audio feature data obtained by feature extraction on the audio input data.
[0248] In some embodiments, the second communication device can extract the audio feature data by feature extraction on the audio input data based on the requirement of the first task.
[0249] In some embodiments, the second communication device can include a feature extraction module, which extracts the audio feature data by feature extraction on the audio input data based on the requirement of the first task.
[0250] The audio feature data can include audio data related to the first task.
[0251] The requirement of the first task can include a requirement for any one or more of the audio data, the audio feature data, and a feature extraction manner; the requirement for the audio data can include a requirement for content, type, format, etc., and the requirement for the audio feature data can include a requirement for type, format, etc., which are not limited here.
[0252] In some embodiments, the first encoding data is a bitstream obtained by encoding the audio feature data.
[0253] In some embodiments, the second communication device can further encode the audio feature data to further compress the audio feature data after extracting the audio feature data from the audio input data by feature extraction, obtain the first encoding data, and send the first encoding data to the first communication device.
[0254] In some embodiments, the second communication device can include a feature extraction module and a feature encoding module, the feature extraction module extracts the audio feature data from the audio input data, and the feature encoding module encodes the audio feature data to obtain the first encoding data, and sends the first encoding data to the first communication device.
[0255] In some embodiments, the second communication device can include a feature extraction module and a feature encoding module, the feature extraction module extracts audio feature data from the audio input data, and the feature encoding module encodes the audio feature data to obtain the first encoded data and sends the first encoded data to the first communication device.
[0256] In some embodiments, the second communication device can flexibly select a corresponding encoding mode as needed when performing the encoding operation.
[0257] In some embodiments, the second communication device can include a feature extraction module, a mode selection module, and a feature encoding module corresponding to each encoding mode, the feature extraction module extracts audio feature data from the audio input data, the mode selection module selects a currently needed encoding mode from a plurality of available encoding modes, and the audio feature data is sent to the feature encoding module corresponding to the selected encoding mode for encoding to obtain the first encoded data.
[0258] In some embodiments, the encoding mode is determined by at least one of the following information: a proportion of the audio feature data to the audio input data; a number of audio channels of the audio input data; a correlation between the audio channels of the audio input data; and configuration information of the first task.
[0259] In some embodiments, the configuration information of the first task includes: encoding information; and related information of the audio feature data required by the first task.
[0260] In some embodiments, the encoding mode can include at least one of the following: a first mode, the first mode being an encoding mode without compressing the audio feature data; a second mode, the second mode being an encoding mode for lossless compression of the audio feature data; and a third mode, the third mode being an encoding mode for lossy compression of the audio feature data.
[0261] In some embodiments, the first mode is used when the proportion of the audio feature data to the audio input data is less than a first threshold value; the second mode is used when the proportion of the audio feature data to the audio input data is greater than or equal to the first threshold value and less than a second threshold value; or the third mode is used when the proportion of the audio feature data to the audio input data is greater than or equal to the second threshold value.
[0262] In some embodiments, the encoding mode includes: an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data.
[0263] In some embodiments, the second mode includes: an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data.
[0264] The third mode includes an encoding mode suitable for single-channel audio data and / or an encoding mode suitable for multi-channel audio data.
[0265] In some embodiments, the method further includes that the second communication device can send encoding information to the first communication device, the encoding information being used to indicate the encoding mode for encoding the audio feature data.
[0266] The encoding information sent by the second communication device to the first communication device can include a mode indication bit (mode_flag), and the encoding mode is indicated by the value of the mode indication bit.
[0267] In some embodiments, the names of information and the like are not limited to the names described in the embodiments, and the terms of "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "symbol", "codebook", "codeword", "codepoint", "bit", "data", "program", "chip", and the like can be replaced with each other.
[0268] In some embodiments, the terms of "time", "time point", "time", "time position", and the like can be replaced with each other, and the terms of "time length", "time period", "time window", "window", "time", and the like can be replaced with each other.
[0269] In some embodiments, the terms of "component carrier (CC)", "cell", "frequency carrier", "carrier frequency", and the like can be replaced with each other.
[0270] In some embodiments, "acquire", "obtain", "get", "receive", "transmit", "bidirectional transmission", "send and / or receive" can be replaced with each other, which can be interpreted as receiving from other subjects, obtaining from protocols, obtaining from higher layers, obtaining by self-processing, and the like.
[0271] In some embodiments, the terms “sending”, “transmitting”, “reporting”, “issuing”, “transferring”, “bidirectional transferring”, “sending and / or receiving”, and the like can be replaced by each other.
[0272] Corresponding to the foregoing embodiments of the decoding and encoding methods, the disclosure also provides embodiments of a terminal and a network device.
[0273] Embodiments of the disclosure also propose a first communication device, comprising: one or more processors; a memory coupled to the processors, the memory having stored thereon executable instructions that, when executed by the processors, cause the terminal to perform the decoding method described in the foregoing embodiments.
[0274] FIG. 5 is a schematic block diagram of an apparatus structure of a first communication device according to an embodiment of the disclosure. As shown in FIG. 5, the first communication device can be a decoding apparatus, and the apparatus comprises a processing module 501 and a transceiver module 502.
[0275] In some embodiments, the transceiver module 502 is configured to receive first encoded data, the first encoded data comprising feature information of audio input data related to a first task; and the processing module 501 is configured to decode the first encoded data to obtain first decoded data, wherein the first decoded data is used to perform the first task based on machine hearing.
[0276] In some embodiments, the audio input data comprises: audio data and label data; and the label data comprises description information of the audio data.
[0277] In some embodiments, the first encoded data is audio feature data obtained by performing feature extraction on the audio input data.
[0278] In some embodiments, the first encoded data is a bitstream obtained by encoding the audio feature data.
[0279] In some embodiments, the transceiver module 502 is configured to receive encoding information, the encoding information being used to indicate an encoding mode for encoding the audio feature data; and the processing module 501 is configured to decode the first encoded data based on a decoding mode corresponding to the encoding mode.
[0280] In some embodiments, the encoding mode is determined by at least one of the following information: a proportion of the audio feature data to the audio input data; a number of audio channels of the audio input data; a correlation between the audio channels of the audio input data; and configuration information of the first task.
[0281] In some embodiments, the configuration information of the first task comprises: the encoding information; and related information of the audio feature data required by the first task.
[0282] In some embodiments, the encoding mode can include at least one of: a first mode, the first mode being an encoding mode without compression of the audio feature data; a second mode, the second mode being an encoding mode with lossless compression of the audio feature data; and a third mode, the third mode being an encoding mode with lossy compression of the audio feature data.
[0283] In some embodiments, the first mode is adopted when the proportion of the audio feature data to the audio input data is less than a first threshold; the second mode is adopted when the proportion of the audio feature data to the audio input data is greater than or equal to the first threshold and less than a second threshold; or the third mode is adopted when the proportion of the audio feature data to the audio input data is greater than or equal to the second threshold.
[0284] In some embodiments, the encoding mode includes: an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data.
[0285] In some embodiments, the second mode includes: an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data.
[0286] In some embodiments, the third mode includes: an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data.
[0287] It should be noted that the modules included in the first communication device are not limited to the modules described in the above embodiments, and can also include other modules, such as a storage module, a display module, etc.
[0288] Embodiments of the present disclosure also propose a second communication device, comprising: one or more processors; a memory coupled to the processors, the memory having stored thereon executable instructions that, when executed by the processors, cause the network device to perform the encoding method described in the above embodiments.
[0289] FIG. 6 is a schematic block diagram of an apparatus structure of a second communication device according to an embodiment of the present disclosure. As shown in FIG. 6, the network device can be an encoding apparatus, and the apparatus includes a processing module 601 and a transceiver module 602.
[0290] In some embodiments, the processing module 601 is configured to encode the audio input data to obtain first encoded data; wherein the first encoded data includes feature information related to a first task in the audio input data; and the transceiver module is configured to send the first encoded data.
[0291] In some embodiments, the audio input data includes: audio data and label data; wherein the label data includes description information of the audio data.
[0292] In some embodiments, the first encoded data is audio feature data obtained by performing feature extraction on the audio input data.
[0293] In some embodiments, the first encoded data is a bitstream obtained by encoding the audio feature data.
[0294] In some embodiments, the transceiver module 602 is further configured to send encoding information, the encoding information being used to indicate an encoding mode for encoding the audio feature data.
[0295] In some embodiments, the encoding mode is determined by at least one of the following: a proportion of the audio feature data to the audio input data; a number of audio channels of the audio input data; a correlation between audio channels of the audio input data; configuration information of the first task.
[0296] In some embodiments, the configuration information of the first task comprises: the encoding information; and related information of the audio feature data required by the first task.
[0297] In some embodiments, the encoding mode can comprise at least one of the following: a first mode, the first mode being an encoding mode without compression of the audio feature data; a second mode, the second mode being an encoding mode for lossless compression of the audio feature data; and a third mode, the third mode being an encoding mode for lossy compression of the audio feature data.
[0298] In some embodiments, the first mode is adopted when the proportion of the audio feature data to the audio input data is less than a first threshold; the second mode is adopted when the proportion of the audio feature data to the audio input data is greater than or equal to the first threshold and less than a second threshold; or the third mode is adopted when the proportion of the audio feature data to the audio input data is greater than or equal to the second threshold.
[0299] In some embodiments, the encoding mode comprises: an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data.
[0300] In some embodiments, the second mode comprises: an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data.
[0301] In some embodiments, the third mode comprises: an encoding mode suitable for single-channel audio data; and / or, an encoding mode suitable for multi-channel audio data.
[0302] It should be noted that the modules included in the second communication device are not limited to the modules described in the above embodiments, and can also include other modules, such as a storage module, a display module, etc.
[0303] For the apparatus embodiment, since it basically corresponds to the method embodiment, the relevant part can be seen from the part of the method embodiment. The apparatus embodiment described above is only illustrative, wherein the modules described as separate components can or can not be physically separated, and the components displayed as modules can or can not be physical modules, i.e., can be located in one place or distributed to multiple network modules. Part or all of the modules can be selected to achieve the purpose of the embodiment according to actual needs. Those skilled in the art can understand and implement it without creative labor.
[0304] The embodiment of the disclosure also proposes a communication device, comprising: one or more processors; a memory coupled to the processor, the memory having stored executable instructions, wherein the executable instructions, when executed by the processor, cause the processor to invoke the executable instructions to cause the communication device to perform the decoding or encoding method described in the optional embodiment.
[0305] The embodiment of the disclosure also proposes a communication system, comprising a first communication device and a second communication device, wherein the first communication device is configured to implement the decoding method described in the optional embodiment, and the second communication device is configured to implement the encoding method described in the optional embodiment.
[0306] The embodiment of the disclosure also proposes a storage medium, the storage medium storing instructions, when the instructions are run on a communication device, causing the communication device to perform the decoding or encoding method described in the optional embodiment.
[0307] The embodiment of the disclosure also proposes an apparatus for implementing any of the above methods, for example, proposes an apparatus, the apparatus comprising units or modules for implementing each step performed by the terminal in any of the above methods. For another example, another apparatus is also proposed, comprising units or modules for implementing each step performed by the network device (such as an access network device, a core network function node, a core network device, etc.) in any of the above methods.
[0308] It should be understood that the division of each unit or module in the above apparatus is only a logical function division, and all or part of them can be integrated into a physical entity or physically separated in actual implementation. In addition, the units or modules in the apparatus can be implemented in the form of processor calling software: for example, the apparatus includes a processor, the processor is connected with a memory, the memory stores instructions, and the processor calls the instructions stored in the memory to realize the functions of any of the above methods or the units or modules of the above apparatus, wherein the processor is a general processor such as a central processing unit (CPU) or a microprocessor, and the memory is a memory in the apparatus or a memory outside the apparatus. Alternatively, the units or modules in the apparatus can be implemented in the form of hardware circuit, and the functions of part or all of the units or modules can be realized by the design of the hardware circuit. The above hardware circuit can be understood as one or more processors; for example, in one implementation, the above hardware circuit is an application-specific integrated circuit (ASIC), and the functions of part or all of the units or modules are realized by the design of the logical relationship between the elements in the circuit; for another example, in another implementation, the above hardware circuit is a programmable logic device (PLD), and a field programmable gate array (FPGA) is taken as an example, which can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by a configuration file, so as to realize the functions of part or all of the units or modules. All units or modules of the above apparatus can be all implemented in the form of processor calling software, or all implemented in the form of hardware circuit, or part implemented in the form of processor calling software and the remaining part implemented in the form of hardware circuit.
[0309] In the embodiments of the present disclosure, the processor is a circuit with signal processing capability. In one implementation, the processor can be a circuit with instruction reading and running capability, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), a digital signal processor (DSP), and the like. In another implementation, the processor can implement certain functions through a logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or reconfigurable. For example, the processor is a hardware circuit implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In the reconfigurable hardware circuit, the processor loads a configuration document to implement the configuration of the hardware circuit. It can be understood that the processor loads instructions to implement the functions of the above part or all units or modules. In addition, the hardware circuit can also be designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), and the like.
[0310] FIG. 7 is a structural schematic diagram of a communication device 7100 according to the embodiments of the present disclosure. The communication device 7100 can be a network device (for example, an access network device, a core network device, and the like), or a terminal (for example, a user equipment, and the like), or a chip, a chip system, or a processor supporting the network device to implement any of the above methods, or a chip, a chip system, or a processor supporting the terminal to implement any of the above methods. The communication device 7100 can be used to implement the methods described in the above method embodiments, and details can be referred to the descriptions in the above method embodiments.
[0311] As shown in FIG. 7, the communication device 7100 includes one or more processors 7101. The processor 7101 can be a general processor or a special-purpose processor, etc., for example, a baseband processor or a central processor. The baseband processor can be used to process communication protocols and communication data, the central processor can be used to control a communication apparatus (e.g., a base station, a baseband chip, a terminal device, a terminal device chip, a DU or a CU, etc.), execute programs, and process data of the programs. The processor 7101 is configured to invoke instructions to enable the communication device 7100 to perform any of the above methods.
[0312] In some embodiments, the communication device 7100 further includes one or more memories 7102 configured to store instructions. Optionally, all or part of the memory 7102 can also be outside the communication device 7100.
[0313] In some embodiments, the communication device 7100 further includes one or more transceivers 7103. When the communication device 7100 includes one or more transceivers 7103, the communication steps in the above methods, such as transmitting and receiving, are performed by the transceiver 7103, and the other steps are performed by the processor 7101.
[0314] In some embodiments, the transceiver can include a receiver and a transmitter, which can be separate or integrated together. Optionally, the terms transceiver, transceiving unit, transceiver, transceiving circuit, etc. can be replaced by each other, the terms transmitter, transmitting unit, transmitter, transmitting circuit, etc. can be replaced by each other, and the terms receiver, receiving unit, receiver, receiving circuit, etc. can be replaced by each other.
[0315] Optionally, the communication device 7100 further includes one or more interface circuits 7104 connected to the memory 7102. The interface circuit 7104 can be used to receive signals from the memory 7102 or other devices, and can be used to send signals to the memory 7102 or other devices. For example, the interface circuit 7104 can read instructions stored in the memory 7102 and send the instructions to the processor 7101.
[0316] The communication device 7100 described in the above embodiments may be a network device or a terminal, but the scope of the communication device 7100 described in this disclosure is not limited thereto, and the structure of the communication device 7100 may not be limited by FIG. 7. The communication device may be a standalone device or a part of a larger device. For example, the communication device may be: (1) a standalone integrated circuit IC, or chip, or chip system or subsystem; (2) a collection of one or more ICs, optionally, the IC collection may also include storage components for storing data and programs; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, terminal device, smart terminal device, cellular phone, wireless device, handheld device, mobile unit, vehicle device, network device, cloud device, artificial intelligence device, etc.; (6) others, etc.
[0317] Figure 8 is a schematic diagram of the structure of chip 8200 according to an embodiment of this disclosure. For cases where the communication device 7100 can be a chip or a chip system, the schematic diagram of chip 8200 shown in Figure 8 can be referenced, but is not limited thereto.
[0318] Chip 8200 includes one or more processors 8201, which are used to invoke instructions to cause chip 8200 to perform any of the above methods.
[0319] In some embodiments, the chip 8200 further includes one or more interface circuits 8202, which are connected to the memory 8203. The interface circuits 8202 can be used to receive signals from the memory 8203 or other devices, and can be used to send signals to the memory.
[0320] 8203 or other devices send signals. For example, interface circuit 8202 can read instructions stored in memory 8203 and send those instructions to processor 8201. Optionally, terms such as interface circuit, interface, transceiver pin, transceiver, etc., can be used interchangeably.
[0321] In some embodiments, chip 8200 further includes one or more memories 8203 for storing instructions. Optionally, all or part of the memories 8203 may be located outside of chip 8200.
[0322] The present disclosure further provides a storage medium having stored instructions which, when executed on the communication device 7100, cause the communication device 7100 to perform any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but is not limited thereto and can also be a storage medium readable by other apparatuses. Optionally, the storage medium can be a non-transitory storage medium, but is not limited thereto and can also be a transitory storage medium.
[0323] The present disclosure further provides a program product which, when executed by the communication device 7100, causes the communication device 7100 to perform any of the above methods. Optionally, the program product is a computer program product.
[0324] The present disclosure further provides a computer program which, when executed on a computer, causes the computer to perform any of the above methods.
Claims
1. A decoding method, comprising: The method comprises: receiving first encoded data, the first encoded data comprising feature information in audio input data related to a first task; decoding the first encoded data to obtain first decoded data, wherein the first decoded data is used to perform the first task based on machine hearing.
2. The method of claim 1, wherein, The audio input data comprises: audio data and label data; wherein the label data comprises description information of the audio data.
3. The method according to claim 1 or 2, characterized in that, The first encoded data is audio feature data obtained after feature extraction on the audio input data; or, The first encoded data is a bitstream obtained after encoding on the audio feature data.
4. The method according to any one of claims 1 to 3, characterized in that, Further comprising: receiving encoding information, the encoding information being used to indicate an encoding mode for encoding the audio feature data; The decoding of the first encoded data comprises: decoding the first encoded data based on a decoding mode corresponding to the encoding mode.
5. The method of claim 4, wherein, The encoding mode is determined by at least one of the following: a proportion of the audio feature data to the audio input data; a number of audio channels of the audio input data; a correlation between audio channels of the audio input data; configuration information of the first task.
6. The method of claim 5, wherein, The configuration information of the first task comprises: the encoding information; related information of the audio feature data required by the first task.
7. The method according to any one of claims 4-6, characterized by, The encoding mode can comprise at least one of the following: a first mode, the first mode being an encoding mode without compression on the audio feature data; a second mode, the second mode being an encoding mode for lossless compression on the audio feature data; a third mode, the third mode being an encoding mode for lossy compression on the audio feature data.
8. The method of claim 7, wherein: in a case where the proportion of the audio feature data to the audio input data is less than a first threshold, the first mode is adopted; in a case where the proportion of the audio feature data to the audio input data is greater than or equal to the first threshold and less than a second threshold, the second mode is adopted; or in a case where the proportion of the audio feature data to the audio input data is greater than or equal to the second threshold, the third mode is adopted.
9. The method of claim 7 or 8, wherein: the second mode comprises: an encoding mode suitable for single-channel audio data; and / or an encoding mode suitable for multi-channel audio data; the third mode comprises: an encoding mode suitable for single-channel audio data; and / or an encoding mode suitable for multi-channel audio data. The method comprises:
10. An encoding method characterized by comprising: encoding audio input data to obtain first encoded data; wherein the first encoded data comprises feature information in the audio input data related to a first task; sending the first encoded data. The audio input data comprises: audio data and label data; wherein the label data comprises description information of the audio data.
11. The method of claim 10, wherein, The first encoded data is audio feature data obtained after feature extraction on the audio input data; or, 12. The method according to claim 10 or 11, characterized in that, The first encoded data is a bitstream obtained by encoding the audio feature data.
13. The method according to any one of claims 10-12, characterized in that, The method further comprises: sending encoding information, the encoding information being used to indicate an encoding mode for encoding the audio feature data.
14. The method according to any one of claims 10-13, characterized in that, The encoding mode is determined by at least one of the following information: a proportion of the audio feature data to the audio input data; a number of audio channels of the audio input data; a correlation between audio channels of the audio input data; configuration information of the first task.
15. The method of claim 14, wherein, The configuration information of the first task comprises: the encoding information; relevant information of the audio feature data required by the first task.
16. The method according to any one of claims 13-15, characterized by, The encoding mode can comprise at least one of the following: a first mode, the first mode being an encoding mode without compression of the audio feature data; a second mode, the second mode being an encoding mode for lossless compression of the audio feature data; a third mode, the third mode being an encoding mode for lossy compression of the audio feature data.
17. The method of claim 16, wherein, in a case where the proportion of the audio feature data to the audio input data is less than a first threshold, the first mode is adopted; in a case where the proportion of the audio feature data to the audio input data is greater than or equal to the first threshold and less than a second threshold, the second mode is adopted; or in a case where the proportion of the audio feature data to the audio input data is greater than or equal to the second threshold, the third mode is adopted.
18. The method of claim 16 or 17, wherein, the second mode comprises: an encoding mode suitable for single-channel audio data; and / or an encoding mode suitable for multi-channel audio data; the third mode comprises: an encoding mode suitable for single-channel audio data; and / or an encoding mode suitable for multi-channel audio data. comprises:
19. A decoding apparatus, comprising: a transceiver module configured to receive first encoded data, the first encoded data comprising feature information related to a first task in audio input data; a processing module configured to decode the first encoded data to obtain first decoded data, wherein the first decoded data is used to perform the first task based on machine hearing. comprises:
20. An encoding apparatus, comprising: a processing module configured to encode audio input data to obtain first encoded data, wherein the first encoded data comprises feature information related to a first task in the audio input data; a transceiver module configured to send the first encoded data. comprises:
21. A communications device, characterized by one or more processors; a memory coupled to the processors, the memory having stored thereon executable instructions that, when executed by the processors, cause the processors to invoke instructions to cause the communication device to perform the decoding method of any one of claims 1-9 and / or the encoding method of any one of claims 10-18. comprises a first communication device and a second communication device, wherein the first communication device is configured to implement the decoding method of any one of claims 1-9, and the second communication device is configured to implement the encoding method of any one of claims 10-18.
22. A communication system, characterized by 23. A storage medium, the storage medium storing instructions, wherein, When the instructions are run on a communication device, they cause the communication device to perform the decoding method of any one of claims 1-9, and / or the encoding method of any one of claims 10-18.
Citation Information
Patent Citations
Video encoding and decoding fusion processing method based on multiple compression systems
CN111314778A
Video encoding method, video encoding device, video decoding method and video decoding device
CN112383778A
Signal coding and decoding method and device, coding equipment, decoding equipment and storage medium
CN114127844A
Signal coding and decoding method and device, and storage medium
CN118043887A
Coding and decoding method and device and storage medium
CN118160286A