Encoding, decoding and encoding-decoding methods, encoding apparatus, decoding apparatus and storage medium
By using encoder and decoder in the audio data encoding process, combined with Huffman encoding, filtering and multi-channel data mixing technologies, the problem of low accuracy of audio data encoding is solved, achieving more efficient machine analysis and resource savings.
Patent Information
- Application Number
- PCT/CN2024/075038
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-31
- Publication Date
- 2025-08-07
AI Technical Summary
In the prior art, the accuracy of audio data encoding is low, resulting in insufficient accuracy of subsequent machine analysis.
The audio data is encoded through the encoding information included in the first signal to ensure the accuracy of encoding. The audio data is encoded and decoded by using encoder and decoder, and the technical means such as Huffman encoding, filtering, core selection and multiplexed data mixing are used to improve the accuracy of encoding and decoding.
Improve the accuracy of audio data encoding, ensure the accuracy of subsequent machine analysis, save transmission resources and enhance the flexibility and stability of the encoding process.
Smart Images

Figure CN2024075038_07082025_PF_FP_ABST
Abstract
Description
Coding and decoding method, device and storage medium Technical Field
[0001] The present disclosure relates to the field of communication technologies, and in particular to a coding and decoding method, device, and storage medium. Background Art
[0002] With the rapid development of multimedia technology, audio data can be encoded and transmitted, and then the encoded data can be decoded to obtain the corresponding audio data, ensuring that the audio data can be transmitted efficiently.
[0003] Summary of the Invention
[0004] The present disclosure solves the problem of low accuracy in encoding audio data. The audio data is encoded by encoding information included in the first signal, thereby ensuring the accuracy of the encoding of the audio data and further ensuring the accuracy of subsequent machine analysis of the audio data.
[0005] The embodiments of the present disclosure provide a coding and decoding method, apparatus, and storage medium.
[0006] According to a first aspect of an embodiment of the present disclosure, a coding method is proposed. The method is performed by an encoder, and the method includes:
[0007] The audio data is encoded based on a first signal to obtain first encoded data, where the first signal includes encoding information for encoding the audio data. The first signal is fed back by a first device, and the audio data is used for machine analysis or machine training by the first device.
[0008] According to a second aspect of an embodiment of the present disclosure, a decoding method is proposed, where the method is performed by a decoder and includes:
[0009] The first encoded data is decoded based on the second signal to obtain audio data, the first signal includes encoding information for encoding the audio data, the first signal is fed back by the first device, and the audio data is used for machine analysis or machine training of the first device.
[0010] According to a third aspect of the embodiments of the present disclosure, a coding and decoding method is proposed, the method comprising:
[0011] The encoder encodes the audio data based on a first signal to obtain first encoded data, where the first signal includes encoding information for encoding the audio data, the first signal is fed back by a first device, and the audio data is used for machine analysis or machine training by the first device;
[0012] The decoder decodes the first encoded data based on the second signal to obtain audio data, where the first signal includes encoding information for encoding the audio data.
[0013] According to a fourth aspect of the embodiments of the present disclosure, an encoding device is provided, including:
[0014] A processing module is used to encode audio data based on a first signal to obtain first encoded data, where the first signal includes encoding information for encoding the audio data, the first signal is fed back by a first device, and the audio data is used for machine analysis or machine training by the first device.
[0015] According to a fifth aspect of the embodiments of the present disclosure, a decoding device is provided, including:
[0016] A processing module is used to decode the first encoded data using a second signal to obtain audio data, wherein the first signal includes encoding information for encoding the audio data, the first signal is fed back by a first device, and the audio data is used by the first device for machine analysis or machine training.
[0017] According to a sixth aspect of the embodiments of the present disclosure, an encoding device is provided, including:
[0018] one or more processors;
[0019] The encoding and decoding device is used to execute any one of the methods described in the first aspect.
[0020] According to a seventh aspect of the embodiments of the present disclosure, a decoding device is provided, including:
[0021] one or more processors;
[0022] The encoding and decoding device is used to execute any method described in the second aspect.
[0023] According to an eighth aspect of the embodiments of the present disclosure, a coding and decoding system is proposed, including:
[0024] An encoder and a decoder, wherein the encoder is configured to implement the encoding and decoding method described in the first aspect, and the decoder is configured to implement the encoding and decoding method described in the second aspect.
[0025] According to a ninth aspect of an embodiment of the present disclosure, a storage medium is proposed, wherein the storage medium stores instructions. When the instructions are executed on a communication device, the communication device executes a method as described in any one of the first aspect or the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings described herein are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the present disclosure. The illustrative embodiments of the embodiments of the present disclosure and their descriptions are used to explain the embodiments of the present disclosure and do not constitute an improper limitation on the embodiments of the present disclosure. In the drawings:
[0027] FIG1 is a schematic diagram of the architecture of a coding and decoding system according to an embodiment of the present disclosure;
[0028] FIG2A is an interactive schematic diagram illustrating a coding and decoding method according to an embodiment of the present disclosure;
[0029] FIG2B is a schematic diagram showing a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0030] FIG2C is a schematic diagram showing a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0031] FIG2D is a schematic diagram illustrating a filtering method according to an embodiment of the present disclosure;
[0032] FIG3A is a schematic diagram showing a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0033] FIG3B is a schematic diagram showing a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0034] FIG4 is a schematic diagram of a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0035] FIG5 is a schematic diagram of a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0036] FIG6 is a schematic diagram of a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0037] FIG7A is a schematic structural diagram of a coding and decoding device proposed in an embodiment of the present disclosure;
[0038] FIG7B is a schematic structural diagram of a coding and decoding device proposed in an embodiment of the present disclosure;
[0039] FIG8A is a schematic structural diagram of a communication device proposed in an embodiment of the present disclosure;
[0040] FIG8B is a schematic diagram of the structure of the chip proposed in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0041] The present disclosure provides a coding and decoding method, device, and storage medium.
[0042] According to a first aspect of an embodiment of the present disclosure, a coding method is proposed. The method is performed by an encoder, and the method includes:
[0043] The audio data is encoded based on a first signal to obtain first encoded data, where the first signal includes encoding information for encoding the audio data. The first signal is fed back by a first device, and the audio data is used for machine analysis or machine training by the first device.
[0044] In the above embodiment, the first signal indicates the coding information involved in encoding the audio data, which solves the problem of low accuracy in encoding the audio data. The audio data is encoded by the coding information included in the first signal, ensuring the accuracy of the encoding of the audio data, and further ensuring the accuracy of subsequent machine analysis of the audio data.
[0045] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0046] The first signal is encoded to obtain second encoded data; the second encoded data is used for at least one of being sent together with the first encoded data or being decoded by a decoder.
[0047] In the above embodiment, the encoder can not only encode audio data, but also encode the first signal, and then send second encoded data. In addition, the second encoded data also supports decoding by the decoder. On the basis of ensuring the decoding of the first signal, the comprehensiveness of the feedback encoded data is also expanded.
[0048] In conjunction with some embodiments of the first aspect, in some embodiments, encoding the first signal to obtain second encoded data includes:
[0049] The first signal is encoded using Huffman coding to obtain the second encoded data.
[0050] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0051] Determining, based on the scene indicated by the first signal, encoding parameters corresponding to the first signal, where the encoding parameters are used to encode the audio data;
[0052] The encoding parameters are encoded to obtain the second encoded data.
[0053] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0054] The first signal is sent.
[0055] In the above embodiment, the first signal may also be sent directly, saving the encoding process and further receiving resource consumption.
[0056] In conjunction with some embodiments of the first aspect, in some embodiments, encoding the audio data based on the first signal to obtain first encoded data includes:
[0057] Perform at least one of the following on the audio data to obtain the first encoded data:
[0058] Pretreatment;
[0059] Signal analysis;
[0060] Nuclear option;
[0061] nuclear coding;
[0062] Multi-channel data mixing.
[0063] In the above embodiment, the audio data is encoded by executing multiple processes to obtain the first encoded data, thereby ensuring the accuracy of the audio data encoding.
[0064] In conjunction with some embodiments of the first aspect, in some embodiments, the first signal includes a first frequency band, the first frequency band is smaller than a second frequency band of the audio data, and the preprocessing includes:
[0065] The audio data is filtered based on the first signal to obtain filtered data, where a frequency band of the filtered data is the first frequency band.
[0066] In the above embodiment, by filtering the audio data to obtain filtered data with a smaller frequency band, the amount of filtered data in subsequent encoding is reduced, thereby saving resource consumption.
[0067] In conjunction with some embodiments of the first aspect, in some embodiments, the preprocessing further includes at least one of the following:
[0068] Divide into multiple frames;
[0069] Time-frequency transformation;
[0070] Time window adjustment;
[0071] Time Domain Noise Shaping;
[0072] Frequency Domain Noise Shaping.
[0073] In conjunction with some embodiments of the first aspect, in some embodiments, the first signal includes a first core module (Core Module), and the core selection includes:
[0074] A third core module is determined based on the second core module selected based on the audio data and the first core module, and the third core module is used to process the audio data.
[0075] In the above embodiment, both the first signal and the encoder can determine the core module. Therefore, the core module for processing the audio data is determined based on the core module indicated by the first signal and the core module determined by the encoder, thereby ensuring the accuracy of the determined core module and further ensuring the accuracy of processing the audio data.
[0076] In conjunction with some embodiments of the first aspect, in some embodiments, determining the third core module based on the second core module selected by the audio data and the first core module includes:
[0077] The first core module is the same as the second core module, and the third core module is determined to be either the first core module or the second core module.
[0078] In conjunction with some embodiments of the first aspect, in some embodiments, determining the third core module based on the second core module selected by the audio data and the first core module includes:
[0079] The first core module is different from the second core module, and the third core module is determined to be the first core module.
[0080] In combination with some embodiments of the first aspect, in some embodiments, the first signal is a fixed signal.
[0081] In the above embodiment, the first signal is fixed information, which ensures the stability of the first signal and further ensures the stability of encoding the audio data.
[0082] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0083] The first signal is received, where the first signal supports real-time adjustment.
[0084] In the above embodiment, the first signal is a signal that can be adjusted in real time, which improves the flexibility of the first signal and further ensures the flexibility of encoding the audio data.
[0085] In conjunction with some embodiments of the first aspect, in some embodiments, the configuration of the encoder includes at least one of the following parameters:
[0086] Input parameters;
[0087] Output parameters;
[0088] Coding rate;
[0089] Sampling rate;
[0090] a starting frequency and an ending frequency for encoding the first coded data;
[0091] Core module identifier.
[0092] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0093] A data stream comprising the first encoded data is sent to a decoder.
[0094] In a second aspect, an embodiment of the present disclosure provides a decoding method, which is performed by an encoder and includes:
[0095] The first encoded data is decoded based on the second signal to obtain audio data, the first signal includes encoding information for encoding the audio data, the first signal is fed back by the first device, and the audio data is used for machine analysis or machine training of the first device.
[0096] In conjunction with some embodiments of the second aspect, in some embodiments, the method further includes:
[0097] The second encoded data is decoded to obtain the second signal.
[0098] In conjunction with some embodiments of the second aspect, in some embodiments, decoding the second encoded data to obtain the first signal includes:
[0099] The second coded data is decoded using Huffman coding to obtain the second signal.
[0100] In conjunction with some embodiments of the second aspect, in some embodiments, the method further includes:
[0101] The second signal is received.
[0102] In conjunction with some embodiments of the second aspect, in some embodiments, decoding the first encoded data based on the first signal includes:
[0103] Performing at least one of the following on the first encoded data to obtain the audio data includes:
[0104] Bitstream demultiplexing;
[0105] nuclear decoding;
[0106] Inverse transform.
[0107] In conjunction with some embodiments of the second aspect, in some embodiments, the core decoding includes:
[0108] The core module determined by the second signal performs core decoding on the audio data.
[0109] In conjunction with some embodiments of the second aspect, in some embodiments, the inverse transform includes:
[0110] Inverse time-frequency transform;
[0111] Inverse time domain noise shaping transform;
[0112] Inverse frequency-domain noise shaping transform.
[0113] In a third aspect, an embodiment of the present disclosure provides a coding and decoding method, the method comprising:
[0114] The encoder encodes the audio data based on a first signal to obtain first encoded data, where the first signal includes encoding information for encoding the audio data, the first signal is fed back by a first device, and the audio data is used for machine analysis or machine training by the first device;
[0115] The decoder decodes the first encoded data based on the second signal to obtain audio data.
[0116] In a fourth aspect, an embodiment of the present disclosure provides an encoding device, wherein the encoding and decoding device includes at least one of a transceiver module and a processing module; wherein the second device is used to execute the optional implementation methods of the first and third aspects.
[0117] In a fifth aspect, an embodiment of the present disclosure provides a decoding device, wherein the encoding and decoding device includes at least one of a transceiver module and a processing module; wherein the encoder is used to execute the optional implementation methods of the second and third aspects.
[0118] In a sixth aspect, an embodiment of the present disclosure provides an encoding device, including:
[0119] one or more processors;
[0120] The encoding and decoding device is used to execute the method described in any one of the first and third aspects.
[0121] In a seventh aspect, an embodiment of the present disclosure provides a decoding device, including:
[0122] one or more processors;
[0123] The encoding and decoding device is used to execute the method described in any one of the second and third aspects.
[0124] In an eighth aspect, an embodiment of the present disclosure provides a storage medium storing first information. When the first information is run on a communication device, the communication device executes a method as described in any one of the first, second and third aspects.
[0125] In a ninth aspect, an embodiment of the present disclosure proposes a program product. When the program product is executed by a communication device, the communication device executes any one of the methods described in the first, second and third aspects.
[0126] In a tenth aspect, an embodiment of the present disclosure proposes a computer program, which, when executed on a communication device, enables the communication device to execute any one of the methods described in the first, second, and third aspects.
[0127] In an eleventh aspect, an embodiment of the present disclosure provides a chip or a chip system, wherein the chip or chip system includes a processing circuit configured to execute any one of the methods described in the first, second, and third aspects.
[0128] It is understandable that the above-mentioned encoder, storage medium, program product, computer program, chip or chip system are all used to execute the method proposed in the embodiment of the present disclosure. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding method and will not be repeated here.
[0129] The present disclosure provides a coding method, apparatus, and storage medium. In some embodiments, the terms coding method, information coding method, and coding method are interchangeable; the terms coding apparatus, information coding apparatus, and indicating apparatus are interchangeable; and the terms information processing system, coding system, and so on are interchangeable.
[0130] The embodiments of the present disclosure are not exhaustive and are merely illustrative of some embodiments, and are not intended to be a specific limitation on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be arbitrarily exchanged. In addition, the optional implementation methods in a certain embodiment can be arbitrarily combined; in addition, the embodiments can be arbitrarily combined. For example, some or all steps of different embodiments can be arbitrarily combined, and a certain embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.
[0131] In each embodiment of the present disclosure, unless otherwise specified or provided for by logic, the terms and / or descriptions between the embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form a new embodiment based on their inherent logical relationships.
[0132] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.
[0133] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular, such as "a", "an", "the", "above", "said", "the", "the", etc., may mean "one and only one", or "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English in translation, the noun following the article may be understood as a singular expression or a plural expression.
[0134] In the embodiments of the present disclosure, “plurality” refers to two or more.
[0135] In some embodiments, the terms "at least one," "one or more," "a plurality of," "multiple," etc. may be used interchangeably.
[0136] In some embodiments, descriptions such as "at least one of A and B," "A and / or B," "A in one case, B in another case," or "in response to one case A, in response to another case B" may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed); and in some embodiments, A and B (both A and B are executed). The above is also applicable when there are more branches such as A, B, and C.
[0137] In some embodiments, "A or B" and other descriptions may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed). The above is also applicable when there are more branches such as A, B, C, etc.
[0138] The prefixes such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different description objects and do not constitute any restriction on the position, order, priority, quantity or content of the description objects. For the statement of the description object, please refer to the description in the context of the claims or embodiments, and no unnecessary restriction should be constituted due to the use of prefixes. For example, if the description object is a "field", the ordinal number before the "field" in the "first field" and the "second field" does not limit the position or order between the "fields". "First" and "second" do not limit whether the "fields" they modify are in the same message, nor do they limit the order of the "first field" and the "second field". For another example, if the description object is a "level", the ordinal number before the "level" in the "first level" and the "second level" does not limit the priority between the "levels". For another example, the number of description objects is not limited by the ordinal number and can be one or more. Taking "first device" as an example, the number of "devices" can be one or more. In addition, the objects modified by different prefixes can be the same or different. For example, if the description object is "device", then the "first device" and the "second device" can be the same device or different devices, and their types can be the same or different; for another example, if the description object is "information", then the "first information" and the "second information" can be the same information or different information, and their contents can be the same or different.
[0139] In some embodiments, “including A,” “comprising A,” “used to indicate A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.
[0140] In some embodiments, terms such as "time / frequency" and "time / frequency domain" refer to the time domain and / or the frequency domain.
[0141] In some embodiments, terms such as "in response to...", "in response to determining...", "in the case of...", "at the time of...", "when...", "if...", "if...", etc. can be used interchangeably.
[0142] In some embodiments, terms such as "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not less than", and "above" can be replaced with each other, and terms such as "less than", "less than or equal to", "not greater than", "less than", "less than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", and "below" can be replaced with each other.
[0143] In some embodiments, devices and equipment can be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. In some cases, they can also be understood as "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "subject", etc.
[0144] In some embodiments, "network" can be interpreted as devices included in the network, such as access network equipment, core network equipment, etc.
[0145] In some embodiments, "access network device (AN device)" may also be referred to as "radio access network device (RAN device)", "base station (BS)", "radio base station", "fixed station", and in some embodiments may also be understood as "node", "access point", "transmission point (TP)", "reception point (RP)", "transmission and / or reception point (TRP)" "panel", "antenna panel", "antenna array", "cell", "macro cell", "small cell", "femto cell", "pico cell", "sector", "cell group", "serving cell", "carrier", "component carrier", "bandwidth part (BWP)", etc.
[0146] In some embodiments, "encoder (terminal)" or "encoder device (terminal device)" may be referred to as "user equipment (encoder)", "user terminal" "mobile station (MS)", "mobile terminal (MT)", subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, etc.
[0147] In some embodiments, obtaining data, information, etc. may comply with the laws and regulations of the country where the data is obtained.
[0148] In some embodiments, data, information, etc. may be obtained with the user's consent.
[0149] In addition, each element, each row, or each column in the table of the embodiment of the present disclosure can be implemented as an independent embodiment, and the combination of any elements, any rows, and any columns can also be implemented as an independent embodiment.
[0150] FIG1 is a schematic diagram of the architecture of a codec system according to an embodiment of the present disclosure. As shown in FIG1 , the method provided in the embodiment of the present disclosure can be applied to a codec system 100, which can include an encoder 101 and a decoder 102. It should be noted that the codec system 100 can also include other devices, and the present disclosure does not limit the devices included in the codec system 100.
[0151] In some embodiments, the encoder 101 and the decoder 102 are both provided in a terminal. In some embodiments, the terminal can be various devices. For example, the terminal can be a mobile phone, a wearable device, an Internet of Things device, a car with communication functions, a smart car, a tablet computer, a computer with wireless transceiver functions, a virtual reality (VR) encoder device, an augmented reality (AR) encoder device, a wireless encoder device in industrial control, a wireless encoder device in self-driving, a wireless encoder device in remote medical surgery, a wireless encoder device in a smart grid, a wireless encoder device in transportation safety, a wireless encoder device in a smart city, and a wireless encoder device in a smart home, but is not limited thereto.
[0152] It can be understood that the encoding and decoding system described in the embodiment of the present disclosure is for the purpose of more clearly illustrating the technical solution of the embodiment of the present disclosure, and does not constitute a limitation on the technical solution proposed in the embodiment of the present disclosure. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution proposed in the embodiment of the present disclosure is also applicable to similar technical problems.
[0153] The following embodiments of the present disclosure may be applied to the codec system 100 shown in FIG1 , or a portion thereof, but are not limited thereto. The entities shown in FIG1 are illustrative only. The codec system may include all or part of the entities shown in FIG1 , or may include other entities outside of FIG1 . The number and form of the entities are arbitrary, and the entities may be physical or virtual. The connection relationship between the entities is illustrative only. The entities may be connected or disconnected, and the connection may be in any manner, including direct or indirect, wired or wireless.
[0154] The embodiments of the present disclosure can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G New Radio (NR), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New Radio Access (NX), Future Generation Radio Access (FX), Global System for Mobile Communications (GSM (registered trademark)), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, Ultra-WideBand (UWB), Bluetooth (registered trademark), Public Land Mobile Network (PLMN) networks, Device-to-Device (D2D) systems, Machine-to-Machine (M2M) systems, Internet of Things (IoT) systems, Vehicle-to-Everything (V2X), systems using other coding and decoding methods, and next-generation systems based on these. In addition, multiple systems can also be combined (for example, a combination of LTE or LTE-A with 5G, etc.) for application.
[0155] In some embodiments, the present disclosure further includes a second device, or may also be referred to as an acquisition device. Optionally, the second device is used to acquire audio data and generate first data of the audio data, and the audio data is described by the first data. Optionally, the second device is an audio sensor device.
[0156] FIG2A is an interactive diagram of a coding and decoding method according to an embodiment of the present disclosure. As shown in FIG2A , the embodiment of the present disclosure relates to a coding and decoding method, which includes:
[0157] Step S2101: The second device collects audio data.
[0158] In some embodiments, the second device refers to a device with a collection function. For example, the second device may also be referred to as a collection device, a collection terminal, etc., which is not limited in the embodiments of the present disclosure. In some embodiments, the audio data refers to audio, or it can also be understood as sound collected by the second device. In some embodiments, the audio data is used to indicate the audio generated by the device during operation.
[0159] In some embodiments, the audio data is used for machine analysis or machine training. Alternatively, it can be understood that machine analysis or machine training of the audio data can determine whether the operating status of the device indicated by the audio data is normal. Alternatively, it can be understood that the audio data indicates the operating status of the device, and whether the device is operating normally can be determined based on the audio data.
[0160] In some embodiments, the embodiments of the present disclosure are applied to automated production and inspection scenarios, where the second device collects audio data of the produced products (components), and then performs machine analysis or machine training on the collected audio data to determine whether the generated products (components) are qualified.
[0161] It should be noted that the embodiments of the present disclosure are described using the role of audio data as an example. Furthermore, after the decoder decodes and obtains the audio data in the following embodiments, the first device analyzes the audio data obtained by the decoder to determine whether the object monitored by the second device is operating normally. In some embodiments, the object can be understood as the second device, the product produced, or other information that the audio data refers to, which is not limited in the embodiments of the present disclosure. In some embodiments, the first device is used to control the second device, which can be understood as the first device being the controller of the second device.
[0162] In some embodiments, the embodiments of the present disclosure are applied in product testing, see Figure 2B, audio data of the product is collected by a second device, the audio data is encoded by an encoder, the encoded data is sent to a decoder, and the decoder decodes the encoded data to determine whether the product corresponding to each audio data is qualified.
[0163] Step S2102: The second device sends audio data to the encoder.
[0164] In some embodiments, the encoder receives audio data sent by the second device. In some embodiments, the second device sends audio data. In some embodiments, the encoder receives audio data.
[0165] In step S2103 , the encoder encodes the audio data based on the first signal to obtain first encoded data.
[0166] In the embodiment of the present disclosure, the first encoded data obtained by encoding the audio data through an encoder and then transmitting it can reduce the amount of transmitted data and save transmission resources.
[0167] In some embodiments, the first signal is sent by a first device for controlling a machine. In some embodiments, the first signal indicates what kind of audio data the first device needs, so the encoder can encode the audio data according to the instruction of the first signal.
[0168] In some embodiments, the first signal is used to indicate parameters for processing the audio data. Alternatively, it can be understood that the first information is used to indicate information for processing the audio data. In some embodiments, the first signal is a signal returned by the first device. The first signal may also be referred to as a feedback signal, a first signal, a processed signal, etc., which is not limited in the present embodiment.
[0169] In some embodiments, the first signal includes encoding information for encoding the audio data. In the disclosed embodiment, the encoding information indicates how to encode the audio data to ensure that the encoded data meets the requirements.
[0170] In some embodiments, encoding the audio data based on the first signal to obtain the first encoded data includes performing at least one of the following on the audio data to obtain the first encoded data:
[0171] (1) Pretreatment;
[0172] In some embodiments, the first signal includes a first frequency band, the first frequency band being smaller than a second frequency band of the audio data, and the preprocessing includes filtering the audio data based on the first signal to obtain filtered data, where the frequency band of the filtered data is the first frequency band. Alternatively, it can be understood that the embodiments of the present disclosure refer to filtering the audio data using an instruction of the first signal during the bandwidth filtering stage of preprocessing to obtain filtered data of the first frequency band.
[0173] Optionally, bandwidth filtering refers to filtering the audio data through a high-pass filter or a band-pass filter to obtain filtered audio data. Bandwidth filtering can reduce the bandwidth of the audio data.
[0174] In some embodiments, pre-processing further comprises at least one of the following:
[0175] 1. Divide into multiple frames;
[0176] Optionally, framing refers to inputting continuous PCM data and a time-varying signal, and then splitting the audio data into multiple audio data of fixed duration using a window function signal. Optionally, the fixed duration can be a value such as 5ms (milliseconds), 10ms, 20ms, or other lengths, which are not limited in the embodiments of the present disclosure. In the embodiments of the present disclosure, framing can ensure that the characteristics of the split audio data are stable within the fixed duration.
[0177] 2. Time-frequency transformation;
[0178] Optionally, time-frequency transformation refers to converting audio data from the time domain to the frequency domain using MDCT or FFT methods. The disclosed embodiment can process signals in different frequency ranges separately, thereby improving processing efficiency.
[0179] 3. Time window adjustment;
[0180] Optionally, the time window adjustment refers to using a short window for transient audio data and a long window for non-transient audio data.
[0181] 4. Time domain noise shaping;
[0182] Optionally, time-domain noise shaping refers to inputting audio data with a shorter time domain or transient audio data, and then filtering the data using a time-domain noise shaping filter to obtain filtered audio data.
[0183] 5. Frequency domain noise shaping.
[0184] Optionally, frequency domain noise shaping refers to filtering frequency domain audio data using frequency domain noise shaping to obtain filtered audio data.
[0185] (2) Signal analysis;
[0186] In some embodiments, signal analysis refers to at least one of bandwidth detection, signal feature analysis, and signal classification. Optionally, signal analysis refers to at least one of bandwidth detection, feature analysis, and classification of audio data.
[0187] (3) Nuclear option;
[0188] In some embodiments, the first signal includes a first core module, and the core selection includes: a second core module selected based on the audio data and a third core module determined by the first core module, and the third core module is used to process the audio data. In the embodiment of the present disclosure, the second core module selected based on the audio data refers to the core module required for use after analyzing the audio data. In some embodiments, the type of audio data can be determined after signal analysis, and then the corresponding core module is determined according to the type of audio data. Optionally, the type of audio data and the core module are in a one-to-one correspondence.
[0189] In some embodiments, determining a third core module based on the second core module selected by the audio data and the first core module includes: if the first core module is identical to the second core module, and determining the third core module to be either the first core module or the second core module. In the disclosed embodiment, if the first core module is identical to the second core module, it means that the same core module is used regardless of which one is used, and therefore the third core module is determined to be either the first core module or the second core module.
[0190] In some embodiments, determining a third core module based on the second core module selected based on the audio data and the first core module includes: determining the third core module to be the first core module because the first core module is different from the second core module. In some embodiments, the priority of the core module indicated by the first signal is higher than the priority of the core module selected based on the audio data. Alternatively, it can also be understood that the priority of the core module indicated by the first signal is higher than the priority of the core module determined based on signal analysis.
[0191] In some embodiments, determining a third core module based on the second core module selected based on the audio data and the first core module includes: determining the third core module to be the second core module because the first core module is different from the second core module. In some embodiments, the priority of the core module indicated by the first signal is lower than the priority of the core module selected based on the audio data. Alternatively, it can also be understood that the priority of the core module indicated by the first signal is higher than the priority of the core module determined based on signal analysis.
[0192] In some embodiments, determining a third core module based on the second core module selected by the audio data and the first core module includes: if the first core module is different from the second core module, determining the third core module to be either the first core module or the second core module. In some embodiments, the priority of the core module indicated by the first signal is not differentiated from the priority of the core module selected based on the audio data, and one of the core modules determined by the two methods is randomly selected.
[0193] It should be noted that, in the embodiment of the present disclosure, when performing core selection, the encoder will encode the identifier of the core module indicated by the first signal, so as to facilitate transmission of the encoded identifier of the core module indicated by the first signal.
[0194] In some embodiments, based on the scene indicated by the first signal, encoding parameters corresponding to the first signal are determined, the encoding parameters are used to encode the audio data, and the encoding parameters are encoded to obtain the second encoded data. In the disclosed embodiments, the encoder stores encoding parameters corresponding to each scene. After the encoder receives the first signal, it can determine the encoding parameters corresponding to the first signal, and then the encoder encodes the encoding parameters of the first signal for subsequent transmission.
[0195] Optionally, the encoding parameter corresponding to the first signal includes an identifier of a core module corresponding to the core selection, and the encoder may encode the identifier of the core module to obtain encoded second encoded data.
[0196] (4) nuclear coding;
[0197] In some embodiments, kernel selection and kernel encoding refer to the process of selecting an encoding kernel and performing encoding using the selected encoding kernel.
[0198] (5) Multi-channel data mixing.
[0199] In some embodiments, mixing of multiple data channels refers to carrying all the data encoded by the above process into one data stream for transmission.
[0200] In some embodiments, the first signal is a fixed signal. In the disclosed embodiments, the first signal is pre-configured and will not be adjusted subsequently. In some embodiments, the first signal may also be referred to as a time-invariant signal. In some embodiments, the first signal is a signal that will not be adjusted.
[0201] It should be noted that the preconfiguration of the first signal in the embodiments of the present disclosure is completed during the initialization configuration of the encoder. In some embodiments, the initialization configuration of the encoder refers to configuring the encoder's parameters. In some embodiments, the encoder configuration includes at least one of the following parameters: input parameters; output parameters; encoding rate; sampling rate; starting frequency and ending frequency for encoding the first encoded data; and core module identifier. In some embodiments, the preconfiguration of the first signal can also be understood as configuring at least one of the parameters of the aforementioned encoder.
[0202] In some embodiments, based on the scene indicated by the first signal, encoding parameters corresponding to the first signal are determined, the encoding parameters are used to encode the audio data, and the encoding parameters are encoded to obtain the second encoded data. In the disclosed embodiments, the encoder stores encoding parameters corresponding to each scene. After the encoder receives the first signal, it can determine the encoding parameters corresponding to the first signal, and then the encoder encodes the encoding parameters of the first signal for subsequent transmission.
[0203] In some embodiments, after receiving the first signal, the encoder needs to first analyze the first signal, then determine the encoding parameters corresponding to the first signal, and then perform subsequent operations. Alternatively, it can be understood that the encoder needs to first convert the first signal to obtain the encoding parameters corresponding to the first signal.
[0204] In some embodiments, the steps in the above embodiment are shown in FIG2C , where audio data is input, and then preprocessing, signal analysis, kernel selection, kernel encoding, and stream mixing are performed on the audio data based on the first signal, and a data stream is output.
[0205] In some embodiments, when filtering audio data, referring to FIG. 2D , the first frequency band is 0 kHz-24 kHz, and the second frequency band obtained by filtering is 10 kHz-20 kHz.
[0206] In some embodiments, the method further includes: receiving a first signal, wherein the first signal supports real-time adjustment. In the disclosed embodiment, the first device may send the first signal to the encoder to update the first signal, thereby ensuring that encoding can be performed based on the updated first signal and ensuring encoding accuracy.
[0207] In some embodiments, the encoder configuration includes at least one of the following parameters:
[0208] Input parameters;
[0209] Output parameters;
[0210] Coding rate;
[0211] Sampling rate;
[0212] a starting frequency and an ending frequency for encoding the first coded data;
[0213] Core module identifier.
[0214] Step S2104: The encoder encodes the first signal to obtain second encoded data.
[0215] In the embodiment of the present disclosure, the encoder not only needs to encode the audio data and transmit the encoded data, but also needs to encode the first signal and transmit the second encoded data obtained by encoding.
[0216] In some embodiments, the second coded data is used for at least one of being transmitted together with the first coded data or being decoded by a decoder. In some embodiments, the first signal is encoded using Huffman coding to obtain the second coded data.
[0217] In some embodiments, based on the scene indicated by the first signal, encoding parameters corresponding to the first signal are determined, the encoding parameters are used to encode the audio data, and the encoding parameters are encoded to obtain the second encoded data. In the disclosed embodiments, the encoder stores encoding parameters corresponding to each scene. After the encoder receives the first signal, it can determine the encoding parameters corresponding to the first signal, and then the encoder encodes the encoding parameters of the first signal for subsequent transmission.
[0218] In some embodiments, the encoding parameters used when processing audio data may include the above-mentioned parameters during preprocessing, parameters during kernel selection, etc., which are not limited in the embodiments of the present disclosure.
[0219] In some embodiments, step S2104 in the embodiments of the present disclosure is an optional step, and in another embodiment, step S2104 may not be performed.
[0220] Step S2105: The encoder carries the first encoded data and the second encoded data in a data stream and sends it to the decoder.
[0221] In some embodiments, if the encoder encodes the audio data to obtain first encoded data and encodes the first signal to obtain second encoded data, the first encoded data and the second encoded data are carried in a data stream and sent to the decoder. In some embodiments, if the encoder encodes the first signal to obtain second encoded data, the second encoded data are carried in a data stream and sent to the decoder. In some embodiments, if the encoder encodes the audio data to obtain first encoded data, the first encoded data are carried in a data stream and sent to the decoder.
[0222] It should be noted that the embodiment of the present disclosure is described by taking the example of the encoder carrying the first encoded data and the second encoded data in the data stream. In another embodiment, the encoder may not perform step S2104, that is, the encoder does not encode the first signal. Correspondingly, the encoder carries the first encoded data and the first signal in the data stream and sends it to the decoder. Alternatively, the encoder may directly send the first signal. Alternatively, it can also be understood that the embodiment of the present disclosure does not perform step S2104, and step S2105 is replaced by: the encoder carries the first encoded data and the first signal in the data stream and sends it to the decoder.
[0223] Step S2106: The decoder decodes the first encoded data based on the second signal to obtain audio data.
[0224] In some embodiments, the second signal in the disclosed embodiments is formed by combining the first signal and parameters encoded by the encoder. In some embodiments, the second signal in the disclosed embodiments is a signal obtained by encoding and decoding the first signal. In some embodiments, the decoded second signal includes additional content compared to the first signal. In some embodiments, the second signal may include the same information as the first signal.
[0225] In some embodiments, a decoder receives a data stream comprising first coded data and second coded data. Accordingly, the decoder needs to decode the first coded data and the second coded data. In some embodiments, a decoder receives a data stream comprising first coded data and a first signal. The decoder then decodes the first coded data.
[0226] The method further includes: decoding the second coded data to obtain the second signal. In some embodiments, decoding the second coded data to obtain the first signal includes: decoding the second coded data using Huffman coding to obtain the second signal.
[0227] In some embodiments, the method further includes: receiving a second signal. Correspondingly, the terminal decodes the first encoded data based on the second signal.
[0228] In some embodiments, decoding the first encoded data based on the first signal includes:
[0229] Performing at least one of the following on the first encoded data to obtain the audio data includes:
[0230] Bitstream demultiplexing;
[0231] nuclear decoding;
[0232] Inverse transform.
[0233] It should be noted that the inverse transformation refers to the opposite process of the encoding in the above embodiment.
[0234] In some embodiments, the core decodes, including:
[0235] The core module determines that the second signal indicates to core decode the audio data.
[0236] In some embodiments, the inverse transform comprises:
[0237] Inverse time-frequency transform;
[0238] Inverse time domain noise shaping transform;
[0239] Inverse frequency-domain noise shaping transform.
[0240] It should be noted that the process of decoding the encoded data by the decoder in the embodiment of the present disclosure is the opposite process of the encoding process in the above embodiment.
[0241] In some embodiments, the names of information, etc. are not limited to the names described in the embodiments, and terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "symbol", "codeword", "codebook", "codeword", "codepoint", "bit", "data", "program", and "chip" can be used interchangeably.
[0242] In some embodiments, terms such as "uplink", "uplink", "physical uplink" can be interchangeable with each other, and terms such as "downlink", "downlink", "physical downlink" can be interchangeable with each other, and terms such as "side", "sidelink", "side communication", "sidelink communication", "direct connection", "direct link", "direct communication", "direct link communication" can be interchangeable with each other.
[0243] In some embodiments, "obtain", "get", "get", "receive", "transmit", "bidirectional transmission", "send and / or receive" can be interchangeable, and can be interpreted as receiving from other entities, obtaining from protocols, obtaining from higher layers, obtaining by self-processing, autonomous implementation, etc.
[0244] In some embodiments, terms such as "send", "transmit", "report", "download", "transmit", "bidirectional transmission", "send and / or receive" can be used interchangeably.
[0245] In some embodiments, terms such as "moment", "time point", "time", and "time position" can be replaced with each other, and terms such as "duration", "period", "time window", "window", and "time" can be replaced with each other.
[0246] In some embodiments, terms such as "certain", "preset", "preset", "setting", "indicated", "a certain", "any", and "first" can be interchangeable. "Specific A", "preset A", "preset A", "setting A", "indicated A", "a certain A", "any A", and "first A" can be interpreted as A pre-specified in a protocol, etc., or as A obtained through setting, configuration, or indication, etc., or as specific A, a certain A, any A, or first A, etc., but not limited to this.
[0247] The encoding and decoding method involved in the embodiment of the present disclosure may include at least one of steps S2101 to S2106. For example, step S2101 can be implemented as an independent embodiment, step S2102 can be implemented as an independent embodiment, step S2103 can be implemented as an independent embodiment, step S2104 can be implemented as an independent embodiment, step S2105 can be implemented as an independent embodiment, step S2106 can be implemented as an independent embodiment, step S2101 and step S2102 can be implemented as independent embodiments, step S2103 and step S2104 can be implemented as independent embodiments, step S2105 and step S2106 can be implemented as independent embodiments, step S2101, step S2102, step S2103, and step S2104 can be implemented as independent embodiments, step S2101, step S2102, step S2105, and step S2106 can be implemented as independent embodiments, but is not limited to this.
[0248] In some embodiments, step S2101 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0249] In some embodiments, step S2102 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0250] In some embodiments, step S2103 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0251] In some embodiments, step S2104 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0252] In some embodiments, step S2105 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0253] In some embodiments, step S2106 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0254] In some embodiments, step S2101 and step S2102 are optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0255] In some embodiments, step S2103 and step S2104 are optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0256] In some embodiments, step S215 and step S2106 are optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0257] In some embodiments, reference may be made to other optional implementations described before or after the description corresponding to FIG. 2A .
[0258] FIG3A is a flow chart of a coding and decoding method according to an embodiment of the present disclosure, which is applied to an encoder. As shown in FIG3A , the embodiment of the present disclosure relates to a coding and decoding method, which includes:
[0259] In step S3101 , the encoder encodes the audio data based on the first signal to obtain first encoded data.
[0260] The optional implementation of step S3101 can be found in step S2103 of FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.
[0261] Step S3102: The encoder encodes the first signal to obtain second encoded data.
[0262] Optional implementations of step S3102 may refer to step S2104 in FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.
[0263] Step S3103: The encoder carries the first encoded data and the second encoded data in a data stream and sends it to the decoder.
[0264] The optional implementation of step S3103 can be found in step S2105 of FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.
[0265] The encoding and decoding method involved in the embodiments of the present disclosure may include at least one of steps S3101 to S3103. For example, step S3101 may be implemented as an independent embodiment, step S3102 may be implemented as an independent embodiment, and step S3103 may be implemented as an independent embodiment, or at least two steps may be combined, but the present invention is not limited thereto.
[0266] In some embodiments, step S3101 is optional, step S3102 is optional, and step S3103 is optional, and one or more of these steps may be omitted or replaced in different embodiments, but the present invention is not limited thereto.
[0267] FIG3B is a flow chart of a coding and decoding method according to an embodiment of the present disclosure, which is applied to an encoder. As shown in FIG3B , the embodiment of the present disclosure relates to a coding and decoding method, which includes:
[0268] In step S3201 , the encoder encodes the audio data based on the first signal to obtain first encoded data.
[0269] Optional implementations of step S3201 may refer to step S2103 in FIG. 2A , step S3101 in FIG. 3A , and other related parts in the embodiments involved in FIG. 2A and FIG. 3A , which will not be described in detail here.
[0270] FIG4 is a flow chart of a coding and decoding method according to an embodiment of the present disclosure, which is applied to a decoder. As shown in FIG4 , the embodiment of the present disclosure relates to a coding and decoding method, which includes:
[0271] Step S4101: The decoder decodes the first encoded data based on the second signal to obtain audio data.
[0272] In some embodiments, the audio data is used for machine analysis, and the first signal includes encoding information for encoding the audio data.
[0273] Optional implementations of step S4101 may refer to step S2106 in FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.
[0274] FIG5 is a flow chart of a coding and decoding method according to an embodiment of the present disclosure. As shown in FIG6 , an embodiment of the present disclosure relates to a coding and decoding method, which includes:
[0275] Step S5101: The encoder encodes the audio data based on the first signal to obtain first encoded data.
[0276] In some embodiments, the audio data is used for machine analysis, and the first signal includes encoding information for encoding the audio data.
[0277] Step S5102: The decoder decodes the first encoded data based on the second signal to obtain audio data.
[0278] In some embodiments, the audio data is used for machine analysis, and the first signal includes encoding information for encoding the audio data.
[0279] In some embodiments, the above method may include the methods of the embodiments of the above communication system side, the second device side, the encoder side, the decoder side, etc., which will not be repeated here.
[0280] FIG6 is a flow chart of a coding and decoding method according to an embodiment of the present disclosure. As shown in FIG6 , the embodiment of the present disclosure relates to a coding and decoding method, and the method includes:
[0281] Step S6101: The encoder encodes the data to obtain encoded data.
[0282] In some embodiments, encoding includes pre-processing, encoding, and core selection of a feedback signal from a terminal, and the feedback signal is time-varying or time-invariant.
[0283] In some embodiments, the bandwidth range used by the terminal is included in the encoded feedback signal. The minimum frequency used by the terminal is 10 kHz, and the maximum frequency is 20 kHz. The entire encoded frequency range of the signal is 0 Hz-24 kHz, as shown in the figure below. After bandpass filtering based on the feedback signal, the encoded frequency range is reduced to only 10 kHz-20 kHz, significantly reducing the amount of encoded data.
[0284] In some embodiments, the coding core selection stage includes a feedback signal from the terminal behavior. In some embodiments, the coding feedback signal includes information that task1 is a terminal task. The coding core selection module determines the coding core based on the feedback signal and the results of the signal analysis module.
[0285] In some embodiments, among the multiple conditions of the core selection module, the feedback signal has the highest priority. The feedback signal has a higher priority than the signal classification priority output by the signal analysis.
[0286] In some embodiments, the feedback signal from the terminal may be time-varying or time-invariant.
[0287] In some embodiments, the configuration of the encoder is typically performed during the initialization phase. A command line format may be used, as shown below:
[0288] Input file name, output file name, bit rate: encoding rate (bps), input signal sampling rate (kHz), target frequency range. Where A is the start frequency and B is the end frequency, in Hz. Name of the encoding kernel.
[0289] In some embodiments, the decoder further performs decoding and demultiplexing of the bitstream to obtain the audio compression data and metadata information, and decodes the metadata information to obtain core selection information indicating the corresponding decoding core.
[0290] In some embodiments, the decoded data of the audio signal is then passed through a post-processing module to obtain a final decoded audio signal. The post-processing module performs one or more operations on the audio signal, such as inverse time-frequency transform, inverse time-domain noise shaping transform, and inverse frequency-domain noise shaping transform.
[0291] In the embodiments of the present disclosure, some or all of the steps and their optional implementations may be arbitrarily combined with some or all of the steps in other embodiments, or may be arbitrarily combined with the optional implementations of other embodiments.
[0292] The present disclosure also provides an apparatus for implementing any of the above methods. For example, a device is provided that includes units or modules for implementing each step performed by an encoder in any of the above methods. For another example, another device is provided that includes units or modules for implementing each step performed by a decoder in any of the above methods.
[0293] It should be understood that the division of the various units or modules in the above device is merely a division of logical functions. In actual implementation, they may be fully or partially integrated into a physical entity, or they may be physically separated. In addition, the units or modules in the device may be implemented in the form of a processor calling software: for example, the device includes a processor, the processor is connected to a memory, and the memory stores instructions. The processor calls the instructions stored in the memory to implement any of the above methods or implement the functions of the various units or modules of the above device, wherein the processor is, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory is a memory within the device or a memory outside the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits, and the functions of some or all of the units or modules can be realized by designing the hardware circuits. The above-mentioned hardware circuits can be understood as one or more processors; for example, in one implementation, the above-mentioned hardware circuit is an application-specific integrated circuit (ASIC), which realizes the functions of some or all of the above units or modules by designing the logical relationship of the components in the circuit; for example, in another implementation, the above-mentioned hardware circuit can be realized by a programmable logic device (PLD). Taking a field programmable gate array (FPGA) as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by configuring the configuration file, thereby realizing the functions of some or all of the above units or modules. All units or modules of the above devices can be realized in the form of software called by the processor, or in the form of hardware circuits, or in part by the form of software called by the processor, and the rest by hardware circuits.
[0294] In the embodiments of the present disclosure, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationship of the hardware circuit. The logical relationship of the above-mentioned hardware circuit is fixed or reconfigurable. For example, the processor is a hardware circuit implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and implementing the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), etc.
[0295] Figure 7A is a structural diagram of the encoding and decoding device proposed in an embodiment of the present disclosure. As shown in Figure 7A, the encoding and decoding device 7100 may include: at least one of a transceiver module 7101, a processing module 7102, etc. In some embodiments, the processing module 7102 is used to encode audio data based on a first signal to obtain first encoded data, and the audio data is used for machine analysis, and the first signal includes encoding information for encoding the audio data. Optionally, the above-mentioned transceiver module 7101 is used to execute at least one of the communication steps such as sending and / or receiving performed by the encoder in any of the above methods (such as step S2101 but not limited to this), which will not be repeated here. Optionally, the above-mentioned processing module is used to execute at least one of the other steps performed by the encoder in any of the above methods, which will not be repeated here.
[0296] Optionally, the processing module 7102 is used to execute at least one of the communication steps such as processing performed by the encoder in any of the above methods, which will not be repeated here.
[0297] FIG7B is a schematic diagram of the structure of the encoding and decoding device proposed in an embodiment of the present disclosure. As shown in FIG7B , the encoding and decoding device 7200 may include: at least one of a transceiver module 7201 and a processing module 7202. In some embodiments, the processing module 7202 is used to decode the first encoded data based on the second signal to obtain audio data, wherein the audio data is used for machine analysis, and the first signal includes encoding information for encoding the audio data. Optionally, the above-mentioned transceiver module is used to perform at least one of the communication steps such as sending and / or receiving performed by the decoder in any of the above methods, which will not be repeated here.
[0298] Optionally, the processing module 7202 is used to execute at least one of the communication steps such as processing performed by the decoder in any of the above methods, which will not be repeated here.
[0299] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module, and the transmitting module and the receiving module may be separate or integrated. Optionally, the transceiver module may be interchangeable with the transceiver.
[0300] In some embodiments, the processing module can be a single module or can include multiple submodules. Optionally, the multiple submodules respectively execute all or part of the steps required to be executed by the processing module. Optionally, the processing module can be interchangeable with the processor.
[0301] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module, and the transmitting module and the receiving module may be separate or integrated. Optionally, the transceiver module may be interchangeable with the transceiver.
[0302] In some embodiments, the processing module can be a single module or can include multiple submodules. Optionally, the multiple submodules respectively execute all or part of the steps required to be executed by the processing module. Optionally, the processing module can be interchangeable with the processor.
[0303] Figure 8A is a schematic diagram of the structure of a communication device 8100 proposed in an embodiment of the present disclosure. Communication device 8100 can be a second device, an encoder, a decoder, or a chip, a chip system, or a processor that supports the second device, encoder, or decoder in implementing any of the above methods. Communication device 8100 can be used to implement the methods described in the above method embodiments. For details, please refer to the description of the above method embodiments.
[0304] As shown in Figure 8A, communication device 8100 includes one or more processors 8101. Processor 8101 can be a general-purpose processor or a dedicated processor, such as a baseband processor or a central processing unit. The baseband processor can be used to process communication protocols and communication data, while the central processing unit can be used to control the codec device, execute programs, and process program data. Communication device 8100 is configured to perform any of the above methods.
[0305] In some embodiments, the communication device 8100 further includes one or more memories 8102 for storing instructions. Optionally, all or part of the memories 8102 may be located outside the communication device 8100.
[0306] In some embodiments, the communication device 8100 further includes one or more transceivers 8103. When the communication device 8100 includes one or more transceivers 8103, the transceiver 8103 performs at least one of the communication steps of sending and / or receiving in the above method.
[0307] In some embodiments, a transceiver may include a receiver and / or a transmitter. The receiver and transmitter may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, and transceiver circuit may be used interchangeably; the terms transmitter, transmitting unit, transmitter, and transmitting circuit may be used interchangeably; and the terms receiver, receiving unit, receiver, and receiving circuit may be used interchangeably.
[0308] In some embodiments, the communication device 8100 may include one or more interface circuits 8104. Optionally, the interface circuit 8104 is connected to the memory 8102. The interface circuit 8104 may be configured to receive signals from the memory 8102 or other devices, and may be configured to send signals to the memory 8102 or other devices. For example, the interface circuit 8104 may read instructions stored in the memory 8102 and send the instructions to the processor 8101.
[0309] The communication device 8100 described in the above embodiments may be a second device, a decoder, or an encoder, but the scope of the communication device 8100 described in the present disclosure is not limited thereto, and the structure of the communication device 8100 may not be limited by FIG. 8A. The communication device may be an independent device or may be part of a larger device. For example, the communication device may be: 1) an independent integrated circuit IC, or a chip, or a chip system or subsystem; (2) a collection of one or more ICs, optionally, the above IC collection may also include a storage component for storing data or programs; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, an encoder device, an intelligent encoder device, a cellular phone, a wireless device, a handheld device, a mobile unit, an in-vehicle device, a decoder, a cloud device, an artificial intelligence device, etc.; (6) others, etc.
[0310] FIG8B is a schematic diagram of the structure of a chip 8200 according to an embodiment of the present disclosure. If the communication device 8100 can be a chip or a chip system, please refer to the schematic diagram of the structure of the chip 8200 shown in FIG8B , but the present disclosure is not limited thereto.
[0311] The chip 8200 includes one or more processors 8201 , and the chip 8200 is configured to execute any of the above methods.
[0312] In some embodiments, the chip 8200 further includes one or more interface circuits 8202. Optionally, the interface circuit 8202 is connected to the memory 8203. The interface circuit 8202 can be used to receive signals from the memory 8203 or other devices, and can be used to send signals to the memory 8203 or other devices. For example, the interface circuit 8202 can read instructions stored in the memory 8203 and send the instructions to the processor 8201.
[0313] In some embodiments, the interface circuit 8202 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processor 8201 performs at least one of the other steps.
[0314] In some embodiments, terms such as interface circuit, interface, transceiver pin, and transceiver may be used interchangeably.
[0315] In some embodiments, the chip 8200 further includes one or more memories 8203 for storing instructions. Alternatively, all or part of the memories 8203 may be outside the chip 8200.
[0316] The present disclosure also proposes a storage medium having instructions stored thereon, which, when executed on the communication device 8100, causes the communication device 8100 to execute any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but is not limited thereto, and may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but is not limited thereto, and may also be a temporary storage medium.
[0317] The present disclosure also provides a program product, which, when executed by the communication device 8100, enables the communication device 8100 to perform any of the above methods. Optionally, the program product is a computer program product.
[0318] The present disclosure also proposes a computer program, which, when executed on a computer, causes the computer to perform any one of the above methods.
Claims
1. A coding method, characterized in that: The method is performed by an encoder, and includes: The audio data is encoded based on a first signal to obtain first encoded data, where the first signal includes encoding information for encoding the audio data. The first signal is fed back by a first device, and the audio data is used for machine analysis or machine training by the first device.
2. The method according to claim 1, characterized in that The method further comprises: The first signal is encoded to obtain second encoded data; the second encoded data is used for at least one of being sent together with the first encoded data or being decoded by a decoder.
3. The method according to claim 2, characterized in that The encoding of the first signal to obtain second encoded data includes: The first signal is encoded using Huffman coding to obtain the second encoded data.
4. The method according to claim 2, characterized in that The method further comprises: Determining, based on the scene indicated by the first signal, encoding parameters corresponding to the first signal, where the encoding parameters are used to encode the audio data; The encoding parameters are encoded to obtain the second encoded data.
5. The method according to claim 1, characterized in that The method further comprises: The first signal is sent.
6. The method according to any one of claims 1 to 5, characterized in that: The encoding of the audio data based on the first signal to obtain first encoded data includes: Perform at least one of the following on the audio data to obtain the first encoded data: Pretreatment; Signal analysis; Nuclear option; nuclear coding; Multi-channel data mixing.
7. The method according to claim 6, characterized in that The first signal includes a first frequency band, the first frequency band is smaller than a second frequency band of the audio data, and the preprocessing includes: The audio data is filtered based on the first signal to obtain filtered data, where a frequency band of the filtered data is the first frequency band.
8. The method according to claim 7, characterized in that The pre-processing further comprises at least one of the following: Divide into multiple frames; Time-frequency transformation; Time window adjustment; Time Domain Noise Shaping; Frequency Domain Noise Shaping.
9. The method according to claim 6, characterized in that The first signal includes a first core module Core Module, and the core selection includes: A third core module is determined based on the second core module selected based on the audio data and the first core module, and the third core module is used to process the audio data.
10. The method according to claim 9, characterized in that The determining of the third core module based on the second core module selected based on the audio data and the first core module includes: The first core module is the same as the second core module, and the third core module is determined to be either the first core module or the second core module.
11. The method according to claim 9, characterized in that The determining of the third core module based on the second core module selected based on the audio data and the first core module includes: The first core module is different from the second core module, and the third core module is determined to be the first core module.
12. The method according to any one of claims 1 to 11, characterized in that: The first signal is a fixed signal.
13. The method according to any one of claims 1 to 11, characterized in that The method further comprises: The first signal is received, where the first signal supports real-time adjustment.
14. The method according to any one of claims 1 to 13, characterized in that: The configuration of the encoder includes at least one of the following parameters: Input parameters; Output parameters; Coding rate; Sampling rate; a starting frequency and an ending frequency for encoding the first coded data; Core module identifier.
15. The method according to any one of claims 1 to 14, characterized in that: The method further comprises: A data stream comprising the first encoded data is sent to a decoder.
16. A decoding method, characterized in that: The method is performed by a decoder, and comprises: The first encoded data is decoded based on the second signal to obtain audio data, and the audio data is used for machine analysis. The first signal includes encoding information for encoding the audio data. The first signal is fed back by the first device, and the audio data is used for machine analysis or machine training by the first device.
17. The method according to claim 16, characterized in that The method further comprises: The second encoded data is decoded to obtain the second signal.
18. The method according to claim 17, characterized in that The decoding the second coded data to obtain the first signal includes: The second coded data is decoded using Huffman coding to obtain the second signal.
19. The method according to claim 17, wherein The decoding of the second coded data to obtain the second signal includes: The second encoded data is decoded to obtain decoding parameters, where the decoding parameters are used to decode the first encoded data.
20. The method according to claim 16, wherein The method further comprises: The second signal is received.
21. The method according to any one of claims 16 to 20, characterized in that The decoding of the first coded data based on the first signal includes: Performing at least one of the following on the first encoded data to obtain the audio data includes: Bitstream demultiplexing; nuclear decoding; Inverse transform.
22. The method according to claim 21, characterized in that The core decoding includes: The core module determined by the second signal instructs core decoding of the audio data.
23. The method according to claim 21, characterized in that The inverse transformation comprises: Inverse time-frequency transform; Inverse time domain noise shaping transform; Inverse frequency-domain noise shaping transform.
24. A coding and decoding method, characterized in that: The method comprises: The encoder encodes the audio data based on the first signal to obtain first encoded data, where the audio data is used for machine analysis. The first signal includes encoding information for encoding the audio data. The first signal is fed back by the first device. The audio data is used for machine analysis or machine training by the first device. The decoder decodes the first encoded data based on the second signal to obtain audio data, where the audio data is used for machine analysis. The first signal includes encoding information for encoding the audio data.
25. An encoding device, characterized in that The encoding device comprises: A processing module is used to encode audio data based on a first signal to obtain first encoded data, where the audio data is used for machine analysis. The first signal includes encoding information for encoding the audio data. The first signal is fed back by a first device, and the audio data is used for machine analysis or machine training by the first device.
26. A decoding device, characterized in that: The decoding device comprises: A processing module is used to decode the first encoded data using a second signal to obtain audio data, where the audio data is used for machine analysis. The first signal includes encoding information for encoding the audio data. The first signal is fed back by a first device, and the audio data is used for machine analysis or machine training by the first device.
27. An encoding device, characterized in that The encoding device comprises: one or more processors; The processor is configured to execute the encoding and decoding method according to any one of claims 1 to 15.
28. A decoding device, characterized in that: The decoding device comprises: one or more processors; The processor is configured to execute the encoding and decoding method according to any one of claims 16 to 23.
29. A coding and decoding system, characterized in that: The invention comprises an encoder and a decoder, wherein the encoder is configured to implement the encoding method according to any one of claims 1 to 15, and the encoder is configured to implement the decoding method according to any one of claims 16 to 23.
30. A storage medium storing instructions, characterized in that: When the instruction is executed on a communication device, the communication device is caused to execute the encoding method according to any one of claims 1 to 15, or execute the decoding method according to any one of claims 16 to 23.
Citation Information
Patent Citations
Audio transmission method and electronic equipment
CN113314133A
Audio signal processing method and device, equipment and storage medium
CN116348952A
Audio coding method and device, storage medium and computer equipment
CN116580716A
Coding and decoding method and device and storage medium
CN118160286A
Method and apparatus for controlling enhancement of low-bitrate coded audio
WO2020047298A1