Encoding, decoding and encoding-decoding methods, encoding apparatus, decoding apparatus and storage medium
By adjusting the band range of audio data to match the encoding frequency, the problem of the band range not matching the encoder is solved, and more accurate encoding and stable data transmission is achieved, suitable for machine analysis and training.
Patent Information
- Application Number
- PCT/CN2024/075045
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-31
- Publication Date
- 2025-08-07
AI Technical Summary
The frequency band range of the audio data does not match the encoder's encoding frequency, resulting in inaccurate encoding effect.
By adjusting the band range of the audio data to match it with the encoded frequency, the band range of the audio data is moved to the second band range by using the encoder and encoded based on the encoding bandwidth, and the decoder moves the band range of the encoded data back to the original band range.
Improves encoding accuracy and data transmission stability, ensuring that audio data can be accurately decoded in machine analysis and machine training.
Smart Images

Figure CN2024075045_07082025_PF_FP_ABST
Abstract
Description
Coding and decoding method, device and storage medium Technical Field
[0001] The present disclosure relates to the field of communication technologies, and in particular to a coding and decoding method, device, and storage medium. Background Art
[0002] With the rapid development of multimedia technology, audio data can be encoded and transmitted, and then the encoded data can be decoded to obtain the corresponding audio data, ensuring that the audio data can be transmitted efficiently.
[0003] Summary of the Invention
[0004] The present disclosure solves the problem of mismatch between the frequency band range of audio data and the encoding frequency of the encoder. By adjusting the frequency band range of the audio data so that the adjusted frequency band range matches the encoding frequency for encoding the second audio data, the accuracy of the encoding of the audio data is ensured and the encoding effect is improved.
[0005] The embodiments of the present disclosure provide a coding and decoding method, apparatus, and storage medium.
[0006] According to a first aspect of an embodiment of the present disclosure, a coding method is proposed. The method is performed by an encoder, and the method includes:
[0007] Acquire first audio data, and move a first frequency band of the first audio data to a second frequency band to obtain second audio data, wherein the first audio data is used for machine analysis / machine training, the second frequency band partially overlaps with or does not overlap with the first frequency band, and the second frequency band is determined by an available encoding bandwidth;
[0008] The second audio data is encoded based on a coding bandwidth corresponding to the second audio data to obtain first encoded data.
[0009] According to a second aspect of an embodiment of the present disclosure, a decoding method is proposed, where the method is performed by a decoder and includes:
[0010] Decoding the first encoded data to obtain second audio data;
[0011] The second frequency band range of the second audio data is moved to the first frequency band range to obtain first audio data, where the first audio data is used for machine analysis / machine training, the second frequency band range partially overlaps or does not overlap with the first frequency band range, and the second frequency band range is determined by the available encoding bandwidth.
[0012] According to a third aspect of the embodiments of the present disclosure, a coding and decoding method is proposed, the method comprising:
[0013] An encoder acquires first audio data, and moves a first frequency band of the first audio data to a second frequency band to obtain second audio data, wherein the first audio data is used for machine analysis / machine training, and the second frequency band partially overlaps or does not overlap with the first frequency band, and the second frequency band is determined by an available encoding bandwidth;
[0014] The encoder encodes the second audio data based on a coding bandwidth corresponding to the second audio data to obtain first coded data;
[0015] The decoder decodes the first encoded data to obtain second audio data;
[0016] The decoder moves the second frequency band range of the second audio data to the first frequency band range to obtain first audio data.
[0017] According to a fourth aspect of the embodiments of the present disclosure, an encoding device is provided, including:
[0018] a processing module, configured to obtain first audio data, and shift a first frequency band of the first audio data to a second frequency band to obtain second audio data, wherein the first audio data is used for machine analysis / machine training, the second frequency band partially overlaps or does not overlap with the first frequency band, and the second frequency band is determined by an available encoding bandwidth;
[0019] The processing module is configured to encode the second audio data based on an encoding bandwidth corresponding to the second audio data to obtain first encoded data.
[0020] According to a fifth aspect of the embodiments of the present disclosure, a decoding device is provided, including:
[0021] a processing module, configured to decode the first encoded data to obtain second audio data;
[0022] The processing module is used to move the second frequency band range of the second audio data to the first frequency band range to obtain first audio data, where the first audio data is used for machine analysis / machine training, the second frequency band range partially overlaps or does not overlap with the first frequency band range, and the second frequency band range is determined by the available encoding bandwidth.
[0023] According to a sixth aspect of the embodiments of the present disclosure, an encoding device is provided, including:
[0024] one or more processors;
[0025] The encoding and decoding device is used to execute any one of the methods described in the first aspect.
[0026] According to a seventh aspect of the embodiments of the present disclosure, a decoding device is provided, including:
[0027] one or more processors;
[0028] The encoding and decoding device is used to execute any method described in the second aspect.
[0029] According to an eighth aspect of the embodiments of the present disclosure, a coding and decoding system is proposed, including:
[0030] An encoder and a decoder, wherein the encoder is configured to implement the encoding and decoding method described in the first aspect, and the decoder is configured to implement the decoding method described in the second aspect.
[0031] According to a ninth aspect of an embodiment of the present disclosure, a storage medium is proposed, wherein the storage medium stores instructions. When the instructions are executed on a communication device, the communication device executes a method as described in any one of the first aspect or the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The drawings described herein are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the present disclosure. The illustrative embodiments of the embodiments of the present disclosure and their descriptions are used to explain the embodiments of the present disclosure and do not constitute an improper limitation on the embodiments of the present disclosure. In the drawings:
[0033] FIG1 is a schematic diagram of the architecture of a coding and decoding system according to an embodiment of the present disclosure;
[0034] FIG2A is an interactive schematic diagram illustrating a coding and decoding method according to an embodiment of the present disclosure;
[0035] FIG2B is a schematic diagram showing a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0036] FIG2C is a schematic diagram showing a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0037] FIG2D is a schematic diagram illustrating a coding and decoding method according to an embodiment of the present disclosure;
[0038] FIG2E is a schematic diagram illustrating a coding and decoding method according to an embodiment of the present disclosure;
[0039] FIG2F is a schematic diagram illustrating a coding and decoding method according to an embodiment of the present disclosure;
[0040] FIG3A is a schematic diagram showing a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0041] FIG3B is a schematic diagram showing a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0042] FIG4 is a schematic diagram of a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0043] FIG5 is a schematic diagram of a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0044] FIG6 is a schematic diagram of a flow chart of a coding and decoding method according to an embodiment of the present disclosure;
[0045] FIG7A is a schematic structural diagram of a coding and decoding device proposed in an embodiment of the present disclosure;
[0046] FIG7B is a schematic structural diagram of a coding and decoding device proposed in an embodiment of the present disclosure;
[0047] FIG8A is a schematic structural diagram of a communication device proposed in an embodiment of the present disclosure;
[0048] FIG8B is a schematic diagram of the structure of the chip proposed in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0049] The present disclosure provides a coding and decoding method, device, and storage medium.
[0050] According to a first aspect of an embodiment of the present disclosure, a coding method is proposed. The method is performed by an encoder, and the method includes:
[0051] Acquire first audio data, and move a first frequency band of the first audio data to a second frequency band to obtain second audio data, wherein the first audio data is used for machine analysis / machine training, the second frequency band partially overlaps with or does not overlap with the first frequency band, and the second frequency band is determined by an available encoding bandwidth;
[0052] The second audio data is encoded based on a coding bandwidth corresponding to the second audio data to obtain first encoded data.
[0053] In the above embodiment, the problem of mismatch between the frequency band range of audio data and the encoding frequency of the encoder is solved. By adjusting the frequency band range of the audio data so that the adjusted frequency band range matches the encoding frequency for encoding the second audio data, the accuracy of encoding the audio data is ensured and the encoding effect is improved.
[0054] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0055] The first encoded data is sent to a decoder, where the first encoded data is used for decoding by the decoder.
[0056] In the above embodiment, after the encoder completes encoding, it sends the first encoded data obtained by encoding to the decoder, and then the decoder decodes the first encoded data so that the decoder can obtain audio data and ensure the stability of data transmission.
[0057] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0058] generating first information, wherein the first information is used to describe parameters for shifting the first audio data;
[0059] The first information is encoded to obtain second encoded data.
[0060] In the above embodiment, since the frequency of the first audio data is shifted, the parameters of the shift are recorded through the generated first information, thereby ensuring the accuracy of the recorded parameters and further ensuring the accuracy of subsequent decoding.
[0061] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0062] The second encoded data is sent to a decoder, where the second encoded data is used for decoding by the decoder.
[0063] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0064] downsampling the second audio data to obtain third audio data;
[0065] The encoding of the second audio data to obtain first encoded data includes:
[0066] The third audio data is encoded to obtain the first encoded data.
[0067] In the above embodiment, by downsampling the audio data, the sampling rate of the data can be reduced, thereby reducing the amount of data and ensuring the accuracy of data processing.
[0068] In conjunction with some embodiments of the first aspect, in some embodiments, downsampling the second audio data to obtain the third audio data includes:
[0069] The second audio data is downsampled according to a sampling rate that is a multiple of the encoding rate to obtain the third audio data.
[0070] In the above embodiment, the audio data is downsampled at a sampling rate that matches the encoding rate, thereby ensuring that the downsampled data meets the encoding requirements, thereby ensuring the accuracy of the encoding.
[0071] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0072] generating second information, where the second information is used to describe a sampling bandwidth of the downsampling;
[0073] The second information is encoded to obtain third encoded data.
[0074] In the above embodiment, since the audio data is sampled, the sampling frequency is recorded by the generated second information, thereby ensuring the accuracy of the recorded parameters and further ensuring the accuracy of subsequent decoding.
[0075] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0076] The third encoded data is sent to a decoder, and the third encoded data is used for decoding by the decoder.
[0077] In conjunction with some embodiments of the first aspect, in some embodiments, the first audio data includes a plurality of data, and encoding the second audio data to obtain the first encoded data includes:
[0078] performing multi-channel redundancy removal processing on the plurality of second audio data;
[0079] The plurality of second audio data after multi-channel redundancy removal processing is encoded to obtain the first encoded data.
[0080] In the above embodiment, since there are multi-channel audio data, there is redundancy between the multi-channel audio data. Therefore, by performing de-redundancy processing on the multi-channel audio data, the redundancy between the audio data between channels can be eliminated, and the accuracy of encoding the data between multiple channels can be guaranteed.
[0081] In a second aspect, an embodiment of the present disclosure provides a decoding method, which is performed by an encoder and includes:
[0082] Decoding the first encoded data to obtain second audio data;
[0083] The second frequency band range of the second audio data is moved to the first frequency band range to obtain first audio data, where the first audio data is used for machine analysis / machine training, the second frequency band range partially overlaps or does not overlap with the first frequency band range, and the second frequency band range is determined by the available encoding bandwidth.
[0084] In conjunction with some embodiments of the second aspect, in some embodiments, the method further includes:
[0085] The first encoded data sent by the encoder is received, where the first encoded data is used for decoding by the decoder.
[0086] In conjunction with some embodiments of the second aspect, in some embodiments, the method further includes:
[0087] receiving second encoded data sent by the encoder;
[0088] The second encoded data is decoded to obtain first information, where the first information is used to describe the first frequency band range shift and the second frequency band range.
[0089] In conjunction with some embodiments of the second aspect, in some embodiments, decoding the first encoded data to obtain the second audio data includes:
[0090] decoding the first encoded data to obtain third audio data;
[0091] The third audio data is up-sampled to obtain the second audio data.
[0092] In conjunction with some embodiments of the second aspect, in some embodiments, the method further includes:
[0093] receiving third encoded data sent by the encoder;
[0094] The third encoded data is decoded to obtain second information, where the second information is used to describe a sampling bandwidth of the downsampling.
[0095] In conjunction with some embodiments of the second aspect, in some embodiments, the second audio data includes a plurality of data, and decoding the first encoded data to obtain the second audio data includes:
[0096] performing multi-channel processing on the plurality of first encoded data;
[0097] The plurality of first encoded data after multi-channel processing is decoded to obtain the second audio data.
[0098] In a third aspect, an embodiment of the present disclosure provides a coding and decoding method, the method comprising:
[0099] The encoder moves a first frequency band range of the first audio data to a second frequency band range to obtain second audio data, where the first audio data is used for machine analysis / machine training, the second frequency band range partially overlaps with or does not overlap with the first frequency band range, and the second frequency band range is determined by an available encoding bandwidth;
[0100] The encoder encodes the second audio data based on a coding bandwidth corresponding to the second audio data to obtain first coded data;
[0101] The decoder decodes the first encoded data to obtain second audio data;
[0102] The decoder moves the second frequency band range of the second audio data to the first frequency band range to obtain first audio data.
[0103] In a fourth aspect, an embodiment of the present disclosure provides an encoding device, wherein the encoding and decoding device includes at least one of a transceiver module and a processing module; wherein the first device is used to execute the optional implementation methods of the first and third aspects.
[0104] In a fifth aspect, an embodiment of the present disclosure provides a decoding device, wherein the encoding and decoding device includes at least one of a transceiver module and a processing module; wherein the encoder is used to execute the optional implementation methods of the second and third aspects.
[0105] In a sixth aspect, an embodiment of the present disclosure provides an encoding device, including:
[0106] one or more processors;
[0107] The encoding and decoding device is used to execute the method described in any one of the first and third aspects.
[0108] In a seventh aspect, an embodiment of the present disclosure provides a decoding device, including:
[0109] one or more processors;
[0110] The encoding and decoding device is used to execute the method described in any one of the second and third aspects.
[0111] In an eighth aspect, an embodiment of the present disclosure provides a storage medium storing first information. When the first information is run on a communication device, the communication device executes a method as described in any one of the first, second and third aspects.
[0112] In a ninth aspect, an embodiment of the present disclosure proposes a program product. When the program product is executed by a communication device, the communication device executes any one of the methods described in the first, second and third aspects.
[0113] In a tenth aspect, an embodiment of the present disclosure proposes a computer program, which, when executed on a communication device, enables the communication device to execute any one of the methods described in the first, second, and third aspects.
[0114] In an eleventh aspect, an embodiment of the present disclosure provides a chip or a chip system, wherein the chip or chip system includes a processing circuit configured to execute any one of the methods described in the first, second, and third aspects.
[0115] It is understandable that the above-mentioned encoder, storage medium, program product, computer program, chip or chip system are all used to execute the method proposed in the embodiment of the present disclosure. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding method and will not be repeated here.
[0116] The present disclosure provides a coding method, apparatus, and storage medium. In some embodiments, the terms coding method, information coding method, and coding method are interchangeable; the terms coding apparatus, information coding apparatus, and indicating apparatus are interchangeable; and the terms information processing system, coding system, and so on are interchangeable.
[0117] The embodiments of the present disclosure are not exhaustive and are merely illustrative of some embodiments, and are not intended to be a specific limitation on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be arbitrarily exchanged. In addition, the optional implementation methods in a certain embodiment can be arbitrarily combined; in addition, the embodiments can be arbitrarily combined. For example, some or all steps of different embodiments can be arbitrarily combined, and a certain embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.
[0118] In each embodiment of the present disclosure, unless otherwise specified or provided for by logic, the terms and / or descriptions between the embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form a new embodiment based on their inherent logical relationships.
[0119] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.
[0120] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular, such as "a", "an", "the", "above", "said", "the", "the", etc., may mean "one and only one", or "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English in translation, the noun following the article may be understood as a singular expression or a plural expression.
[0121] In the embodiments of the present disclosure, “plurality” refers to two or more.
[0122] In some embodiments, the terms "at least one," "one or more," "a plurality of," "multiple," etc. may be used interchangeably.
[0123] In some embodiments, descriptions such as "at least one of A and B," "A and / or B," "A in one case, B in another case," or "in response to one case A, in response to another case B" may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed); and in some embodiments, A and B (both A and B are executed). The above is also applicable when there are more branches such as A, B, and C.
[0124] In some embodiments, "A or B" and other descriptions may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed). The above is also applicable when there are more branches such as A, B, C, etc.
[0125] The prefixes such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different description objects and do not constitute any restriction on the position, order, priority, quantity or content of the description objects. For the statement of the description object, please refer to the description in the context of the claims or embodiments, and no unnecessary restriction should be constituted due to the use of prefixes. For example, if the description object is a "field", the ordinal number before the "field" in the "first field" and the "second field" does not limit the position or order between the "fields". "First" and "second" do not limit whether the "fields" they modify are in the same message, nor do they limit the order of the "first field" and the "second field". For another example, if the description object is a "level", the ordinal number before the "level" in the "first level" and the "second level" does not limit the priority between the "levels". For another example, the number of description objects is not limited by the ordinal number and can be one or more. Taking "first device" as an example, the number of "devices" can be one or more. In addition, the objects modified by different prefixes can be the same or different. For example, if the description object is "device", then the "first device" and the "second device" can be the same device or different devices, and their types can be the same or different; for another example, if the description object is "information", then the "first information" and the "second information" can be the same information or different information, and their contents can be the same or different.
[0126] In some embodiments, “including A,” “comprising A,” “used to indicate A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.
[0127] In some embodiments, terms such as "time / frequency" and "time / frequency domain" refer to the time domain and / or the frequency domain.
[0128] In some embodiments, terms such as "in response to...", "in response to determining...", "in the case of...", "at the time of...", "when...", "if...", "if...", etc. can be used interchangeably.
[0129] In some embodiments, terms such as "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not less than", and "above" can be replaced with each other, and terms such as "less than", "less than or equal to", "not greater than", "less than", "less than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", and "below" can be replaced with each other.
[0130] In some embodiments, devices and equipment can be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. In some cases, they can also be understood as "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "subject", etc.
[0131] In some embodiments, "network" can be interpreted as devices included in the network, such as access network equipment, core network equipment, etc.
[0132] In some embodiments, "access network device (AN device)" may also be referred to as "radio access network device (RAN device)", "base station (BS)", "radio base station", "fixed station", and in some embodiments may also be understood as "node", "access point", "transmission point (TP)", "reception point (RP)", "transmission and / or reception point (TRP)" "panel", "antenna panel", "antenna array", "cell", "macro cell", "small cell", "femto cell", "pico cell", "sector", "cell group", "serving cell", "carrier", "component carrier", "bandwidth part (BWP)", etc.
[0133] In some embodiments, "encoder (terminal)" or "encoder device (terminal device)" may be referred to as "user equipment (encoder)", "user terminal" "mobile station (MS)", "mobile terminal (MT)", subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, etc.
[0134] In some embodiments, obtaining data, information, etc. may comply with the laws and regulations of the country where the data is obtained.
[0135] In some embodiments, data, information, etc. may be obtained with the user's consent.
[0136] In addition, each element, each row, or each column in the table of the embodiment of the present disclosure can be implemented as an independent embodiment, and the combination of any elements, any rows, and any columns can also be implemented as an independent embodiment.
[0137] FIG1 is a schematic diagram of the architecture of a codec system according to an embodiment of the present disclosure. As shown in FIG1 , the method provided in the embodiment of the present disclosure can be applied to a codec system 100, which can include an encoder 101 and a decoder 102. It should be noted that the codec system 100 can also include other devices, and the present disclosure does not limit the devices included in the codec system 100.
[0138] In some embodiments, the encoder 101 and the decoder 102 are both provided in a terminal. In some embodiments, the terminal can be various devices. For example, the terminal can be a mobile phone, a wearable device, an Internet of Things device, a car with communication functions, a smart car, a tablet computer, a computer with wireless transceiver functions, a virtual reality (VR) encoder device, an augmented reality (AR) encoder device, a wireless encoder device in industrial control, a wireless encoder device in self-driving, a wireless encoder device in remote medical surgery, a wireless encoder device in a smart grid, a wireless encoder device in transportation safety, a wireless encoder device in a smart city, and a wireless encoder device in a smart home, but is not limited thereto.
[0139] It can be understood that the encoding and decoding system described in the embodiment of the present disclosure is for the purpose of more clearly illustrating the technical solution of the embodiment of the present disclosure, and does not constitute a limitation on the technical solution proposed in the embodiment of the present disclosure. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution proposed in the embodiment of the present disclosure is also applicable to similar technical problems.
[0140] The following embodiments of the present disclosure may be applied to the codec system 100 shown in FIG1 , or a portion thereof, but are not limited thereto. The entities shown in FIG1 are illustrative only. The codec system may include all or part of the entities shown in FIG1 , or may include other entities outside of FIG1 . The number and form of the entities are arbitrary, and the entities may be physical or virtual. The connection relationship between the entities is illustrative only. The entities may be connected or disconnected, and the connection may be in any manner, including direct or indirect, wired or wireless.
[0141] The embodiments of the present disclosure can be applied to Long Term Evolution (LTE), LTE-Advanced (LTE-A), LTE-Beyond (LTE-B), SUPER 3G, IMT-Advanced, 4th generation mobile communication system (4G), 5th generation mobile communication system (5G), 5G New Radio (NR), Future Radio Access (FRA), New-Radio Access Technology (RAT), New Radio (NR), New Radio Access (NX), Future Generation Radio Access (FX), Global System for Mobile Communications (GSM (registered trademark)), CDMA2000, Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, Ultra-WideBand (UWB), Bluetooth (registered trademark), Public Land Mobile Network (PLMN) networks, Device-to-Device (D2D) systems, Machine-to-Machine (M2M) systems, Internet of Things (IoT) systems, Vehicle-to-Everything (V2X), systems using other coding and decoding methods, and next-generation systems based on these. In addition, multiple systems can also be combined (for example, a combination of LTE or LTE-A with 5G, etc.) for application.
[0142] In some embodiments, the present disclosure further includes a first device, or may also be referred to as an acquisition device. Optionally, the first device is used to acquire audio data and generate first data of the audio data, and the audio data is described by the first data. Optionally, the first device is an audio sensor device.
[0143] FIG2A is an interactive diagram of a coding and decoding method according to an embodiment of the present disclosure. As shown in FIG2A , the embodiment of the present disclosure relates to a coding and decoding method, which includes:
[0144] Step S2101: The first device collects first audio data.
[0145] In some embodiments, the first device refers to a device with a collection function. For example, the first device may also be referred to as a collection device, a collection terminal, etc., which is not limited in the embodiments of the present disclosure. In some embodiments, the first audio data refers to audio, or it can also be understood as sound collected by the first device. In some embodiments, the first audio data is used to indicate the audio generated by the device during operation.
[0146] In some embodiments, the first audio data is used for machine analysis or machine training. Alternatively, it can be understood that machine analysis or machine training of the first audio data can determine whether the operating status of the device indicated by the first audio data is normal. Alternatively, it can be understood that the first audio data indicates the operating status of the device, and whether the device is operating normally can be determined based on the first audio data.
[0147] In some embodiments, the embodiments of the present disclosure are applied to automated production and inspection scenarios, where the first device collects first audio data of the produced product (component), and then performs machine analysis or machine training on the collected first audio data to determine whether the generated product (component) is qualified.
[0148] It should be noted that the embodiment of the present disclosure uses the role of the first audio data as an example for explanation. Furthermore, after the decoder decodes and obtains the first audio data in the following embodiment, the second device analyzes the first audio data obtained by the decoder to determine whether the object monitored by the first device is operating normally. In some embodiments, the object can be understood as the first device, the product produced, or other information that the first audio data refers to, and the embodiment of the present disclosure is not limited thereto. In some embodiments, the second device is used to control the first device, and it can be understood that the second device is a controller of the first device.
[0149] In some embodiments, the embodiments of the present disclosure are applied in product testing, see Figure 2B, the first audio data of the product is collected by a first device, the first audio data is encoded by an encoder, the encoded data is sent to a decoder, and the decoder decodes the encoded data to determine whether the product corresponding to each piece of first audio data is qualified.
[0150] Step S2102: The first device sends first audio data to the encoder.
[0151] In some embodiments, the encoder receives first audio data sent by the first device. In some embodiments, the first device sends the first audio data. In some embodiments, the encoder receives the first audio data.
[0152] In some embodiments, after the encoder receives the first audio data sent by the first device, the encoder obtains the first audio data.
[0153] In step S2103 , the encoder moves the first frequency band of the first audio data to a second frequency band to obtain second audio data.
[0154] In some embodiments, the second frequency band partially overlaps or does not overlap with the first frequency band. In some embodiments, the second frequency band partially overlaps with the first frequency band. In some embodiments, the second frequency band does not overlap with the first frequency band. Optionally, if the second frequency band does not overlap with the first frequency band, the maximum frequency of the second frequency band is less than the minimum frequency of the first frequency band. Optionally, if the second frequency band overlaps with the first frequency band, the maximum frequency of the second frequency band is less than the maximum frequency of the first frequency band, and the maximum frequency of the second frequency band is greater than the minimum frequency of the first frequency band. In some embodiments, the second frequency band is determined by the available encoding bandwidth. Optionally, the available encoding bandwidth refers to the bandwidth that can be used to encode the audio data. Optionally, the available encoding bandwidth refers to the effective frequency range of the audio data when encoding the audio data. Optionally, the available encoding bandwidth is related to the encoding frequency when encoding the audio data. Optionally, the available encoding bandwidth is no greater than half the encoding frequency when encoding the audio data. In some embodiments, the second frequency band matches the encoding frequency used to encode the second audio data.
[0155] For example, if the first frequency band is 30kHz (kilohertz)-40kHz, the first frequency band can be moved to a second frequency band, and the second frequency band is 0kHz-10kHz. Alternatively, if the first frequency band is 30kHz (kilohertz)-40kHz, the first frequency band can be moved to a second frequency band, and the second frequency band is 25kHz-35kHz. It should be noted that the embodiments of the present disclosure only illustrate the frequency bands by way of example. Other frequency bands may also be used, and the embodiments of the present disclosure do not limit them.
[0156] In some embodiments, the encoding frequency used by the encoder to encode the second audio data is used to indicate the frequency used by the encoder when encoding the second audio data. In some embodiments, the encoding frequency used by the encoder to encode the second audio data can also be understood as the frequency used by the core module in the encoder when performing core encoding on the second audio data.
[0157] It should be noted that the encoder in the embodiments of the present disclosure may be preconfigured with a preset encoding frequency, wherein the encoding frequency used by the encoder to encode the second audio data may be different from the preset encoding frequency. Alternatively, it can be understood that when encoding the second audio data, the encoder adjusts the preset encoding frequency and uses the adjusted encoding frequency for encoding.
[0158] It should be noted that, in the embodiment of the present disclosure, the encoder adjusting the frequency band of the audio data is also part of encoding the audio data. Optionally, in the embodiment of the present disclosure, adjusting the frequency band of the audio data can also be understood as part of preprocessing in encoding.
[0159] In some embodiments, an encoder generates first information describing parameters for shifting first audio data, and encodes the first information to produce second encoded data. In embodiments of the present disclosure, the encoder also encodes the frequency band range of the shift and transmits it to the decoder, so that the decoder can decode the encoded data based on the frequency band range shifted by the encoder.
[0160] In some embodiments, the encoder sends the second encoded data to the decoder, and the second encoded data is used for decoding by the decoder.
[0161] In some embodiments, encoding the audio data by the encoder may include at least one of the following:
[0162] (1) Pretreatment;
[0163] In some embodiments, the preprocessing includes the above step S2103. In addition, the preprocessing includes at least one of the following:
[0164] 1. Divide into multiple frames;
[0165] Optionally, framing refers to inputting continuous PCM (Pulse Code Modulation) data and a time-varying signal, and then splitting the audio data into multiple audio data of fixed duration using a window function signal. Optionally, the fixed duration can be a value such as 5ms (milliseconds), 10ms, 20ms, or other lengths, which are not limited in the embodiments of the present disclosure. In the embodiments of the present disclosure, framing can ensure that the characteristics of the split audio data are stable within the fixed duration.
[0166] 2. Time-frequency transformation;
[0167] Optionally, time-frequency transform refers to converting audio data from the time domain to the frequency domain using MDCT (Modified Discrete Cosine Transform) or FFT (Fast Fourier Transform). The disclosed embodiments can process signals in different frequency bands separately, thereby increasing processing efficiency.
[0168] 3. Time window adjustment;
[0169] Optionally, the time window adjustment refers to using a short window for transient audio data and a long window for non-transient audio data.
[0170] 4. Time domain noise shaping;
[0171] Optionally, time-domain noise shaping refers to inputting audio data with a shorter time domain or transient audio data, and then filtering the data using a time-domain noise shaping filter to obtain filtered audio data.
[0172] 5. Frequency domain noise shaping.
[0173] Optionally, frequency domain noise shaping refers to filtering frequency domain audio data using frequency domain noise shaping to obtain filtered audio data.
[0174] It should be noted that step S2103 in the disclosed embodiment is a different preprocessing process from the other preprocessing processes. Alternatively, it can be understood that the encoder first preprocesses the audio data using at least one of the five steps above, then performs the following signal analysis process on the preprocessed audio data, and then performs step S2103 above to preprocess the audio data after signal analysis.
[0175] (2) Signal analysis;
[0176] In some embodiments, signal analysis refers to at least one of bandwidth detection, signal feature analysis, and signal classification. Optionally, signal analysis refers to at least one of bandwidth detection, feature analysis, and classification of audio data.
[0177] (3) Nuclear option;
[0178] In some embodiments, core selection refers to selecting a core module that matches the audio data, performing core encoding on the audio data, and thereby obtaining encoded data. In some embodiments, the encoder analyzes the audio data to determine the core module that performs core encoding on the audio data.
[0179] It should be noted that, in the embodiment of the present disclosure, when performing core selection, the encoder will encode the identifier of the core module to be used, so as to facilitate transmission of the encoded identifier of the core module.
[0180] (4) nuclear coding;
[0181] In some embodiments, kernel selection and kernel encoding refer to the process of selecting an encoding kernel and performing encoding using the selected encoding kernel.
[0182] (5) Multi-channel data mixing.
[0183] In some embodiments, mixing of multiple data channels refers to carrying all the data encoded by the above process into one data stream for transmission.
[0184] It should be noted that the embodiments of the present disclosure involve frequency band ranges. In some embodiments, the frequency band ranges and the frequency ranges are the same concepts, that is, the frequency band ranges can also be called the frequency ranges, which is not limited in the embodiments of the present disclosure.
[0185] Step S2104: The encoder encodes the second audio data based on the encoding bandwidth corresponding to the second audio data to obtain first encoded data.
[0186] In the embodiment of the present disclosure, the encoder encodes the second audio data according to the encoding frequency to facilitate the transmission of the encoded first encoded data.
[0187] In some embodiments, the above step S2104 can actually be understood as a core encoding process in the encoding process.
[0188] It should be noted that, in the embodiment of the present disclosure, the encoder downsamples the second audio data before encoding the second audio data, and then encodes the third audio data obtained after the downsampling.
[0189] In some embodiments, the second audio data is downsampled to obtain third audio data, and the third audio data is encoded to obtain the first encoded data. In some embodiments, the second audio data is downsampled at a sampling rate that is a multiple of the encoding bandwidth to obtain the third audio data. In some embodiments, the encoding bandwidth can be understood as the encoding bandwidth used by the encoder to encode the second audio data. In some embodiments, the encoding bandwidth can be understood as the encoding bandwidth used by the core module when the encoder employs the core module to encode the second audio data.
[0190] Optionally, the multiple of the coding bandwidth can be understood as 1 times, 2 times or other values, which is not limited in the embodiments of the present disclosure.
[0191] In some embodiments, the encoder generates second information that describes the downsampled sampling bandwidth, and encodes the second information to obtain third encoded data. In the disclosed embodiment, the encoder also encodes the downsampled sampling bandwidth used and sends it to the decoder so that the decoder can decode the encoded data based on the sampling bandwidth used by the encoder. Optionally, the downsampled sampling bandwidth includes the bandwidth before sampling and the bandwidth after sampling.
[0192] In some embodiments, the encoder sends the third encoded data to the decoder, and the third encoded data is used for decoding by the decoder. In some embodiments, the decoder receives the third encoded data sent by the encoder.
[0193] It should be noted that downsampling in the disclosed embodiments can be understood as a step in preprocessing the audio data in the above embodiments. In some embodiments, the downsampling and step S2103 are performed in the same preprocessing process. Alternatively, it can be understood that the five steps in the above embodiments are not performed in the same step as step S2103 and the downsampling process.
[0194] It should be noted that if the first audio data includes multiple data, the encoder will move the first frequency band range of each first audio data to the second frequency band range to obtain second audio data, and then perform multi-channel de-redundancy processing on the multiple second audio data, and encode the multiple second audio data after multi-channel de-redundancy processing to obtain first encoded data.
[0195] In some embodiments, de-redundancy processing can be understood as taking a reference audio data as a benchmark, sequentially obtaining the degree of difference between the second audio data and the reference audio data, and then encoding the second heterodyne data based on the degree of difference to obtain first encoded data.
[0196] In some embodiments, the redundancy removal process may be understood as adjusting the phases and energies of the plurality of second audio data so as to perform encoding processing on the adjusted second audio data to obtain first encoded data.
[0197] It should be noted that, in the embodiment of the present disclosure, when there are multiple first audio data, the audio data can also be downsampled to obtain downsampled audio data, and then the multiple downsampled audio data can be encoded to obtain first encoded data.
[0198] In some embodiments, referring to FIG2C , after audio data is input, the audio data is preprocessed by the preprocessing module 1, and then the preprocessed audio data is subjected to signal analysis, and then the audio data after signal analysis is preprocessed by the preprocessing module 2, and then the audio data is subjected to core encoding by the core module. During this period, first information and second information are also generated, and then the first information and the second information are encoded, and then the encoded audio data and information are mixed and carried in the data stream. Among them, the preprocessing module 1 is used to perform at least one of 1-5 in the above embodiments. The preprocessing module 2 is used to perform at least one of the above steps S2103 or downsampling.
[0199] Step S2105: The encoder carries the first encoded data in a data stream and sends it to the decoder.
[0200] In the embodiment of the present disclosure, the encoder sends data to the decoder by sending a data stream to the decoder. In some embodiments, the encoder needs to send multiple data to the decoder, so the encoder can carry the data to be sent in the data stream and send it to the decoder.
[0201] In some embodiments, the encoder sends the first encoded data to the decoder, and the first encoded data is used for decoding by the decoder. In some embodiments, the encoder carries the first encoded data in a data stream and sends the data stream to the decoder.
[0202] It should be noted that, in the embodiment of the present disclosure, the encoder involves first coded data, second coded data and third coded data that need to be transmitted. The encoder can carry the first coded data, second coded data and third coded data in a data stream and send it to the decoder.
[0203] It should be noted that the embodiments of the present disclosure include steps such as frequency shifting, downsampling, and core encoding of audio data. Frequency shifting of audio data involves a frequency band range, downsampling of audio data involves a frequency band range, and core encoding also involves a coding rate. Therefore, the frequency of each method can be specified. See Table 1, which shows the relationship between the second frequency band range, the downsampling bandwidth, and the core encoding frequency:
[0204] Table 1
[0205] In some embodiments, the downsampling bandwidth is determined by the core coding frequency and the frequency / band range of the shifted audio signal. Assuming the frequency / band range of the shifted audio signal is 6 kHz, and the core coding frequency set includes 8 kHz, 16 kHz, 32 kHz, and 48 kHz, the smallest core coding frequency greater than 6*2 kHz, i.e., 16 kHz, is selected.
[0206] If the maximum core coding frequency is NkHz and the frequency / band range of the shifted audio signal is N / 2+B (where B>0), the downsampling bandwidth is determined only by the maximum core coding frequency (whose value is N). In this case, the shifted audio signal will lose some high-frequency components.
[0207] Step S2106: The decoder decodes the first encoded data to obtain second audio data.
[0208] In some embodiments, the decoder receives first encoded data sent by the encoder, and the first encoded data is used for decoding by the decoder.
[0209] In some embodiments, the decoder receives the second encoded data sent by the encoder, decodes the second encoded data, and obtains first information, where the first information is used to describe the first frequency band range shift and the second frequency band range.
[0210] In some embodiments, the decoder decodes the first encoded data to obtain third audio data, and upsamples the third audio data to obtain second audio data.
[0211] In some embodiments, the decoder receives the third encoded data sent by the encoder, decodes the third encoded data, and obtains second information, where the second information is used to describe a down-sampling sampling bandwidth.
[0212] In some embodiments, the second audio data includes multiple first encoded data, the decoder performs multi-channel processing on the multiple first encoded data, and decodes the multiple first encoded data after multi-channel processing to obtain the second audio data.
[0213] It should be noted that the second coded data and the third coded data in the embodiment of the present disclosure are carried in one information, that is, the second coded data and the third coded data can be generated and encoded at the same time without having to be encoded separately.
[0214] In step S2107 , the decoder moves the second frequency band of the second audio data to the first frequency band to obtain the first audio data.
[0215] In the disclosed embodiment, the encoder adjusts the frequency band of the audio data when encoding it. Therefore, the decoder needs to adjust the frequency band of the audio data back when decoding it. Therefore, the decoder moves the second frequency band of the second audio data to the first frequency band to obtain the first audio data. The first audio data is the audio data before being encoded by the encoder.
[0216] It should be noted that the audio data in the embodiments of the present disclosure involves two processes, encoding and decoding, wherein the encoding and decoding processes may cause partial loss of the audio data, but the resulting loss will not affect the substantial content of the audio data.
[0217] In some embodiments, decoding the first encoded data includes:
[0218] Performing at least one of the following on the first encoded data to obtain the audio data includes:
[0219] Bitstream demultiplexing;
[0220] nuclear decoding;
[0221] Inverse transform.
[0222] It should be noted that the inverse transformation refers to the opposite process of the encoding in the above embodiment.
[0223] In some embodiments, the core decodes, including:
[0224] The audio data is kernel-decoded based on the determined kernel module.
[0225] In some embodiments, the inverse transform comprises:
[0226] Inverse time-frequency transform;
[0227] Inverse time domain noise shaping transform;
[0228] Inverse frequency-domain noise shaping transform.
[0229] It should be noted that the process of decoding the encoded data by the decoder in the embodiment of the present disclosure is the opposite process of the encoding process in the above embodiment.
[0230] In some embodiments, referring to FIG2D , after a data stream is input, the data stream is demultiplexed to obtain first information, second information, and encoded data. The encoded data is then decoded by the core decoding module and then processed by post-processing module 2 and post-processing module 1 to obtain audio data. Optionally, post-processing module 2 is configured to perform step S2107 and the upsampling process described above. Post-processing module 1 is configured to perform processes such as inverse transformation.
[0231] In some embodiments, referring to FIG2E , after the input data stream is demultiplexed, core decoding is performed, and the second frequency band range of the audio data obtained by upsampling is 0kHz-10kHz, and then the second frequency band range is adjusted to the first frequency band range of 30kHz-40kHz to obtain the adjusted audio data.
[0232] It should be noted that the embodiment of the present disclosure involves the processing of audio data of multiple channels. See Figure 2F. If audio data of two channels is input, for each audio data, the audio data is frequency band adjusted and down-sampled, and then the processed audio data is subjected to multi-channel de-redundancy processing, and then the audio data of each channel is kernel-encoded separately, and then mixed to obtain the data stream.
[0233] In some embodiments, the names of information, etc. are not limited to the names described in the embodiments, and terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", "field", "symbol", "symbol", "codeword", "codebook", "codeword", "codepoint", "bit", "data", "program", and "chip" can be used interchangeably.
[0234] In some embodiments, terms such as "uplink", "uplink", "physical uplink" can be interchangeable with each other, and terms such as "downlink", "downlink", "physical downlink" can be interchangeable with each other, and terms such as "side", "sidelink", "side communication", "sidelink communication", "direct connection", "direct link", "direct communication", "direct link communication" can be interchangeable with each other.
[0235] In some embodiments, "obtain", "get", "get", "receive", "transmit", "bidirectional transmission", "send and / or receive" can be interchangeable, and can be interpreted as receiving from other entities, obtaining from protocols, obtaining from higher layers, obtaining by self-processing, autonomous implementation, etc.
[0236] In some embodiments, terms such as "send", "transmit", "report", "download", "transmit", "bidirectional transmission", "send and / or receive" can be used interchangeably.
[0237] In some embodiments, terms such as "moment", "time point", "time", and "time position" can be replaced with each other, and terms such as "duration", "period", "time window", "window", and "time" can be replaced with each other.
[0238] In some embodiments, terms such as "certain", "preset", "preset", "setting", "indicated", "some", "any", and "first" can be interchangeable. "Specific A", "preset A", "preset A", "setting A", "indicated A", "some A", "any A", and "first A" can be interpreted as A pre-specified in a protocol, etc., or as A obtained through setting, configuration, or indication, etc., or as specific A, some A, any A, or first A, etc., but not limited to this.
[0239] The encoding and decoding method involved in the embodiment of the present disclosure may include at least one of steps S2101 to S2107. For example, step S2101 can be implemented as an independent embodiment, step S2102 can be implemented as an independent embodiment, step S2103 can be implemented as an independent embodiment, step S2104 can be implemented as an independent embodiment, step S2105 can be implemented as an independent embodiment, step S2106 can be implemented as an independent embodiment, step S2107 can be implemented as an independent embodiment, steps S2101 and S2102 can be implemented as independent embodiments, steps S2103 and S2107 can be implemented as an independent embodiment. 2104 can be implemented as an independent embodiment, step S2105 and step S2106 can be implemented as independent embodiments, step S2101, step S2102, step S2103, and step S2104 can be implemented as independent embodiments, step S2101, step S2102, step S2105, and step S2106 can be implemented as independent embodiments, step S2103, step S2104, step S2105, and step S2106 can be implemented as independent embodiments, but are not limited to this.
[0240] In some embodiments, step S2101 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0241] In some embodiments, step S2102 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0242] In some embodiments, step S2103 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0243] In some embodiments, step S2104 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0244] In some embodiments, step S2105 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0245] In some embodiments, step S2106 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0246] In some embodiments, step S2107 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0247] In some embodiments, step S2101 and step S2102 are optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0248] In some embodiments, step S2103 and step S2104 are optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0249] In some embodiments, step S215 and step S2106 are optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0250] In some embodiments, reference may be made to other optional implementations described before or after the description corresponding to FIG. 2A .
[0251] FIG3A is a flow chart of a coding and decoding method according to an embodiment of the present disclosure, which is applied to an encoder. As shown in FIG3A , the embodiment of the present disclosure relates to a coding and decoding method, which includes:
[0252] In step S3101 , the encoder moves the first frequency band of the first audio data to a second frequency band to obtain second audio data.
[0253] Optional implementations of step S3101 may refer to step S2103 in FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.
[0254] Step S3102: The encoder encodes the second audio data according to the encoding frequency to obtain first encoded data.
[0255] The optional implementation of step S3102 can be found in step S2104 of FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.
[0256] Step S3103: The encoder carries the first encoded data in a data stream and sends it to the decoder.
[0257] The optional implementation of step S3103 can be found in step S2105 of FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.
[0258] The encoding and decoding method involved in the embodiments of the present disclosure may include at least one of steps S3101 to S3103. For example, step S3101 may be implemented as an independent embodiment, step S3102 may be implemented as an independent embodiment, and step S3103 may be implemented as an independent embodiment, or at least two steps may be combined, but the present invention is not limited thereto.
[0259] In some embodiments, step S3101 is optional, step S3102 is optional, and step S3103 is optional, and one or more of these steps may be omitted or replaced in different embodiments, but the present invention is not limited thereto.
[0260] FIG3B is a flow chart of a coding and decoding method according to an embodiment of the present disclosure, which is applied to an encoder. As shown in FIG3B , the embodiment of the present disclosure relates to a coding and decoding method, which includes:
[0261] In step S3201 , the encoder moves the first frequency band of the first audio data to a second frequency band to obtain second audio data.
[0262] Optional implementations of step S3201 may refer to step S2103 in FIG. 2A , step S3101 in FIG. 3A , and other related parts in the embodiments involved in FIG. 2A and FIG. 3A , which will not be described in detail here.
[0263] Step S3202: The encoder encodes the second audio data according to the encoding frequency to obtain first encoded data.
[0264] Optional implementations of step S3202 may be referred to step S2104 in FIG. 2A , step S3102 in FIG. 3A , and other related parts in the embodiments involved in FIG. 2A and FIG. 3A , which will not be described in detail here.
[0265] FIG4 is a flow chart of a coding and decoding method according to an embodiment of the present disclosure, which is applied to a decoder. As shown in FIG4 , the embodiment of the present disclosure relates to a coding and decoding method, which includes:
[0266] Step S4101: The decoder decodes the first encoded data to obtain second audio data.
[0267] Optional implementations of step S4101 may refer to step S2106 in FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.
[0268] In step S4102 , the decoder moves the second frequency band of the second audio data to the first frequency band to obtain the first audio data.
[0269] The optional implementation of step S4102 can be found in step S2107 of FIG. 2A and other related parts of the embodiment involved in FIG. 2A , which will not be described in detail here.
[0270] FIG5 is a flow chart of a coding and decoding method according to an embodiment of the present disclosure. As shown in FIG6 , an embodiment of the present disclosure relates to a coding and decoding method, which includes:
[0271] Step S5101: The encoder moves the first frequency band range of the first audio data to the second frequency band range to obtain second audio data.
[0272] In some embodiments, the first audio data is used for machine analysis / machine training, the second frequency band range partially overlaps or does not overlap with the first frequency band range, and the second frequency band range matches the encoding frequency for encoding the second audio data.
[0273] Step S5102: The encoder encodes the second audio data according to the encoding frequency to obtain first encoded data.
[0274] Step S5103: The decoder decodes the first encoded data to obtain second audio data.
[0275] Step S5104: The decoder moves the second frequency band range of the second audio data to the first frequency band range to obtain the first audio data.
[0276] In some embodiments, the above method may include the methods of the embodiments of the above communication system side, the first device side, the encoder side, the decoder side, etc., which will not be repeated here.
[0277] FIG6 is a flow chart of a coding and decoding method according to an embodiment of the present disclosure. As shown in FIG6 , the embodiment of the present disclosure relates to a coding and decoding method, and the method includes:
[0278] In step S6101 , the encoder adjusts the frequency band of the data to obtain data after the frequency band is adjusted.
[0279] In some embodiments, the input signal to the encoder has only spectral components between 30-40 kHz.
[0280] In some embodiments, the encoder performs mathematical operations to shift the spectrum from 30-40 kHz to 0-10 kHz.
[0281] In some embodiments, the shifted audio signal is downsampled to adapt to the encoding sampling rate of the core encoding module. The sampling rate satisfies the Nyquist sampling theorem, that is, the sampling rate is greater than twice the highest effective frequency.
[0282] In some embodiments, core encoding is then performed and the encoded data is mixed into a bitstream.
[0283] In some embodiments, the side information of the frequency shift value 30 kHz and the downsampling rate is encoded by a metadata encoding module.
[0284] In some embodiments, the decoder decouples the bitstream to obtain the encoded data. The encoded data is first decoded by the core decoding module and then upsampled and frequency-shifted.
[0285] In some embodiments, multi-channel audio data input is also provided.
[0286] In some embodiments, the audio data of each channel is bandwidth-shifted and down-sampled before being multi-channel encoded to reduce inter-channel redundancy.
[0287] In the embodiments of the present disclosure, some or all of the steps and their optional implementations may be arbitrarily combined with some or all of the steps in other embodiments, or may be arbitrarily combined with the optional implementations of other embodiments.
[0288] The present disclosure also provides an apparatus for implementing any of the above methods. For example, a device is provided that includes units or modules for implementing each step performed by an encoder in any of the above methods. For another example, another device is provided that includes units or modules for implementing each step performed by a decoder in any of the above methods.
[0289] It should be understood that the division of the various units or modules in the above device is merely a division of logical functions. In actual implementation, they may be fully or partially integrated into a physical entity, or they may be physically separated. In addition, the units or modules in the device may be implemented in the form of a processor calling software: for example, the device includes a processor, the processor is connected to a memory, and the memory stores instructions. The processor calls the instructions stored in the memory to implement any of the above methods or implement the functions of the various units or modules of the above device, wherein the processor is, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory is a memory within the device or a memory outside the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits, and the functions of some or all of the units or modules can be realized by designing the hardware circuits. The above-mentioned hardware circuits can be understood as one or more processors; for example, in one implementation, the above-mentioned hardware circuit is an application-specific integrated circuit (ASIC), which realizes the functions of some or all of the above units or modules by designing the logical relationship of the components in the circuit; for example, in another implementation, the above-mentioned hardware circuit can be realized by a programmable logic device (PLD). Taking a field programmable gate array (FPGA) as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by configuring the configuration file, thereby realizing the functions of some or all of the above units or modules. All units or modules of the above devices can be realized in the form of software called by the processor, or in the form of hardware circuits, or in part by the form of software called by the processor, and the rest by hardware circuits.
[0290] In the embodiments of the present disclosure, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationship of the hardware circuit. The logical relationship of the above-mentioned hardware circuit is fixed or reconfigurable. For example, the processor is a hardware circuit implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and implementing the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), etc.
[0291] Figure 7A is a structural diagram of the encoding and decoding device proposed in an embodiment of the present disclosure. As shown in Figure 7A, the encoding and decoding device 7100 may include: at least one of a transceiver module 7101, a processing module 7102, etc. In some embodiments, the processing module 7102 is used to move the first frequency band range of the first audio data to the second frequency band range to obtain second audio data, the first audio data is used for machine analysis / machine training, the second frequency band range partially overlaps or does not overlap with the first frequency band range, and the second frequency band range matches the encoding frequency for encoding the second audio data; the second audio data is encoded according to the encoding frequency to obtain the first encoded data. Optionally, the above-mentioned transceiver module 7101 is used to perform at least one of the communication steps such as sending and / or receiving performed by the encoder in any of the above methods (for example, step S2101 but not limited to this), which will not be repeated here. Optionally, the above-mentioned processing module is used to perform at least one of the other steps performed by the encoder in any of the above methods, which will not be repeated here.
[0292] Optionally, the processing module 7102 is used to execute at least one of the communication steps such as processing performed by the encoder in any of the above methods, which will not be repeated here.
[0293] Figure 7B is a structural diagram of the encoding and decoding device proposed in an embodiment of the present disclosure. As shown in Figure 7B, the encoding and decoding device 7200 may include: at least one of a transceiver module 7201, a processing module 7202, etc. In some embodiments, the processing module 7202 is used to decode the first encoded data to obtain second audio data; move the second frequency band range of the second audio data to the first frequency band range to obtain first audio data, the first audio data is used for machine analysis / machine training, the second frequency band range partially overlaps or does not overlap with the first frequency band range, and the second frequency band range matches the encoding frequency of the encoder. Optionally, the above-mentioned transceiver module is used to perform at least one of the communication steps such as sending and / or receiving performed by the decoder in any of the above methods, which will not be repeated here.
[0294] Optionally, the processing module 7202 is used to execute at least one of the communication steps such as processing performed by the decoder in any of the above methods, which will not be repeated here.
[0295] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module, and the transmitting module and the receiving module may be separate or integrated. Optionally, the transceiver module may be interchangeable with the transceiver.
[0296] In some embodiments, the processing module can be a single module or can include multiple submodules. Optionally, the multiple submodules respectively execute all or part of the steps required to be executed by the processing module. Optionally, the processing module can be interchangeable with the processor.
[0297] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module, and the transmitting module and the receiving module may be separate or integrated. Optionally, the transceiver module may be interchangeable with the transceiver.
[0298] In some embodiments, the processing module can be a single module or can include multiple submodules. Optionally, the multiple submodules respectively execute all or part of the steps required to be executed by the processing module. Optionally, the processing module can be interchangeable with the processor.
[0299] Figure 8A is a schematic diagram of the structure of a communication device 8100 proposed in an embodiment of the present disclosure. Communication device 8100 can be a first device, an encoder, a decoder, or a chip, a chip system, or a processor that supports the first device, encoder, or decoder in implementing any of the above methods. Communication device 8100 can be used to implement the methods described in the above method embodiments. For details, please refer to the description of the above method embodiments.
[0300] As shown in Figure 8A, communication device 8100 includes one or more processors 8101. Processor 8101 can be a general-purpose processor or a dedicated processor, such as a baseband processor or a central processing unit. The baseband processor can be used to process communication protocols and communication data, while the central processing unit can be used to control the codec device, execute programs, and process program data. Communication device 8100 is configured to perform any of the above methods.
[0301] In some embodiments, the communication device 8100 further includes one or more memories 8102 for storing instructions. Optionally, all or part of the memories 8102 may be located outside the communication device 8100.
[0302] In some embodiments, the communication device 8100 further includes one or more transceivers 8103. When the communication device 8100 includes one or more transceivers 8103, the transceiver 8103 performs at least one of the communication steps of sending and / or receiving in the above method.
[0303] In some embodiments, a transceiver may include a receiver and / or a transmitter. The receiver and transmitter may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, and transceiver circuit may be used interchangeably; the terms transmitter, transmitting unit, transmitter, and transmitting circuit may be used interchangeably; and the terms receiver, receiving unit, receiver, and receiving circuit may be used interchangeably.
[0304] In some embodiments, the communication device 8100 may include one or more interface circuits 8104. Optionally, the interface circuit 8104 is connected to the memory 8102. The interface circuit 8104 may be configured to receive signals from the memory 8102 or other devices, and may be configured to send signals to the memory 8102 or other devices. For example, the interface circuit 8104 may read instructions stored in the memory 8102 and send the instructions to the processor 8101.
[0305] The communication device 8100 described in the above embodiments may be a first device, a decoder, or an encoder, but the scope of the communication device 8100 described in the present disclosure is not limited thereto, and the structure of the communication device 8100 may not be limited by FIG. 8A. The communication device may be an independent device or may be part of a larger device. For example, the communication device may be: 1) an independent integrated circuit IC, or a chip, or a chip system or subsystem; (2) a collection of one or more ICs, optionally, the above IC collection may also include a storage component for storing data or programs; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, an encoder device, an intelligent encoder device, a cellular phone, a wireless device, a handheld device, a mobile unit, an in-vehicle device, a decoder, a cloud device, an artificial intelligence device, etc.; (6) others, etc.
[0306] FIG8B is a schematic diagram of the structure of a chip 8200 according to an embodiment of the present disclosure. If the communication device 8100 can be a chip or a chip system, please refer to the schematic diagram of the structure of the chip 8200 shown in FIG8B , but the present disclosure is not limited thereto.
[0307] The chip 8200 includes one or more processors 8201 , and the chip 8200 is configured to execute any of the above methods.
[0308] In some embodiments, the chip 8200 further includes one or more interface circuits 8202. Optionally, the interface circuit 8202 is connected to the memory 8203. The interface circuit 8202 can be used to receive signals from the memory 8203 or other devices, and can be used to send signals to the memory 8203 or other devices. For example, the interface circuit 8202 can read instructions stored in the memory 8203 and send the instructions to the processor 8201.
[0309] In some embodiments, the interface circuit 8202 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processor 8201 performs at least one of the other steps.
[0310] In some embodiments, terms such as interface circuit, interface, transceiver pin, and transceiver may be used interchangeably.
[0311] In some embodiments, the chip 8200 further includes one or more memories 8203 for storing instructions. Alternatively, all or part of the memories 8203 may be outside the chip 8200.
[0312] The present disclosure also proposes a storage medium having instructions stored thereon, which, when executed on the communication device 8100, causes the communication device 8100 to execute any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but is not limited thereto, and may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but is not limited thereto, and may also be a temporary storage medium.
[0313] The present disclosure also provides a program product, which, when executed by the communication device 8100, enables the communication device 8100 to perform any of the above methods. Optionally, the program product is a computer program product.
[0314] The present disclosure also proposes a computer program, which, when executed on a computer, causes the computer to perform any one of the above methods.
Claims
1. A coding method, characterized in that: The method is performed by an encoder, and includes: Acquire first audio data, and move a first frequency band of the first audio data to a second frequency band to obtain second audio data, wherein the first audio data is used for machine analysis / machine training, the second frequency band partially overlaps with or does not overlap with the first frequency band, and the second frequency band is determined by an available encoding bandwidth; The second audio data is encoded based on a coding bandwidth corresponding to the second audio data to obtain first encoded data.
2. The method according to claim 1, characterized in that The method further comprises: The first encoded data is sent to a decoder, where the first encoded data is used for decoding by the decoder.
3. The method according to claim 1 or 2, characterized in that The method further comprises: generating first information, wherein the first information is used to describe parameters for shifting the first audio data; The first information is encoded to obtain second encoded data.
4. The method according to claim 3, characterized in that The method further comprises: The second encoded data is sent to a decoder, where the second encoded data is used for decoding by the decoder.
5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: downsampling the second audio data to obtain third audio data; The encoding of the second audio data to obtain first encoded data includes: The third audio data is encoded to obtain the first encoded data.
6. The method according to claim 5, characterized in that The downsampling of the second audio data to obtain the third audio data includes: The second audio data is downsampled according to a sampling rate that is a multiple of the encoding bandwidth to obtain the third audio data.
7. The method according to claim 5 or 6, characterized in that The method further comprises: generating second information, where the second information is used to describe a sampling bandwidth of the downsampling; The second information is encoded to obtain third encoded data.
8. The method according to claim 7, characterized in that The method further comprises: The third encoded data is sent to a decoder, and the third encoded data is used for decoding by the decoder.
9. The method according to any one of claims 1 to 8, characterized in that: The first audio data includes a plurality of pieces, and encoding the second audio data to obtain the first encoded data includes: performing multi-channel redundancy removal processing on the plurality of second audio data; The plurality of second audio data after multi-channel redundancy removal processing is encoded to obtain the first encoded data.
10. A decoding method, characterized in that: The method is performed by a decoder, and comprises: Decoding the first encoded data to obtain second audio data; The second frequency band range of the second audio data is moved to the first frequency band range to obtain first audio data, where the first audio data is used for machine analysis / machine training, the second frequency band range partially overlaps or does not overlap with the first frequency band range, and the second frequency band range is determined by the available encoding bandwidth.
11. The method according to claim 10, characterized in that The method further comprises: The first encoded data sent by the encoder is received, where the first encoded data is used for decoding by the decoder.
12. The method according to claim 10 or 11, characterized in that The method further comprises: receiving second encoded data sent by the encoder; The second encoded data is decoded to obtain first information, where the first information is used to describe the first frequency band range shift and the second frequency band range.
13. The method according to any one of claims 10 to 12, characterized in that: The decoding of the first encoded data to obtain the second audio data includes: decoding the first encoded data to obtain third audio data; The third audio data is up-sampled to obtain the second audio data.
14. The method according to claim 13, characterized in that The method further comprises: receiving third encoded data sent by the encoder; The third encoded data is decoded to obtain second information, where the second information is used to describe a sampling bandwidth of the downsampling.
15. The method according to any one of claims 10 to 14, characterized in that: The second audio data includes a plurality of items, and decoding the first encoded data to obtain the second audio data includes: performing multi-channel processing on the plurality of first encoded data; The plurality of first encoded data after multi-channel processing is decoded to obtain the second audio data.
16. A coding and decoding method, characterized in that: The method comprises: The encoder moves a first frequency band range of the first audio data to a second frequency band range to obtain second audio data, where the first audio data is used for machine analysis / machine training, the second frequency band range partially overlaps or does not overlap with the first frequency band range, and the second frequency band range matches an encoding frequency used to encode the second audio data; The encoder encodes the second audio data according to the encoding frequency to obtain first encoded data; The decoder decodes the first encoded data to obtain second audio data; The decoder moves the second frequency band range of the second audio data to the first frequency band range to obtain first audio data.
17. An encoding device, characterized in that: The encoding device comprises: a processing module, configured to move a first frequency band of the first audio data to a second frequency band to obtain second audio data, wherein the first audio data is used for machine analysis / machine training, the second frequency band partially overlaps or does not overlap with the first frequency band, and the second frequency band is determined by an available encoding bandwidth; The processing module is configured to encode the second audio data based on an encoding bandwidth corresponding to the second audio data to obtain first encoded data.
18. A decoding device, characterized in that: The decoding device comprises: a processing module, configured to decode the first encoded data to obtain second audio data; The processing module is used to move the second frequency band range of the second audio data to the first frequency band range to obtain first audio data, where the first audio data is used for machine analysis / machine training, the second frequency band range partially overlaps or does not overlap with the first frequency band range, and the second frequency band range is determined by the available encoding bandwidth.
19. An encoding device, characterized in that: The encoding device comprises: one or more processors; The processor is configured to execute the encoding method according to any one of claims 1 to 9.
20. A decoding device, characterized in that: The decoding device comprises: one or more processors; The processor is configured to execute the encoding and decoding method according to any one of claims 10 to 15.
21. A coding and decoding system, characterized in that: The invention comprises an encoder and a decoder, wherein the encoder is configured to implement the encoding method according to any one of claims 1 to 9, and the encoder is configured to implement the decoding method according to any one of claims 10 to 15.
22. A storage medium storing instructions, characterized in that: When the instruction is executed on a communication device, the communication device is caused to execute the encoding method according to any one of claims 1 to 9, or execute the decoding method according to any one of claims 10 to 15.
Citation Information
Patent Citations
Optical-sampling-based radio frequency measuring method and measuring device
CN102981048A
Transmission method and system of sound channel information
CN104702343A
Voice signal processing method and related device and system
CN105869653A
Audio signal coding compression and transmission method and electronic equipment
CN113129911A
Audio processing method and device, equipment and storage medium
CN117351943A