Coding method and device, decoding method and device, and storage medium

By splitting the characteristic information of the audio signal and coding the core group encoding, the audio signal encoding problem with diverse frequency ranges is solved, and a more accurate and reasonable encoding process is achieved, reducing the problems of information loss and excessive bit flow.

WO2025145384A1PCT designated stage expired Publication Date: 2025-07-10BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/070597
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-04
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

The prior art lacks effective encoding methods to process audio signals with a variety of frequency ranges, especially infrasonic waves, audible sound waves and ultrasonic signals, resulting in information loss or bit flow too much during the encoding process, which cannot meet actual needs.

Method used

By splitting the audio signal into audio components of different frequency ranges according to the characteristic information of the audio signal, and encode each component using a corresponding encoding core group, metadata is generated to indicate the encoding core group information and split information, and a bit stream is output for easy decoding.

Benefits of technology

Improves the accuracy and rationality of encoding, ensures that audio signals in different frequency ranges can be properly processed, reduces information loss and optimizes the bitstream size.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024070597_10072025_PF_FP_ABST
    Figure CN2024070597_10072025_PF_FP_ABST
Patent Text Reader

Abstract

A coding method and device, a decoding method and device, and a storage medium. The coding method comprises: on the basis of characteristic information of an audio signal, splitting the audio signal into at least one group of audio components (S2101, S3101); on the basis of a coding kernel set corresponding to each group of audio components among the at least one group of audio components, separately coding the at least one group of audio components(S2102, S3102), the coding kernel set having a mapping relationship with the audio components; and, on the basis of a coding result corresponding to the at least one group of audio components and metadata, outputting a bit stream (S2103, S3103), the metadata being used for indicating information about the coding kernel set that executes coding and split information corresponding to the audio signal being split into the at least one group of audio components. The coding method splits an input audio signal on the basis of characteristic information, and separately codes the split audio components, thereby helping to select a proper coding method on the basis of characteristics of different audio components, and improving coding accuracy and rationality. In addition, writing in metadata during a coding process facilitates accurate decoding at a decoding end.
Need to check novelty before this filing date? Find Prior Art

Description

Coding and decoding method, device and storage medium Technical Field

[0001] The present disclosure relates to the field of signal processing technology, and in particular to an encoding and decoding method, device, and storage medium. Background Art

[0002] Sound waves can be classified into different types based on their frequency. Sound waves with frequencies below 20 Hz are called infrasound waves; those with frequencies between 20 Hz and 20 kHz are called audible sound waves; those with frequencies above 20 kHz are called ultrasonic waves; and those with frequencies above 500 kHz are called megasonic waves. Audible sound waves include audio signals audible to the human ear, such as music or speech. Audio encoding can be performed on these audio signals based on the frequency range audible to the human ear and the masking effect.

[0003] Summary of the Invention

[0004] With the development of the automation field, there is a lack of effective encoding methods for signals with diverse frequency ranges.

[0005] Embodiments of the present disclosure provide an encoding and decoding method, device, and storage medium.

[0006] In a first aspect, an embodiment of the present disclosure provides an encoding method, including:

[0007] Splitting the audio signal into at least one group of audio components according to characteristic information of the audio signal;

[0008] encoding the at least one group of audio components according to a coding core group corresponding to each audio component in the at least one group of audio components, wherein the coding core group and the audio component have a mapping relationship;

[0009] A bitstream is output according to the encoding result and metadata corresponding to the at least one group of audio components, wherein the metadata is used to indicate the encoding core group information performing the encoding and the splitting information corresponding to the splitting of the audio signal into the at least one group of audio components.

[0010] In a second aspect, an embodiment of the present disclosure provides a decoding method, including:

[0011] Determining, based on the bitstream, metadata and an encoding result corresponding to at least one group of audio components, wherein the metadata is used to indicate information about a coding core group on which encoding is performed and information corresponding to splitting the audio signal into the at least one group of audio components;

[0012] Determining a decoding core group corresponding to the at least one group of audio components according to the metadata;

[0013] Outputting a decoding result according to the decoding core group corresponding to the at least one group of audio components.

[0014] In a third aspect, an embodiment of the present disclosure provides an encoding device, including:

[0015] a processing module, configured to split the audio signal into at least one group of audio components according to characteristic information of the audio signal;

[0016] The processing module is further configured to encode the at least one group of audio components respectively according to the encoding core group corresponding to each audio component in the at least one group of audio components, wherein the encoding core group and the audio component have a mapping relationship;

[0017] The processing module is further configured to output a bitstream based on the encoding results and metadata corresponding to the at least one group of audio components, wherein the metadata is configured to indicate encoding core group information for performing encoding and splitting information corresponding to splitting the audio signal into at least one group of audio components.

[0018] In a fourth aspect, an embodiment of the present disclosure provides a decoding device, including:

[0019] a processing module, configured to determine, based on the bitstream, metadata and an encoding result corresponding to at least one group of audio components, wherein the metadata is used to indicate information about a coding core group on which encoding is performed and information corresponding to splitting the audio signal into at least one group of audio components;

[0020] The processing module is further configured to determine, based on the metadata, a decoding core group corresponding to the at least one group of audio components;

[0021] The processing module is further configured to output a decoding result according to the decoding core group corresponding to the at least one group of audio components.

[0022] In a fifth aspect, an embodiment of the present disclosure provides a communication device, including:

[0023] one or more processors;

[0024] The communication device is used to execute the method described in the first aspect or the second aspect.

[0025] In a sixth aspect, an embodiment of the present disclosure provides a storage medium, wherein the storage medium stores instructions, wherein:

[0026] When the instruction is executed on a communication device, the communication device is caused to execute the method according to the first aspect or the second aspect.

[0027] In a seventh aspect, an embodiment of the present disclosure provides a communication system, including: an encoding device and a decoding device, wherein:

[0028] The encoding device is used to perform the method according to the first aspect;

[0029] The decoding device is used to execute the method described in the second aspect.

[0030] In the disclosed embodiments, the input audio signal can be split based on feature information, and the split audio components can be encoded using corresponding coding core groups. This facilitates the selection of appropriate encoding methods based on the characteristics of different audio components, improving the accuracy and rationality of encoding. Furthermore, metadata is written into the encoding process to facilitate accurate decoding at the decoding end. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following drawings required for describing the embodiments are introduced. The following drawings are merely some embodiments of the present disclosure and do not impose specific limitations on the protection scope of the present disclosure.

[0032] FIG1 is a schematic diagram of an architecture provided according to an embodiment of the present disclosure;

[0033] FIG2a is a schematic flow chart of a method according to an embodiment of the present disclosure;

[0034] Figures 2b to 2h are process flow charts provided according to an embodiment of the present disclosure;

[0035] FIG3 is a schematic diagram of a method according to an embodiment of the present disclosure;

[0036] FIG4 is a schematic diagram of a processing flow according to an embodiment of the present disclosure;

[0037] FIG5a is a schematic structural diagram of an encoding device according to an embodiment of the present disclosure;

[0038] FIG5b is a schematic structural diagram of a decoding device according to an embodiment of the present disclosure;

[0039] FIG6a is a schematic diagram of a communication device according to an embodiment of the present disclosure;

[0040] FIG6 b is a schematic diagram of a communication device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0041] Embodiments of the present disclosure provide an encoding and decoding method, device, and storage medium.

[0042] In a first aspect, an embodiment of the present disclosure provides an encoding method, including:

[0043] Splitting the audio signal into at least one group of audio components according to characteristic information of the audio signal;

[0044] encoding the at least one group of audio components according to a coding core group corresponding to each audio component in the at least one group of audio components, wherein the coding core group and the audio component have a mapping relationship;

[0045] A bitstream is output according to the encoding result and metadata corresponding to the at least one group of audio components, wherein the metadata is used to indicate the encoding core group information performing the encoding and the splitting information corresponding to the splitting of the audio signal into the at least one group of audio components.

[0046] In the above embodiment, the input audio signal can be split based on feature information, and the split audio components can be encoded using corresponding coding core groups. This facilitates the selection of appropriate encoding methods based on the characteristics of different audio components, improving the accuracy and rationality of encoding. In addition, metadata is written during the encoding process to facilitate accurate decoding at the decoding end.

[0047] In conjunction with the embodiments of the first aspect, in some embodiments, the feature information includes at least one of the following:

[0048] Frequency band range;

[0049] Signal type;

[0050] The frequency band range includes the infrasound range, the audible sound range and the ultrasonic range, and the signal type is used to indicate the signal type of the signal within each frequency band range.

[0051] In the above embodiment, the audio signal is split according to its frequency band range and / or signal type, which is conducive to selecting a suitable coding core group for each audio component.

[0052] In combination with the embodiments of the first aspect, in some embodiments, in different encoding core groups, each encoding core group includes at least one encoding core, wherein the metadata is further used to indicate information of the encoding core that performs encoding.

[0053] In the above embodiment, the coding core group can be divided into different coding cores, so that a suitable coding core can be selected within the coding core group to perform coding processing on the corresponding audio component.

[0054] In conjunction with the embodiments of the first aspect, in some embodiments, outputting a bitstream according to encoding results and metadata corresponding to at least one group of audio components includes:

[0055] Encoding metadata to obtain encoding parameters;

[0056] Output the bitstream based on the encoding result and encoding parameters.

[0057] In the above embodiment, encoding the metadata and then writing it into the bitstream can effectively balance or control the bits occupied by the metadata.

[0058] In conjunction with the embodiments of the first aspect, in some embodiments, encoding the at least one group of audio components according to the coding core group corresponding to each audio component in the at least one group of audio components includes:

[0059] In the encoding core group corresponding to each group of audio components, an encoding core that performs encoding in at least one encoding core is determined, and encoding of the corresponding audio component is performed based on the encoding core that performs encoding.

[0060] In the above embodiment, for each coding core group, a suitable coding core therein may be used to encode the corresponding audio component, in order to improve the coding performance.

[0061] In conjunction with the embodiments of the first aspect, in some embodiments, the encoding core group includes a task core, and the task core is used to encode audio signals having the same audio characteristic parameters and / or the same processing tasks, wherein the audio characteristic parameters are extracted from the audio signals;

[0062] The input of the task core includes multiple frequency band ranges.

[0063] In the above embodiment, task cores may be configured in the encoding core group so that signals with the same task characteristics can be processed, thereby improving processing efficiency.

[0064] In conjunction with the embodiments of the first aspect, in some embodiments, the method further includes:

[0065] The audio components in different frequency ranges are encoded according to the task kernel.

[0066] In the above embodiment, audio components in different frequency ranges can be processed based on the task core, so that signals with the same task characteristics can be processed more efficiently.

[0067] In conjunction with the embodiments of the first aspect, in some embodiments, the method further includes:

[0068] The coding core group corresponding to each group of audio components is determined according to the user configuration information.

[0069] In the above embodiment, the audio components are encoded in combination with user configuration to meet the user's needs.

[0070] In conjunction with the embodiments of the first aspect, in some embodiments, the split information includes at least one of the following:

[0071] The number or groups of audio components;

[0072] The starting frequency of the audio component;

[0073] The stop frequency of the audio component.

[0074] In combination with the embodiments of the first aspect, in some embodiments, the audio signal is a signal obtained by pre-processing the input signal.

[0075] In the above embodiments, various possible pre-processings may be performed on the input signal to improve the subsequent encoding efficiency.

[0076] In a second aspect, an embodiment of the present disclosure provides a decoding method, including:

[0077] Determining, based on the bitstream, metadata and an encoding result corresponding to at least one group of audio components, wherein the metadata is used to indicate information about a coding core group on which encoding is performed and information corresponding to splitting the audio signal into the at least one group of audio components;

[0078] Determining a decoding core group corresponding to the at least one group of audio components according to the metadata;

[0079] Outputting a decoding result according to the decoding core group corresponding to the at least one group of audio components.

[0080] In the above embodiment, the encoding core group for encoding performed by the encoding end is determined based on metadata, so that a suitable decoding core group is selected for accurate decoding.

[0081] In conjunction with the embodiments of the second aspect, in some embodiments, the audio components are separated based on feature information of the audio signal, where the feature information includes at least one of the following:

[0082] Frequency band range;

[0083] Signal type;

[0084] The frequency band range includes the infrasound range, the audible sound range and the ultrasonic range, and the signal type is used to indicate the signal type of the signal within each frequency band range.

[0085] In combination with the embodiments of the second aspect, in some embodiments, in different decoding core groups, each decoding core group includes at least one decoding core, wherein the metadata is also used to indicate information of the encoding core that performs encoding.

[0086] In conjunction with the embodiments of the second aspect, in some embodiments, determining metadata according to the bitstream includes:

[0087] Obtaining encoding parameters according to a bitstream, where the encoding parameters are obtained by encoding the metadata;

[0088] The encoding parameters are decoded to obtain the metadata.

[0089] In conjunction with the embodiments of the second aspect, in some embodiments, the method further includes:

[0090] A decoding core that performs decoding in a decoding core group corresponding to each group of audio components is determined according to the metadata.

[0091] In conjunction with the embodiments of the second aspect, in some embodiments, the decoding core group includes a task core, wherein the task core is used to decode audio signals having the same audio characteristic parameters and / or the same processing tasks, wherein the audio characteristic parameters are extracted from the audio signals;

[0092] The input of the task core includes multiple frequency bands.

[0093] In conjunction with the embodiments of the second aspect, in some embodiments, the method further includes:

[0094] The decoding core group corresponding to each group of audio components is determined according to the user configuration information.

[0095] In combination with the embodiments of the second aspect, in some embodiments, the audio signal is a signal obtained by pre-processing the input signal.

[0096] In conjunction with the embodiments of the second aspect, in some embodiments, the split information includes at least one of the following:

[0097] The number or groups of audio components;

[0098] The starting frequency of the audio component;

[0099] The stop frequency of the audio component.

[0100] In a third aspect, an embodiment of the present disclosure provides an encoding device, including:

[0101] a processing module, configured to split the audio signal into at least one group of audio components according to characteristic information of the audio signal;

[0102] The processing module is further configured to encode the at least one group of audio components respectively according to the encoding core group corresponding to each audio component in the at least one group of audio components, wherein the encoding core group and the audio component have a mapping relationship;

[0103] The processing module is further configured to output a bitstream based on the encoding results and metadata corresponding to the at least one group of audio components, wherein the metadata is configured to indicate encoding core group information for performing encoding and splitting information corresponding to splitting the audio signal into at least one group of audio components.

[0104] In a fourth aspect, an embodiment of the present disclosure provides a decoding device, including:

[0105] a processing module, configured to determine, based on the bitstream, metadata and an encoding result corresponding to at least one group of audio components, wherein the metadata is used to indicate information about a coding core group on which encoding is performed and information corresponding to splitting the audio signal into at least one group of audio components;

[0106] The processing module is further configured to determine, based on the metadata, a decoding core group corresponding to the at least one group of audio components;

[0107] The processing module is further configured to output a decoding result according to the decoding core group corresponding to the at least one group of audio components.

[0108] In a fifth aspect, an embodiment of the present disclosure provides a communication device, including:

[0109] one or more processors;

[0110] The communication device is used to execute the method described in the first aspect or the second aspect.

[0111] In a sixth aspect, an embodiment of the present disclosure provides a storage medium, wherein the storage medium stores instructions, wherein:

[0112] When the instruction is executed on a communication device, the communication device is caused to execute the method according to the first aspect or the second aspect.

[0113] In a seventh aspect, an embodiment of the present disclosure provides a communication system, including: an encoding device and a decoding device, wherein:

[0114] The encoding device is used to perform the method according to the first aspect;

[0115] The decoding device is used to execute the method described in the second aspect.

[0116] The embodiments of the present disclosure are not exhaustive and are merely illustrative of some embodiments, and are not intended to be a specific limitation on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be arbitrarily exchanged. In addition, the optional implementation methods in a certain embodiment can be arbitrarily combined; in addition, the embodiments can be arbitrarily combined. For example, some or all steps of different embodiments can be arbitrarily combined, and a certain embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.

[0117] In each embodiment of the present disclosure, unless otherwise specified or provided for by logic, the terms and / or descriptions between the embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form a new embodiment based on their inherent logical relationships.

[0118] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.

[0119] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular, such as "a", "an", "the", "above", "said", "the", "the", etc., may mean "one and only one", or "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English in translation, the noun following the article may be understood as a singular expression or a plural expression.

[0120] In the embodiments of the present disclosure, “plurality” refers to two or more.

[0121] In some embodiments, the terms "at least one," "one or more," "a plurality of," "multiple," etc. may be used interchangeably.

[0122] In some embodiments, descriptions such as "at least one of A and B," "A and / or B," "A in one case, B in another case," or "in response to one case A, in response to another case B" may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed); and in some embodiments, A and B (both A and B are executed). The above is also applicable when there are more branches such as A, B, and C.

[0123] In some embodiments, "A or B" and other descriptions may include the following technical solutions depending on the situation: in some embodiments, A (A is executed independently of B); in some embodiments, B (B is executed independently of A); in some embodiments, execution is selected from A and B (A and B are selectively executed). The above is also applicable when there are more branches such as A, B, C, etc.

[0124] The prefixes such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different description objects and do not constitute any restriction on the position, order, priority, quantity or content of the description objects. For the statement of the description object, please refer to the description in the context of the claims or embodiments, and no unnecessary restriction should be constituted due to the use of prefixes. For example, if the description object is a "field", the ordinal number before the "field" in the "first field" and the "second field" does not limit the position or order between the "fields". "First" and "second" do not limit whether the "fields" they modify are in the same message, nor do they limit the order of the "first field" and the "second field". For another example, if the description object is a "level", the ordinal number before the "level" in the "first level" and the "second level" does not limit the priority between the "levels". For another example, the number of description objects is not limited by the ordinal number and can be one or more. Taking "first device" as an example, the number of "devices" can be one or more. In addition, the objects modified by different prefixes can be the same or different. For example, if the description object is "device", then the "first device" and the "second device" can be the same device or different devices, and their types can be the same or different; for another example, if the description object is "information", then the "first information" and the "second information" can be the same information or different information, and their contents can be the same or different.

[0125] In some embodiments, “including A,” “comprising A,” “used to indicate A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.

[0126] In some embodiments, terms such as "in response to...", "in response to determining...", "in the case of...", "at the time of...", "when...", "if...", "if...", etc. can be used interchangeably.

[0127] In some embodiments, terms such as "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not less than", and "above" can be replaced with each other, and terms such as "less than", "less than or equal to", "not greater than", "less than", "less than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", and "below" can be replaced with each other.

[0128] In some embodiments, devices and equipment can be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. In some cases, they can also be understood as "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "subject", etc.

[0129] In some embodiments, obtaining data, information, etc. may comply with the laws and regulations of the country where the data is obtained.

[0130] In some embodiments, data, information, etc. may be obtained with the user's consent.

[0131] In addition, each element, each row, or each column in the table of the embodiment of the present disclosure can be implemented as an independent embodiment, and the combination of any elements, any rows, and any columns can also be implemented as an independent embodiment.

[0132] Figure 1 is a schematic diagram of an architecture provided according to an embodiment of the present disclosure. As shown in Figure 1, system 100 may include an encoding device 101 and a decoding device 102. Encoding device 101 may be deployed on a server, and decoding device 102 may be deployed on a server or in a more powerful terminal product. The embodiments of the present disclosure do not limit the specific structures of encoding device 101 or decoding device 102. The structural descriptions in the following embodiments are for illustrative purposes only.

[0133] In the disclosed embodiments, a lossy coding algorithm can be used to process audible audio signals. However, in machine audio applications, lossy coding algorithms can lose valid audio components and fail to meet practical requirements. When encoding using a lossless coding algorithm, the input frequency range for audible audio signals is much larger than the audible audio signal range, resulting in a larger bitstream and a lower compression ratio (e.g., 1:2).

[0134] In some embodiments, both lossy and lossless encoding methods process the input audio signal as a whole. When the frequency range of the input audio signal is large, if high resolution is used to process the signal as a whole, high-frequency signals (such as ultrasonic signals) will generate a larger bitstream data; whereas, if low resolution is used to process the signal as a whole, low-frequency signals (such as infrasonic signals) will be significantly distorted.

[0135] It is necessary to provide an efficient coding method to process signals with rich frequency ranges.

[0136] FIG2a is a schematic diagram of an interactive process of an encoding and decoding method according to an embodiment of the present disclosure. As shown in FIG2a, an embodiment of the present disclosure relates to an encoding and decoding method, the method comprising:

[0137] In step S2101 , the encoding device 101 splits an audio signal into at least one group of audio components according to characteristic information of the audio signal.

[0138] In some embodiments, the audio signal is a signal obtained by pre-processing the input signal.

[0139] Optionally, a pre-processing module may be provided in the encoding device 101. The pre-processing module may perform one or more of the following pre-processing operations on the input signal:

[0140] Framing, time-frequency transformation, window length decision, long and short frame decision, time domain noise shaping, frequency domain noise shaping, etc.

[0141] In some embodiments, the feature information includes at least one of the following:

[0142] Frequency band range (frequency band range analysis);

[0143] signal type;

[0144] The frequency band range includes the infrasound range, the audible sound range and the ultrasonic range, and the signal type is used to indicate the signal type of the signal within each frequency band range.

[0145] In some embodiments, the frequency band range or effective frequency band range of the audio signal includes one or more types.

[0146] Optionally, as shown in Figures 2b to 2c, a signal analyzer may be provided in the encoding device 101 to analyze the frequency band range of the audio signal and perform bandwidth detection and band type analysis on the audio components of each frequency band range.

[0147] In one example, the signal analyzer analyzes the audio signal to include at least one of an infrasonic audio component, an audible sound audio component, and an ultrasonic audio component according to different frequency bands of the audio signal, such as the audio signal includes audio components of the above three frequency bands at the same time.

[0148] Alternatively, infrasound signals can be used in areas such as oceanography and strata. They can be used to explore deep-buried mineral deposits, determine the distribution of hot and cold air masses in the stratosphere, detect hidden dangers in operating machinery, and predict natural phenomena such as tsunamis, storms, volcanic eruptions, and magnetic storms. Therefore, infrasound can be used to detect weather, earthquakes, and predict typhoons or tsunamis.

[0149] Ultrasonic signals, with their excellent directionality, strong penetration, easy acquisition of concentrated sound energy, and long propagation distances in water, can be used for distance measurement, speed measurement, cleaning, welding, and stone crushing. Ultrasonic signals can also be applied to ultrasonic welding, ultrasonic chemistry, ultrasonic cleaning, ultrasonic machining (such as drilling, engraving, and polishing), ultrasonic therapy, ultrasonic surgery, ultrasonic cosmetic surgery, or ultrasonic motors and levitation. Ultrasonic frequencies used in medical diagnosis range from 1 to 5 MHz.

[0150] Optionally, the audible sound wave signal includes music or speech that is audible or hearable to human ears.

[0151] Alternatively, audio signals in the medical field, such as brain waves, have a frequency range different from that of sound waves audible to the human ear. Brain waves refer to the electrical oscillations generated by the activity of nerve cells in the human brain. Brain waves can be divided into five categories based on frequency: beta waves (conscious, 14-30Hz), alpha waves (bridge consciousness, 8-14Hz), theta waves (subconscious, 4-8Hz), delta waves (unconscious, below 4Hz) and gamma waves (focusing on something, above 30Hz). The combination of these consciousnesses forms a person's internal and external behavior, emotions and learning performance. The audio encoding algorithm used for the human ear compresses brain wave signals, resulting in the loss of audio components.

[0152] In another example, based on the frequency band analysis in the previous example, the signal analyzer can further analyze the signal type of each frequency band audio component. For example, audible sound wave audio components include silent frames, narrowband frames, wideband frames, ultra-wideband frames, and full-band frames; or audible sound wave audio components include speech frames, music frames, noise frames, etc. The signal type can also include transient signals, quasi-periodic signals, etc.

[0153] Optionally, a splitting module (or Bandwidth splitting module) 2202 is further provided in the encoding device 101. Based on the analysis result of the audio analyzer 2201, the splitting module 2202 divides the audio signal into three parts according to the frequency band analysis result: infrasonic audio component, audible sound audio component and ultrasonic audio component.

[0154] Optionally, the infrasonic audio component and the ultrasonic audio component are used for machine listening (for machine listening or machine listening), and the audible sound audio component includes a part audible to the human ear (for human listening or human listening) or a part used for machine listening.

[0155] In step S2102 , the encoding device 101 encodes the at least one group of audio components according to the encoding core group corresponding to each audio component in the at least one group of audio components.

[0156] In some embodiments, the coding core groups have a mapping relationship with the audio components.

[0157] In one example, as shown in Figures 2b to 2d, the audio components of the audio signal segmentation include: infrasound audio components, audible sound audio components, and ultrasonic audio components. The coding core groups set by the encoding device 101 include an infrasound core group, an audible sound core group, and an ultrasonic core group. There is no limit on the number of each core group. For example, the audible sound core group can be divided into an audible sound core group for human ears and an audible sound core group for machine hearing.

[0158] In this example, any core group can encode the corresponding audio component. For example, the infrasound core group encodes the infrasound audio component, the audible sound core group encodes the audible sound audio component, and the ultrasonic core group encodes the ultrasonic audio component.

[0159] In some embodiments, the encoding core group corresponding to each group of audio components may also be determined according to user configuration information (encoder config).

[0160] Optionally, as shown in FIG2e , the user may specify to use an encoding core group for machine monitoring for encoding.

[0161] In some embodiments, in different encoding core groups, each encoding core group includes at least one encoding core, wherein the metadata is further used to indicate information about the encoding core that performs encoding.

[0162] Optionally, each encoding core group may be divided into different encoding cores according to application scenarios.

[0163] In one example, as shown in Figures 2b to 2d, the infrasound core group includes: a human disease detection coding core and a natural disaster detection coding core, etc. The audible sound wave core group for the human ear includes: a speech coding core and a music coding core, etc., or the audible sound wave core group for machine hearing includes: a speech recognition coding core and an audio classification coding core, etc. The ultrasonic core group includes: an imaging coding core and an object detection coding core, etc. In some embodiments, in the coding core group corresponding to each group of audio components, a coding core that performs encoding is determined in at least one coding core, and encoding of the corresponding audio component is performed based on the coding core that performs encoding.

[0164] Optionally, in combination with the description of the above embodiments, the coding core group corresponding to each group of audio components may include one or more coding cores. During the coding process, one of the coding cores may be selected to perform coding processing on the group of audio components according to the application scenario.

[0165] For example, based on the description of the above example and as shown in Figures 2b to 2d, the input audio signal is split into an infrasonic audio component, an audible sound audio component for the human ear, and an ultrasonic audio component. For the infrasonic audio component, if its application scenario is human disease detection, the human disease detection coding core can be used to encode the infrasonic audio component. For the audible sound audio component, if its application scenario is speech, the speech coding core can be used to encode the audible sound audio component. For the ultrasonic audio component, if its application scenario is imaging, the imaging coding core can be used to encode the ultrasonic audio component.

[0166] In some embodiments, the encoding core group includes a task core, and the task core is used to encode audio signals having the same audio characteristic parameters and / or the same processing tasks, wherein the audio characteristic parameters are extracted from the audio signals;

[0167] The input of the task core includes multiple frequency band ranges.

[0168] Optionally, a coding core group may include only one or more task cores, or one or more coding cores, or at least one task core and at least one coding core. In this embodiment, a coding core group may include both task cores and coding cores.

[0169] Optionally, the task core encodes the audio signal, which may be encoding a group or part of the audio components after the audio signal is split.

[0170] Optionally, a task core can implement or integrate the functions of multiple encoding cores. For example, a task core can process signals within the frequency bands that can be processed by multiple encoding cores. For example, an audio signal is split into multiple audio components. If the frequency bands of more than one audio component are within the range that can be processed by the task core, then the more than one audio component can be encoded by the task core to simplify the code and data processing when multiple encoding cores are used for processing. Alternatively, if the frequency band of the input audio signal is within the range that can be processed by the task core, that is, the audio signal corresponds to one audio component, then the task core can be used to encode the audio signal.

[0171] Optionally, the same audio characteristic parameters may refer to the same compression characteristics or the same audio characteristics, for example, the audio characteristic parameters are Mel-Frequency Cepstrum (MFCC), and the task core may process the audio signal or audio component with the same MFCC. In one example, as shown in reference figure 2f, the audible sound wave core group for machine hearing includes three coding cores: a speech recognition coding core, an audio classification coding core, and a speaker identification coding core. The above three coding cores correspond to the same MFCC, and the audio components processed by the three coding cores satisfy the same audio characteristic parameters. Therefore, the audio components corresponding to the above three coding cores can be processed by the MFCC task core to realize the functions of multiple cores and simplify code and data processing.

[0172] Optionally, taking the terminal device as an example, the same processing task may refer to the same terminal task. For example, if several audio components after the audio signal is split or the entire audio signal (corresponding to one audio component) is used for sound recognition, the task core can be used to process the several audio components or audio signals. For example, for the first processing, the audio signal or audio component of the frequency band range FRA can be input to the task core for speech recognition; for the second processing, the audio signal or audio component of the frequency band range FRB can be input to the task core for speaker recognition; thus, the task core can process the audio signals or audio components of FRA and FRB. Optionally, audio components of different frequency ranges are encoded according to the task core. For example, as shown in Figures 2f to 2g, the task core acts as a general core to process the tasks of different encoding cores.

[0173] In some embodiments, the encoding methods used by each encoding core group may be different, including lossy encoding methods or lossless encoding methods.

[0174] Optionally, the lossy coding method can perform frequency screening on the input audio signal according to the frequency range audible to the human ear and the masking effect, and perform audio encoding so that the human ear is not easily aware of the distortion. For example, the Advanced Audio Coding (AAC) compression algorithm in the Moving Picture Experts Group (MPEG) standard. The AAC compression algorithm uses a psychoacoustic model to achieve lossy compression. The compression algorithm considers the importance of audio components based on the sensitive frequency bands and masking effects of the human ear, and compresses signal components that are not obvious to the human ear. Another example is the Enhanced Voice Services (EVS) algorithm of the 3rd Generation Partnership Project (3GPP). The algorithm considers the importance of audio components based on the sensitive frequency bands and masking effects of the human ear, and compresses signal components that are not obvious to the human ear.

[0175] Optionally, a lossless encoding method may be used, such as the Scalable to Lossless (SLS) algorithm in MPEG-4.

[0176] Optionally, the embodiment of the present disclosure may adopt a lossy coding method for encoding.

[0177] In step S2103 , the encoding device 101 outputs a bitstream according to the encoding results and metadata corresponding to at least one group of audio components.

[0178] Optionally, the metadata is used to indicate information about a coding core group performing encoding and splitting information corresponding to the audio signal.

[0179] In some embodiments, after the signal is split or divided in step S2101 and the encoding core group is selected in step S2102, metadata is generated. The metadata indicates audio segmentation information or splitting information and information about the encoding core group used for encoding. The metadata is written to the bitstream so that the decoding device 102 can select an appropriate decoding core group based on the metadata and can determine how to merge audio components.

[0180] Optionally, the split information includes at least one of the following:

[0181] The number or groups of audio components;

[0182] The starting frequency of the audio component;

[0183] The stop frequency of the audio component.

[0184] Based on the split information, the decoding end can perform adaptive merging processing.

[0185] In some embodiments, this step may include:

[0186] Encoding metadata to obtain encoding parameters;

[0187] Output the bitstream based on the encoding result and encoding parameters.

[0188] Alternatively, the encoding of the metadata may be a linear mapping, such as writing the encoding or identification of the selected encoding core directly into the bitstream.

[0189] The metadata may also be encoded using Huffman coding. For example, the metadata may be encoded using the following method:

[0190] There are N coding cores for human monitoring and 1 core for machine monitoring. The N+1 coding cores have different usage frequencies, and the N coding cores are categorized based on their usage frequencies. For example, during metadata encoding, the two most frequently used coding cores are encoded with 1 bit, while the third most frequently used coding core is encoded with more bits. Thus, different coding cores have corresponding numbers of bits based on their usage frequencies. Based on the selection results of different coding cores, the corresponding number of bits is used for encoding. If the most frequently used coding core is selected, 1 bit can be used to encode the corresponding metadata encoding parameters. This effectively balances or controls the bits occupied by metadata or encoding parameters based on probability during the actual encoding process.

[0191] In step S2104 , the decoding device 102 determines encoding results corresponding to metadata and at least one group of audio components according to the bitstream.

[0192] Optionally, the decoding device 102 demultiplexes the bit stream to obtain the encoding result, ie, the audio compression data and metadata.

[0193] Optionally, the decoding device 102 determines metadata according to the bitstream, which may include:

[0194] The decoding device 102 obtains encoding parameters according to the bit stream, where the encoding parameters are obtained by encoding the metadata;

[0195] The decoding device 102 decodes the encoding parameters to obtain metadata.

[0196] Step S2105 : The decoding device 102 determines a decoding core group corresponding to at least one group of audio components according to the metadata.

[0197] Optionally, the metadata is decoded to obtain the coding core group and splitting information corresponding to the splitting of the audio signal into at least one group of audio components, so as to facilitate the decoding device 102 to determine the corresponding decoding core group.

[0198] Optionally, the split information includes at least one of the following:

[0199] The number or groups of audio components;

[0200] The starting frequency of the audio component;

[0201] The stop frequency of the audio component.

[0202] Optionally, the decoding device 102 may decode the encoding parameters to obtain metadata.

[0203] Optionally, there is a correspondence or mapping relationship between the decoding core group and the encoding core group. In different decoding core groups, each decoding core group includes at least one decoding core, wherein the metadata is also used to indicate information of the encoding core that performs encoding.

[0204] For example, as shown in Figure 2h, the decoding core group includes: a decoding core group for human hearing; and a decoding core group for machine monitoring. The decoding core group for human hearing includes speech decoding cores, music decoding cores, and noise decoding cores, while the decoding core group for machine monitoring includes task 1 decoding cores and task 2 decoding cores.

[0205] Optionally, the decoding core group includes a task core, which is used to decode audio signals with the same audio characteristic parameters and / or the same processing tasks, wherein the audio characteristic parameters are extracted from the audio signals; wherein the input of the task core includes multiple frequency band ranges.

[0206] Optionally, the decoding device 102 determines, according to the metadata, a decoding core in the decoding core group corresponding to each group of audio components and a decoding core that performs decoding.

[0207] Step S2106 : The decoding device 102 outputs a decoding result according to the decoding core group corresponding to the at least one group of audio components.

[0208] Optionally, the decoding result, i.e., decoded data of the audio signal, is passed through a post-processing module to obtain a final decoded audio signal. The post-processing module performs one or more of the following operations on the audio signal: inverse time-frequency transform, inverse time-domain noise shaping transform, and inverse frequency-domain noise shaping transform.

[0209] In some embodiments, the names of information, etc. are not limited to the names described in the embodiments, and terms such as "information", "message", "signal", "signaling", "report", "configuration", "indication", "instruction", "command", "channel", "parameter", "domain", and "field" can be used interchangeably.

[0210] In some embodiments, "obtain", "get", "get", "receive", "transmit", "bidirectional transmission", "send and / or receive" can be interchangeable, and can be interpreted as receiving from other entities, obtaining from protocols, obtaining from higher layers, obtaining by self-processing, autonomous implementation, etc.

[0211] In some embodiments, terms such as "certain", "preset", "preset", "setting", "indicated", "a certain", "any", and "first" can be interchangeable. "Specific A", "preset A", "preset A", "setting A", "indicated A", "a certain A", "any A", and "first A" can be interpreted as A pre-specified in a protocol, etc., or as A obtained through setting, configuration, or indication, etc., or as specific A, a certain A, any A, or first A, etc., but not limited to this.

[0212] The method involved in the embodiment of the present disclosure may include at least one of steps S2101 to S2106. For example, steps S2101 to S2103 may be implemented as independent embodiments, and steps S2104 to S2106 may be implemented as independent embodiments, but are not limited thereto.

[0213] In some embodiments, reference may be made to other optional implementations described before or after the description corresponding to FIG. 2 a .

[0214] FIG3 is a flow chart of an encoding method according to an embodiment of the present disclosure. As shown in FIG3 , the embodiment of the present disclosure relates to an encoding method, which includes:

[0215] In step S3101 , the encoding device 101 splits an audio signal into at least one group of audio components according to characteristic information of the audio signal.

[0216] Optionally, the implementation of step S3101 can refer to the optional implementation of step S2101, which will not be repeated here.

[0217] In some embodiments, the feature information includes at least one of the following:

[0218] Frequency band range;

[0219] Signal type;

[0220] The frequency band range includes the infrasound range, the audible sound range and the ultrasonic range, and the signal type is used to indicate the signal type of the signal within each frequency band range.

[0221] In some embodiments, the audio signal is a signal obtained by pre-processing the input signal.

[0222] In step S3102 , the encoding device 101 encodes the at least one group of audio components according to the encoding core group corresponding to each audio component in the at least one group of audio components.

[0223] Optionally, the implementation of step S3102 can refer to the optional implementation of step S2102, which will not be repeated here.

[0224] Optionally, the coding core group and the audio component have a mapping relationship.

[0225] In some embodiments, encoding the at least one group of audio components according to the coding kernel group corresponding to each audio component in the at least one group of audio components includes:

[0226] In the encoding core group corresponding to each group of audio components, an encoding core that performs encoding in at least one encoding core is determined, and encoding of the corresponding audio component is performed based on the encoding core that performs encoding.

[0227] In some embodiments, the encoding core group includes a task core, and the task core is used to encode audio signals having the same audio characteristic parameters and / or the same processing tasks, wherein the audio characteristic parameters are extracted from the audio signals;

[0228] The input of the task core includes multiple frequency bands.

[0229] In some embodiments, the method further comprises:

[0230] The audio components in different frequency ranges are encoded according to the task kernel.

[0231] In some embodiments, the method further comprises:

[0232] The coding core group corresponding to each group of audio components is determined according to the user configuration information.

[0233] In step S3103 , the encoding device 101 outputs a bitstream according to the encoding results and metadata corresponding to at least one group of audio components.

[0234] Optionally, the implementation of step S3103 can refer to the optional implementation of step S2103, which will not be repeated here.

[0235] Optionally, the metadata is used to indicate information of a coding core group for performing encoding and splitting information corresponding to splitting the audio signal into at least one group of audio components.

[0236] Optionally, the split information includes at least one of the following:

[0237] The number or groups of audio components;

[0238] The starting frequency of the audio component;

[0239] The stop frequency of the audio component.

[0240] In some embodiments, in different encoding core groups, each encoding core group includes at least one encoding core, wherein the metadata is further used to indicate information about the encoding core that performs encoding.

[0241] In some embodiments, outputting a bitstream according to encoding results and metadata corresponding to at least one group of audio components includes:

[0242] Encoding metadata to obtain encoding parameters;

[0243] Output the bitstream based on the encoding result and encoding parameters.

[0244] In some embodiments, reference may be made to other optional implementations described before or after the description corresponding to FIG. 3 .

[0245] FIG4 is a flow chart of a decoding method according to an embodiment of the present disclosure. As shown in FIG4 , an embodiment of the present disclosure relates to a decoding method, which includes:

[0246] In step S4101 , the decoding device 101 determines encoding results corresponding to metadata and at least one group of audio components according to a bitstream.

[0247] Optionally, the implementation of step S4101 can refer to the optional implementation of step S2104, which will not be repeated here.

[0248] Optionally, the metadata is used to indicate information of a coding core group for performing encoding and splitting information corresponding to splitting the audio signal into at least one group of audio components.

[0249] Optionally, the split information includes at least one of the following:

[0250] The number or groups of audio components;

[0251] The starting frequency of the audio component;

[0252] The stop frequency of the audio component.

[0253] In some embodiments, the audio components are separated based on characteristic information of the audio signal, where the characteristic information includes at least one of the following:

[0254] Frequency band range;

[0255] Signal type;

[0256] The frequency band range includes the infrasound range, the audible sound range and the ultrasonic range, and the signal type is used to indicate the signal type of the signal within each frequency band range.

[0257] In some embodiments, the method further comprises:

[0258] Decode the encoding parameters to obtain metadata.

[0259] In some embodiments, the audio signal is a signal obtained by pre-processing the input signal.

[0260] Step S4102: The decoding device 101 determines a decoding core group corresponding to at least one group of audio components according to the metadata.

[0261] Optionally, the implementation of step S4102 can refer to the optional implementation of step S2105, which will not be repeated here.

[0262] In some embodiments, in different decoding core groups, each decoding core group includes at least one decoding core, wherein the metadata is further used to indicate information about the encoding core that performs encoding.

[0263] In some embodiments, the method further comprises:

[0264] A decoding core that performs decoding in a decoding core group corresponding to each group of audio components is determined according to the metadata.

[0265] Optionally, metadata is determined based on the bitstream, including:

[0266] Obtaining encoding parameters according to a bitstream, where the encoding parameters are obtained by encoding the metadata;

[0267] The encoding parameters are decoded to obtain the metadata.

[0268] Step S4103 : The decoding device 101 outputs a decoding result according to the decoding core group corresponding to at least one group of audio components.

[0269] Optionally, the implementation of step S4103 can refer to the optional implementation of step S2106, which will not be repeated here.

[0270] In some embodiments, optionally, the decoding core group includes a task core, wherein the task core is used to decode audio signals having the same audio characteristic parameters and / or the same processing tasks, wherein the audio characteristic parameters are extracted from the audio signals;

[0271] The input of the task core includes multiple frequency bands.

[0272] In some embodiments, the method further comprises:

[0273] The decoding core group corresponding to each group of audio components is determined according to the user configuration information.

[0274] In some embodiments, reference may be made to other optional implementations described before or after the description corresponding to FIG. 4 .

[0275] In the method of the embodiment of the present disclosure, an audio signal with a frequency range exceeding the audible frequency range is split into different frequency components, such as infrasound, audible sound, ultrasonic wave, etc., and encoded using different coding cores. Metadata is generated based on signal splitting and coding core selection, and the metadata is written into a bitstream, which is used in the decoder to select a corresponding decoding core and combine the separated audio components. In this embodiment, different coding cores can be used according to the sound frequency range, different application scenarios, different audio task specifics or different audio characteristics. If the task characteristics are the same in some application scenarios, the same coding core can be used. In signal splitting, bandwidth detection is used to obtain the audio range, and signal feature analysis and classification are used to obtain audio features.

[0276] To facilitate understanding of the present disclosure, some specific examples are listed below:

[0277] Example 1:

[0278] The audio signal is a pre-processed signal of the input audio signal. The effective frequency band of the signal includes one or more sound wave ranges, that is, it may include infrasound, audible sound and ultrasonic sound at the same time.

[0279] Optionally, the preprocessing module performs one or more of the following operations on the audio signal: framing, time-frequency transformation, window length determination, long frame and short frame determination, time domain noise shaping, frequency domain noise shaping, etc.

[0280] Optionally, a signal analyzer analyzes the audio input signal. The signal analyzer includes bandwidth detection and frequency band type analysis for each frequency band. The frequency band analysis results include: infrasound, audible sound, and ultrasonic waves. The frequency band type analysis for each frequency band further analyzes the frequency band. For example, the frequency band analysis results for audible sound waves include: silent frame, narrowband frame, wideband frame, ultra-wideband frame, and full-band frame.

[0281] Optionally, the signal analyzer also includes signal type analysis within each frequency band. For example, the signal type analysis results within the audible sound wave range include: speech frames, music frames, noise frames, etc. The signal analyzer can analyze signal characteristics such as transient signals, quasi-periodic signals, etc.

[0282] Optionally, the bandwidth segmentation module divides the audio signal into three parts according to the result of bandwidth detection: an infrasound component, an audible sound component, and an ultrasonic component.

[0283] Optionally, after splitting the signal and core selection, metadata is generated and needs to be written to the bitstream to let the decoder know which decoding core to select and how to merge the audio components.

[0284] Optionally, the decoding end demultiplexes the bitstream to obtain compressed audio data and metadata information. The metadata information is decoded to obtain core selection information indicating the corresponding decoding core. Decoding cores fall into two categories: decoding cores for human hearing and decoding cores for machine listening. Human hearing decoding cores include speech decoding cores, music decoding cores, and noise decoding cores. Machine task decoding cores include task 1 decoding cores, task 2 decoding cores, and so on.

[0285] Optionally, the decoded data of the audio signal is passed through a post-processing module to obtain a final decoded audio signal. The post-processing module performs one or more of the following operations on the audio signal: inverse time-frequency transform, inverse time-domain noise shaping transform, inverse frequency-domain noise shaping transform, etc.

[0286] Example 2:

[0287] The audio signal is a pre-processed signal of the input audio signal. The effective frequency band of the signal includes one or more sound wave ranges, that is, it may include infrasound, audible sound and ultrasonic sound at the same time.

[0288] Optionally, the preprocessing module performs one or more of the following operations on the audio signal: framing, time-frequency transformation, window length determination, long frame and short frame determination, time domain noise shaping, frequency domain noise shaping, etc.

[0289] Optionally, a signal analyzer analyzes the audio input signal. The signal analyzer performs bandwidth detection and frequency band type analysis for each frequency band. Frequency band analysis results include: infrasound, audible sound, and ultrasonic waves. Frequency band type analysis for each frequency band further analyzes the frequency band. For example, frequency band analysis results for audible sound waves include: silent frame, narrowband frame, wideband frame, ultra-wideband frame, and full-band frame.

[0290] Optionally, the bandwidth segmentation module separates the audio signal into three components based on the bandwidth detection results: infrasound, audible sound, and ultrasonic sound. After segmenting the signal, metadata is generated and written to the bitstream to let the decoder know how to combine the audio components.

[0291] Optionally, the decoding end demultiplexes the bitstream to obtain the audio compression data and metadata information. The metadata information is decoded to obtain the audio segmentation auxiliary information. Bandwidth combination is performed based on the audio segmentation auxiliary information.

[0292] Example 3:

[0293] The audio signal is a pre-processed signal of the input audio signal. The effective frequency band of the signal includes one or more sound wave ranges, that is, it may include infrasound, audible sound and ultrasonic sound at the same time.

[0294] Optionally, the preprocessing module performs one or more of the following operations on the audio signal: framing, time-frequency transformation, window length determination, long and short frame determination, time domain noise shaping, frequency domain noise shaping, etc.

[0295] Optionally, there are different groups of encoding cores, each with one or more encoding cores.

[0296] Optionally, the coding core groups are divided according to frequency ranges, such as four core groups, including an infrasound core group, an audible sound core group for human hearing, an audible sound core group for machine hearing, and an ultrasonic core group.

[0297] The different cores within a group are divided according to application scenarios. For example, each core group contains many different cores. The infrasound core group includes cores such as human disease detection and natural disaster detection. The audible sound core group for humans includes cores such as speech and music. The audible sound core group for machines includes cores such as speech recognition and audio classification. The ultrasound core group includes cores such as imaging and object detection.

[0298] Optionally, after core group and core selection, metadata will be generated and needs to be written to the bitstream to let the decoder know which core to decode the bitstream.

[0299] Optionally, the decoding end demultiplexes the bitstream to obtain the audio compression data and metadata information. The metadata information is decoded to obtain the audio core group and core information. A decoding core is selected based on the metadata.

[0300] Optionally, the three task encoding cores in the audible sound core group for machine hearing, such as the speech recognition core, audio classification core, and speaker recognition core, all use the same features (MFCC). Therefore, the task encoding cores are combined into one MFCC task core because they use the same task features.

[0301] Alternatively, audio segmentation is an option for the encoding core processing. For example, if the audio signal only has an ultrasonic component, the audio signal does not need to be separated.

[0302] Optionally, regarding the selection of the coding core, in addition to the adaptive selection based on frequency range and signal characteristics, the user can also specify the coding core through coding configuration, for example, configuration 0 indicates for manual monitoring, and configuration 1 indicates for machine monitoring.

[0303] Example 4:

[0304] The audio signal is a preprocessed signal of the input audio signal.

[0305] Optionally, the preprocessing module performs one or more of the following operations on the audio signal: framing, time-frequency transformation, window length determination, long and short frame determination, time domain noise shaping, frequency domain noise shaping, etc.

[0306] Optionally, there are different encoding core groups, each group having one or more encoding cores.

[0307] Optionally, the coding core groups are divided according to frequency ranges, such as four core groups, including an infrasound core group, an audible sound core group for human hearing, an audible sound core group for machine hearing, and an ultrasonic core group.

[0308] Optionally, the different cores within a group are divided based on signal characteristics. For example, each core group can contain many different cores. The infrasound core group includes quasi-periodic signal cores and transient signal cores. The audible sound core group for human hearing includes speech cores and harmonic cores. The audible sound core group for machine hearing includes speech cores and harmonic cores. The ultrasonic core group includes quasi-periodic signal cores and transient signal cores.

[0309] Alternatively, identical features can be combined into a common core group, and the signal information can be used as input to the core encoding.

[0310] Among them, the signal information includes the signal target, which can be 0 for manual monitoring and 1 for machine monitoring;

[0311] The signal information also includes a signal frequency range, such as 0 for infrasound, 1 for audible sound, and 2 for ultrasonic wave, or the signal frequency range can also be expressed as a minimum frequency value (Hz): 20, a maximum frequency value (Hz): 20000.

[0312] The embodiments of the present disclosure further provide an apparatus for implementing any of the above methods. For example, an apparatus is provided, and the apparatus includes units or modules for implementing each step performed by the communication device in any of the above methods.

[0313] It should be understood that the division of the various units or modules in the above device is merely a division of logical functions. In actual implementation, they may be fully or partially integrated into a physical entity, or they may be physically separated. In addition, the units or modules in the device may be implemented in the form of a processor calling software: for example, the device includes a processor, the processor is connected to a memory, and the memory stores instructions. The processor calls the instructions stored in the memory to implement any of the above methods or implement the functions of the various units or modules of the above device, wherein the processor is, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory is a memory within the device or a memory outside the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits, and the functions of some or all of the units or modules can be realized by designing the hardware circuits. The above-mentioned hardware circuits can be understood as one or more processors; for example, in one implementation, the above-mentioned hardware circuit is an application-specific integrated circuit (ASIC), which realizes the functions of some or all of the above units or modules by designing the logical relationship of the components in the circuit; for example, in another implementation, the above-mentioned hardware circuit can be realized by a programmable logic device (PLD). Taking a field programmable gate array (FPGA) as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by configuring the configuration file, thereby realizing the functions of some or all of the above units or modules. All units or modules of the above devices can be realized in the form of software called by the processor, or in the form of hardware circuits, or in part by the form of software called by the processor, and the rest by hardware circuits.

[0314] In the embodiments of the present disclosure, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationship of the hardware circuit. The logical relationship of the above-mentioned hardware circuit is fixed or reconfigurable. For example, the processor is a hardware circuit implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and implementing the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), etc.

[0315] Figure 5a is a schematic diagram of the structure of the encoding device proposed in an embodiment of the present disclosure. As shown in Figure 5a, the encoding device 5100 may include: at least one of a transceiver module 5101, a processing module 5102, etc. In some embodiments, the processing module 5102 is used to split the audio signal into at least one group of audio components based on the characteristic information of the audio signal; the processing module is also used to encode the at least one group of audio components according to the encoding core group corresponding to each group of audio components in the at least one group of audio components, wherein the encoding core group and the audio component have a mapping relationship; the processing module is also used to output a bitstream according to the encoding result corresponding to the at least one group of audio components and metadata, wherein the metadata is used to indicate the encoding core group information for performing the encoding and the splitting information corresponding to the splitting of the audio signal into at least one group of audio components.

[0316] Optionally, the transceiver module 5101 is used to perform at least one of the communication steps such as sending and / or receiving performed by the encoding device 5100 in any of the above methods, which will not be described in detail here. Optionally, the processing module is used to perform at least one of the other steps performed by the encoding device 5100 in any of the above methods, which will not be described in detail here.

[0317] FIG5b is a schematic diagram of the structure of a decoding device proposed in an embodiment of the present disclosure. As shown in FIG5b, the decoding device 5200 may include: a transceiver module 5201, a processing module 5202, etc. In some embodiments, the processing module 5202 is used to determine metadata and encoding results corresponding to at least one group of audio components based on the bitstream, wherein the metadata is used to indicate encoding core group information for performing encoding and splitting information corresponding to the splitting of the audio signal into at least one group of audio components; the processing module is further used to determine the decoding core group corresponding to the at least one group of audio components based on the metadata; the processing module is further used to output a decoding result based on the decoding core group corresponding to the at least one group of audio components.

[0318] Optionally, the transceiver module 5201 is configured to execute at least one of the communication steps, such as sending and / or receiving, performed by the decoding device 5200 in any of the above methods, and will not be described in detail here. Optionally, the processing module is configured to execute at least one of the other steps performed by the decoding device 5200 in any of the above methods, and will not be described in detail here.

[0319] Figure 6a is a schematic diagram of the structure of a communication device 6100 proposed in an embodiment of the present disclosure. Communication device 6100 can be a terminal (e.g., user equipment), or a chip, chip system, or processor that supports implementation of any of the above methods. Communication device 6100 can be used to implement the methods described in the above method embodiments. For details, please refer to the description of the above method embodiments.

[0320] As shown in Figure 6a, the communication device 6100 includes one or more processors 6101. The processor 6101 can be a general-purpose processor or a dedicated processor, for example, a baseband processor or a central processing unit. The baseband processor can be used to process the communication protocol and communication data, and the central processing unit can be used to control the communication device (such as a base station, a baseband chip, a terminal device, a terminal device chip, a DU or a CU, etc.), execute programs, and process program data. Optionally, the communication device 6100 is used to perform any of the above methods. Optionally, one or more processors 6101 are used to call instructions to enable the communication device 6100 to perform any of the above methods.

[0321] In some embodiments, the communication device 6100 further includes one or more transceivers 6102. When the communication device 6100 includes one or more transceivers 6102, the transceiver 6102 performs at least one of the communication steps, such as sending and / or receiving, in the above-described method, and the processor 6101 performs at least one of the other steps. In an optional embodiment, the transceiver may include a receiver and / or a transmitter, and the receiver and transmitter may be separate or integrated. Optionally, the terms transceiver, transceiver unit, transceiver, transceiver circuit, interface circuit, and interface may be used interchangeably; the terms transmitter, transmitting unit, transmitter, and transmitting circuit may be used interchangeably; and the terms receiver, receiving unit, receiver, and receiving circuit may be used interchangeably.

[0322] In some embodiments, the communication device 6100 further includes one or more memories 6103 for storing data. Alternatively, all or part of the memories 6103 may be located outside the communication device 6100. In alternative embodiments, the communication device 6100 may include one or more interface circuits 6104. Optionally, the interface circuits 6104 are connected to the memory 6102 and may be configured to receive data from the memory 6102 or other devices, or to send data to the memory 6102 or other devices. For example, the interface circuits 6104 may read data stored in the memory 6102 and send the data to the processor 6101.

[0323] A communication device may be an independent device or a part of a larger device. For example, the communication device may be: 1) an independent integrated circuit (IC), or a chip, or a chip system or subsystem; (2) a collection of one or more ICs, optionally including a storage component for storing data or programs; (3) an ASIC, such as a modem; (4) a module that can be embedded in other devices; (5) a receiver, a terminal device, an intelligent terminal device, a cellular phone, a wireless device, a handheld device, a mobile unit, an in-vehicle device, a network device, a cloud device, an artificial intelligence device, etc.; (6) others, etc.

[0324] FIG6b is a schematic diagram of the structure of a chip 6200 according to an embodiment of the present disclosure. If the communication device 6100 can be a chip or a chip system, reference can be made to the schematic diagram of the structure of the chip 6200 shown in FIG6b , but the present disclosure is not limited thereto.

[0325] The chip 6200 includes one or more processors 6201. The chip 6200 is configured to execute any of the above methods.

[0326] In some embodiments, chip 6200 further includes one or more interface circuits 6202. Terms such as interface circuit, interface, and transceiver pins may be used interchangeably. In some embodiments, chip 6200 further includes one or more memories 6203 for storing data. Alternatively, all or part of memory 6203 may be located external to chip 6200. Optionally, interface circuit 6202 is connected to memory 6203 and may be used to receive data from memory 6203 or other devices, or may be used to send data to memory 6203 or other devices. For example, interface circuit 6202 may read data stored in memory 6203 and send the data to processor 6201.

[0327] In some embodiments, the interface circuit 6202 performs at least one of the communication steps, such as sending and / or receiving, in the above-described method. For example, the interface circuit 6202 performing the communication steps, such as sending and / or receiving, in the above-described method means that the interface circuit 6202 performs data exchange between the processor 6201, the chip 6200, the memory 6203, or the transceiver device. In some embodiments, the processor 6201 performs at least one of the other steps.

[0328] The modules and / or devices described in various embodiments, such as virtual devices, physical devices, and chips, can be arbitrarily combined or separated according to circumstances. Optionally, some or all steps can also be performed collaboratively by multiple modules and / or devices, which is not limited here.

[0329] The present disclosure also proposes a storage medium having instructions stored thereon. When the instructions are executed on the communication device 6100, the communication device 6100 executes any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but is not limited thereto and may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but is not limited thereto and may also be a transient storage medium.

[0330] The present disclosure also provides a program product, which, when executed by the communication device 6100, enables the communication device 6100 to perform any of the above methods. Optionally, the program product is a computer program product.

[0331] The present disclosure also proposes a computer program, which, when executed on a computer, causes the computer to perform any one of the above methods. Industrial Applicability

[0332] The input audio signal is split based on feature information, and the split audio components are encoded using corresponding coding core groups. This helps select the appropriate encoding method based on the characteristics of different audio components, improving encoding accuracy and rationality. In addition, metadata is written during the encoding process to facilitate accurate decoding on the decoder.

Claims

1. A coding method, the method comprising: Splitting the audio signal into at least one group of audio components according to the characteristic information of the audio signal; Encoding the at least one group of audio components respectively according to the corresponding coding kernel groups of each group of audio components in the at least one group of audio components, wherein there is a mapping relationship between the coding kernel groups and the audio components; Outputting a bitstream according to the coding results corresponding to the at least one group of audio components and the metadata, wherein the metadata is used to indicate the coding kernel group information for performing encoding and the splitting information corresponding to splitting the audio signal into at least one group of audio components.

2. The method according to claim 1, wherein The characteristic information includes at least one of the following: Frequency band range; Signal type; Wherein, the frequency band range includes infrasonic range, audible sound range and ultrasonic range, and the signal type is used to indicate the signal type of the signal within each frequency band range.

3. The method according to claim 1, wherein, In different coding kernel groups, each coding kernel group includes at least one coding kernel, wherein the metadata is further used to indicate the coding kernel information for performing encoding.

4. The method according to claim 3, wherein The outputting a bitstream according to the coding results corresponding to the at least one group of audio components and the metadata includes: Encoding the metadata to obtain coding parameters; Outputting a bitstream according to the coding results and the coding parameters.

5. The method according to claim 3, wherein, The encoding the at least one group of audio components respectively according to the corresponding coding kernel groups of each group of audio components in the at least one group of audio components includes: In the coding kernel group corresponding to each group of audio components, determining the coding kernel for performing encoding among the at least one coding kernel, and performing encoding on the corresponding audio component based on the coding kernel for performing encoding.

6. The method according to claim 1 or 3, wherein, The coding kernel group includes task kernels, and the task kernels are used to encode audio signals with the same audio characteristic parameters and / or the same processing tasks, wherein the audio characteristic parameters are extracted from the audio signals; Wherein, the input of the task kernel includes multiple frequency band ranges.

7. The method according to claim 1, wherein, The method further includes: Determining the coding kernel group corresponding to each group of audio components according to the user configuration information.

8. The method according to any one of claims 1 to 7, wherein, The splitting information includes at least one of the following: The number or number of groups of audio components; The starting frequency of the audio component; The ending frequency of the audio component.

9. The method according to any one of claims 1 to 7, wherein, The audio signal is a signal after preprocessing the input signal.

10. A decoding method, the method comprising: Determining the metadata and the coding results corresponding to at least one group of audio components according to the bitstream, wherein the metadata is used to indicate the coding kernel group information for performing encoding and the splitting information corresponding to splitting the audio signal into at least one group of audio components; Determining the decoding kernel group corresponding to the at least one group of audio components according to the metadata; Outputting a decoding result according to the decoding kernel group corresponding to the at least one group of audio components.

11. The method according to claim 10, wherein, The audio components are split based on the characteristic information of the audio signal, and the characteristic information includes at least one of the following: Frequency band range; Signal type; Among them, the frequency band range includes the infrasonic wave range, the audible sound wave range, and the ultrasonic wave range, and the signal type is used to indicate the signal type of the signal in each frequency band range.

12. The method according to claim 10, wherein, In different decoding core groups, each decoding core group includes at least one decoding core, and wherein the metadata is further used to indicate the encoding core information for performing encoding.

13. As claimed in claim 12, wherein, The determining the metadata according to the bitstream includes: Obtaining encoding parameters from the bitstream, where the encoding parameters are obtained by encoding the metadata; Decoding the encoding parameters to obtain the metadata.

14. The method according to claim 12, wherein The method further includes: Determining, according to the metadata, the decoding core in the decoding core group corresponding to each group of the audio components that performs decoding.

15. The method according to claim 10 or 12, wherein, The decoding core group includes task cores, and the task cores are used to decode audio signals with the same audio characteristic parameters and / or the same processing tasks, where the audio characteristic parameters are extracted from the audio signals; Among them, the input of the task core includes multiple frequency band ranges.

16. The method according to claim 10, wherein, The method further includes: Determining the decoding core group corresponding to each group of the audio components according to the user configuration information.

17. The method according to any one of claims 10 to 16, wherein, The splitting information includes at least one of the following: The number or number of groups of audio components; The starting frequency of the audio component; The ending frequency of the audio component.

18. An encoding device, comprising: A processing module, configured to split the audio signal into at least one group of audio components according to the characteristic information of the audio signal; The processing module is further configured to encode each group of the at least one group of audio components respectively according to the encoding core group corresponding to each group of the at least one group of audio components, where there is a mapping relationship between the encoding core group and the audio component; The processing module is further configured to output a bitstream according to the encoding result corresponding to the at least one group of audio components and the metadata, where the metadata is used to indicate the encoding core group information for performing encoding and the splitting information corresponding to the process of splitting the audio signal into audio components.

19. A decoding device, comprising: A processing module, configured to determine the metadata and the encoding result corresponding to at least one group of audio components according to the bitstream, where the metadata is used to indicate the encoding core group information for performing encoding and the splitting information corresponding to the process of splitting the audio signal into audio components; The processing module is further configured to determine the decoding core group corresponding to the at least one group of audio components according to the metadata; The processing module is further configured to output a decoding result according to the decoding core group corresponding to the at least one group of audio components.

20. A communication device, comprising: One or more processors; Among them, the communication device is configured to execute the method according to any one of claims 1 to 9 or any one of claims 10 to 17.

21. A storage medium, where the storage medium stores instructions, and wherein, When the instructions run on the communication device, the communication device is caused to execute the method according to any one of claims 1 to 9 or any one of claims 10 to 17.

22. A communication system, comprising: An encoding device and a decoding device, wherein, The encoding device is configured to execute the method according to any one of claims 1 to 9; The decoding device is used to perform the method according to any one of claims 10 to 17.

Citation Information

Patent Citations

  • Audio scene encoder, audio scene decoder and related methods using hybrid encoder / decoder spatial analysis

    CN112074902A

  • Method and system for coding metadata in audio streams and for efficient bitrate allocation to audio streams coding

    CN114072874A

  • Signal coding and decoding method and device, coding equipment, decoding equipment and storage medium

    CN114127844A

  • Audio processing method and device, computer equipment, storage medium and program product

    CN115130569A

  • Method and apparatus for encoding / decoding mpeg-4 bsac audio bitstream having auxillary information

    CN1684523A