Audio processing method, decoding method, coding method, device, equipment and medium

By dividing the original audio into high-frequency and low-frequency sub-audio and using deep learning models for feature extraction, the problem of low degree of audio processing in the prior art is solved, and audio performance improvement and encoding and decoding efficiency optimization are achieved.

CN120020945APending Publication Date: 2025-05-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311546039.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

In the prior art, the degree of processing refinement is low in the process and it is difficult to improve audio performance.

Method used

The target audio is determined by dividing the original audio into sub-audios of the high-frequency area and the low-frequency area, and performing feature extraction processing respectively. The method includes encoding and decoding the trained deep learning model to extract its features.

Benefits of technology

It improves the degree of refinement of audio processing, improves audio performance, and optimizes audio encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020945A_ABST
    Figure CN120020945A_ABST
Patent Text Reader

Abstract

The invention provides an audio processing method, an audio decoding method, an audio coding method, an audio processing device, electronic equipment and a computer readable storage medium, and is applied to the technical field of audio processing. The method comprises the following steps: dividing an original audio in a frequency domain, and converting a division result into a time domain, so that the original audio is divided into sub-audios belonging to a high-frequency region and sub-audios belonging to a low-frequency region. And further, feature extraction is performed on the sub-audios belonging to the high-frequency region and the sub-audios belonging to the low-frequency region in a targeted manner, so that the target audio of the original audio can be determined. Compared with the prior art in which the original audio is directly coded and decoded, the audio processing scheme provided by the invention can improve the fine processing degree of the audio and is beneficial to improving the audio performance. Furthermore, compared with encoding and decoding of the original audio, the effect of encoding and decoding of the target audio corresponding to the original audio is improved, and the audio encoding and decoding efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio processing, and in particular, to an audio processing method, an audio decoding method, an audio encoding method, an audio processing device, an electronic device, and a computer-readable storage medium. Background Art

[0002] In the related art, when processing audio, it is often based on the original audio. For example, for audio compression processing, specifically, in the related art, the audio is directly compressed in the following encoding formats, including Pulse Code Modulation (PCM) format, Windows Media Audio (WMA) format, Advanced Audio Coding (AAC) format, Moving Picture Experts Group Audio Layer III (MP3) format, etc. Thus, by encoding the audio, the compression of the audio file can be achieved, thereby reducing the data volume and further reducing the storage and transmission costs of the audio.

[0003] In the related art, the original audio is directly processed, such as performing the encoding process of the above-mentioned encoding formats. However, the related art has the problem of low refinement degree in audio processing, which is not conducive to improving audio performance. Summary of the Invention

[0004] The present application provides an audio processing method, an audio decoding method, an audio encoding method, an audio processing device, an electronic device, and a computer-readable storage medium, which at least to a certain extent improve the refinement degree of audio processing and are conducive to improving audio performance.

[0005] In a first aspect, the present application provides an audio processing method, which includes: dividing the original audio into a plurality of sub-audios, where the frequency domain intervals corresponding to the plurality of sub-audios are different, at least one of the plurality of sub-audios has a frequency domain interval belonging to the high-frequency region, and at least one of the plurality of sub-audios has a frequency domain interval belonging to the low-frequency region; and determining the target audio corresponding to the original audio by respectively performing feature extraction processing on the sub-audios belonging to the high-frequency region and the sub-audios belonging to the low-frequency region.

[0006] In some embodiments, based on the above solution, determining the target audio corresponding to the original audio includes: determining the sub-audios with frequency domain intervals belonging to the high-frequency region among the multiple sub-audios as high-frequency sub-audios, and determining the sub-audios with frequency domain intervals belonging to the low-frequency region among the multiple sub-audios as low-frequency sub-audios; performing a first feature extraction process on the high-frequency sub-audios to obtain restored sub-audios corresponding to the high-frequency sub-audios; merging the restored sub-audios and the low-frequency sub-audios to obtain a merged audio; and performing a second feature extraction process on the merged audio to obtain the target audio corresponding to the original audio.

[0007] In some embodiments, based on the above solution, performing a first feature extraction process on the high-frequency sub-audios to obtain the restored sub-audios corresponding to the high-frequency sub-audios includes: inputting the i-th high-frequency sub-audio into the i-th first encoding and decoding model to perform an encoding process on the i-th high-frequency sub-audio through the i-th first encoding and decoding model, and performing a decoding process on the encoding result of the i-th high-frequency sub-audio, and the i-th first encoding and decoding model outputs the i-th restored sub-audio corresponding to the i-th high-frequency sub-audio, where i is a positive integer not greater than the number of high-frequency sub-audios; wherein, the first encoding and decoding model is a trained deep learning model.

[0008] In some embodiments, based on the above solution, the first encoding and decoding model includes: P encoding units and Q decoding units, where both P and Q are positive integers; wherein, the downsampling multiples corresponding to the P encoding units are the same as the upsampling multiples corresponding to the Q decoding units.

[0009] In some embodiments, based on the above solution, performing a second feature extraction process on the merged audio to obtain the target audio corresponding to the original audio includes: inputting the merged audio into a second encoding and decoding model to perform an encoding process on the merged audio through the second encoding and decoding model, and performing a decoding process on the encoding result of the merged audio, and the second encoding and decoding model outputs the target audio corresponding to the original audio; wherein, the second encoding and decoding model is a trained deep learning model.

[0010] In some embodiments, based on the above solution, after dividing the original audio into multiple sub-audios, the method further includes: performing personalized processing on the target sub-audio, where the target sub-audio corresponds to a target frequency domain interval, and the personalized processing is processing for the target frequency domain interval.

[0011] In some embodiments, based on the above solution, the process of decoding the i-th high-frequency sub-audio by the i-th first codec model includes: performing personalized processing on the i-th high-frequency sub-audio, where the i-th high-frequency sub-audio corresponds to the i-th frequency domain interval, and the personalized processing is processing for the i-th frequency domain interval.

[0012] In some embodiments, based on the above solution, the division of the original audio into multiple sub-audios includes: converting the original audio to the frequency domain to obtain a target frequency domain signal; determining M-1 frequency points according to the distribution of the target frequency domain signal, where M is an integer greater than 1; dividing the target frequency domain signal into M frequency domain intervals through the M-1 frequency points; and converting the M sub-audio signals corresponding to the M frequency domain intervals to the time domain respectively to obtain M sub-audios.

[0013] In some embodiments, based on the above solution, after determining the target audio corresponding to the original audio, the method further includes: encoding the target audio by an audio sending end to obtain a target bitstream; and sending the target bitstream to an audio receiving end by the audio sending end for decoding the target bitstream at the audio decoding end.

[0014] The audio processing method provided by the embodiments of the present application divides the original audio into sub-audios belonging to the high-frequency region and sub-audios belonging to the low-frequency region, and respectively extracts features from the sub-audios belonging to the high-frequency region and the low-frequency region, which can improve the degree of refined processing of the audio and is beneficial to improving the audio performance. Further, compared with encoding and decoding the original audio, the effect of encoding and decoding the target audio corresponding to the original audio is improved, which is beneficial to improving the audio encoding and decoding efficiency.

[0015] In a second aspect, the present application provides an audio processing device, which includes: a sub-audio determination module and a feature extraction module;

[0016] Among them, the sub-audio determination module is used to divide the original audio into multiple sub-audios, where the frequency domain intervals corresponding to the multiple sub-audios are different, at least one of the multiple sub-audios corresponds to a frequency domain interval belonging to the high-frequency region, and at least one of the multiple sub-audios corresponds to a frequency domain interval belonging to the low-frequency region; and the feature extraction module is used to determine the target audio corresponding to the original audio by respectively performing feature extraction processing on the sub-audios belonging to the high-frequency region and the sub-audios belonging to the low-frequency region.

[0017] In some embodiments, based on the foregoing solution, the feature extraction module includes: a screening unit, a first feature extraction unit, a merging unit, and a second feature extraction unit;

[0018] Among them, the above screening unit is used to: determine the sub-audio with the frequency domain interval belonging to the above high-frequency region among the above multiple sub-audios as the high-frequency sub-audio, and determine the sub-audio with the frequency domain interval belonging to the above low-frequency region among the above multiple sub-audios as the low-frequency sub-audio; the above first feature extraction unit is used to: perform a first feature extraction process on the above high-frequency sub-audio to obtain the restored sub-audio corresponding to the above high-frequency sub-audio; the above merging unit is used to: merge the above restored sub-audio and the above low-frequency sub-audio to obtain a merged audio; and, the above second feature extraction unit is used to: perform a second feature extraction process on the above merged audio to obtain the target audio corresponding to the above original audio.

[0019] In some embodiments, based on the foregoing solution, the above first feature extraction unit is specifically used to: input the i-th high-frequency sub-audio into the i-th first encoding and decoding model, so as to perform an encoding process on the i-th high-frequency sub-audio through the i-th first encoding and decoding model, and perform a decoding process on the encoding result of the i-th high-frequency sub-audio, and the i-th first encoding and decoding model outputs the i-th restored sub-audio corresponding to the i-th high-frequency sub-audio, where i takes a positive integer not greater than the number of the above high-frequency sub-audios; among them, the above first encoding and decoding model is a trained deep learning model.

[0020] In some embodiments, based on the foregoing solution, the above first encoding and decoding model includes: P encoding units and Q decoding units, where P and Q take positive integer values; among them, the downsampling factor corresponding to the above P encoding units is the same as the upsampling factor corresponding to the above Q decoding units.

[0021] In some embodiments, based on the foregoing solution, the above second feature extraction unit is specifically used to: input the above merged audio into the second encoding and decoding model, so as to perform an encoding process on the above merged audio through the second encoding and decoding model, and perform a decoding process on the encoding result of the above merged audio, and the second encoding and decoding model outputs the target audio corresponding to the above original audio; among them, the above second encoding and decoding model is a trained deep learning model.

[0022] In some embodiments, based on the foregoing solution, the above audio processing device further includes: a personalized processing module;

[0023] Among them, the above personalized processing module is used to: after the above sub-audio determination module divides the original audio into multiple sub-audios, perform personalized processing on the target sub-audio, where the above target sub-audio corresponds to a target frequency domain interval, and the above personalized processing is processing for the above target frequency domain interval.

[0024] In some embodiments, based on the foregoing solution, the first feature extraction unit is further specifically configured to: perform personalized processing on the i-th high-frequency sub-audio, where the i-th high-frequency sub-audio corresponds to the i-th frequency domain interval, and the personalized processing is processing for the i-th frequency domain interval.

[0025] In some embodiments, based on the foregoing solution, the sub-audio determination module is specifically configured to: convert the original audio to the frequency domain to obtain a target frequency domain signal; determine M-1 frequency points according to the distribution of the target frequency domain signal, where M is an integer greater than 1; divide the target frequency domain signal into M frequency domain intervals through the M-1 frequency points; and convert the M sub-audio signals corresponding to the M frequency domain intervals to the time domain respectively to obtain M sub-audios.

[0026] In some embodiments, based on the above solution, the audio processing device further includes: an encoding module and a sending module; wherein, the encoding module is configured to: after the feature extraction module determines the target audio corresponding to the original audio, encode the target audio through an audio sending end to obtain a target bitstream; and the sending module is configured to: send the target bitstream to an audio receiving end through the audio sending end for decoding the target bitstream at the audio decoding end.

[0027] The audio processing device provided by the embodiments of the present application divides the original audio into sub-audios belonging to the high-frequency region and sub-audios belonging to the low-frequency region, and respectively performs feature extraction on the sub-audios belonging to the high-frequency region and the low-frequency region in a targeted manner, which can improve the degree of refined processing of the audio and is beneficial to improving the audio performance. Further, compared with encoding and decoding the original audio, the effect of encoding and decoding the target audio corresponding to the original audio is better, which is thus beneficial to improving the audio encoding and decoding efficiency.

[0028] In a third aspect, the present application provides an audio decoding method, which is applied to a decoder, and the method includes: decoding the bitstream corresponding to the original audio to obtain a target audio; where the target audio is determined by processing the original audio according to the method provided in the first aspect.

[0029] Due to the refined processing of the audio in the preprocessing process before audio encoding, during the audio transmission process, it will be beneficial to improve the encoding effect and decoding effect corresponding to the audio encoder at the audio transmitting end and the audio decoder at the audio receiving end respectively, and further beneficial to improve the encoding and decoding efficiency during the transmission process between audio terminals. Also, due to the improvement of the audio fidelity performance, that is, the distortion degree can be reduced, which is thus beneficial to the encoding and decoding effect during the audio transmission process, and further beneficial to improving the audio encoding and decoding efficiency during the audio transmission process.

[0030] Fourth aspect, the present application provides a decoder, which includes: a decoding module; wherein, the above decoding module is used to decode the code stream corresponding to the original audio to obtain the target audio; wherein, the above target audio is determined by processing the original audio according to the method provided in the above first aspect.

[0031] Due to the refined processing of the audio during the preprocessing process before audio encoding, during the audio transmission process, it will be beneficial to improve the encoding effect and decoding effect corresponding to the audio encoder at the audio transmitting end and the audio decoder at the audio receiving end respectively, and thus beneficial to improving the encoding and decoding efficiency during the transmission process between terminals in the audio. Also, due to the improvement of the audio fidelity performance, that is, the distortion degree can be reduced, which is beneficial to the encoding and decoding effect during the audio transmission process, and thus beneficial to improving the audio encoding and decoding efficiency during the audio transmission process.

[0032] Fifth aspect, the present application provides an audio encoding method, which includes: obtaining a target audio, wherein the above target audio is determined by processing the original audio according to the method provided in the above first aspect; and encoding the above target audio to obtain the code stream corresponding to the above original audio.

[0033] Due to the refined processing of the audio during the preprocessing process before audio encoding, during the audio transmission process, it will be beneficial to improve the encoding effect and decoding effect corresponding to the audio encoder at the audio transmitting end and the audio decoder at the audio receiving end respectively, and thus beneficial to improving the encoding and decoding efficiency during the transmission process between terminals in the audio. Also, due to the improvement of the audio fidelity performance, that is, the distortion degree can be reduced, which is beneficial to the encoding and decoding effect during the audio transmission process, and thus beneficial to improving the audio encoding and decoding efficiency during the audio transmission process.

[0034] Sixth aspect, the present application provides an encoder, which includes: an obtaining module and an encoding module;

[0035] Wherein, the above obtaining module is used to obtain a target audio, wherein the above target audio is determined by processing the original audio according to the method provided in the above first aspect or any one of its embodiments; and the above encoding module is used to encode the above target audio to obtain the code stream corresponding to the above original audio.

[0036] Due to the refined processing of the audio during the preprocessing process before audio encoding, during the audio transmission process, it will be beneficial to improve the encoding effect and decoding effect corresponding to the audio encoder at the audio transmitting end and the audio decoder at the audio receiving end respectively, and thus beneficial to improving the encoding and decoding efficiency during the transmission process between terminals in the audio. Also, due to the improvement of the audio fidelity performance, that is, the distortion degree can be reduced, which is beneficial to the encoding and decoding effect during the audio transmission process, and thus beneficial to improving the audio encoding and decoding efficiency during the audio transmission process.

[0037] In a seventh aspect, an electronic device is provided, including a processor and a memory; the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the audio processing method provided in the first aspect or any one of its embodiments, or to execute the audio decoding method provided in the third aspect, or to execute the audio encoding method provided in the fifth aspect.

[0038] In an eighth aspect, a chip is provided for implementing the audio processing method provided in the first aspect or any one of its embodiments; specifically, the chip includes: a processor for calling and running a computer program from a memory, such that a device installed with the chip executes the audio processing method provided in the first aspect, or executes the audio decoding method provided in the third aspect, or executes the audio encoding method provided in the fifth aspect.

[0039] In a ninth aspect, a computer-readable storage medium is provided for storing a computer program, and the computer program causes a computer to execute the audio processing method provided in the first aspect or any one of its embodiments, or to execute the audio decoding method provided in the third aspect, or to execute the audio encoding method provided in the fifth aspect.

[0040] In a tenth aspect, a computer program product is provided, including computer program instructions, and the computer program instructions cause a computer to execute the audio processing method provided in the first aspect or any one of its embodiments, or to execute the audio decoding method provided in the third aspect, or to execute the audio encoding method provided in the fifth aspect.

[0041] In an eleventh aspect, a computer program is provided, which when running on a computer, causes the computer to execute the audio processing method provided in the first aspect or any one of its embodiments, or to execute the audio decoding method provided in the third aspect, or to execute the audio encoding method provided in the fifth aspect.

[0042] In summary, since audio at different frequencies has different characteristics, in the data audio processing solution provided by the embodiments of the present application, the original audio is divided in the frequency domain, and the division result is converted to the time domain, so as to divide the original audio into sub-audio belonging to the high-frequency region and sub-audio belonging to the low-frequency region. Further, feature extraction is respectively performed on the sub-audio belonging to the high-frequency region and the sub-audio belonging to the low-frequency region in a targeted manner to determine the target audio of the original audio. The audio processing method provided by the embodiments of the present application can improve the degree of refined processing of audio, which is beneficial to improving audio performance. Further, compared with encoding and decoding the original audio, the effect of encoding and decoding the target audio corresponding to the original audio is improved, which is beneficial to improving the audio encoding and decoding efficiency. Brief Description of the Drawings

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention and this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention and this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0044] Figure 1 It is a schematic flowchart of encoding and decoding the original audio in the related art;

[0045] Figure 2 It is a schematic flowchart of encoding and decoding the original audio provided by the embodiment of this application;

[0046] Figure 3 It is a schematic flowchart of the audio processing method provided by the embodiment of this application;

[0047] Figure 4 It is a schematic flowchart of dividing the original audio into sub-audios provided by the embodiment of this application;

[0048] Figure 5 It is a schematic flowchart of determining the target audio for encoding according to the sub-audio provided by the embodiment of this application;

[0049] Figure 6 It is a schematic flowchart of the method for determining the target audio according to the sub-audio provided by the embodiment of this application;

[0050] Figure 7 It is a schematic flowchart of determining the target audio for encoding by combining a deep learning model provided by the embodiment of this application;

[0051] Figure 8 It is a schematic structural diagram of the first codec provided by the embodiment of this application;

[0052] Figure 9 It is another schematic flowchart of determining the target audio by combining a deep learning model provided by the embodiment of this application;

[0053] Figure 10 It is a schematic flowchart of the audio encoding method provided by the embodiment of this application;

[0054] Figure 11 It is a schematic flowchart of the audio decoding method provided by the embodiment of this application;

[0055] Figure 12 It is a schematic block diagram of the audio processing device provided by the embodiment of this application;

[0056] Figure 13Schematic block diagram of the encoder provided by an embodiment of the present application;

[0057] Figure 14 Schematic block diagram of the decoder provided by an embodiment of the present application;

[0058] Figure 15 Schematic block diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners

[0059] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0060] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In the embodiments of the present invention of the present application, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, but B can also be determined according to A and / or other information. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In the description of the embodiments of the present application, unless otherwise specified, "a plurality of" means two or more than two.

[0061] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit including the functions of the module or unit.

[0062] The technical solution proposed in the embodiments of the present application can be applied to the field of audio processing technology, specifically in the field of audio coding and decoding technology. For example, during the end-to-end audio transmission process, before encoding the original audio at the audio sending end, fine processing of the original audio is beneficial to improving audio performance. Further, compared with encoding and decoding the original audio, it can improve the audio coding and decoding effect, and thus is beneficial to improving the audio coding and decoding efficiency.

[0063] Figure 1 It is a schematic flowchart of encoding and decoding the original audio in the related art. Refer to Figure 1 , in the end-to-end audio coding and decoding scheme based on deep learning provided by the related art, both the encoder 112 and the decoder 122 are trained deep learning models. Specifically, at the audio sending end 110, the original audio 111 is directly input into the encoder 112. Specifically, the encoder 112 can perform a non-linear transformation on the original audio 111 through an encoding network to obtain a corresponding first latent variable. Further, a quantization operation can be performed on the above first latent variable through a residual vector quantizer (RVQ). Further, the quantization result is generated into a binary code stream 101. At the audio receiving end 120, the decoder 122 first recovers the value of the first latent variable from the binary code stream 101, then performs an inverse quantization operation on the recovered first latent variable, and inputs it into the decoder 122 for non-linear transformation to obtain the reconstructed audio 121.

[0064] As can be seen from the above, in the audio coding scheme provided by the related art, the related art does not consider the obvious differences between the high-frequency and low-frequency components in the original audio 111, but directly inputs the original audio 111 into the encoder 112 for encoding processing, which is not conducive to improving the coding effect. Exemplarily, since the high-frequency sub-audio contains relatively rich audio feature information, the related art does not specifically process the high-frequency sub-audio, which is not conducive to improving the audio fidelity. In addition, the related art does not distinguish between the high-frequency and low-frequency information in the original audio 111, and thus cannot perform processing related to high-frequency or low-frequency audio specifically, such as enhancing high-frequency classification and reducing low-frequency noise, which is not conducive to improving the auditory effect of the audio.

[0065] The solution provided by the embodiments of the present application can solve the above technical problems. Exemplarily, Figure 2 It is a schematic flowchart of encoding and decoding the original audio provided by the embodiments of the present application. Refer to Figure 2 , in the embodiments of the present application, before encoding the above original audio, preprocessing of the encoder is first performed on the original audio 211. Among them, the audio preprocessing process before encoding will be in Figure 3The corresponding embodiments are described in detail. Further, the target audio 214 processed as above is input into the encoder 212; specifically, the encoder 212 can perform a nonlinear transformation on the target audio 214 through the coding network to obtain the corresponding second latent variable. The second latent variable is then quantized. For example, the second latent variable can be quantized by a residual-based vector quantizer RVQ. Further, the quantization result is used to generate a binary code stream 201. At the audio receiving end 220, the decoder 222 first recovers the value of the second latent variable from the binary code stream 201, and then performs an inverse quantization operation on the recovered second latent variable, and inputs it into the decoder 222 for nonlinear transformation to obtain the reconstructed audio 221.

[0066] It can be seen that in the data audio processing scheme provided by the embodiment of the present application, the original audio is not directly encoded and decoded, but pre-processed before encoding. Specifically, before audio encoding (for example, encoding the original audio), the original audio is divided in the frequency domain, and the division result is converted to the time domain, so that the original audio is divided into sub-audio belonging to the high-frequency area and sub-audio belonging to the low-frequency area. Further, the sub-audio belonging to the high-frequency area and the sub-audio belonging to the low-frequency area are respectively feature extracted to determine the target audio of the above-mentioned original audio. Among them, the target audio is an audio that can be used for audio encoding. Compared with the related art that directly encodes and decodes the original audio without distinguishing between high and low frequencies, the audio processing scheme provided by the embodiment of the present application realizes the refined processing of the audio by dividing the original audio in the frequency domain and extracting features from the sub-audio belonging to the high-frequency area and the low-frequency area, which is conducive to improving the audio performance. Further, compared with encoding and decoding the original audio, it is conducive to improving the encoding and decoding effect of the audio, and then it is conducive to improving the audio encoding and decoding efficiency. At the same time, the embodiment of the present application distinguishes the high-frequency and low-frequency information in the original audio, so that targeted processing related to high-frequency audio or low-frequency audio can be performed, such as enhancing high-frequency classification, reducing low-frequency noise, etc., which is beneficial to improving the auditory effect of the audio.

[0067] References Figure 2, where the audio sending end 210 can be a terminal or a server, and the audio receiving end 220 can be a terminal. The audio sending end 210 and the audio receiving end 220 can be connected through a network. Among them, the network can be a wired communication link, a wireless communication link, an optical fiber cable, etc., and the embodiments of the present application do not limit this here. For example, it can be a communication medium of various connection types that can provide a communication link between a terminal and a server. In addition, the above-mentioned terminal can be a computer, a smart phone, a tablet, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a wearable smart device, a medical device, etc. The device is often configured with a display device, and the display device can also be a display, a display screen, a touch screen, etc. The touch screen can also be a touch control screen, a touch control panel, etc. However, it is not limited to this. The above-mentioned server can be a cloud server, and specifically can provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms; in addition, the server can also be an independent physical server, or a server cluster or distributed system composed of multiple physical servers.

[0068] The technical solutions of the embodiments of the present application will be described in detail below through some embodiments. These several embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0069] Figure 3 It is a schematic flowchart of the audio processing method P300 provided by the embodiments of the present application. Among them, the execution subject of the method P300 can be a server or a terminal, or, as Figure 13 shown in the electronic device, or can also be the audio sending end 210 as Figure 2 shown. It can be understood that if the execution subject of P300 is not the audio sending end 210, the audio sending end 210 can obtain the above-mentioned target audio 214 from other electronic devices to perform encoding operations on it. Refer to Figure 3 , the method P300 includes: S310 and S320.

[0070] In S310, the original audio is divided into multiple sub-audios. Among them, the frequency domain intervals corresponding to the multiple sub-audios are different, at least one of the multiple sub-audios corresponds to a frequency domain interval belonging to the high-frequency region, and at least one of the multiple sub-audios corresponds to a frequency domain interval belonging to the low-frequency region.

[0071] Among them, the above-mentioned original audio is the audio to be encoded and processed. For example, for creative music, the background music in film and television works, etc., in order to facilitate storage and transmission, it generally needs to be encoded and compressed. Therefore, the above-mentioned creative music to be encoded and compressed, the background music in film and television works can be used as the above-mentioned original audio.

[0072] In an exemplary embodiment, the above-mentioned original audio is a time-domain signal. Exemplarily, referring to Figure 2 the original audio 211, the target audio 224, and the reconstructed audio 221 in [reference], are all represented as time-domain diagrams of the audio, that is, the horizontal axis represents time and the vertical axis represents the amplitude of the audio signal. In other embodiments, the above-mentioned original signal may also be a frequency-domain signal, and the embodiments of the present application do not limit this.

[0073] There are obvious differences in the frequency range and audio characteristics between higher-frequency signals and lower-frequency signals. Higher-frequency signals have advantages such as a large capacity transmission ability and strong anti-interference ability, while having disadvantages such as a short signal propagation distance, high energy consumption, being affected by attenuation, multi-path effects, and poor penetration ability. Therefore, the embodiments of the present application will divide the original audio in the frequency domain to determine multiple sub-audios belonging to different frequency-domain intervals. At least one sub-audio corresponding to the frequency-domain interval among the obtained multiple sub-audios belongs to the high-frequency region, and at least one sub-audio corresponding to the frequency-domain interval among the obtained multiple sub-audios belongs to the low-frequency region. It should be noted that the threshold for distinguishing the above-mentioned high-frequency region and low-frequency region can be determined according to actual needs and the frequency-domain distribution of the original audio. The embodiments of the present application do not set a fixed threshold, and different thresholds can be set for different original audios to distinguish the high-frequency region and the low-frequency region. Furthermore, the characteristics of higher-frequency signals and lower-frequency signals can be considered separately to extract their audio features.

[0074] Specifically, taking the original audio as a time-domain signal as an example, the specific implementation manner of dividing it into sub-audios includes: S310-1 - S310-4.

[0075] In S310-1, the original audio is converted to the frequency domain to obtain a target frequency-domain signal.

[0076] In an exemplary embodiment, referring to Figure 4 , the original audio in the time domain is converted into a target frequency-domain signal in the frequency domain. Among them, the conversion method can be the method of Fourier transform. It can be understood that any method of converting audio from the time domain to the frequency domain can be adopted, and the embodiments of the present application do not limit this.

[0077] In S310-2, M - 1 frequency points are determined according to the distribution of the above-mentioned target frequency-domain signal.

[0078] Among them, M takes an integer greater than 1.

[0079] In an exemplary embodiment, the frequency range can be divided by a preset threshold. For example, two frequency points for dividing the frequency domain range are set, namely a high threshold and a low threshold. Thus, the part of the target frequency domain signal above the high threshold is divided into the first range, the part of the target frequency domain signal below the low threshold is divided into the third range, and the part above the low threshold and below the high threshold is divided into the second range. By means of presetting the threshold, the division of the target frequency domain signal can be efficiently realized.

[0080] In an exemplary embodiment, since the distributions of different original audio signals in the frequency domain are different, the interval values for dividing the sub-audio can be determined according to the distributions of different original audio signals in the frequency domain. In addition, the target frequency domain signal is divided into M ranges, where the specific value of M can be determined according to the distribution of the signal and the actual requirements for the fineness. For example, if the distribution of the target frequency domain signal is relatively wide, a larger number of frequency domain ranges can be divided, that is, the value of M is larger; on the contrary, if the distribution of the target frequency domain signal is relatively narrow, a smaller number of frequency domain ranges can be divided, that is, the value of M is smaller. For example, if the actual requirement for the fineness is higher, a larger number of frequency domain ranges can be divided, that is, the value of M is larger; on the contrary, if the actual requirement for the fineness is lower, a smaller number of frequency domain ranges can be divided, that is, the value of M is smaller. In this embodiment, the corresponding threshold is adjusted according to the high-frequency and low-frequency distributions of the original audio signal, so that signal distributions can be ensured in both the high-frequency and low-frequency domains. By performing personalized division on different original audio signals, it is beneficial to improve the subsequent feature extraction efficiency.

[0081] Exemplarily, after converting the time-domain audio signal to the frequency domain, the frequency domain diagram of the original audio can be obtained. The threshold for division can be determined according to the frequency distribution shown in the diagram to ensure that audio signals are distributed in each frequency domain range.

[0082] Exemplarily, the threshold can also be calculated by adapting to the frequency distribution of the original audio. For example, the clustering algorithm can be adopted. Taking the adoption of K-means clustering and determining two thresholds as an example, the specific implementation method includes:

[0083] 1) Randomly select 2 data sample points from the above-mentioned target frequency domain signal as the centroids (initial clustering centers).

[0084] 2) Measure the distance from each data point in the above-mentioned target frequency domain signal to each centroid and assign it to the class of the nearest centroid.

[0085] 3) Recalculate the centroids of the obtained classes.

[0086] 4) Iterate steps 2) - 3) until the distance from each data point to the new centroid is equal to the distance from each data point to the original centroid; or until the distance from each data point to the new centroid is less than the distance from each data point to the original centroid. The ordinate of the center point of the distance between the two centroids is the threshold value.

[0087] In S310-3, the above target frequency-domain signal is divided into M frequency-domain intervals through the above M-1 frequency points; and in S310-4, the M sub-audio signals corresponding to the above M frequency-domain intervals are respectively transformed into the time domain to obtain M sub-audios.

[0088] Exemplarily, Figure 4 is a flowchart showing the process of dividing an original audio into sub-audios provided by an embodiment of the present application. Refer to Figure 4 , after determining M-1 frequency points in S310-2, the above target frequency-domain signal can be divided into M frequency-domain intervals. Thus, the frequency-domain signals corresponding to the 1st frequency-domain interval, the 2nd frequency-domain interval, the Nth (N is less than M) frequency-domain interval, …… the (M-1)th frequency-domain interval, and the Mth frequency-domain interval can be obtained. Then, the frequency-domain signals corresponding to the above respective frequency-domain intervals are respectively subjected to inverse Fourier transform, and then M sub-audios can be obtained.

[0089] As described above, the embodiment of the present application does not set a fixed threshold for distinguishing high-frequency regions and low-frequency regions. However, for the convenience of distinguishing whether the frequency-domain intervals of sub-audios belong to the high-frequency region or the low-frequency region, the embodiment of the present application can determine the above threshold value among the frequency points used for dividing the frequency-domain intervals to be used for distinguishing high-frequency regions and low-frequency regions. Suppose the original audio A is divided into three sub-audios, that is, it includes two frequency points a1 and a2 (a1 is greater than a2) for dividing the frequency-domain intervals; in one embodiment, a1 can be used as the threshold for dividing high-frequency regions and low-frequency regions, that is, the sub-audio with the largest frequency-domain interval value is determined as the high-frequency sub-audio, and the two sub-audios with smaller frequency-domain interval values belong to the low-frequency region and can be determined as low-frequency sub-audios; in another embodiment, a2 can also be used as the threshold for dividing high-frequency regions and low-frequency regions, that is, the two sub-audios with larger frequency-domain interval values are determined as high-frequency sub-audios, and the sub-audio with the smallest frequency-domain interval value belongs to the low-frequency region and can be determined as the low-frequency sub-audio. At the same time, it can be seen that in the embodiment of the present application, among the multiple sub-audios obtained by dividing one original audio, the number of sub-audios belonging to the high-frequency region is at least one, and the specific number is not limited.

[0090] Continuing to refer to Figure 3 , in S320, by respectively performing feature extraction processing on the sub-audios belonging to the high-frequency region and the sub-audios belonging to the low-frequency region, the target audio corresponding to the above original audio is determined.

[0091] Among them, Figure 5 is a schematic flow chart of determining a target audio for encoding according to a sub-audio provided by an embodiment of the present application. Refer to Figure 5 , the multiple sub-audios determined in S310 are input into a trained deep learning model for feature extraction. Exemplarily, the trained deep learning model is any network model that can be used for audio feature extraction, and the specific network model of the deep learning model is not limited in the embodiments of the present application. In an exemplary embodiment, the above deep learning model performs an encoding process and a decoding process on the audio. Specifically, since the embodiment of the present application provides a preprocessing solution before audio encoding, and when the original audio to be preprocessed is a time-domain signal, the determined target audio should also be a time-domain signal. Therefore, the above feature extraction includes an encoding process for obtaining audio features and a decoding process for restoring to an audio format. Among them, regarding Figure 5 the trained deep learning model and the feature extraction process based on it will be described in detail in subsequent embodiments.

[0092] In Figure 3 the corresponding embodiment of the provided method P300, the original audio is not directly encoded and decoded, but preprocessed before encoding. Specifically, the original audio is divided in the frequency domain, and the division result is converted to the time domain, so as to divide the original audio into sub-audios belonging to the high-frequency region and sub-audios belonging to the low-frequency region. Further, feature extraction is respectively performed on the sub-audios belonging to the high-frequency region and the low-frequency region to determine the target audio of the above original audio. Among them, the target audio is an audio that can be used for audio encoding. Compared with the related art that directly encodes and decodes the original audio without distinguishing between high and low frequencies, the audio processing solution provided by the embodiment of the present application realizes refined processing of the audio by dividing the original audio in the frequency domain and separately extracting features of sub-audios in different frequency domain intervals, which is beneficial to improving audio performance. Further, compared with encoding and decoding the original audio, it is beneficial to improve the encoding and decoding effect, and thus beneficial to improving the audio encoding and decoding efficiency.

[0093] Further, in the case where the above original audio needs to be transmitted, the target audio processed by the above audio processing embodiment provided by the present application can be audio-encoded by an audio transmitter. As described above, due to the refined processing of the audio, it will also be beneficial to improve the encoding effect and decoding effect corresponding to the audio transmitter and the audio receiver respectively, and thus beneficial to improving the encoding and decoding efficiency during the transmission process between audio terminals.

[0094] Meanwhile, the embodiments of the present application distinguish the high-frequency and low-frequency information in the original audio, so that processing related to high-frequency audio or low-frequency audio can be carried out targeted, such as enhancing high-frequency classification, reducing low-frequency noise, etc., which is beneficial to improving the auditory effect of the audio.

[0095] Figure 6 It is a schematic flowchart of method P400 for determining a target audio according to a sub-audio provided by an embodiment of the present application. Refer to Figure 6 , the execution subject of method P400 can be a server or a terminal, similar to the execution subject of method P300. Or, as Figure 13 shown in the electronic device, or it can also be an audio sending end 210 as Figure 2 shown. It can be understood that if the execution subject of P400 is not the audio sending end 210, the audio sending end 210 can obtain the above-mentioned target audio 214 from other electronic devices for encoding operations. Refer to Figure 6 , method P400 includes: S410 - S440.

[0096] In S410, among the multiple sub-audios obtained by dividing the original audio, the sub-audio whose frequency domain interval belongs to the high-frequency region is determined as the high-frequency sub-audio, and the sub-audio whose frequency domain interval belongs to the low-frequency region is determined as the low-frequency sub-audio.

[0097] Since the signal belonging to the high-frequency region is a relatively high-frequency signal, and the period of the relatively high-frequency signal is short, more information can be transmitted per unit time. Therefore, the relatively high-frequency signal carries more audio information. Therefore, in the embodiments of the present application, separate feature extraction processing is performed on the sub-audio belonging to the high-frequency region. Exemplarily refer to Figure 5 , if M sub-audios are arranged in descending order of frequency domain interval, and the frequency domain intervals of the first N sub-audios belong to the high-frequency region, then the first N sub-audios can be denoted as high-frequency sub-audios.

[0098] It can be understood that among the multiple sub-audios obtained by dividing the above-mentioned original audio, the frequency domain interval either belongs to the above-mentioned high-frequency region or belongs to the above-mentioned low-frequency region. Then, among the multiple sub-audios obtained by dividing the original audio, the sub-audio whose frequency domain interval belongs to the high-frequency region is determined as the high-frequency sub-audio; among the multiple sub-audios obtained by dividing the original audio again, the sub-audio whose frequency domain interval belongs to the above-mentioned low-frequency region can be determined as the low-frequency sub-audio.

[0099] In an exemplary embodiment, after the original audio is divided into M sub-audios as described above, not only can higher-frequency and lower-frequency sub-audios be obtained, but also the sub-audios in different frequency ranges can be processed individually. For example, suppressing low-frequency noise, enhancing the sub-audio in a certain frequency range, etc.; for another example, in the process of adding an end-to-end audio codec based on deep learning to a symphony, before encoding it, the symphony is first transformed into the frequency domain and split into multiple sub-audios according to the divided frequency domain intervals; if the symphony includes the sounds of violins and cellos, and if it is desired to enhance the sound of the violins, the sub-audio corresponding to the lower audio range (corresponding to the sound of the violins) can be enhanced. In other embodiments, after the original audio is split into multiple sub-audios, processing such as adjusting the drumbeat and bass can also be performed. It can be seen that the embodiments of the present application distinguish the high-frequency and low-frequency information in the original audio, so that processing related to high-frequency audio or low-frequency audio can be carried out specifically, which is beneficial to improving the auditory effect of the audio.

[0100] In S420, the above-mentioned high-frequency sub-audios are respectively subjected to a first feature extraction process to obtain the restored sub-audios corresponding to the above-mentioned high-frequency sub-audios.

[0101] In an exemplary embodiment, Figure 7 is a schematic flowchart of a process for determining a target audio for encoding by combining a deep learning model provided by an embodiment of the present application. Refer to Figure 7 , the trained deep learning model 500 specifically includes N first codecs and a second codec. Among them, the above-mentioned N first codecs are trained according to audio samples in their respectively corresponding frequency ranges. For example, the first codec A corresponding to the frequency range [4 kHz, 5 kHz] is trained according to audio samples with a frequency of [4 kHz, 5 kHz]. Therefore, the trained first codec A can more accurately and specifically extract the features of the high-frequency sub-audios belonging to [4 kHz, 5 kHz].

[0102] Refer to Figure 7 , the first N high-frequency sub-audios are respectively input into the corresponding first codecs for a first feature extraction process to obtain N restored sub-audios. Specifically, the i-th high-frequency sub-audio is input into the i-th first codec model, so as to perform an encoding process on the i-th high-frequency sub-audio through the above-mentioned i-th first codec model, and perform a decoding process on the encoding result of the i-th high-frequency sub-audio. The above-mentioned i-th first codec model outputs the i-th restored sub-audio corresponding to the i-th high-frequency sub-audio, where i takes a positive integer not greater than N.

[0103] In an exemplary embodiment, Figure 8Schematic diagram of the structure of the first codec provided by the embodiments of the present application. It should be noted that the codec provided by the embodiments of the present application can be deployed on the same device, and the input sub-audio is encoded and decoded by the same device. For example, referring to Figure 8 The codec shown in is deployed on the electronic device 80.

[0104] Referring to Figure 8 , exemplarily, the above first codec model includes: P encoding units, where P is a positive integer. Among them, the above encoding unit includes a plurality of convolutional layers to extract features of the corresponding sub-audio through the convolutional layers. Specifically, a high-frequency sub-audio 802 obtained by dividing the original audio is input to the first encoding unit. After feature extraction corresponding to the P encoding units respectively, the hidden vector X corresponding to the high-frequency sub-audio 802 is output by the P encoding units. Further, quantization 820 processing is performed. For example, vector quantization processing can be performed through a residual-based vector quantizer RVQ to generate a binary code stream 804 corresponding to the high-frequency sub-audio 802.

[0105] Continuing to refer to Figure 8 , exemplarily, the above first codec model further includes: Q decoding units, where Q is a positive integer. Among them, the above decoding unit includes a plurality of transposed convolutional layers to perform upsampling through the transposed convolutional layers to restore the corresponding sub-audio. Specifically, the above binary code stream 804 is dequantized to obtain the corresponding latent variable, and then the latent variable is used as the input of the decoding unit to decode the final restored sub-audio 806.

[0106] As described above, using a targeted deep learning model (the first codec) to obtain features of the high-frequency sub-audio is beneficial to maintaining the audio information, improving the audio fidelity, and enhancing the audio performance. It should be noted that inside a codec as shown in Figure 8 , the downsampling ratio of the above P encoding units is the same as the upsampling ratio of the above Q decoding units to ensure the restoration effect of the input sub-audio.

[0107] In an exemplary embodiment, the above first codec may include a network structure for performing targeted processing. For example, Figure 8 The frequency range corresponding to the first codec shown in is [4 kHz, 5 kHz]. If it is necessary to enhance the high-frequency sub-audio with a frequency range of [4 kHz, 5 kHz] according to actual needs, then in the model structure as shown in Figure 8 , a network structure for enhancement processing can be set, so as to be able to perform personalized processing on the sub-audio in this audio range specifically, and then play a role in fine audio processing, which is beneficial to improving the auditory effect of the audio.

[0108] It can be understood that the above embodiments provide a structure of the first codec. In addition, the first codec may further include other network structures, which are not limited in the embodiments of the present application.

[0109] Continuing to refer to Figure 6 , in S430, the above restored sub-audio and the above low-frequency sub-audio are merged to obtain a merged audio. And, in S440, a second feature extraction process is performed on the above merged audio to obtain the target audio corresponding to the above original audio.

[0110] Exemplarily, referring to Figure 7 , after the first feature extraction process is respectively performed on N high-frequency sub-audios, N restored sub-audios are obtained. Further, the N restored sub-audios and the above low-frequency sub-audio are added in the time domain to obtain a merged audio. Exemplarily, the above low-frequency sub-audio specifically refers to (M - N) sub-audios that have not undergone the above first feature extraction process. Specifically, the merging can be performed according to the timestamps of each sub-audio. When there are overlapping timestamps between at least two sub-audios, the corresponding audio is superimposed at the overlapping timestamps; in the case where there are no overlapping timestamps between sub-audios, the relevant sub-audios are merged according to the time sequence before and after. Further, the above merged audio is input into the second codec model, and through the second codec model, the above merged audio can be fused to obtain the target audio corresponding to the above original audio.

[0111] In an exemplary embodiment, the network structure of the above second codec may be the same as or similar to the network structure of the first codec. Specifically, it may include an encoding unit and a decoding unit. Among them, the number of convolutional layers in the encoding unit of the second codec can be increased or decreased according to the complexity of the information to be processed. For example, if the complexity of the information of the merged audio processed by the above second codec is higher than the complexity of the information of the corresponding high-frequency sub-audio processed by a certain first encoder, then the number of convolutional layers in the encoding unit of the first codec is greater than the number of convolutional layers in the encoding unit of the first codec. Similarly, the number of convolutional layers in the decoding unit of the second codec can be increased or decreased according to the complexity of the information to be processed. Among them, in the second codec, the downsampling ratio of the encoding unit is the same as the upsampling ratio of the decoding unit to ensure the restoration effect of the input audio.

[0112] Through the embodiment provided by P400, the high-frequency sub-audios whose frequency domain intervals belong to the high-frequency region are separately subjected to feature extraction (the first feature extraction process), so that rich audio feature information contained in the high-frequency sub-audios can be extracted; and then the target audio of the above original audio is obtained by fusing each sub-audio through the second feature extraction process. Further, audio encoding can be performed on the target audio. For example, it can be performed through, such as Figure 2The encoder 212 in it encodes the target audio. Compared with the related art where the original audio is directly encoded and decoded without distinguishing between high and low frequencies, the audio processing solution provided by the embodiments of the present application can not only achieve refined processing of the audio, but also help improve the audio fidelity performance, that is, it can reduce the distortion degree, thereby helping to improve the encoding and decoding effects in the above audio processing process, and further helping to improve the audio encoding and decoding efficiency in the above audio processing process. Further, during the audio transmission process, the target audio processed by the audio processing embodiments provided by the present application is encoded. As mentioned above, due to the refined processing of the audio, it will also help improve the encoding effects and decoding effects corresponding to the audio transmitting end and the audio receiving end respectively, and further help improve the encoding and decoding efficiency during the transmission process between the audio terminals. Also, due to the improvement of the audio fidelity performance, that is, the distortion degree can be reduced, which helps the encoding and decoding effects during the audio transmission process, and further helps to improve the audio encoding and decoding efficiency during the audio transmission process.

[0113] The above provides an overall introduction to the audio processing method provided by the embodiments of the present application. The following further introduces the audio processing method provided by the embodiments of the present application through specific embodiments.

[0114] Figure 9 This is another flowchart showing the determination of the target audio in combination with a deep learning model provided by the embodiments of the present application. The following combines Figure 9 to further introduce the audio processing method provided by the embodiments of the present application.

[0115] Refer to Figure 9 , in the embodiment shown in this figure, after converting the original audio 900 to the frequency domain, a frequency point (e.g., 5.5 kHz) for dividing the frequency domain interval is determined, so that the frequency domain signal corresponding to the original audio 900 can be divided into two parts: one part is the high-frequency signal with a frequency greater than 5.5 kHz, and the other part is the low-frequency signal with a frequency less than or equal to 5.5 kHz. After converting the two parts of the signals to the time domain respectively, the sub-audio 902 and the sub-audio 904 are obtained. Since the frequency domain signal of the original audio 900 is divided into two frequency domain intervals, for the convenience of processing, the above frequency point 5.5 kHz can be used as the division threshold between the high-frequency region and the low-frequency region, that is, the above sub-audio 902 is the high-frequency sub-audio, and the sub-audio 904 can also be denoted as the low-frequency sub-audio.

[0116] In an exemplary embodiment, if it is necessary to suppress low-frequency noise, the above low-frequency sub-audio 904 can be processed specifically. It can be seen that through the refined frequency division processing of the original audio, it is also possible to perform personalized and targeted processing on the sub-audios in different frequency domain intervals, which is beneficial to improving the audio quality.

[0117] Since the above high-frequency sub-audio 902 contains more audio feature information compared to the low-frequency sub-audio 904, in the embodiments of the present application, the high-frequency sub-audio 902 is separately subjected to feature extraction. Specifically, the trained deep learning model 90 includes a first codec model 92 and a second codec model 94. Among them, the specific structures of the two codec sub-models can be referred to Figure 8 as shown; it can be understood that the number of convolutional layers for upsampling in the first codec model 92 can be the same as the number of convolutional layers for upsampling in the second codec model 94. Similarly, the number of transposed convolutional layers for downsampling in the first codec model 92 can be the same as the number of transposed convolutional layers for downsampling in the second codec model 94. The above high-frequency sub-audio 902 is input into the trained first codec model 92 to perform a non-linear transformation on the high-frequency sub-audio 902 through the encoding unit of the first codec model 92 to obtain the corresponding high-frequency latent variable. Further, a quantization operation is performed on the above high-frequency latent variable. For example, the above high-frequency latent variable can be quantized through a residual-based vector quantizer RVQ. Then, the quantization result of the high-frequency latent variable is generated into a corresponding binary bitstream. Then, the decoding unit of the first codec model 92 restores the value of the above high-frequency latent variable from the above binary bitstream, and then the restored high-frequency latent variable is dequantized and non-linearly transformed to obtain a restored sub-audio 906 of the high-frequency sub-audio 902.

[0118] Refer to Figure 9 , and the restored sub-audio 906 and the low-frequency sub-audio 904 are merged in the time domain to obtain a merged audio 908. In order to obtain the audio corresponding to the above original audio, the merged audio 908 can be input into the trained second codec model 94 to realize the feature fusion of the restored sub-audio 906 and the low-frequency sub-audio 904. Specifically, the merged audio 908 is non-linearly transformed through the encoding unit of the second codec model 94 to obtain the corresponding merged latent variable. Further, the above merged latent variable can be quantized through a residual-based vector quantizer RVQ. Then, the quantization result of the merged latent variable is generated into a corresponding binary bitstream. Then, the decoding unit of the second codec model 94 restores the value of the above merged latent variable from the above binary bitstream, and then the restored merged latent variable is dequantized and non-linearly transformed to obtain a target audio 910. Among them, the target audio 910 can be used for audio encoding.

[0119] Compared with directly encoding the original audio 900 in the related art, the audio processing solution provided in the embodiments of the present application can not only improve the degree of refined processing of the audio, but also help to improve the audio fidelity performance, that is, it can reduce the distortion degree, thereby helping to improve the audio codec effect and further helping to improve the audio codec efficiency.

[0120] The above introduced the method embodiments for preprocessing the original audio before audio encoding. Next, the audio encoding method and the audio decoding method provided by the embodiments of the present application will be introduced.

[0121] Figure 10 FIG. 5 is a schematic flowchart of the audio encoding method P500 provided by the embodiments of the present application. Among them, the execution subject of the method P500 may be an encoder, and exemplarily, it may be an audio sending end 210 as shown in Figure 2 FIG. Figure 10 . The method P500 includes: S510 - S530.

[0122] In S510, the original audio is divided into multiple sub - audios. Among them, the frequency - domain intervals corresponding to the multiple sub - audios are different, at least one of the multiple sub - audios corresponds to a frequency - domain interval belonging to the high - frequency region, and at least one of the multiple sub - audios corresponds to a frequency - domain interval belonging to the low - frequency region;

[0123] Among them, the specific implementation manner of S510 is the same as that of S310, and will not be elaborated here.

[0124] In S520, by respectively performing feature extraction processing on the sub - audios belonging to the high - frequency region and the sub - audios belonging to the low - frequency region, the target audio corresponding to the original audio is determined;

[0125] Among them, the specific implementation manner of S520 is the same as that of S320, and will not be elaborated here.

[0126] In S530, the target audio is encoded to obtain the bitstream corresponding to the original audio.

[0127] In the end - to - end audio encoding and decoding process based on deep learning, first, according to the distribution of the original audio in the frequency domain, it is divided into multiple sub - audios corresponding to different frequency - domain intervals. Then, through the embodiments provided by P400, the high - frequency sub - audios whose frequency - domain intervals belong to the high - frequency region are separately subjected to feature extraction (the first feature extraction processing), so that rich audio feature information contained in the high - frequency sub - audios can be extracted; then, through the second feature extraction processing, each sub - audio is fused to obtain the target audio of the original audio. Further, the target audio can be subjected to audio encoding, for example, through Figure 2The encoder 212 in [it] performs audio encoding on the target audio. Specifically, the encoder 212 can perform a non-linear transformation on the target audio 214 through an encoding network to obtain a corresponding second latent variable. Further, a quantization operation can be performed on the above-mentioned second latent variable through a residual-based vector quantizer RVQ. Further, the quantization result is used to generate a binary bitstream 201. At the audio receiving end 220, the decoder 222 first recovers the value of the second latent variable from the binary bitstream 201, then performs an inverse quantization operation on the recovered second latent variable, and inputs it into the decoder 222 for non-linear transformation to obtain a reconstructed audio 221.

[0128] In the audio encoding method provided in this embodiment, due to the refined processing of the audio during the preprocessing before audio encoding, and the improvement of audio fidelity brought by separately extracting features from high-frequency sub-audios whose frequency domain intervals belong to the high-frequency region, during the audio transmission process, it will be beneficial to improve the encoding effect and decoding effect corresponding to the encoder at the audio transmitting end and the decoder at the audio receiving end respectively, and thus beneficial to improving the encoding and decoding efficiency during the transmission process between terminals in the audio. Also, due to the improvement of audio fidelity performance, that is, the distortion degree can be reduced, which is beneficial to the encoding and decoding effect during the audio transmission process, and thus beneficial to improving the audio encoding and decoding efficiency during the audio transmission process.

[0129] Figure 11 It is a schematic flowchart of the audio decoding method P600 provided by an embodiment of this application. Among them, the execution subject of the method P600 can be a decoder. Exemplarily, it can be the audio receiving end 220 as shown in Figure 2 shown. Referring to Figure 11 , the method P600 includes: S610 and S620.

[0130] In S610, a target audio is obtained, where the above-mentioned target audio is determined according to an audio processing method for processing the original audio. And, in S620, the above-mentioned target audio is encoded to obtain a bitstream corresponding to the above-mentioned original audio.

[0131] In the audio encoding method provided by the embodiments of the present application, as described above, the original audio is not directly encoded and decoded. Instead, preprocessing is performed on the original audio before encoding, and then audio encoding is performed on the obtained target audio. Since, in the process of determining the above target audio, feature extraction is separately performed on the high-frequency sub-audio whose frequency domain interval belongs to the high-frequency region (the first feature extraction process, which can refer to the embodiments provided in P400), rich audio feature information contained in the high-frequency sub-audio can be extracted. Furthermore, performing audio encoding on the target audio processed by the above audio encoding embodiments of the present application is beneficial to improving the fidelity of the audio and the encoding efficiency. Thus, decoding an audio with high fidelity can also reduce the decoding complexity and improve the decoding efficiency.

[0132] As described above in conjunction with Figures 2 to 11 , the method embodiments of the present application have been described in detail. Below, in conjunction with Figures 12 - 14 , the apparatus embodiments of the present application will be described in detail.

[0133] Figure 12 FIG. is a schematic block diagram of an audio processing apparatus provided by an embodiment of the present application. As Figure 12 shown, the audio processing apparatus 1200 includes: a sub-audio determination module 1210 and a feature extraction module 1220;

[0134] Among them, the above sub-audio determination module 1210 is used to divide the original audio into multiple sub-audios, where the frequency domain intervals corresponding to the multiple sub-audios are different, at least one of the multiple sub-audios corresponds to a frequency domain interval belonging to the high-frequency region, and at least one of the multiple sub-audios corresponds to a frequency domain interval belonging to the low-frequency region; and the above feature extraction module 1220 is used to determine the target audio corresponding to the original audio by performing feature extraction processing on the multiple sub-audios, where the target audio is used for audio encoding.

[0135] In some embodiments, based on the foregoing solution, the above feature extraction module 1220 includes: a screening unit, a first feature extraction unit, a merging unit, and a second feature extraction unit;

[0136] Among them, the above screening unit is used to: determine the sub-audio with the frequency domain interval belonging to the high-frequency region among the above-mentioned multiple sub-audios as the high-frequency sub-audio, and determine the sub-audio with the frequency domain interval belonging to the above-mentioned low-frequency region among the above-mentioned multiple sub-audios as the low-frequency sub-audio; the above first feature extraction unit is used to: perform a first feature extraction process on the above high-frequency sub-audio to obtain the restored sub-audio corresponding to the above high-frequency sub-audio; the above merging unit is used to: merge the above restored sub-audio and the above low-frequency sub-audio to obtain a merged audio; and, the above second feature extraction unit is used to: perform a second feature extraction process on the above merged audio to obtain the target audio corresponding to the above original audio.

[0137] In some embodiments, based on the foregoing solution, the above first feature extraction unit is specifically used to: input the i-th high-frequency sub-audio into the i-th first encoding and decoding model, so as to perform an encoding process on the i-th high-frequency sub-audio through the i-th first encoding and decoding model, and perform a decoding process on the encoding result of the i-th high-frequency sub-audio, and the i-th first encoding and decoding model outputs the i-th restored sub-audio corresponding to the i-th high-frequency sub-audio, where i takes a positive integer not greater than the number of the above high-frequency sub-audios; among them, the above first encoding and decoding model is a trained deep learning model.

[0138] In some embodiments, based on the foregoing solution, the above first encoding and decoding model includes: P encoding units and Q decoding units, where P and Q take positive integer values; among them, the downsampling multiple corresponding to the above P encoding units is the same as the upsampling multiple corresponding to the above Q decoding units.

[0139] In some embodiments, based on the foregoing solution, the above second feature extraction unit is specifically used to: input the above merged audio into a second encoding and decoding model, so as to perform an encoding process on the above merged audio through the second encoding and decoding model, and perform a decoding process on the encoding result of the above merged audio, and the second encoding and decoding model outputs the target audio corresponding to the above original audio; among them, the above second encoding and decoding model is a trained deep learning model.

[0140] In some embodiments, based on the foregoing solution, the above audio processing device further includes: a personalized processing module;

[0141] Among them, the above personalized processing module is used to: after the above sub-audio determination module 1210 divides the original audio into M sub-audios, perform personalized processing on the target sub-audio, where the above target sub-audio corresponds to a target frequency domain interval, and the above personalized processing is processing for the above target frequency domain interval.

[0142] In some embodiments, based on the above scheme, the first feature extraction unit is further specifically used to: perform personalized processing on the i-th high-frequency sub-audio, wherein the i-th high-frequency sub-audio corresponds to the i-th frequency domain interval, and the personalized processing is processing for the i-th frequency domain interval.

[0143] In some embodiments, based on the above scheme, the sub-audio determination module 1210 is specifically used to: convert the original audio into the frequency domain to obtain a target frequency domain signal; determine M-1 frequency points according to the distribution of the target frequency domain signal; divide the target frequency domain signal into M frequency domain intervals through the M-1 frequency points; and convert the M sub-audio signals corresponding to the M frequency domain intervals into the time domain respectively to obtain M sub-audios.

[0144] In some embodiments, based on the above scheme, the audio processing device 1200 further includes: an encoding module and a sending module; wherein the encoding module is used to: after the feature extraction module performs feature extraction processing on the multiple sub-audios to determine the target audio corresponding to the original audio, encode the target audio through the audio transmitting end to obtain a target code stream; and the sending module is used to: send the target code stream to the audio receiving end through the audio transmitting end, so that the target code stream is decoded at the audio decoding end.

[0145] In the data audio processing device provided in the embodiment of the present application, the original audio is not directly encoded and decoded, but pre-processed before encoding. Specifically, before audio encoding (for example, encoding the original audio), the original audio is divided in the frequency domain, and the division result is converted to the time domain, so that the original audio is divided into sub-audio belonging to the high-frequency area and sub-audio belonging to the low-frequency area. Further, the sub-audio belonging to the high-frequency area and the sub-audio belonging to the low-frequency area are respectively feature extracted to determine the target audio of the above-mentioned original audio. Among them, the target audio is an audio that can be used for audio encoding. Compared with the related art that directly encodes and decodes the original audio without distinguishing between high and low frequencies, the audio processing scheme provided in the embodiment of the present application realizes the refined processing of the audio by dividing the original audio in the frequency domain and extracting the features of the sub-audio in different frequency domain intervals separately, which is conducive to improving the audio performance. Further, compared with encoding and decoding the original audio, it is conducive to improving the encoding and decoding effect, and then it is conducive to improving the audio encoding and decoding efficiency. At the same time, the embodiment of the present application distinguishes the high-frequency and low-frequency information in the original audio, so that targeted processing related to high-frequency audio or low-frequency audio can be performed, such as enhancing high-frequency classification, reducing low-frequency noise, etc., which is beneficial to improving the auditory effect of the audio.

[0146] It should be understood that the embodiments of the audio processing apparatus and the embodiments of the audio processing method can correspond to each other. Similar descriptions can refer to the embodiments of the audio processing method. To avoid repetition, they will not be elaborated here. Specifically, Figure 12 The illustrated audio processing apparatus can execute the embodiments of the above audio processing method, and the foregoing and other operations and / or functions of each module in the audio processing apparatus are respectively for implementing the embodiments of the above audio processing method. For the sake of brevity, they will not be elaborated here.

[0147] Figure 13 It is a schematic block diagram of an encoder provided by an embodiment of the present application. As Figure 13 shown, the encoder 1300 includes: a sub-audio determination module 1310, a feature extraction module 1320, and an encoding module 1330;

[0148] Among them, the above-mentioned sub-audio determination module 1210 is used to divide the original audio into multiple sub-audios, where the frequency domain intervals corresponding to the multiple sub-audios are different, at least one of the multiple sub-audios corresponds to a frequency domain interval belonging to the high-frequency region, and at least one of the multiple sub-audios corresponds to a frequency domain interval belonging to the low-frequency region; the above-mentioned feature extraction module 1220 is used to determine the target audio corresponding to the original audio by performing feature extraction processing on the multiple sub-audios; and the above-mentioned encoding module 1330 is used to perform encoding processing on the target audio to obtain the bitstream corresponding to the original audio.

[0149] In the end-to-end audio encoding and decoding process based on deep learning, the above encoder first divides the original audio into multiple sub-audios with different frequency domain intervals according to the distribution of the original audio in the frequency domain. Then, through the embodiment provided by P400, the high-frequency sub-audios whose frequency domain intervals belong to the high-frequency region are separately subjected to feature extraction (the first feature extraction processing), so that rich audio feature information contained in the high-frequency sub-audios can be extracted; and then the sub-audios are fused through the second feature extraction processing to obtain the target audio of the original audio. Further, audio encoding can be performed on the target audio. Due to the refined processing of the audio in the preprocessing process before audio encoding and the improvement of audio fidelity brought by separately performing feature extraction on the high-frequency sub-audios whose frequency domain intervals belong to the high-frequency region, in the audio transmission process, it will be beneficial to improve the encoding effect and decoding effect corresponding to the audio transmitting end and the audio receiving end respectively, and further beneficial to improve the encoding and decoding efficiency in the transmission process between audio terminals. Also, due to the improvement of audio fidelity performance, that is, the distortion degree can be reduced, which is beneficial to the encoding and decoding effect in the audio transmission process, and further beneficial to improve the audio encoding and decoding efficiency in the audio transmission process.

[0150] It should be understood that the encoder embodiments and the audio encoding method embodiments can correspond to each other. Similar descriptions can refer to the audio encoding method embodiments. To avoid repetition, they will not be elaborated here. Specifically, Figure 13 The encoder shown can execute the embodiments of the above audio encoding method, and the foregoing and other operations and / or functions of each module in the encoder are respectively for implementing the embodiments of the audio encoding method. For the sake of brevity, they will not be elaborated here.

[0151] Figure 14 is a schematic block diagram of a decoder provided in an embodiment of the present application. As Figure 14 shown, the decoder 1300 includes: a target audio acquisition module 1410 and a decoding module 1420;

[0152] Among them, the above-mentioned target audio acquisition module 1410 is used to acquire target audio, where the above-mentioned target audio is determined by processing the original audio according to an audio processing method; and, the above-mentioned decoding module 1420 is used as a decoding module to decode the bitstream corresponding to the original audio to obtain the target audio.

[0153] As described above, the decoder provided in the embodiments of the present application does not directly encode and decode the original audio, but first performs preprocessing before encoding the original audio, and then performs audio encoding on the obtained target audio. Since in the process of determining the above-mentioned target audio, feature extraction is separately performed on the high-frequency sub-audio whose frequency domain interval belongs to the high-frequency region (the first feature extraction process, which can refer to the embodiment provided in P400), rich audio feature information contained in the high-frequency sub-audio can be extracted. Furthermore, audio encoding is performed on the target audio processed by the above audio encoding embodiments of the present application, which is beneficial to improving the audio fidelity and the audio encoding efficiency. Thus, decoding the audio with high fidelity can also reduce the decoding complexity and improve the decoding efficiency.

[0154] It should be understood that the decoder embodiments and the audio decoding method embodiments can correspond to each other. Similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, Figure 14 The decoder shown can execute the embodiments of the above audio decoding method, and the foregoing and other operations and / or functions of each module in the decoder are respectively for implementing the embodiments of the audio decoding method. For the sake of brevity, they will not be elaborated here.

[0155] In the above, the device according to the embodiments of the present application has been described from the perspective of functional modules in combination with the accompanying drawings. It should be understood that the functional modules can be implemented in the form of hardware, can also be implemented by instructions in the form of software, and can also be implemented by a combination of hardware and software modules. Specifically, each step of the method embodiments in the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or can be executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0156] Figure 15 FIG. is a schematic block diagram of an electronic device provided in an embodiment of the present application. Figure 15 The electronic device 1500 can be the above audio processing device and can be used to execute the above audio processing method. The electronic device 1500 can also be the above encoder and can execute the above audio encoding method. Executing the above audio processing method can also be the above decoder and can execute the above audio decoding method.

[0157] As Figure 15 shown, the electronic device 1500 may include:

[0158] A memory 1510 and a processor 1520. The memory 1510 is used to store a computer program 1530 and transmit the program code 33 to the processor 1520. In other words, the processor 1520 can call and run the computer program 1530 from the memory 1510 to implement the method in the embodiments of the present application.

[0159] For example, when the above electronic device is the above audio processing device, the processor 1520 can be used to execute the steps in the above audio processing method according to the instructions in the computer program 1530. For another example, when the above electronic device is the above encoder, the processor 1520 can be used to execute the steps in the above audio encoding method according to the instructions in the computer program 1530. For another example, when the above electronic device is the above decoder, the processor 1520 can be used to execute the steps in the above audio decoding method according to the instructions in the computer program 1530.

[0160] In some embodiments of the present application, the processor 1520 may include but is not limited to:

[0161] General-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like.

[0162] In some embodiments of the present application, the memory 1510 includes, but is not limited to:

[0163] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0164] In some embodiments of the present application, the computer program 1530 can be divided into one or more modules, and the one or more modules are stored in the memory 1510 and executed by the processor 1520 to complete the audio processing method provided by the present application, or to complete the audio decoding method provided by the present application, or to complete the audio encoding method provided by the present application. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 1530 in the electronic device.

[0165] Such asFigure 15 As shown, the electronic device 1500 may further include:

[0166] A transceiver 1540, which may be connected to the processor 1520 or the memory 1510.

[0167] Among them, the processor 1520 can control the transceiver 1540 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 1540 may include a transmitter and a receiver. The transceiver 1540 may further include an antenna, and the number of antennas may be one or more.

[0168] It should be understood that the various components in the electronic device 1500 are connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.

[0169] According to one aspect of the present application, a computer storage medium is provided, on which a computer program is stored. When the computer program is executed by the computer, the computer can execute the method of the above method embodiment. Or rather, the embodiment of the present application further provides a computer program product containing instructions. When the instructions are executed by the computer, the computer executes the method of the above method embodiment.

[0170] According to another aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method of the above method embodiment.

[0171] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0172] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0173] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.

[0174] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, each functional module can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0175] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. An audio processing method, characterized in that: The method comprises: Divide the original audio into a plurality of sub-audios, wherein the frequency domain intervals corresponding to the plurality of sub-audios are different, the frequency domain interval corresponding to at least one of the plurality of sub-audios belongs to a high frequency region, and the frequency domain interval corresponding to at least one of the plurality of sub-audios belongs to a low frequency region; The target audio corresponding to the original audio is determined by performing feature extraction processing on the sub-audio belonging to the high-frequency region and the sub-audio belonging to the low-frequency region respectively.

2. The method according to claim 1, characterized in that The determining the target audio corresponding to the original audio includes: Determine the sub-audio whose frequency domain interval among the multiple sub-audio frequencies belongs to the high frequency region as a high frequency sub-audio, and determine the sub-audio whose frequency domain interval among the multiple sub-audio frequencies belongs to the low frequency region as a low frequency sub-audio; Performing a first feature extraction process on the high-frequency sub-audio to obtain a restored sub-audio corresponding to the high-frequency sub-audio; Merging the restored sub-audio and the low-frequency sub-audio to obtain a merged audio; Perform a second feature extraction process on the combined audio to obtain a target audio corresponding to the original audio.

3. The method according to claim 2, characterized in that The performing a first feature extraction process on the high frequency sub-audio to obtain a restored sub-audio corresponding to the high frequency sub-audio includes: Inputting the ith high frequency sub-audio into the ith first codec model, performing an encoding process on the ith high frequency sub-audio through the ith first codec model, and performing a decoding process on the encoding result of the ith high frequency sub-audio, the ith first codec model outputting the ith restored sub-audio corresponding to the ith high frequency sub-audio, where i is a positive integer not greater than the number of the high frequency sub-audio; Among them, the first encoding and decoding model is a trained deep learning model.

4. The method according to claim 3, characterized in that The first coding and decoding model includes: P coding units and Q decoding units, where P and Q are both positive integers; The downsampling multiples corresponding to the P encoding units are the same as the upsampling multiples corresponding to the Q decoding units.

5. The method according to claim 2, characterized in that: The performing a second feature extraction process on the combined audio to obtain a target audio corresponding to the original audio includes: Inputting the merged audio into a second codec model, so as to perform an encoding process on the merged audio through the second codec model, and perform a decoding process on the encoding result of the merged audio, and the second codec model outputs a target audio corresponding to the original audio; Among them, the second encoding and decoding model is a trained deep learning model.

6. The method according to claim 3, characterized in that The encoding process of the i-th high frequency sub-audio by using the i-th first encoding and decoding model comprises: The i-th high frequency sub-audio is personalizedly processed, wherein the i-th high frequency sub-audio corresponds to the i-th frequency domain interval, and the personalized processing is processing for the i-th frequency domain interval.

7. The method according to any one of claims 1 to 5, characterized in that After dividing the original audio into a plurality of sub-audios, the method further comprises: The target sub-audio is personalizedly processed, wherein the target sub-audio corresponds to a target frequency domain interval, and the personalized processing is processing for the target frequency domain interval.

8. The method according to any one of claims 1 to 5, characterized in that The step of dividing the original audio into a plurality of sub-audios comprises: Convert the original audio to the frequency domain to obtain the target frequency domain signal; Determine M-1 frequency points according to the distribution of the target frequency domain signal, where M is an integer greater than 1; Dividing the target frequency domain signal into M frequency domain intervals through the M-1 frequency points; The M sub-audio signals corresponding to the M frequency domain intervals are respectively converted into the time domain to obtain M sub-audio signals.

9. The method according to any one of claims 1 to 5, characterized in that: After determining the target audio corresponding to the original audio, the method further includes: Encode the target audio through the audio sending end to obtain the target code stream; The target code stream is sent to the audio receiving end through the audio sending end, so that the target code stream is decoded at the audio decoding end.

10. An audio decoding method, characterized in that: Applied to a decoder, the method comprises: Decode the code stream corresponding to the original audio to obtain the target audio; The target audio is determined by processing the original audio according to the method described in any one of claims 1 to 9.

11. An audio encoding method, characterized in that: Applied to an encoder, the method comprises: Acquire target audio, wherein the target audio is determined by processing original audio according to the method according to any one of claims 1 to 9; The target audio is encoded to obtain a code stream corresponding to the original audio.

12. An audio processing device, characterized in that: The device comprises: A sub-audio determining module, configured to divide the original audio into a plurality of sub-audios, wherein the frequency domain intervals corresponding to the plurality of sub-audios are different, the frequency domain interval corresponding to at least one of the plurality of sub-audios belongs to a high frequency region, and the frequency domain interval corresponding to at least one of the plurality of sub-audios belongs to a low frequency region; The feature extraction module is used to determine the target audio corresponding to the original audio by performing feature extraction processing on the sub-audio belonging to the high-frequency area and the sub-audio belonging to the low-frequency area respectively.

13. An electronic device comprising a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program to implement the audio processing method as described in any one of claims 1 to 9 above, or to execute the computer program to implement the audio decoding method as described in claim 10 above, or to execute the computer program to implement the audio encoding method as described in claim 11 above.

14. A computer-readable storage medium, characterized in that: For storing computer programs; The computer program enables a computer to execute the audio processing method as described in any one of claims 1 to 9 above, or the computer program enables a computer to execute the computer program to implement the audio decoding method as described in claim 10 above, or the computer program enables a computer to execute the computer program to implement the audio encoding method as described in claim 11 above.