Audio processing methods, devices, electronic equipment and media

CN115798507BActive Publication Date: 2026-08-11VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-10
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]本申请实施例的目的是提供一种音频处理方法、装置、电子设备及介质,能够解决音频处理过程中出现信息损失的问题

Benefits of technology

[0012]在本申请实施例中,电子设备可以先获取第一音频的N个频段中每个频段内出现的第一音频信号的第一时间段,N为正整数;接着,针对每个上述第一时间段,在一个第一时间段所处频段对应的傅里叶窗口内对上述第一时间段内的上述第一音频信号进行处理,得到每个上述第一时间段对应的第一音频信号特征信息,不同频段对应不同的傅里叶窗口;然后,将每个上述第一时间段对应的第一音频信号特征信息输入生成网络模型进行超清分辨率处理,得到每个上述第一时间段对应的第二音频信号;最后,基于该第二音频信号,得到第二音频。如此,由于电子设备在处理N个频段对应的N个第一音频信号时,是为不同频段匹配不同的傅里叶窗口,并在每个频段对应的傅里叶窗口内对该频段对应的第一音频信号进行处理,因此,提高了电子设备对上述N个第一音频信号的处理精细程度,从而避免了音频处理过程中出现信息损失。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115798507B_ABST
    Figure CN115798507B_ABST
Patent Text Reader

Abstract

This application discloses an audio processing method, apparatus, electronic device, and medium, belonging to the field of artificial intelligence technology. The audio processing method includes: acquiring a first time period of a first audio signal appearing in each of N frequency bands of a first audio signal, where N is a positive integer; for each of the first time periods, processing the first audio signal within a Fourier window corresponding to the frequency band of the first time period to obtain first audio signal feature information corresponding to each of the first time periods, with different Fourier windows corresponding to different frequency bands; inputting the first audio signal feature information corresponding to each of the first time periods into a generator network model for ultra-high-definition resolution processing to obtain a second audio signal corresponding to each of the first time periods; and obtaining a second audio signal based on the second audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to an audio processing method, apparatus, electronic device, and medium. Background Technology

[0002] With the development of smart audio devices, people's demand for sound quality is getting higher and higher. The sound quality of lossy audio formats can no longer meet the needs. Therefore, it is necessary to transform low-quality audio into high-quality audio.

[0003] In related technologies, when improving the sound quality of low-quality audio, it is necessary to first acquire the audio signal of the low-quality audio, and then process all the acquired audio signals as a whole to obtain the audio signal of high-quality audio, so as to generate high-quality audio.

[0004] However, the above scheme processes all audio signals as a whole within a fixed-size Fourier window. This fixed-size Fourier window cannot accurately meet the processing requirements of all audio signals, resulting in a low level of precision in the processing of the audio signals and thus information loss during the audio processing. Summary of the Invention

[0005] The purpose of this application is to provide an audio processing method, apparatus, electronic device, and medium that can solve the problem of information loss during audio processing.

[0006] In a first aspect, embodiments of this application provide an audio processing method, the method comprising: acquiring a first time period of a first audio signal appearing in each of N frequency bands of a first audio, where N is a positive integer; for each of the first time periods, processing the first audio signal within a Fourier window corresponding to the frequency band of the first time period to obtain first audio signal feature information corresponding to each of the first time periods, wherein different frequency bands correspond to different Fourier windows; inputting the first audio signal feature information corresponding to each of the first time periods into a generator network model for ultra-high definition resolution processing to obtain a second audio signal corresponding to each of the first time periods; and obtaining a second audio based on the second audio signal.

[0007] Secondly, embodiments of this application provide an audio processing apparatus, which includes an acquisition module and a processing module, wherein: the acquisition module is used to acquire a first time period of a first audio signal appearing in each of N frequency bands of a first audio signal, where N is a positive integer; the processing module is further used to process the first audio signal within the first time period within a Fourier window corresponding to the frequency band of the first time period for each first time period acquired by the acquisition module, to obtain first audio signal feature information corresponding to each first time period, where different frequency bands correspond to different Fourier windows; the processing module is further used to input the first audio signal feature information corresponding to each first time period into a generator network model for ultra-high definition resolution processing, to obtain a second audio signal corresponding to each first time period; the processing module is further used to obtain a second audio signal based on the second audio signal.

[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0011] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0012] In this embodiment, the electronic device first acquires a first time period of the first audio signal appearing in each of the N frequency bands of the first audio, where N is a positive integer. Then, for each of the first time periods, the first audio signal within the frequency band corresponding to that first time period is processed within a Fourier window to obtain the first audio signal feature information corresponding to each first time period. Different frequency bands correspond to different Fourier windows. Next, the first audio signal feature information corresponding to each of the first time periods is input into a generator network model for ultra-high-definition resolution processing to obtain a second audio signal corresponding to each of the first time periods. Finally, based on the second audio signal, a second audio signal is obtained. Thus, because the electronic device matches different Fourier windows for different frequency bands when processing the N first audio signals corresponding to the N frequency bands, and processes the first audio signal corresponding to each frequency band within the Fourier window, the processing precision of the N first audio signals by the electronic device is improved, thereby avoiding information loss during audio processing. Attached Figure Description

[0013] Figure 1 This is a schematic flowchart of an audio processing method provided in an embodiment of this application;

[0014] Figure 2 This is one of the audio signal spectrum diagrams of an audio processing method provided in this application embodiment;

[0015] Figure 3 This is the second audio signal spectrum diagram of an audio processing method provided in the embodiments of this application;

[0016] Figure 4 This is the third audio signal spectrum diagram of an audio processing method provided in this application embodiment;

[0017] Figure 5 This is one of the schematic diagrams of the generative network model structure provided in the embodiments of this application;

[0018] Figure 6 This is a schematic diagram of the subpixel convolution operation provided in an embodiment of this application;

[0019] Figure 7 This is the second schematic diagram of the generative network model structure provided in the embodiments of this application;

[0020] Figure 8 This is a schematic diagram of the structure of an audio processing device provided in an embodiment of this application;

[0021] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0022] Figure 10This is a hardware schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0024] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0025] The audio processing method, apparatus, electronic device, and medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0026] In related technologies, when electronic devices improve the sound quality of low-quality audio, they first process the original low-quality audio signal within a fixed-size Fourier window to obtain the feature information of the original audio signal; then, they further process this feature information to obtain the processed audio signal feature information; finally, they perform restoration processing on the processed audio signal feature information to obtain the desired high-quality audio. However, this audio processing method processes the audio signal within a fixed-size window. If the size of this fixed window does not match the audio signal, it will result in a lower level of precision in the audio signal processing, leading to information loss during the audio processing.

[0027] In the audio processing method, apparatus, electronic device, and medium provided in this application embodiment, when improving the sound quality of audio, the electronic device can first acquire a first time period of the first audio signal appearing in each of the N frequency bands of the first audio, where N is a positive integer; then, for each of the first time periods, the first audio signal within the first time period is processed within a Fourier window corresponding to the frequency band of the first time period to obtain the first audio signal feature information corresponding to each of the first time periods, with different Fourier windows corresponding to different frequency bands; then, the first audio signal feature information corresponding to each of the first time periods is input into a generator network model for ultra-high definition resolution processing to obtain a second audio signal corresponding to each of the first time periods; finally, a second audio is obtained based on the second audio signal. Thus, since the electronic device matches different Fourier windows for different frequency bands and processes the first audio signal corresponding to each frequency band within the Fourier window corresponding to each frequency band when processing the N first audio signals corresponding to the N frequency bands, the processing precision of the electronic device for the N first audio signals is improved, thereby avoiding information loss during audio processing.

[0028] The audio processing method provided in this embodiment can be executed by an audio processing device, which can be an electronic device, or a control module or processing module within the electronic device. The following description uses an electronic device as an example to illustrate the technical solution provided in this application embodiment.

[0029] This application provides an audio processing method, such as... Figure 1 As shown, the audio processing method may include the following steps 201 to 204:

[0030] Step 201: The electronic device acquires the first time period in each of the N frequency bands of the first audio signal.

[0031] Where N is a positive integer.

[0032] In the embodiments of this application, one frequency band corresponds to one first audio signal.

[0033] In this embodiment of the application, the first audio can be lossy audio in an electronic device that requires sound quality enhancement.

[0034] It should be noted that the method for enhancing audio quality in this embodiment is to increase the audio bitrate; the bitrate can be the bit rate, which is used to indicate the number of bits transmitted per unit time.

[0035] It is understandable that the more bits transmitted per unit of time, the higher the audio quality; in other words, the higher the audio bitrate, the higher the audio quality.

[0036] For example, the aforementioned lossy audio can be audio from a music file or audio from a video file.

[0037] In this embodiment of the application, the first audio signal is used to indicate the frequency information of the first audio signal.

[0038] In this embodiment of the application, the first time period is used to indicate the duration of the first audio signal within a frequency band.

[0039] In this embodiment of the application, the electronic device may take the first time period of the first audio signal from the start time point to the end time point of the frequency band as the first time period of the first audio signal.

[0040] Optionally, in this embodiment of the application, step 201 above, "the electronic device acquires the first time period in each of the N frequency bands of the first audio signal", may include the following steps 201a and 201c:

[0041] Step 201a: The electronic device inputs the audio signal of the first audio into N filters for filtering processing to obtain N first audio signals corresponding to N frequency bands.

[0042] In this embodiment, the filtering frequency thresholds of the above N filters are different.

[0043] In this embodiment of the application, the electronic device can set multiple filters with different filtering frequency thresholds, and then use the multiple filters with different filtering frequency thresholds to filter the audio signal of the first audio signal to obtain the first audio signal with multiple different frequencies.

[0044] Example 1, taking the above N filters with filtering frequency thresholds of 8kHz, 12kHz, and 16kHz as an example. The electronic device can filter the audio signal through the 8kHz filter, the 12kHz filter, and the 16kHz filter respectively to obtain audio signal 1 with a frequency band of 8kHz to 12kHz, audio signal 2 with a frequency band of 12kHz to 16kHz, and signal 3 with a frequency band above 16kHz.

[0045] In this way, by setting filters with different filtering frequency thresholds, the first audio signal can be divided into multiple frequency bands, which facilitates the subsequent processing of each frequency band and further improves the processing precision of the first audio signal.

[0046] Step 201b: For each first audio signal, the electronic device processes the first audio signal within a first Fourier window to obtain the first time period in which the first audio signal appears in each frequency band.

[0047] In this embodiment of the application, the window size of the first Fourier window is fixed.

[0048] In this embodiment of the application, the processing of the first audio signal includes performing a Fourier transform on the first audio signal.

[0049] In this embodiment, the electronic device can perform Fourier transform on the above N first audio signals in a Fourier window with a fixed window size to obtain the spectrum information corresponding to each first audio signal; then, the start time point and end time point of each first audio signal are read from the spectrum information, and the time between the start time point and the end time point is taken as the first time period of the first audio signal; finally, the first time period of each first audio signal is recorded as a frequency positioning table in a programming language format.

[0050] For example, the programming language format described above can be JavaScript Object Notation (JSON) format.

[0051] Example 2, in conjunction with Example 1, using the aforementioned spectral information as a spectrum diagram, and taking the first Fourier window size of 4092 as an example, after obtaining signals 1, 2, and 3, the electronic device can perform Fourier transforms on signals 1, 2, and 3 respectively within a Fourier transform window of size 4096 to obtain the spectrum corresponding to signal 1. Figure 1 The spectrum corresponding to signal 2 Figure 2 The spectrum corresponding to signal 3 Figure 3 .like Figure 2 As shown, the spectrum corresponding to signal 1 is displayed. Figure 1 Where t7 to t1 is the time period corresponding to signal 1; for example Figure 3 As shown, the spectrum corresponding to signal 2 is displayed. Figure 2 Where t5 to t6 is the time period corresponding to signal 2; for example Figure 4 As shown, the spectrum corresponding to signal 3 is displayed. Figure 3 t1 to t2 and t3 to t4 are the time periods corresponding to signal 3; and t0 to t1 are taken as signal 4 with a frequency of 0kHz to 8kHz.

[0052] In addition, the time periods corresponding to the above signals 1, 2, 3, and 4 are compiled into a frequency positioning table and recorded in JSON format. For example: {"16kHz": [[t1, t2], [t3, t4]], "12kHz": [[t5, t6]], "8kHz": [[t7, t1]], "0kHz": [[t0, t1]]}.

[0053] Step 201c: The electronic device divides each first time period into equal parts to obtain M equal second time periods within each first time period.

[0054] Where M is an integer greater than 1.

[0055] In this embodiment of the application, after obtaining the first time period, the electronic device can divide the first time period into multiple second time periods according to a preset division duration. For the last second time period that does not meet the preset division duration, blank audio signals can be added to supplement the last second time period into the division duration.

[0056] Example 3, combined with Example 2, takes the above "8kHz": [[t7, t1]] as an example, dividing it into equal parts, with the time interval from t7 to t1 being 35 seconds. If the preset division duration is 10 seconds, then t7 to t1 is divided into 4 second time intervals every 10 seconds. For the last second time interval, which is 30 to 35 seconds, a blank audio signal can be added at the end to extend the second time interval to 10 seconds.

[0057] In this way, by dividing the first time period into multiple equal second time periods, the amount of computation required for each audio signal processing is reduced, thereby improving the processing efficiency of the audio signal.

[0058] Step 202: For each first time period, the electronic device processes the first audio signal within the Fourier window corresponding to the frequency band of the first time period to obtain the first audio signal feature information corresponding to each first time period.

[0059] In the embodiments of this application, different frequency bands correspond to different Fourier windows.

[0060] In this embodiment of the application, the above-mentioned processing of the first audio signal within the first time period includes performing a Fourier transform on the first audio signal within the first time period.

[0061] In this embodiment of the application, the first audio signal feature information may include a first audio signal feature vector.

[0062] In this embodiment of the application, the electronic device can perform Fourier transform on the first audio signal in each first time period in N frequency bands in the Fourier window corresponding to its respective frequency band to obtain N first audio feature vectors.

[0063] Example 4, combined with Example 2, shows that an electronic device can perform a Fourier transform on a first audio signal above 16kHz with a window width of 256 to obtain a two-dimensional eigenvector matrix X. 16(129, 6891) Perform a Fourier transform with a window width of 512 on the first audio signal above 12kHz to obtain the two-dimensional eigenvector matrix X. 12 (257, 3446); Perform a Fourier transform on the first audio signal above 8kHz with a window width of 1024 to obtain a two-dimensional feature vector matrix X8(513, 1723); and perform a Fourier transform on the first audio signal above 0kHz with a window width of 2048 to obtain a two-dimensional feature vector matrix X0(1025, 862).

[0064] Step 203: The electronic device inputs the first audio signal feature information corresponding to each first time period into the generator network model for ultra-high definition resolution processing to obtain the second audio signal corresponding to each first time period.

[0065] In this embodiment of the application, the above-described generative network model is used to improve the audio quality. For example, by increasing the audio bitrate, converting low-bitrate audio into high-bitrate audio.

[0066] For example, the second audio signal is an audio signal whose sound quality has been improved by enhancing the first audio signal.

[0067] In the embodiments of this application, the above-mentioned generative network model can be a generative adversarial network (GAN).

[0068] The GAN model provided in the embodiments of this application will be explained below:

[0069] The GAN model provided in this application mainly consists of two parts: a generator and a discriminator. The generator has the following structure: Figure 5 As shown, it consists of 8 modules. Modules 1 to 4 are downsampling modules, and modules 5 to 8 are upsampling modules. Each module in the upsampling module consists of a one-dimensional deconvolutional neural network unit with 512, 512, 862, and 862 channels, corresponding to kernel sizes of 5, 7, 7, and 9, and strides of 1, 1, 1, and consists of sub-pixel convolutional units, normalization units, and function activation units. Each module in the downsampling module consists of a one-dimensional convolutional neural network unit with 862, 512, 512, and 1024 channels, corresponding to kernel sizes of 7, 5, 3, and 3, and strides of 2, 2, 2, 2, and consists of normalization units and function activation units.

[0070] Furthermore, the upsampling module in the generator differs from the downsampling module in that it includes a layer of one-dimensional subpixel convolutional units, suitable for the audio domain, between the one-dimensional deconvolutional neural network unit and the normalization unit. In traditional upsampling methods, to increase the dimensionality of the output features, blank areas are directly padded with zeros. This doesn't actually improve audio quality because zeros are useless information. However, the subpixel convolution method combines individual pixels from multi-channel features into a single pixel for a single feature, and then fills the blank areas with the combined pixel value. This padded value, unlike zeros, contains additional information, thus achieving the purpose of upsampling and improving resolution.

[0071] In this embodiment of the application, after obtaining the above-mentioned N first audio feature vectors, the electronic device can input the N first audio feature vectors into the GAN model for ultra-high resolution processing to obtain N processed first audio feature vectors. Then, the N processed first audio feature vectors are subjected to inverse Fourier transform in their respective Fourier windows to obtain N second audio signals.

[0072] Step 204: The electronic device obtains the second audio signal based on the second audio signal.

[0073] In the audio processing method provided in this application embodiment, the electronic device can first acquire the first time period of the first audio signal appearing in each of the N frequency bands of the first audio, where N is a positive integer; then, for each of the first time periods, the first audio signal within the first time period is processed within a Fourier window corresponding to the frequency band of the first time period to obtain the first audio signal feature information corresponding to each of the first time periods, with different Fourier windows corresponding to different frequency bands; then, the first audio signal feature information corresponding to each of the first time periods is input into a generator network model for ultra-high resolution processing to obtain the second audio signal corresponding to each of the first time periods; finally, the second audio is obtained based on the second audio signal. Thus, since the electronic device matches different Fourier windows for different frequency bands and processes the first audio signal corresponding to each frequency band within the Fourier window corresponding to each frequency band when processing the N first audio signals corresponding to the N frequency bands, the processing precision of the electronic device for the N first audio signals is improved, thereby avoiding information loss during audio processing.

[0074] Optionally, in this embodiment of the application, the above-mentioned generative network model includes a feature extraction layer, a sub-pixel convolutional layer, and an inverse transformation layer; the step 203 above, "the electronic device inputs the feature information of the first audio signal corresponding to each first time period into the generative network model for ultra-high definition resolution processing to obtain the second audio signal corresponding to each first time period", may include the following steps 203a to 203c:

[0075] Step 203a: The electronic device extracts key audio signal feature information from the first audio signal feature information corresponding to each first time period based on the feature extraction layer in the generative network model.

[0076] In this embodiment of the application, the aforementioned key audio signal feature information may be: the first audio signal feature information in which the phase and amplitude satisfy a preset threshold.

[0077] In this embodiment of the application, the feature extraction layer can be the downsampling module.

[0078] For example, the downsampling module described above includes a one-dimensional convolutional neural network unit, a normalization unit, and a function activation unit.

[0079] For example, the aforementioned one-dimensional convolutional neural network unit can obtain the key audio signal feature vector in the aforementioned first audio signal feature vector through convolution calculation.

[0080] For example, the normalization unit can pull the weights of the key audio signal feature vectors back to a standard normal distribution to avoid gradient vanishing and accelerate the convergence of the weights.

[0081] For example, the aforementioned function activation unit can limit the weight range of the aforementioned key audio signal feature vector to a certain range to avoid gradient explosion, which would affect the convergence effect of the aforementioned generative network model.

[0082] Step 203b: The electronic device performs feature augmentation on each key audio signal feature information based on the sub-pixel convolutional layer in the generative network model to obtain the second audio signal feature information corresponding to each first time period.

[0083] In this embodiment, the subpixel convolutional layer is used to fill in the blank areas in the audio signal feature vector to increase the dimension of the audio signal feature vector.

[0084] In this embodiment of the application, the aforementioned second audio signal feature information can be a second audio signal feature vector.

[0085] In this embodiment, the subpixel convolutional layer may be included in the upsampling module of the generative network model.

[0086] In this embodiment of the application, the subpixel convolutional layer can be a subpixel convolutional unit.

[0087] For example, the upsampling module described above includes a one-dimensional deconvolutional neural network unit, a subpixel convolution unit, a normalization unit, and a function activation unit.

[0088] For example, the aforementioned one-dimensional deconvolutional neural network unit can expand the aforementioned key audio signal feature vector through deconvolution calculation.

[0089] For example, the sub-pixel convolutional unit described above can, for the blank areas in the expanded key audio signal feature vector, combine multiple independent feature vectors in the expanded key audio signal feature vector into a feature vector with the same characteristics according to Formula 1, and then fill the blank areas with this feature vector to generate a second audio signal feature vector. Formula 1 is as follows:

[0090] A sr =f L (A LR )=ρ(W L *f L-1 (A LR )+b L ) Formula 1

[0091] Among them, A sr and A LR These represent the feature information of the first audio in high-resolution and low-resolution spaces, respectively; ρ is the shuffle operator, which can transform a tensor of size H*C·r into a tensor of size r·H*C. Here, H is the height of the feature vector, C is the number of channels in the subpixel convolutional unit, and r is the upsampling coefficient, chosen as 2 here; W L and b L These are the parameter matrix and the offset matrix, respectively.

[0092] For example, the sub-pixel convolutional unit described above can perform the tensor dimension transformation according to Equations 2.1, 2.2, and 2.3, as follows:

[0093] out shape[0] =input shape[0] Formula 2.1

[0094] out shape[1] =input shape[1] *r Formula 2.2

[0095]

[0096] It should be noted that the above subpixel convolution operation can be performed as follows: Figure 6 As shown, the input special effects vector matrix is ​​first transposed, then the three-dimensional matrix is ​​projected into a higher-dimensional (four-dimensional) space using the dimension increase method, then the higher-dimensional matrix is ​​transposed, and finally the transposed matrix is ​​output according to the above formula 2 to obtain a matrix of b*r·H*C / r.

[0097] For example, combined Figure 5 For the output of the one-dimensional convolutional neural network unit in the first upsampling module (8*128*512), the size of channel 2 can be changed to 128*2=256 using the above formula 2.2; the size of channel 3 can be changed to 512 / 2=256 using the above formula 2.3. Therefore, the feature dimension obtained by one-dimensional sub-pixel unit conversion is (8*256*256).

[0098] For example, the normalization unit can pull the weights of the second audio signal feature vector back to a standard normal distribution to avoid gradient vanishing and accelerate the convergence of the weights.

[0099] For example, the aforementioned function activation unit can limit the weight range of the second audio signal feature vector to a certain range to avoid gradient explosion and affect the convergence effect of the aforementioned generative network model.

[0100] Step 203c: The electronic device performs feature inverse transformation on the feature information of each second audio signal based on the inverse transformation layer in the generative network model to obtain the second audio signal corresponding to each first time period.

[0101] In this embodiment, the inverse transformation layer is used to perform inverse Fourier transform on the audio signal feature information.

[0102] For example, the inverse Fourier transform described above is the inverse of the Fourier transform.

[0103] In this embodiment of the application, after obtaining the second audio signal feature vector, the electronic device can perform inverse Fourier transform on the second audio signal feature vector in the Fourier window corresponding to its respective frequency band to obtain the second audio signal corresponding to each first time period.

[0104] Thus, by inputting the feature vector of the first audio signal into the generation network model, the key audio signal feature vector in the feature vector of the first audio signal is extracted, and the key audio signal feature vector is expanded again, and the independent feature vectors in the expanded feature vector are combined to generate a high-quality second audio signal, which further improves the processing precision of the first audio signal and enhances the sound quality of the first audio.

[0105] The training process of the GAN model provided in the embodiments of this application will be explained below:

[0106] For example, the training process of the above GAN model may include the following steps A1 to A4:

[0107] Step A1: Acquire 3000 lossless music tracks with a sampling rate of 44.1kHz. Compress them using existing MP3 compression methods into 32kbps and 128kbps versions. The 32kbps tracks simulate lossy music in real life, while the 128kbps tracks simulate lossless music. The goal of this GAN model is to generate 128kbps audio from 32kbps audio, achieving a 4x improvement in sound quality.

[0108] Step A2: Perform variable-window Fourier transforms on the 32kHz and 128kHz audio signals respectively. Specifically, the audio signal of the 32kHz audio signal can be passed through filters with filtering frequencies of 8kHz, 12kHz, and 16kHz to generate multiple audio signals from 0 to 8kHz, 8 to 12kHz, 12 to 16kHz, and above 16kHz. Then, perform Fourier transforms on the multiple audio signals with corresponding window sizes to generate a two-dimensional feature matrix X. 16 X 12 X8, X0. Perform the corresponding processing on the same 128 bitrate music to obtain a two-dimensional feature matrix Y of the same dimensions. 16 Y 12 ,Y8,Y0.

[0109] Step A3: Combining Figure 5 Eight samples are selected as a batch, so the dimension of the feature vector matrix input to the GAN model is (8*1025*862). After the first downsampling module, the dimension of the feature vector matrix becomes (8*512*862); after the second downsampling module, the dimension becomes (8*256*512); and after the third and fourth downsampling modules, the dimension becomes (8*64*1024). Then, during the upsampling process, the outputs of the sub-pixel convolutional units and the downsampling modules are concatenated along the second dimension. The first upsampling module is concatenated with the third downsampling module, the second upsampling module with the second downsampling module, and the third upsampling module with the first downsampling module. The final output of the GAN model has the same dimension as the input (8*1025*862).

[0110] like Figure 7The diagram illustrates the discriminator of the GAN model, which consists of three downsampling modules, two fully connected layers (FC), and one LReLU activation function layer. Each downsampling module includes a one-dimensional convolutional neural network unit with 1024, 1024, and 1024 channels, kernel sizes of 7, 5, and 3, and strides of 2, 2, and 2, respectively. The feature vector matrix output from the downsampling modules is then expanded into a one-dimensional feature vector (8*131072*1). This feature vector is then passed through a 2048-node fully connected layer, resulting in an output of dimension (8*1).

[0111] Step A4: Generate a feature matrix X from the above two-dimensional feature matrix X using the generator. ~ The data is input into the discriminator. Then, the feature Y corresponding to the 128kbps music is also input into the discriminator. After 1000 iterations, the discriminator output value stabilizes at 0.5, meaning that the discriminator can no longer distinguish between real data Y and generated data X. ~ At this point, the generator is the final trained model.

[0112] Thus, by using the model trained in this way to process low bitrate audio, signal loss can be greatly reduced during the processing, making the processed audio closer to the original lossless audio.

[0113] Optionally, in this embodiment of the application, step 204 above, "the electronic device obtains the second audio based on the second audio signal," may include the following step 204a:

[0114] Step 204a: The electronic device splices the second audio signals corresponding to each first time segment according to the time sequence of the first audio to obtain the second audio.

[0115] In this embodiment of the application, the time sequence of the first audio is used to indicate the sequential arrangement of the N first audio signals.

[0116] In this embodiment of the application, the electronic device can splice each of the above-mentioned second audio signals in the order of the first audio signals to obtain a continuous second audio signal, i.e., the second audio.

[0117] Example 5, combined with Example 2, allows the electronic device to concatenate the signals 1, 2, 3, and 4 according to the time sequence of their respective time periods, using Formula 3, to obtain the final high-bitrate audio X. Formula 3 is as follows:

[0118]

[0119] Where X represents high-bitrate audio. This involves stitching the pieces along the first dimension according to the time points corresponding to the first time period.

[0120] In this way, the second audio signal, after being improved in sound quality, is spliced ​​together according to the time sequence of the first audio signal to avoid the final second audio signal being inconsistent with the first audio signal.

[0121] Optionally, in this embodiment of the application, in conjunction with step 201c above, step 202, "the electronic device processes the first audio signal within the first time period within a Fourier window corresponding to the frequency band of the first time period for each first time period to obtain the first audio signal feature information corresponding to each first time period," may include the following steps 202a and 202b:

[0122] Step 202a: For each of the M second time periods within each first time period, the electronic device processes the first audio signal within the second time period within a Fourier window corresponding to the frequency band of the second time period to obtain the first audio signal feature information corresponding to each second time period.

[0123] In this embodiment of the application, the above-mentioned processing of the first audio signal includes performing a Fourier transform on the first audio signal.

[0124] In this embodiment of the application, the electronic device can perform Fourier transform on the first audio signals within the M second time periods respectively within the Fourier window corresponding to their respective frequency bands to obtain M first audio signal feature vectors.

[0125] Step 202b: The electronic device splices together the first audio signal feature information corresponding to the M second time periods within each first time period to obtain the first audio signal feature information corresponding to each first time period.

[0126] In this embodiment of the application, after obtaining the above-mentioned M first audio signal feature vectors, the electronic device can splice the M first audio signal feature vectors according to the time order of the first time period before the above-mentioned equal division to generate a first audio signal feature vector, that is, the first audio signal feature information corresponding to the above-mentioned first time period.

[0127] Thus, by performing Fourier transforms on the first audio signals within the second time period after equal division, and then concatenating the feature vectors of the first audio signals corresponding to each second time period, not only is the precision of audio signal processing improved, but the processing efficiency of audio signals is also increased.

[0128] Optionally, the audio processing method provided in this application embodiment can also be applied to cochlear implants, because humans can perceive sound frequencies ranging from approximately 20Hz to 20kHz, while cochlear implants can respond to frequencies up to 8kHz. Therefore, if deployed in a cochlear implant, it can improve the frequency range of the cochlear implant's response for hearing-impaired individuals, resulting in a more realistic hearing experience.

[0129] The audio processing method provided in this application can be executed by an audio processing device. This application uses an audio processing device executing the audio processing method as an example to illustrate the audio processing device provided in this application.

[0130] This application provides an audio processing device, such as... Figure 8 As shown, the audio processing device 400 includes an acquisition module 401 and a processing module 402, wherein: the acquisition module 401 is used to acquire a first time period of the first audio signal appearing in each of the N frequency bands of the first audio, where N is a positive integer; the processing module 402 is further used to process the first audio signal within the first time period within a Fourier window corresponding to the frequency band of the first time period for each first time period acquired by the acquisition module 401, to obtain the first audio signal feature information corresponding to each first time period, where different frequency bands correspond to different Fourier windows; the processing module 402 is further used to input the first audio signal feature information corresponding to each first time period into a generator network model for ultra-high resolution processing, to obtain a second audio signal corresponding to each first time period; the processing module 402 is further used to obtain a second audio based on the second audio signal.

[0131] Optionally, in this embodiment of the application, the above-mentioned generative network model includes a feature extraction layer, a subpixel convolutional layer, and an inverse transformation layer; the above-mentioned processing module 402 is specifically used for: extracting key audio signal feature information from the first audio signal feature information corresponding to each of the first time periods based on the feature extraction layer in the above-mentioned generative network model; performing feature augmentation on each of the key audio signal feature information based on the subpixel convolutional layer in the above-mentioned generative network model to obtain second audio signal feature information corresponding to each of the first time periods; and performing feature inverse transformation on each of the second audio signal feature information based on the inverse transformation layer in the above-mentioned generative network model to obtain second audio signals corresponding to each of the first time periods.

[0132] Optionally, in this embodiment of the application, the processing module 402 is specifically used to splice the second audio signal corresponding to each of the first time periods according to the time sequence of the first audio to obtain the second audio.

[0133] Optionally, in this embodiment of the application, the acquisition module 401 is specifically used to: input the audio signal of the first audio into N filters for filtering processing to obtain N first audio signals corresponding to N frequency bands; process the first audio signal within a first Fourier window for each of the first audio signals to obtain a first time period in each frequency band where the first audio signal appears; divide each of the first time periods into equal parts to obtain M equal second time periods in each of the first time periods, where M is an integer greater than 1.

[0134] Optionally, in this embodiment of the application, the processing module 402 is specifically used to: for each of the M second time periods in each of the first time periods, process the first audio signal in the second time period within a Fourier window corresponding to the frequency band of the second time period to obtain the first audio signal feature information corresponding to each of the second time periods; and concatenate the first audio signal feature information corresponding to the M second time periods in each of the first time periods to obtain the first audio signal feature information corresponding to each of the first time periods.

[0135] In the audio processing apparatus provided in this application embodiment, the audio processing apparatus can first acquire a first time period of the first audio signal appearing in each of the N frequency bands of the first audio, where N is a positive integer; then, for each of the first time periods, the first audio signal within the first time period is processed within a Fourier window corresponding to the frequency band of the first time period to obtain the first audio signal feature information corresponding to each of the first time periods, with different Fourier windows corresponding to different frequency bands; then, the first audio signal feature information corresponding to each of the first time periods is input into a generator network model for ultra-high resolution processing to obtain a second audio signal corresponding to each of the first time periods; finally, a second audio is obtained based on the second audio signal. Thus, since the audio processing apparatus matches different Fourier windows for different frequency bands and processes the first audio signal corresponding to each frequency band within the Fourier window corresponding to each frequency band when processing the N first audio signals corresponding to the N frequency bands, the processing precision of the audio processing apparatus for the N first audio signals is improved, thereby avoiding information loss during audio processing.

[0136] The audio processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.

[0137] The audio processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0138] The audio processing device provided in this application embodiment can achieve... Figures 1 to 7 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0139] Optionally, such as Figure 9 As shown, this application embodiment also provides an electronic device 600, including a processor 601 and a memory 602. The memory 602 stores a program or instructions that can run on the processor 601. When the program or instructions are executed by the processor 601, they implement the various steps of the above-described audio processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0140] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0141] Figure 10 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0142] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.

[0143] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 10 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0144] The processor 110 is configured to acquire a first time period of the first audio signal appearing in each of the N frequency bands of the first audio, where N is a positive integer; the processor 110 is further configured to process the first audio signal within the first time period within a Fourier window corresponding to the frequency band of the first time period for each first time period, to obtain the first audio signal feature information corresponding to each first time period, with different Fourier windows corresponding to different frequency bands; the processor 110 is further configured to input the first audio signal feature information corresponding to each first time period into a generator network model for ultra-high resolution processing, to obtain a second audio signal corresponding to each first time period; the processor 110 is further configured to obtain a second audio based on the second audio signal.

[0145] Optionally, in this embodiment of the application, the above-mentioned generative network model includes a feature extraction layer, a subpixel convolutional layer, and an inverse transformation layer; the processor 110 is specifically used to: extract key audio signal feature information from the first audio signal feature information corresponding to each first time period based on the feature extraction layer in the above-mentioned generative network model; perform feature augmentation on each of the key audio signal feature information based on the subpixel convolutional layer in the above-mentioned generative network model to obtain second audio signal feature information corresponding to each of the first time periods; and perform feature inverse transformation on each of the second audio signal feature information based on the inverse transformation layer in the above-mentioned generative network model to obtain second audio signal corresponding to each first time period.

[0146] Optionally, in this embodiment of the application, the processor 110 is specifically used to splice the second audio signals corresponding to each of the first time periods according to the time sequence of the first audio to obtain the second audio.

[0147] Optionally, in this embodiment of the application, the processor 110 is specifically configured to: input the audio signal of the first audio into N filters for filtering processing to obtain N first audio signals corresponding to N frequency bands; process the first audio signal within a first Fourier window for each of the first audio signals to obtain a first time period in each frequency band where the first audio signal appears; divide each of the first time periods into equal parts to obtain M equal second time periods in each of the first time periods, where M is an integer greater than 1.

[0148] Optionally, in this embodiment of the application, the processor 110 is specifically configured to: process the first audio signal in the second time period within a Fourier window corresponding to the frequency band of the second time period for each of the M second time periods in the first time period, to obtain the first audio signal feature information corresponding to each of the second time periods; and concatenate the first audio signal feature information corresponding to the M second time periods in the first time period to obtain the first audio signal feature information corresponding to each of the first time periods.

[0149] In the electronic device provided in this application embodiment, the electronic device can first acquire the first time period of the first audio signal appearing in each of the N frequency bands of the first audio, where N is a positive integer; then, for each of the first time periods, the first audio signal within the first time period is processed within a Fourier window corresponding to the frequency band of the first time period to obtain the first audio signal feature information corresponding to each of the first time periods, with different Fourier windows corresponding to different frequency bands; then, the first audio signal feature information corresponding to each of the first time periods is input into a generator network model for ultra-high resolution processing to obtain the second audio signal corresponding to each of the first time periods; finally, the second audio is obtained based on the second audio signal. Thus, since the electronic device matches different Fourier windows for different frequency bands when processing the N first audio signals corresponding to the N frequency bands, and processes the first audio signal corresponding to each frequency band within the Fourier window corresponding to each frequency band, the processing precision of the electronic device for the N first audio signals is improved, thereby avoiding information loss during audio processing.

[0150] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.

[0151] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0152] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.

[0153] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described audio processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0154] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0155] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described audio processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0156] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0157] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the audio processing method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0158] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0159] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0160] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An audio processing method, characterized in that, The method includes: Obtain the first time interval in each of the N frequency bands of the first audio signal, where N is a positive integer; For each first time period, the first audio signal within the first time period is processed within the second Fourier window corresponding to the frequency band of the first time period to obtain the first audio signal feature information corresponding to each first time period. Different frequency bands correspond to different second Fourier windows. The size of the second Fourier window corresponding to each frequency band is a fixed value, and the size of the second Fourier window is negatively correlated with the size of the frequency band. The first audio signal feature information corresponding to each first time period is input into the generator network model for ultra-high definition resolution processing to obtain the second audio signal corresponding to each first time period. The second audio signal is used to obtain the second audio signal.

2. The method according to claim 1, characterized in that, The generative network model includes a feature extraction layer, a sub-pixel convolutional layer, and an inverse transformation layer; the step of inputting the feature information of the first audio signal corresponding to each first time period into the generative network model for ultra-high definition resolution processing to obtain the second audio signal corresponding to each first time period includes: Based on the feature extraction layer in the generative network model, key audio signal feature information is extracted from the first audio signal feature information corresponding to each first time period; Based on the sub-pixel convolutional layer in the generative network model, feature augmentation is performed on each key audio signal feature information to obtain the second audio signal feature information corresponding to each first time period. Based on the inverse transformation layer in the generative network model, feature inverse transformation is performed on the feature information of each second audio signal to obtain the second audio signal corresponding to each first time period.

3. The method according to claim 1, characterized in that, The process of obtaining the second audio based on the second audio signal includes: According to the time sequence of the first audio, the second audio signals corresponding to each of the first time periods are spliced ​​together to obtain the second audio.

4. The method according to claim 1, characterized in that, The first time period in which the first audio signal appears in each of the N frequency bands of the first audio is obtained includes: The audio signal of the first audio is input into N filters for filtering processing to obtain N first audio signals corresponding to N frequency bands; For each of the first audio signals, the first audio signal is processed within a first Fourier window to obtain a first time period in which the first audio signal appears in each frequency band; After obtaining the first time period in which the first audio signal appears in each frequency band, the method further includes: Each of the first time periods is divided into equal parts to obtain M equal second time periods within each of the first time periods, where M is an integer greater than 1.

5. The method according to claim 4, characterized in that, For each of the first time periods, the first audio signal within the first time period is processed within a second Fourier window corresponding to the frequency band of the first time period to obtain the first audio signal feature information corresponding to each first time period, including: For each of the M second time periods within each first time period, the first audio signal within the second time period is processed within a second Fourier window corresponding to the frequency band of the second time period to obtain the first audio signal feature information corresponding to each second time period. The first audio signal feature information corresponding to the M second time periods within each first time period is concatenated to obtain the first audio signal feature information corresponding to each first time period.

6. An audio processing apparatus, characterized in that, The device includes: an acquisition module and a processing module, wherein: The acquisition module is used to acquire the first time period of the first audio signal appearing in each of the N frequency bands of the first audio, where N is a positive integer; The processing module is further configured to process the first audio signal within the first time period within a second Fourier window corresponding to the frequency band of the first time period for each first time period acquired by the acquisition module, thereby obtaining the first audio signal feature information corresponding to each first time period. Different frequency bands correspond to different second Fourier windows, and the size of the second Fourier window corresponding to each frequency band is a fixed value, and the size of the second Fourier window is negatively correlated with the size of the frequency band. The processing module is further configured to input the first audio signal feature information corresponding to each first time period into the generator network model for ultra-high resolution processing to obtain the second audio signal corresponding to each first time period. The processing module is further configured to obtain a second audio signal based on the second audio signal.

7. The apparatus according to claim 6, characterized in that, The generative network model includes a feature extraction layer, a subpixel convolutional layer, and an inverse transformation layer; The processing module is specifically used for: Based on the feature extraction layer in the generative network model, key audio signal feature information is extracted from the first audio signal feature information corresponding to each first time period; Based on the sub-pixel convolutional layer in the generative network model, feature augmentation is performed on each key audio signal feature information to obtain the second audio signal feature information corresponding to each first time period. Based on the inverse transformation layer in the generative network model, feature inverse transformation is performed on the feature information of each second audio signal to obtain the second audio signal corresponding to each first time period.

8. The apparatus according to claim 6, characterized in that, The processing module is specifically used to splice the second audio signals corresponding to each of the first time periods according to the time sequence of the first audio to obtain the second audio.

9. The apparatus according to claim 6, characterized in that, The acquisition module is specifically used to input the audio signal of the first audio into N filters for filtering processing to obtain N first audio signals corresponding to N frequency bands; and, for each first audio signal, to process the first audio signal within a first Fourier window to obtain a first time period in which the first audio signal appears in each frequency band. The processing module is further configured to, after obtaining the first time period in which the first audio signal appears in each frequency band, divide each first time period into equal parts to obtain M equal second time periods in each first time period, where M is an integer greater than 1.

10. The apparatus according to claim 9, characterized in that, The processing module is specifically used for: For each of the M second time periods within each first time period, the first audio signal within the second time period is processed within a second Fourier window corresponding to the frequency band of the second time period to obtain the first audio signal feature information corresponding to each second time period. The first audio signal feature information corresponding to the M second time periods within each first time period is concatenated to obtain the first audio signal feature information corresponding to each first time period.

11. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the audio processing method as described in any one of claims 1 to 5.

12. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the audio processing method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Human voice detection method and device, electronic equipment and computer readable storage medium

    CN112967738A

  • Voice enhancement method and device based on neural network, and electronic equipment

    CN113808607A