Audio processing method, electronic device and computer readable storage medium
By applying a pre-trained audio enhancement model in the terminal device, spectrum expansion of the real-time playback audio is solved, and the problem of insufficient audio signal components in the prior art is achieved, and the ultra-high-quality audio playback effect is achieved, which improves the user experience.
Patent Information
- Application Number
- CN202210379551.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-04-12
AI Technical Summary
The prior art is difficult to improve the audio signal components of old or poorly recorded audio in terminal devices in real time, resulting in poor user experience.
By applying a pre-trained audio enhancement model in the terminal device, spectrum expansion is performed on the audio being played, the specific steps include obtaining the audio to be played, inputting the audio enhancement model, and processing high-band modes and low-band phases to generate ultra-high-quality audio.
Real-time audio enhancement of the target audio in the real-time audio playback state is achieved, improving the audio playback effect and user auditory experience.
Smart Images

Figure CN114627882B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing technology, and in particular to an audio processing method, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the development of terminal devices, various applications emerge in an endless stream. When using terminal devices to play songs, for some old songs or songs recorded by devices with poor recording effects, these songs are not rich in audio signal components, and therefore cannot provide users with the ultimate auditory experience, thus affecting the user experience. Summary of the invention
[0003] The embodiments of the present application provide an audio processing method, an electronic device, and a computer-readable storage medium, which can perform spectrum expansion on the audio being played in real time to improve the audio playback effect.
[0004] In a first aspect, an embodiment of the present application discloses an audio processing method, which is applied to a terminal device, the terminal device includes an audio enhancement model, and the method includes:
[0005] In response to an audio enhancement start instruction for a target audio, a first audio to be played is obtained from the target audio; the target audio is the audio in a playing state;
[0006] Inputting the first audio into a pre-trained audio enhancement model to obtain a second audio output by the audio enhancement model;
[0007] The second audio is processed according to the high-frequency band mode of the first audio and the phase of the second audio is corrected according to the low-frequency band phase of the first audio, so as to process the second audio into a third audio; the third audio is an ultra-high-quality audio;
[0008] Replace playing the target audio with playing the third audio.
[0009] In a second aspect, an embodiment of the present application discloses an audio processing device, the device comprising:
[0010] An acquisition module, configured to respond to an audio enhancement start instruction for a target audio and acquire a first audio to be played from the target audio; the target audio is the audio in a playing state;
[0011] A processing module, used for inputting the first audio into a pre-trained audio enhancement model to obtain a second audio output by the audio enhancement model;
[0012] The processing module is further used to process the mode of the second audio according to the high-frequency band mode of the first audio and to correct the phase of the second audio according to the low-frequency band phase of the first audio, so as to process the second audio into a third audio; the third audio is an ultra-high-quality audio;
[0013] The playing module is used to replace the playing target audio with the playing third audio.
[0014] In a third aspect, an embodiment of the present invention provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the audio processing method provided in the first aspect.
[0015] In a fourth aspect, an embodiment of the present invention provides a computer storage medium, characterized in that computer program instructions are stored in the computer storage medium, and when the computer program instructions are executed by a processor, they are used to execute the audio processing method provided in the first aspect.
[0016] In a fifth aspect, an embodiment of the present invention provides a computer program product or a computer program, wherein the computer program product includes a computer program, and the computer program is stored in a computer storage medium; when a processor of a computer device reads the computer instructions from the computer storage medium, the processor executes the audio processing method provided in the first aspect above.
[0017] In an embodiment of the present invention, the terminal device receives and responds to an audio enhancement start instruction for a target audio input in a playback state, and can obtain a first audio to be played from the target audio according to the audio enhancement start instruction, input the first audio into a pre-trained audio enhancement model, obtain a second audio output by the audio enhancement model, and process the mode of the second audio according to the high-frequency band mode of the first audio, and correct the phase of the second audio according to the low-frequency band phase of the first audio, so as to process the second audio into an ultra-high-quality third audio, and replace the playback of the target audio with the playback of the third audio, so that the terminal device can perform audio enhancement processing on the target audio in real time in the real-time playback state of the audio, and thereby obtain an ultra-high-quality third audio to enhance the user's auditory experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0019] Figure 1Ais a schematic diagram of time domain spectrum expansion provided by an embodiment of the present application;
[0020] Figure 1B is a schematic diagram of frequency domain spectrum expansion provided by an embodiment of the present application;
[0021] Figure 1C It is a schematic diagram of a system architecture provided by an embodiment of the present application;
[0022] Figure 2 It is a flowchart of an audio processing method provided in an embodiment of the present application;
[0023] Figure 3 is a schematic diagram of a target audio playback interface provided in an embodiment of the present application;
[0024] Figure 4 is a schematic diagram of another target audio playback interface provided in an embodiment of the present application;
[0025] Figure 5 This is a schematic diagram of another target audio playback interface provided in an embodiment of the present application;
[0026] Fig. 6A is a schematic diagram of an audio setting interface corresponding to a target audio provided in an embodiment of the present application;
[0027] Figure 6B is a schematic diagram of a first audio provided by an embodiment of the present application;
[0028] Figure 6C is a schematic diagram of another first audio provided in an embodiment of the present application;
[0029] Fig.6D is a schematic diagram of another first audio provided in an embodiment of the present application;
[0030] Fig. 6E is a schematic diagram of another first audio provided in an embodiment of the present application;
[0031] Fig. 7A This is a schematic diagram of a post-processing process of a high-frequency mode provided in an embodiment of the present application;
[0032] Figure 7B It is a flowchart of an improved Griffinlim algorithm provided in an embodiment of the present application;
[0033] Fig. 8A is a schematic diagram showing a first audio spectrum diagram and a third audio spectrum diagram provided by an embodiment of the present application;
[0034] Figure 8B is a schematic diagram showing a before-and-after comparison spectrum of an audio clip 1 provided in an embodiment of the present application;
[0035] Figure 8C is another schematic diagram showing the before and after comparison of the frequency spectrum of the audio clip 1 provided in an embodiment of the present application;
[0036] Fig. 9 It is a flowchart of another audio enhancement model generation method provided in an embodiment of the present application;
[0037] Fig.10 It is a schematic diagram of the internal structure of an encoding-decoding architecture provided in an embodiment of the present application;
[0038] Fig.11 It is a flowchart of training through a generative adversarial network provided in an embodiment of the present application;
[0039] Fig.12 It is a schematic diagram of low-band and high-band frequency points provided by an embodiment of the present application;
[0040] Fig.13 It is a schematic diagram of a process of making predictions by generating an adversarial network model provided in an embodiment of the present application;
[0041] Fig.14 is a structural schematic diagram of an audio processing device provided in an embodiment of the present application;
[0042] Fig.15 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] To facilitate understanding, the terms involved in this application are first introduced.
[0044] 1. Super quality (SQ) sound quality
[0045] SQ sound quality can also be called lossless sound quality. If the audio source is a compact disc (CD) quality or higher audio than CD quality, then audio A can be SQ sound quality. Alternatively, if the audio obtained by passing audio A through a lossless encoder is recorded as audio B, then audio B can be SQ sound quality. Alternatively, if audio A or audio B can still be stored in a lossless format after being encoded by a higher quality lossy encoder, the encoded audio can be recorded as audio C, and audio C can be SQ sound quality.
[0046] 2. High quality (HQ) sound quality
[0047] HQ sound quality refers to audio in standard mp3 format (bit rate 320kbps) that has not been converted to a low bit rate. The main technical indicators for judging whether it is HQ sound quality are: (1) the highest tangent of the spectrum reaches above 20K; (2) if there is attenuation near the 20K spectrum, the proportion of the spectrum above 20K must be greater than 25%; (3) if there is attenuation near 18K, the proportion of the spectrum above 18K must be greater than 45%. Among them, the spectrum of the standard mp3 format (320kbps) can reach 20K.
[0048] 3. Low quality (LQ) sound quality
[0049] If the audio sampled at a lower sampling rate (such as below 28kHz) is recorded as audio D, then audio D can be of LQ sound quality. Alternatively, if the audio encoded by a lower bit rate encoder (such as LAME 64bits version 3.99.5 encoder at a bit rate of 80kbps or less) and the spectral height does not reach 14kHz is recorded as audio E, then audio E can be of LQ sound quality.
[0050] 4. Not super high quality sound
[0051] Non-ultra-high quality sound can also be called non-lossless sound. Non-ultra-high quality sound can be HQ sound or LQ sound. The main technical indicators for judging whether it is non-lossless sound can be summarized as follows: (1) The file format is compressed by a lossy encoder; (2) The file format is compressed by a lossless encoder, or is uncompressed, but it can be found that it has been compressed with lossy code and then saved as a lossless format, that is, there are obvious traces of lossy compression; among them, the above traces can be mainly described by the spectrum height and density: the spectrum reaches 20k or more (excluding 20k) and meets a certain ratio, such as more than 10% of effective energy.
[0052] 5. Music spectrum expansion
[0053] Music Bandwidth Extension technology can also be called Music Super Resolution technology. From the perspective of time domain, Figure 1A As shown in , deep neural network (DNN) technology can be used for time domain interpolation to introduce high-frequency details; from the frequency domain, as shown in Figure 1B As shown, DNN technology can be used to repair and reconstruct the lost high-frequency components.
[0054] The following will be combined with the accompanying drawings to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0055] In order to enhance the rhythm and somatosensory feedback of music and improve user experience, the present application embodiment provides an audio processing method. In order to better understand the audio processing method provided by the present application embodiment, the system architecture applied by the audio processing method is first introduced below.
[0056] See also Figure 1C , is a schematic diagram of a system architecture provided in an embodiment of the present application, and the audio processing method proposed in the present application can be executed through the system architecture. The system architecture includes a terminal device 101 and a server 102. The terminal device 101 and the server 102 are connected to each other by wired or wireless means. It should be noted that the above-mentioned terminal device 101 can be provided with an audio enhancement model to process the audio being played through the audio enhancement model, so as to obtain ultra-high-quality audio.
[0057] The terminal device 101 is a device with wireless communication function, which can be a smart phone, a tablet computer, a smart wearable device, a personal computer, etc., in which an application client, such as an audio player software, can be run. In some embodiments of the present application, the terminal device can also be a device with transceiver function, such as a chip system. The chip system can include a chip and other discrete devices, which is not limited in the embodiments of the present application.
[0058] The terminal device 101 can obtain data such as video, audio, text, etc. from the server 102. Among them, the server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and basic cloud computing services such as big data and artificial intelligence platforms. The database in the server 102 can be a local database of the server 102, or it can be a cloud database that the server 102 can access, and this application does not limit this. It should be noted that the above system architecture is described by taking a terminal device 101 and a server 102 as an example, and the number of terminal devices and the number of servers do not constitute a limitation on this application.
[0059] The audio processing method provided in the embodiment of the present application is further described in detail below:
[0060] See also Figure 2 , Figure 2 is a flowchart of an audio processing method provided by an embodiment of the present application. The audio processing method can be performed by a terminal device (such as the above Figure 1C The method may also be executed by a server, or by a server and a terminal together (e.g., the server obtains the target audio played by the terminal, performs sound quality enhancement on the first audio in the target audio to obtain the second audio, and determines a third audio of ultra-high quality based on the first audio and the second audio, and returns the third audio to the terminal for playback). For ease of understanding, Figure 2 The embodiment is described by taking the method executed by a terminal device as an example. The audio processing method may at least include the following steps S201 to S204:
[0061] S201. In response to an audio enhancement start instruction for a target audio, obtain a first audio to be played from the target audio.
[0062] Among them, the target audio is the audio in the playing state. The target audio can be non-ultra-high quality audio, that is, non-lossless audio, such as audio with HQ sound quality or LQ sound quality. The target audio can be a song, a recording, etc., and this application does not limit this. The above-mentioned target audio in the playing state can be understood as the terminal device is playing the target audio, for example, the music application or playback module (such as a radio or a recorder) of the terminal device is playing the target audio. This application is explained by taking the music application on the terminal device playing the target audio as an example.
[0063] The above audio enhancement start instruction can be used to start the audio processing flow for the target audio, that is, to execute steps S201 to S204 for the target audio. It should be noted that the user can input the audio enhancement start instruction for the target audio in the above music application, and accordingly, the terminal device receives the enhancement start instruction.
[0064] The terminal device may receive an audio enhancement start instruction for the target audio input in one of the following two interfaces. Specifically:
[0065] Interface 1: target audio playback interface. That is, the terminal device can receive an audio enhancement start instruction for the target audio input on the target audio playback interface.
[0066] The target audio playback interface may be an interface when a music application on a terminal device plays the target audio; Figure 3 As shown, Figure 3 The target audio playback interface is shown. Figure 3In the playback interface shown, different sound qualities and the corresponding file sizes of each sound quality can be listed below the target audio, such as: standard quality 2.0M, HQ quality 7.1M, SQ quality 18.5M, etc.; an audio enhancement start button can be provided below each sound quality option to start the audio enhancement processing flow. Optionally, the user can choose among the three different sound qualities mentioned above, so that the terminal device can play the target audio according to the sound quality selected by the user. Optionally, the user can also set automatic selection of sound quality, so that the terminal device can automatically select the currently applicable sound quality based on factors such as the current network conditions. Figure 3 As shown, before the audio enhancement button is turned on, the target audio can be played at HQ sound quality. Optionally, the sound quality option and the audio enhancement on button are located below the target audio for example only and do not constitute a limitation of the present application.
[0067] It should be noted that in Figure 3 In the target audio playback interface shown, the user can click the audio enhancement start button for the target audio to start the audio enhancement processing flow for the target audio, that is, execute steps S201 to S204.
[0068] Interface 2: Audio setting interface corresponding to the target audio. That is, the terminal device can receive an audio enhancement start instruction for the target audio input in the audio setting interface corresponding to the target audio.
[0069] It should be noted that, if there is no audio enhancement on button on the target audio playback interface, the terminal device can jump to the audio setting interface corresponding to the target audio by clicking on the target audio playback interface. Figure 4 As shown, Figure 4 If there is no audio enhancement on button in the playback interface of the target audio in the video, the terminal device can jump to the audio setting interface corresponding to the target audio after receiving a click operation on the audio enhancement button; or Figure 5 As shown, Figure 5 If there is no audio enhancement on button on the playback interface of the target audio, the terminal device can jump to the audio setting interface corresponding to the target audio after receiving the audio setting operation for the target audio input (such as by right-clicking and selecting the audio setting button from the menu). This application does not impose any restrictions on this.
[0070] like Fig. 6A As shown, Fig. 6A The audio setting interface corresponding to the target audio is shown. Fig. 6A In the audio setting interface corresponding to the target audio shown, the user can click the audio enhancement start button to start the audio enhancement process for the target audio.
[0071] When the terminal device plays the target audio, it can receive the audio enhancement start instruction input by the user for the target audio. In other words, when the terminal device plays the target audio, it can receive the audio enhancement start instruction input by the user for the target audio to expand the spectrum of the target audio, thereby improving the audio quality and further improving the user experience.
[0072] In order to distinguish the audio after the sound quality enhancement, the audio before the sound quality enhancement may be referred to as the first audio. The first audio may be the audio obtained by the terminal device from the target audio, such as Figure 6B As shown, the first audio may be a short audio obtained from the target audio (the first audio is obtained according to a preset duration); optionally, Figure 6C As shown, the first audio may also be a longer audio obtained from the target audio (the first audio is obtained from the time point when the audio enhancement start instruction is received to the end time point of the target audio), and this application does not impose any restrictions on this.
[0073] It should be noted that, when receiving the audio enhancement start instruction, the terminal device may use the time point of receiving the audio enhancement start instruction as the starting time point to obtain the first audio of the preset duration from the starting time point. Figure 6B or Figure 6C As shown, if the target audio has played 30 seconds of audio data when the terminal device receives the audio enhancement start instruction, the terminal device can use the 30th second as the starting time point to obtain the above-mentioned first audio from the 30th second.
[0074] In an optional implementation, the start time point for obtaining the first audio is the time point when the audio enhancement start instruction is received, and the end time point is the end time point of the target audio. In this case, if the preset duration for obtaining the first audio is shorter than the remaining playback duration of the target audio, multiple first audios can be obtained from the target audio. In addition, since the duration of different audios may be different, the number of first audios obtained by the terminal device from different audios may also be different. For example: assuming that the total duration of audio 1 is 3 minutes, if the duration of the first audio is 5 seconds, the terminal device can obtain up to 36 first audios from audio 1; assuming that the total duration of audio 2 is 1 minute, if the duration of the first audio is 5 seconds, the terminal device can obtain up to 12 first audios from audio 2.
[0075] Optionally, for different audios, the duration of the first audio obtained by the terminal device may also be different. For example: assuming that the total duration of audio 3 is 3 minutes, for this audio 3, if the duration of the first audio obtained by the terminal device is 5 seconds, then at most 36 first audios can be obtained from this audio 3; assuming that the total duration of audio 4 is 3 minutes, for this audio 4, if the duration of the first audio obtained by the terminal device is 10 seconds, then at most 18 first audios can be obtained from this audio 4. This application does not limit the number of first audios obtained by the terminal device.
[0076] In one implementation, the first audio may include one or more audio segments. As shown in FIG6D , the first audio may include an audio segment; Fig. 6E As shown, the first audio may also include multiple audio segments, which is not limited in the present application.
[0077] It should be noted that, when the first audio includes one audio segment, the number of the first audio acquired by the terminal device may be the same as the number of the audio segments. Fig.6D As shown, the total duration of the target audio is 3 minutes, and the first audio (such as called the first audio 1) includes an audio segment. If the duration of the audio segment is 5 seconds, then when the target audio has played 30 seconds of audio data, the terminal device can obtain 30 audio segments (i.e., 30 first audio 1s) from the target audio.
[0078] Optionally, when the first audio includes multiple audio segments, the number of the first audio acquired by the terminal device may be different from the number of the audio segments. Fig. 6E As shown, the total duration of the target audio is 3 minutes, and the first audio (referred to as the first audio 2) includes all audio data starting from the 30th second of the target audio. The first audio 2 includes multiple audio clips. If the duration of the audio clip is 5 seconds, when the target audio has played 30 seconds of audio data, the terminal device can obtain 1 first audio 2 from the target audio, and the first audio 2 includes 30 audio clips.
[0079] It should be noted that an audio clip may include the minimum amount of audio data required for the terminal device to perform audio enhancement processing, so that the terminal device can perform audio enhancement processing after obtaining the audio clip. For example, if the minimum amount of audio data that the terminal device can perform audio enhancement processing is 5 milliseconds of audio, then an audio clip may include at least 5 milliseconds of audio. Among them, the minimum amount of audio data for audio enhancement processing is 5 milliseconds of audio for example only and does not constitute a limitation of this application.
[0080] The above audio clips may also be referred to as audio data of unit duration, that is, the first audio may include one or more audio data of unit duration. Optionally, the above audio clip may also be referred to as audio data of a batch size. The embodiments of the present application are all described by taking audio clips as examples, which does not constitute a limitation on the present application. Optionally, the above audio clips may also be in frames, which is not limited in the present application.
[0081] After receiving the audio enhancement start instruction, the terminal device can obtain one or more audio segments included in the first audio from the target audio according to the audio enhancement start instruction, so as to divide the target audio with a longer duration being played into one or more audio segments with a shorter duration, so that after obtaining an audio segment, the audio enhancement processing can be started quickly, and the processed audio segment can be obtained more quickly.
[0082] S202: Input the first audio into a pre-trained audio enhancement model to obtain a second audio output by the audio enhancement model.
[0083] The second audio may be the audio output by the audio enhancement model. The second audio may be the audio after the first audio has been processed by audio enhancement, and the second audio may contain more signal components than the first audio. It should be noted that the second audio may correspond to the first audio, that is, if the first audio includes one audio segment, the second audio also includes one audio segment; if the first audio includes multiple audio segments, the second audio also includes multiple audio segments.
[0084] In one implementation, the audio enhancement model is a model obtained by training audio samples through a generative adversarial network (GAN), and the audio samples are ultra-high-quality audio. How to generate an audio enhancement model will be described in Fig. 9 The embodiments shown are described in detail and will not be repeated here.
[0085] S203, processing the mode of the second audio according to the high-frequency band mode of the first audio and correcting the phase of the second audio according to the low-frequency band phase of the first audio, so as to process the second audio into a third audio; the third audio is an ultra-high-quality audio.
[0086] The third audio may be an audio after high-frequency modulus post-processing and phase correction based on the first audio and the second audio. That is, the third audio may be obtained after the modulus and phase of the second audio output by the audio enhancement model are optimized.
[0087] In one implementation, based on the high-band mode of the first audio and the high-band mode of the second audio, the mode of the second audio is post-processed to obtain the full-band mode of the second audio; according to the low-band phase of the first audio, the phase of the second audio is corrected to obtain the full-band phase of the second audio; the second audio is processed based on the full-band mode and the full-band phase to determine the third audio.
[0088] The high-frequency mode post-processing process can be found in Fig. 7A Schematic diagram shown. Specifically, the terminal device can compare the high-frequency band mode of the first audio with the high-frequency band mode of the second audio to determine the high-frequency band mode with greater energy of the high-frequency frequency point, so as to adopt the high-frequency band mode with greater energy as the high-frequency band mode after post-processing of the high-frequency mode. For example: if the energy of the high-frequency band mode of the first audio is higher than the energy of the high-frequency band mode of the second audio, the high-frequency band mode of the first audio is adopted as the high-frequency band mode after post-processing of the high-frequency mode. Through the post-processing process of the high-frequency mode, the original high-frequency frequency point can be retained as much as possible when the original high-frequency frequency point has greater energy, thereby achieving the goal of not damaging the characteristics of the original audio. Optionally, the above-mentioned post-processing process of the high-frequency mode is similar to the post-processing process of the mode for the image, which will not be repeated here.
[0089] In one implementation, the low-frequency band phase of the first audio is mirrored to obtain a mirror phase; the mirror phase is calculated using a speech signal reconstruction algorithm to obtain a calculated phase; and based on the calculated phase, the phase of the second audio is corrected to obtain a full-band phase of the second audio.
[0090] The above-mentioned speech signal reconstruction algorithm can be an improved Griffinlim algorithm, and the specific process of the improved Griffinlim algorithm can be as follows: Figure 7B Specifically, in the embodiment of the present application, the low-frequency band phase of the first audio may be mirrored, and the mirrored phase may be used as the initial phase in the improved Griffinlim algorithm, so that the calculated phase is more accurate.
[0091] It should be noted that the number of iterations for the phase correction processing of the second audio can be manually set. Although the more iterations, the more accurate the result, in order to reduce the amount of calculation of the terminal device, it can generally be manually set to 1 or 2 iterations, and this application does not limit this. Optionally, for other steps corresponding to the improved Griffinlim algorithm, please refer to the relevant steps of the Griffinlim algorithm, which will not be repeated in this application.
[0092] The terminal device can determine the third audio based on the second audio, the full-band mode after post-processing of the high-frequency mode, and the full-band phase after phase correction by the improved Griffinlim algorithm. Specifically, the terminal device can combine the above-mentioned second audio, full-band mode and full-band phase based on the Euler formula, and restore the time domain signal through the inverse short-time Fourier transform (ISTFT), which is not limited in this application. By optimizing the mode and phase of the second audio, the details in the processed third audio can be made more accurate, thereby improving the user's listening experience.
[0093] It should be noted that after the terminal device obtains the third audio, it can cache the third audio so that the terminal device can obtain and play the third audio. Optionally, the terminal device can cache the third audio to the memory of the terminal device, or cache the third audio to other storage space, which is not limited in this application.
[0094] In one implementation, the terminal device can display the frequency spectra of the first audio and the third audio on the audio setting interface.
[0095] It should be noted that, in order to simplify the description, the following text will use "the third audio is ultra-high quality audio after audio enhancement processing" as an example. The "audio enhancement processing" can mean "audio enhancement processing and modulus and phase optimization processing", which does not constitute a limitation on the present application.
[0096] As can be seen from the above content, the first audio may be the audio before audio enhancement processing, and the third audio may be the ultra-high quality audio after audio enhancement processing. Fig. 8A As shown, after turning on the enhancement processing button in the audio setting interface, the terminal device can also display the spectrum diagram of the audio before audio enhancement (i.e., the first audio) and the spectrum diagram of the audio after audio enhancement (i.e., the third audio) in the audio setting interface, so as to show the user the comparison effect before and after turning on audio enhancement. Fig. 8A The placement positions of the first audio and third audio spectrum diagrams are for example only and do not constitute a limitation to the present application.
[0097] Optionally, the above-mentioned before-after comparison effect diagram of audio enhancement may be exemplary, for example, the spectrum height before audio enhancement is turned on may be a spectrum of about 12K, and the spectrum height after audio enhancement is turned on may be a spectrum of about 22K. It is understandable that the legends corresponding to different audios may be the same or different to show the user a general effect, and do not represent the actual spectrum of the first audio and the third audio.
[0098] Optionally, the before-after comparison effect diagram of the audio enhancement may be an example of an actual legend generated during the audio enhancement process. For example, the spectrum diagram before audio enhancement is turned on may be the spectrum of the first audio being played, and the spectrum diagram after audio enhancement is turned on may be the spectrum of the third audio after audio enhancement processing. It is understandable that the before-after comparison effect diagrams corresponding to different audios may be different, so as to show the user the spectrum of the first audio and the third audio during the actual processing process.
[0099] In one implementation, the terminal device can display a first spectrum and a second spectrum of a target audio segment currently being processed by the audio enhancement model; the first spectrum is the spectrum of the target audio segment before it is input into the audio enhancement model, and the second spectrum is the spectrum of the target audio segment after it is input into the audio enhancement model, and the target audio segment is an audio segment being processed among multiple audio segments.
[0100] Among them, the above-mentioned first spectrum can be the spectrum of the target audio segment before audio enhancement processing (that is, before the above-mentioned input audio enhancement model), and the above-mentioned second spectrum can be the spectrum of the target audio segment after audio enhancement processing (that is, after the above-mentioned input audio enhancement model), and the present application does not impose any restrictions on this.
[0101] It should be noted that, when the first audio includes multiple audio segments (ie, multiple audio segments), the before-after comparison effect diagram can be used to show the frequency spectrum of the audio segment being processed. Figure 8B As shown, assuming that the first audio includes audio segment 1, audio segment 2 and audio segment 3, and the terminal device is performing audio enhancement processing on audio segment 1 (i.e., the above-mentioned target audio segment), the before-after comparison effect diagram may show the before-after comparison spectrum when processing audio segment 1. Optionally, if the terminal device starts to process audio segment 2 (i.e., the above-mentioned target audio segment), the before-after comparison effect diagram may show the before-after comparison spectrum when processing audio segment 2, which is not limited in the present application.
[0102] Optionally, assuming that the first audio includes audio segment 1, audio segment 2, and audio segment 3, the before-after comparison effect diagram may first show the before-after comparison spectrum of audio segment 1, then show the before-after comparison spectrum of audio segment 2, and finally show the before-after comparison spectrum of audio segment 3, so as to display the before-after comparison spectrum of all audio segments included in the first audio. Figure 8C As shown, the figure first shows the before and after comparison frequency spectra of audio clip 1 by way of example.
[0103] S204: Replace the target audio with the third audio.
[0104] After the terminal device obtains the third audio, it can play the third audio for the user. Optionally, when the terminal device obtains the first audio, it can pause the playback of the first audio, and when the third audio is obtained, it can cache and play the third audio in a timely manner to enhance the user's auditory experience. Among them, the interval time from pausing the first audio to playing the third audio is short, that is, the time for the first audio to be input into the audio enhancement model for processing is short, and the audio enhancement model can quickly process one or more audio clips included in the first audio to reduce the user's waiting time.
[0105] In one implementation, the terminal device deletes the cached third audio when the playing of the third audio is completed.
[0106] As can be seen from the foregoing, the third audio may include one or more audio clips. When the terminal device plays the third audio, it may play the audio clips in sequence, and when one audio clip is played, the cache of the audio clip is deleted. Exemplarily, assuming that the third audio includes multiple audio clips, such as audio clip 4, audio clip 5, and audio clip 6, and audio clip 4 is played, the terminal device may delete the cached audio clip 4. Optionally, the terminal device may also delete the cached second audio after the second audio is played, that is, after audio clip 4, audio clip 5, and audio clip 6 are all played, and this application does not impose any restrictions on this. By deleting the audio clips that have been played, the cache can be cleared, thereby freeing up storage space to store audio clips that the terminal device subsequently performs audio enhancement processing on.
[0107] In one implementation, when the terminal device receives a playback progress bar dragging instruction or an audio switching instruction for the target audio input, the terminal device deletes the cached third audio.
[0108] Among them, the play progress bar dragging instruction can be used to drag the play progress of the audio, such as dragging the audio to 1 minute and 30 seconds; the audio switching instruction can be used to switch the audio being played, such as switching the audio 1 being played to audio 3.
[0109] Since the terminal device performs audio enhancement processing on the currently playing audio in real time, when the terminal device receives a play progress bar dragging instruction or an audio switching instruction for the target audio input, the audio that the terminal device needs to perform audio enhancement processing has changed; that is, the audio that has been previously processed and cached will no longer be played, so the terminal device can delete the cached audio to clear the cache, and then use the cache to store the audio that has been processed after the play progress bar dragging instruction or the audio switching instruction has been executed and has undergone audio enhancement processing.
[0110] In one implementation, when the terminal device receives an audio enhancement shut-down instruction for the target audio input, it stops inputting the audio in the target audio into the audio enhancement model.
[0111] The audio enhancement off instruction can be used to turn off the audio enhancement processing flow. When the terminal device receives the audio enhancement off instruction for the target audio input, it can stop the audio enhancement processing for the target audio, that is, stop inputting the audio in the target audio (such as the first audio) into the audio enhancement model, thereby ending the above audio enhancement processing process.
[0112] It should be noted that, when the audio enhancement processing effect for the target audio is not ideal, or when the audio enhancement processing for the target audio fails, the terminal device can receive the above-mentioned music enhancement off instruction to reduce the energy consumption of the terminal device. It is understandable that when the terminal device receives the audio enhancement off instruction, the target audio will be played with the original sound quality.
[0113] It should also be noted that the target audio may also be ultra-high quality audio, i.e., lossless audio, such as SQ quality audio. The terminal device can make up for the lack of high frequencies of the target audio by performing audio enhancement processing on the ultra-high quality target audio, thereby making the enhanced audio sound louder and more detailed.
[0114] In an embodiment of the present application, the terminal device receives an audio enhancement start instruction input for audio (such as target audio) in a playing state, and can obtain the first audio from the target audio according to the audio enhancement start instruction, so as to input the first audio into the audio enhancement model to obtain the audio output by the audio enhancement model (such as the second audio mentioned above), and determine the third audio after the modulus and phase are optimized based on the first audio and the second audio, so as to cache and play the third audio, so that the terminal device can perform audio enhancement processing on the audio in the real-time playback state of the audio, and thus obtain ultra-high-quality audio to enhance the user's auditory experience.
[0115] See also Fig. 9 , Fig. 9 is a flow chart of a method for generating an audio enhancement model provided in an embodiment of the present application. The method for generating an audio enhancement model can be performed by a terminal device (the terminal device can be Figure 2 The terminal device of the embodiment shown may also be other terminal devices) or a server. Figure 2 The illustrated embodiment is executed by the terminal device 1, the audio enhancement model generation method is executed by the terminal device 2 or the server, and the audio enhancement model generated by the terminal device 2 or the server can be deployed in the terminal device 1. For ease of understanding and distinction, Fig. 9The illustrated embodiment is described by taking the method executed by the terminal device 2 as an example. The audio enhancement model generation method may at least include the following steps S901 to S902:
[0116] S901: Generate an initial architecture of an audio enhancement model.
[0117] The initial architecture of the audio enhancement model may include one or more algorithms required by the audio enhancement model. Since the audio enhancement model can obtain the mapping relationship of the high frequency band from the low frequency band through model learning, an encoder-decoder network model may be used as the initial architecture of the audio enhancement model.
[0118] Optional, Fig.10 The internal structure diagram of the encoder-decoder architecture. Fig.10 As shown in the figure, the encoder-decoder architecture can include lightweight algorithms such as depthwise separable convolution (DWconv2D), split super-resolution module (SplitSRBlock) and sub-pixel convolution (SubPixel2D). The specific calculation process of several lightweight algorithms will be briefly described below.
[0119] 1. DWconv2D
[0120] Among them, DWconv2D is the collective name of depthwise (DW) convolution and pointwise (PW) convolution. Depthwise convolution is different from conventional convolution operation. One convolution kernel of depthwise convolution is responsible for one channel, and one channel is convolved by only one convolution kernel; while each convolution kernel of conventional convolution operates each channel of the input at the same time. Taking image processing as an example, for a 5×5 pixel, three-channel color input image (that is, the size of the input image is 5×5×3), depthwise convolution can be operated in a two-dimensional plane, and the number of its convolution kernels is the same as the number of channels in the previous layer (channels and convolution kernels correspond one to one); therefore, a three-channel image can generate 3 feature images after operation.
[0121] The operation of point-by-point convolution is very similar to that of regular convolution, except that the size of the convolution kernel of point-by-point convolution is 1×1×M, where M is the depth of the previous layer. Therefore, the image of the previous step can be weighted combined in the depth direction through point-by-point convolution to generate a new feature image. In point-by-point convolution, the number of feature images is the same as the number of filters.
[0122] 2. SplitSRBlock
[0123] Based on DWconv2D, a new end-to-end mobile super-resolution system SplitSR is proposed. Split convolution divides the input features along the depth channel at a certain ratio (the ratio is adjustable) to reduce the amount of calculation and memory loss, thereby accelerating the inference process. Specifically, the input features can be separated along the depth channel, one part participates in the DWconv2D calculation, and the other part does not participate in any calculation (that is, feature retention); then the features of the two parts are merged and spliced according to the depth channel. This algorithm can reduce the amount of calculation and retain some features to the next layer, so that the decoding layer can also obtain more primary features.
[0124] 3. SubPixel2D
[0125] SubPixel2D is an algorithm that combines upsampling and convolution operations. The algorithm can act on low-resolution features so that the low-resolution features can obtain high-resolution features through the algorithm. This algorithm can reduce the risk of introducing too many factors when using inverse convolution as an upsampling method. SubPixel2D can rearrange each channel of each pixel into an r*r area to correspond to an r*r sub-block in the high-resolution image, that is, a feature image of size 1*H*W can be rearranged into a high-resolution image of size 1*rH*rW. If the feature size is reorganized by a four-dimensional vector, it can be rearranged from [B,H,W, r*r*C] to [B,rH,rW,C]. Although this algorithm is called sub-pixel convolution, it does not actually have a convolution operation.
[0126] In an embodiment of the present application, a streamlined encoder-decoder network model is used to generate the initial architecture of the audio enhancement model, so that the number of network layers and the number of input frames can be reduced; wherein, the encoder-decoder network model uses a special encoding (encoder) network structure unit to reduce the amount of calculation and model size; and uses a special decoding (decoder) module for upsampling to reduce the memory usage unit, so that when the terminal device 1 runs the audio enhancement model, it can run in real time while occupying extremely low central processing unit (CPU) resources and memory consumption.
[0127] S902: Train the audio samples through a generative adversarial network (GAN) to obtain an audio enhancement model, wherein the audio samples may be ultra-high quality audio.
[0128] It should be noted that the above GAN training may include a generator and a discriminator. Through GAN training, the predicted audio generated by the generator can be made indistinguishable from true or false. The generator can use the lightweight encoder-decoder architecture mentioned in the above steps, and the discriminator can use a binary classification network model structure (such as VGG-like).
[0129] like Fig.11 As shown, Fig.11 The flowchart of the GAN training is shown. Specifically, the terminal device 2 can obtain a fourth audio from the audio sample, where the fourth audio is a low-frequency band audio of the audio sample; and input the fourth audio into the encoder-decoder architecture to obtain a fifth audio; the fifth audio can be the audio obtained through the GAN training.
[0130] It should be noted that when the audio sample is obtained, the terminal device 2 can extract the short-time Fourier transform (STFT) feature of the audio sample, take the modulus of the STFT feature, and then take the logarithm to obtain the modulus logarithm of the audio sample; wherein fft_length=2048, hop_length=256. The modulus logarithm size of the audio sample can be expressed as [T, 1024], wherein T can be the length of a feature sequence (i.e., the first audio mentioned above), and the size of the T value is related to the duration of the audio sample. Optionally, if the modulus logarithm is converted to a fixed frame length of 32 frames, the modulus logarithm size of the audio sample can be expressed as [X, 32, 1024], wherein X=T / 32.
[0131] Since the lightweight DWconv2D is used as the convolution operation unit in the encoder-decoder architecture, a four-dimensional input and output can be constructed. For example, if the second dimension is expanded to 1 by default, the logarithmic size of the above audio sample can be expressed as [X, 1, 32, 464], that is, the input of the model during the training process can be [X, 1, 32, 464], and each value can be expressed as [batch size, fixed value 1, frame length, number of low-frequency band points]. Among them, the number of low-frequency band points 464 can correspond to the spectrum height of 10K, and its calculation method can be: 464 = 2048 * 10K (spectrum height) / 44.1K (sampling rate). It can be understood that the output of the model during the training process can be [X, 1, 32, 580], and each value can be expressed as [batch size, fixed value 1, frame length, number of high-frequency band points]. The calculation method of the high-frequency band frequency point number 580 can be: 580 = (1024-464) + 20; 1024 can correspond to the 44.1K sampling rate full-band 22.05K spectrum height, and 20 can correspond to the overlap frequency point number. This value is not fixed and can be adjusted to prevent sudden changes in the calculation process. Fig.12 , Fig.12 A schematic diagram showing low-band and high-band frequency points.
[0132] Optionally, the loss functions of the Generator and Discriminator in GAN can be:
[0133]
[0134]
[0135] in, can represent the loss function of the discriminator; E can represent the cross entropy loss function; x can represent an audio sample, D can represent a discriminator, and D(x) can represent placing the audio sample in the discriminator for discrimination; z can represent the input of the audio enhancement model, which is the low-band audio of the audio sample; G can represent a generator, and G(z) can represent predicting and generating the low-band audio input to the audio enhancement model, and the generated result is full-band audio; D(G(z)) can represent placing the generated full-band audio in the discriminator for discrimination; It can represent the loss function of the generator; D (discriminator) can be trained first, and then G (generator) can be trained, and the two can compete with each other until convergence. It should be noted that in addition to using the above loss function, two additional loss functions can be introduced to strengthen the prediction of low-frequency band features to high-frequency band features, as shown below. The total L G The loss function can be:
[0136]
[0137]
[0138]
[0139] Among them, L G It can represent the total loss function, L LSD It can represent the loss function of the least significant difference (LSD), also known as L2 loss; L l1pixcel It can be expressed as L1 loss, also known as L1 norm loss; the above λ 1 and λ 2 can represent weight hyperparameters, which can generally be set to 10 and 0.1 respectively; L can represent time (i.e. the horizontal axis in the spectrum graph), K can represent frequency (i.e. the vertical axis in the spectrum graph), and X HR Can represent high frequency band, X HR (l,k) can represent the location of a specific frequency point in the high frequency band, X SR can represent the generated high frequency band, X SR (l,k) can represent the location of a specific frequency point in the generated high frequency band. It can be seen that L LSD The loss function can be obtained by taking the square root of the difference between the target value (high frequency band) and the model output value (generated high frequency band); L l1pixcel The loss function can be obtained by taking the absolute value of the difference between the target value (high frequency band) and the model output value (generated high frequency band) to obtain the error.
[0140] In the embodiment of the present application, by adopting the GAN training model instead of directly using the general generation model (such as auto encoder) training method, the model can be trained more fully, so as to learn more high-frequency band details from the low-frequency band, so that the high-frequency band can generate more high-frequency details, thereby more realistically restoring the real high-frequency features.
[0141] Optionally, when the terminal device 2 performs prediction generation through the trained audio enhancement model, the modulus and phase of the audio after the audio enhancement can be post-processed. Fig.13As shown, input 44.1K sampling rate audio (such as the above audio sample), then calculate the STFT feature, get the corresponding modulus and phase, take the logarithm of the modulus to get the logarithmic modulus, cut the logarithmic modulus to the spectrum height 10K (such as the fourth audio above), input the trained generator and high-frequency modulus post-processing module, you can get a full-band modulus, the generated full-band modulus and the low-band phase mirror to get the full-band phase, through the Euler formula, and use ISTFT, you can get a 44.1K sampling rate time domain waveform, further sent to the improved Griffinlim algorithm to correct the phase, you can get the time domain waveform after modulus and phase post-processing (that is Fig.13 ed audio after prediction).
[0142] It should be noted that when the trained audio enhancement model is used for prediction generation, the input audio sample can be an ultra-high-quality audio (i.e. Fig.13 After the audio sample is subjected to audio enhancement processing (including modulus and phase optimization processing), the resulting audio can still be the predicted ultra-high quality audio (i.e., audio with a sampling rate of 44.1K). The predicted audio can make up for the high-frequency loss of the audio sample in the frequency domain, and can make the audio sample fluctuate faster in the time domain, thereby making the predicted audio sound louder and more detailed.
[0143] Optionally, when the audio enhancement model is used for prediction and generation, the high-frequency mode post-processing process and phase correction process used can be referred to in the above Figure 2 The detailed description of the high-frequency mode post-processing and phase correction corresponding to S203 in the corresponding embodiment will not be repeated in this application.
[0144] In an embodiment of the present application, by deploying a lightweight algorithm when generating the initial architecture of the audio enhancement model, the amount of calculation and the model size can be reduced; and by training ultra-high-quality audio samples through GAN, an audio enhancement model can be obtained, thereby learning more high-frequency band details from the low-frequency band; and then optimizing the modulus and phase of the audio predicted and generated by the audio enhancement model can make the prediction result more accurate, thereby improving the audio playback effect.
[0145] Based on the above audio processing method, an embodiment of the present invention provides an audio processing device. Fig.14 , is a schematic diagram of the structure of an audio processing device provided by an embodiment of the present invention. The audio processing device 1400 can run the following units:
[0146] The acquisition unit 1401 is used to respond to the audio enhancement start instruction for the target audio and acquire the first audio to be played from the target audio; the target audio is the audio in the playing state;
[0147] The processing unit 1402 is configured to input the first audio into a pre-trained audio enhancement model to obtain a second audio output by the audio enhancement model;
[0148] The processing unit 1402 is further configured to process the mode of the second audio according to the high-frequency band mode of the first audio and to correct the phase of the second audio according to the low-frequency band phase of the first audio, so as to process the second audio into a third audio; the third audio is an ultra-high-quality audio;
[0149] The playing unit 1403 is used to replace the playing target audio with playing the third audio.
[0150] In one embodiment, the audio processing device further includes a determination unit 1404. The processing unit 1402 is further used to perform high-frequency modulus post-processing on the modulus of the second audio based on the high-frequency band modulus of the first audio and the high-frequency band modulus of the second audio to obtain the full-band modulus of the second audio; the processing unit 1402 is further used to perform phase correction on the phase of the second audio according to the low-frequency band phase of the first audio to obtain the full-band phase of the second audio; the determination unit 1404 is used to process the second audio based on the full-band modulus and the full-band phase to determine the third audio.
[0151] In one embodiment, the processing unit 1402 is further used to mirror the low-frequency band phase of the first audio to obtain a mirror phase; the processing unit 1402 is further used to operate the mirror phase using a speech signal reconstruction algorithm to obtain a calculated phase; the processing unit 1402 is further used to perform phase correction on the phase of the second audio according to the calculated phase to obtain a full-band phase of the second audio.
[0152] In one embodiment, the processing unit 1402 is further configured to delete the cached third audio when receiving a drag instruction or a switch instruction for the target audio input.
[0153] In one embodiment, the processing unit 1402 is further configured to delete the cached third audio when the playing of the third audio is completed.
[0154] In one embodiment, the processing unit 1402 is further configured to stop inputting the audio in the target audio into the audio enhancement model upon receiving an audio enhancement shut-down instruction for the target audio input.
[0155] In one embodiment, the audio processing device further includes a communication unit 1405. The communication unit 1405 is used to receive an audio enhancement start instruction for the target audio input on the target audio playback interface.
[0156] In one embodiment, the communication unit 1405 is further configured to receive an audio enhancement start instruction for the target audio input on an audio setting interface corresponding to the target audio.
[0157] In one embodiment, the audio processing device further includes a display unit 1406. The display unit 1406 is used to display the frequency spectrum of the first audio and the third audio on the audio setting interface.
[0158] In one embodiment, the above-mentioned display unit 1406 is also used to display a first spectrum and a second spectrum of a target audio segment currently being processed by the audio enhancement model; the first spectrum is the spectrum of the target audio segment before it is input into the audio enhancement model, and the second spectrum is the spectrum of the target audio segment after it is input into the audio enhancement model, and the target audio segment is an audio segment being processed among multiple audio segments.
[0159] In one embodiment, the audio enhancement model is a model obtained by training audio samples through a generative adversarial network, and the audio samples are ultra-high-quality audio.
[0160] According to one embodiment of the present invention, Figure 2 The steps involved in the audio processing method shown can be Fig.14 The audio processing device shown in the figure is executed by each unit. For example, Figure 2 Step S201 can be performed by Fig.14 The acquisition unit 1401 in the audio processing device 1400 shown in FIG. 1 is used to perform step S202. Fig.14 The processing unit 1402 in the audio processing device 1400 shown in FIG. 1 is used to perform step S204. Fig.14 The playing unit 1403 in the audio processing device 1400 shown is used to execute the above.
[0161] According to another embodiment of the present invention, Fig.14 The various units in the audio processing device shown can be separately or completely combined into one or several other units to form, or one (some) of the units can be further divided into multiple functionally smaller units to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present invention. The above-mentioned units are divided based on logical functions. In practical applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present invention, other units can also be included based on the audio processing device. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.
[0162] According to another embodiment of the present invention, the program can be executed by running a program on a general computing device such as a computer including a central processing unit (CPU), a random access memory medium (RAM), a read-only memory medium (ROM), and other processing elements and storage elements. Figure 2 A computer program (including program code) for each step involved in the corresponding method shown in FIG. Fig.14 The audio processing device shown in the embodiment of the present invention is used to implement the audio processing method of the embodiment of the present invention. The computer program can be recorded on a computer storage medium, for example, and loaded into the above-mentioned computing device through the computer storage medium and run therein.
[0163] To summarize, by receiving an audio enhancement start instruction input for audio in a playing state (such as a target audio), the terminal device can respond to the audio enhancement start instruction, obtain the first audio to be played from the target audio, input the first audio into a pre-trained audio enhancement model, obtain the audio output by the audio enhancement model (such as the second audio), and determine an ultra-high-quality third audio based on the first audio and the second audio to play the third audio, so that the terminal device can perform audio enhancement processing on the audio in a real-time playback state, and thereby obtain ultra-high-quality audio to enhance the user's auditory experience.
[0164] Based on the above-mentioned embodiments of the audio processing method and the audio processing device, an embodiment of the present invention further provides an electronic device, which may correspond to the above-mentioned terminal device. Please refer to Figure 15, which is a structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 1500 may at least include: a processor 1501, an input interface 1502, an output interface 1503, and a computer storage medium 1504, which may be connected via a bus or other means.
[0165] The computer storage medium 1504 may be stored in the memory 1505 of the electronic device 1500. The computer storage medium 1504 is used to store a computer program, the computer program includes program instructions, and the processor 1501 is used to execute the program instructions stored in the computer storage medium 1504. The processor 1501 (or CPU (Central Processing Unit)) is the computing core and control core of the electronic device, which is suitable for implementing one or more instructions, specifically suitable for loading and executing:
[0166] In response to an audio enhancement start instruction for a target audio, a first audio to be played is obtained from the target audio; the target audio is the audio in a playing state; the first audio is input into a pre-trained audio enhancement model to obtain a second audio output by the audio enhancement model; the mode of the second audio is processed according to the high-frequency band mode of the first audio and the phase of the second audio is corrected according to the low-frequency band phase of the first audio to process the second audio into a third audio; the third audio is an ultra-high-quality audio; and the playing of the target audio is replaced by the playing of the third audio.
[0167] In one embodiment, the processor 1501 is further used to perform high-frequency module post-processing on the module of the second audio based on the high-frequency module of the first audio and the high-frequency module of the second audio to obtain a full-band module of the second audio; the processor 1501 is further used to perform phase correction on the phase of the second audio according to the low-frequency band phase of the first audio to obtain the full-band phase of the second audio; the processor 1501 is further used to process the second audio using the full-band module and the full-band phase to determine the third audio.
[0168] In one embodiment, the processor 1501 is further used to mirror the low-frequency band phase of the first audio to obtain a mirror phase; the processor 1501 is further used to operate the mirror phase using a speech signal reconstruction algorithm to obtain a calculated phase; the processor 1501 is further used to perform phase correction on the phase of the second audio according to the calculated phase to obtain a full-band phase of the second audio.
[0169] In one embodiment, the processor 1501 deletes the cached third audio when receiving a drag instruction or a switch instruction for the target audio input.
[0170] In one embodiment, the processor 1501 deletes the cached third audio when the playing of the third audio is completed.
[0171] In one embodiment, upon receiving an audio enhancement off instruction for a target audio input, the processor 1501 stops inputting audio in the target audio into the audio enhancement model.
[0172] In one embodiment, the processor 1501 receives an audio enhancement start instruction for a target audio input on a playback interface of the target audio.
[0173] In one embodiment, the processor 1501 receives an audio enhancement activation instruction for the target audio input on an audio setting interface corresponding to the target audio.
[0174] In one embodiment, the processor 1501 displays the frequency spectra of the first audio and the third audio on the audio setting interface.
[0175] In one embodiment, the processor 1501 displays a first spectrum and a second spectrum of a target audio segment currently being processed by an audio enhancement model; the first spectrum is a spectrum of the target audio segment before the target audio segment is input into the audio enhancement model, and the second spectrum is a spectrum of the target audio segment after the target audio segment is input into the audio enhancement model, and the target audio segment is an audio segment being processed among multiple audio segments.
[0176] In one embodiment, the audio enhancement model is a model obtained by performing phase correction on a generative adversarial network model, and the generative adversarial network model is a model obtained by training an audio sample through a generative adversarial network, and the audio sample is ultra-high-quality audio.
[0177] In summary, the electronic device receives an audio enhancement start instruction for a target audio input, where the target audio is an audio in a playing state; and responds to the audio enhancement start instruction to obtain a first audio to be played from the target audio; inputs the first audio into a pre-trained audio enhancement model to obtain a second audio output by the audio enhancement model, and determines a third audio of ultra-high quality audio based on the first audio and the second audio; thereby playing the third audio. It should be understood that the electronic device can obtain an ultra-high quality third audio by inputting the first audio into the audio enhancement model, thereby improving the user experience.
[0178] In the above embodiments, the description of each embodiment has its own emphasis. For the part that is not described in detail in a certain embodiment, please refer to the relevant description of other embodiments. The technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server or a network device, etc., specifically a processor in a computer device) to perform all or part of the steps of the above methods of each embodiment of the present application. Among them, the aforementioned storage medium may include: U disk, mobile hard disk, magnetic disk, optical disk, read-only memory (English: Read-Only Memory, abbreviated: ROM) or random access memory (English: Random Access Memory, abbreviated: RAM) and other media that can store program codes.
[0179] Those of ordinary skill in the art will appreciate that the units and steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0180] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer storage medium or transmitted through a computer storage medium. The computer instructions can be transmitted from a website site, a computer, a server or a data center to another website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer storage medium can be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integration. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.
[0181] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. An audio processing method, It is characterized in that The method is applied to a terminal device, and the method comprises: In response to an audio enhancement start instruction for a target audio, obtaining a first audio to be played from the target audio; the target audio is an audio in a playing state; Inputting the first audio into a pre-trained audio enhancement model to obtain a second audio output by the audio enhancement model; Based on the high-frequency band mode of the first audio and the high-frequency band mode of the second audio, performing high-frequency mode post-processing on the mode of the second audio to obtain a full-band mode of the second audio; According to the low-frequency band phase of the first audio, the phase of the second audio is corrected to obtain the full-frequency band phase of the second audio; Based on the full-band mode and the full-band phase, the second audio is processed into a third audio; the third audio is ultra-high quality audio; The playing of the target audio is replaced by the playing of the third audio.
2. The method according to claim 1, It is characterized in that The method of performing phase correction on the phase of the second audio according to the low-frequency band phase of the first audio to obtain the full-frequency band phase of the second audio includes: Performing mirror processing on the low-frequency band phase of the first audio to obtain a mirror phase; Using a speech signal reconstruction algorithm to calculate the mirror phase to obtain a calculated phase; The phase of the second audio is corrected according to the calculated phase to obtain the full-band phase of the second audio.
3. The method according to claim 1, It is characterized in that The audio enhancement model is a model obtained by training audio samples through a generative adversarial network, and the audio samples are ultra-high-quality audio.
4. The method according to any one of claims 1 to 3, It is characterized in that The method further comprises: When a drag instruction or a switch instruction for the target audio input is received or when the third audio is finished playing, the cached third audio is deleted.
5. The method according to any one of claims 1 to 3, It is characterized in that Before responding to the audio enhancement start instruction for the target audio, receiving the audio enhancement start instruction for the target audio input includes: On the playback interface of the target audio or the audio setting interface corresponding to the target audio, an audio enhancement start instruction for the target audio input is received.
6. The method according to claim 5, It is characterized in that The method further comprises: On the audio setting interface, frequency spectra of the first audio and the third audio are displayed.
7. The method according to claim 6, It is characterized in that The first audio includes a plurality of audio segments, the second audio includes audio segments after each of the audio segments is enhanced, and the audio enhancement model processes each of the audio segments in sequence to obtain audio segments after audio enhancement; The displaying of the frequency spectra of the first audio and the third audio comprises: For a target audio segment currently being processed by the audio enhancement model, a first spectrum and a second spectrum of the target audio segment are displayed; wherein the first spectrum is the spectrum of the target audio segment, and the second spectrum is the spectrum of the enhanced audio segment corresponding to the target audio segment.
8. An electronic device, It is characterized in that The method comprises a processor and a memory, wherein the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 7.
10. A computer program product, It is characterized in that The computer program product comprises a computer program, which is stored in a computer storage medium and is suitable for being read by a processor of a computer device and executing the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and apparatus for improving sound quality of speaker
US20240135946A1