Method, apparatus, device and readable storage medium for improving audio generation quality
By combining discrete wavelet transform with time-domain and frequency-domain information in an audio processing method, the problem of inconsistent audio sampling rates across different channels was solved, resulting in higher-quality audio generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA CITIC BANK CO LTD
- Filing Date
- 2022-12-05
- Publication Date
- 2026-05-05
AI Technical Summary
Existing audio generation methods fail to fully utilize the time and frequency domain information of audio in voiceprint recognition, resulting in poor audio generation quality, especially when the audio sampling rates collected from different channels are inconsistent.
By employing discrete wavelet transform combined with time and frequency domain information of audio signals, and through audio preprocessing, signal reconstruction, and fusion models, high-sampling-rate audio is reconstructed to improve the generation quality.
By reconstructing information from both the time and frequency domains, the quality of audio generation is significantly improved, the problem of inconsistent sampling rates of audio collected from different channels is solved, and the generated high-sampling-rate audio is closer to real audio.
Smart Images

Figure CN116013317B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech conversion and recognition technology, and more specifically, to methods, apparatus, devices, and readable storage media for improving the quality of audio generation. Background Technology
[0002] With the development of artificial intelligence technology, voiceprint recognition technology has been widely applied. In the banking industry, voiceprint recognition technology can not only verify user identity but also support the identification of fraudulent applications. In voiceprint recognition applications, audio collected from different channels has different sampling rates. For example, the sampling rate of audio collected from telephone channels is 8kHz, while the sampling rate of audio collected from network channels is 16kHz. To achieve better results in voiceprint recognition models, super-resolution reconstruction methods can be used to reconstruct low-sampling-rate signals into high-sampling-rate signals. Current methods for improving audio generation quality typically use Short-Time Fourier Transform (SFT) to process audio. However, the window length of the SFT is fixed, which can only capture details of the audio at a certain scale and only uses one of the time-domain or frequency-domain information, resulting in insufficient utilization of audio information. Summary of the Invention
[0003] The purpose of this invention is to provide a method, apparatus, device, and readable storage medium for improving audio generation quality, thereby addressing the aforementioned problems. To achieve this objective, the technical solution adopted by this invention is as follows:
[0004] In a first aspect, this application provides a method for improving audio generation quality, comprising: acquiring low-sampling-rate audio, a target audio sampling rate, and an audio processing model, wherein the audio processing method includes an audio preprocessing mathematical model and an audio signal reconstruction mathematical model; calculating an initial high-sampling-rate audio based on the low-sampling-rate audio, the target audio sampling rate, and the audio preprocessing mathematical model; calculating a target audio time-domain signal and target audio wavelet coefficients based on the initial high-sampling-rate audio and the audio signal reconstruction mathematical model; and solving the mathematical model based on the target audio time-domain signal, the target audio wavelet coefficients, and a preset fusion audio signal mathematical model to obtain the target high-sampling-rate audio.
[0005] Secondly, this application also provides an apparatus for improving the quality of generated audio, comprising: a data acquisition module for acquiring low-sampling-rate audio, a target audio sampling rate, and an audio processing model, wherein the audio processing method includes an audio preprocessing mathematical model and an audio signal reconstruction mathematical model; an audio processing module for calculating an initial high-sampling-rate audio based on the low-sampling-rate audio, the target audio sampling rate, and the audio preprocessing mathematical model; an audio analysis module for calculating a target audio time-domain signal and target audio wavelet coefficients based on the initial high-sampling-rate audio and the audio signal reconstruction mathematical model; and an audio reconstruction module for solving the mathematical model to obtain the target high-sampling-rate audio based on the target audio time-domain signal, the target audio wavelet coefficients, and a preset fused audio signal mathematical model.
[0006] Thirdly, this application also provides an apparatus for improving the quality of audio generation, comprising:
[0007] Memory, used to store computer programs;
[0008] A processor for implementing the steps of the method for improving audio generation quality when executing the computer program.
[0009] Fourthly, this application also provides a readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method for improving audio generation quality described above.
[0010] The beneficial effects of this invention are as follows:
[0011] This invention uses discrete wavelet transform instead of short-time Fourier transform to capture multi-scale details of audio signals; this invention reconstructs high-sampling-rate audio by combining the time-domain and frequency-domain information of the audio signal to obtain better high-sampling-rate audio signals, further improving the overall generation quality of audio.
[0012] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a schematic diagram of the method for improving audio generation quality as described in an embodiment of the present invention;
[0015] Figure 2 This is a schematic diagram of the device structure for improving audio generation quality as described in an embodiment of the present invention;
[0016] Figure 3 This is a schematic diagram of the device structure for improving audio generation quality as described in an embodiment of the present invention.
[0017] The diagram is labeled as follows: 1. Data acquisition module; 2. Audio processing module; 21. First calculation unit; 211. First division unit; 212. Second classification unit; 213. Third calculation unit; 22. Second calculation unit; 221. First extraction unit; 222. Fourth calculation unit; 223. Fifth calculation unit; 3. Audio analysis module; 31. Sixth calculation unit; 32. Seventh calculation unit; 33. Eighth calculation unit; 4. Audio reconstruction module; 41. Ninth calculation unit; 42. Tenth calculation unit; 43. Eleventh calculation unit; 800. Lossless compression device for monitoring data; 801. Processor; 802. Memory; 803. Multimedia component; 804. I / O interface; 805. Communication component. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0019] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0020] Example 1:
[0021] This embodiment provides a method for improving the quality of audio generation.
[0022] See Figure 1 The figure shows that the method includes steps S100, S200, S300 and S400.
[0023] S100: Acquire low-sampling-rate audio, target audio sampling rate, and audio processing model. The audio processing methods include an audio preprocessing mathematical model and an audio signal reconstruction mathematical model.
[0024] It should be noted that in this application, the low sampling rate is equal to 8000 Hz, and the target audio sampling rate and high sampling rate are equal to 16000 Hz.
[0025] S200. Based on the low sampling rate audio, the target audio sampling rate, and the audio preprocessing mathematical model, the initial high sampling rate audio is calculated.
[0026] It should be noted that in step S200, the low sampling rate audio is processed by silence removal and interpolation to obtain the initial high sampling rate audio. The duration of the initial high sampling rate audio is the same as that of the low sampling rate audio, and the sampling rate of the initial high sampling rate audio is the same as that of the target audio (e.g., 16,000 sampling points per second). At this time, the initial high sampling rate audio obtained has the problem of over-smoothing and cannot be directly used as the final result to provide a basis for subsequent audio processing.
[0027] S300. Based on the initial high-sampling-rate audio and the audio signal, a mathematical model is reconstructed to calculate the target audio time-domain signal and the target audio wavelet coefficients.
[0028] It should be noted that in step S300, the time-domain signal of the initial high-sampling-rate audio is extracted, and the time-domain signal of the target audio is obtained by calculating the detailed information of the supplementary audio time-domain signal. At the same time, the wavelet coefficients of the target audio are calculated from the time-domain signal of the target audio.
[0029] S400. Based on the target audio time-domain signal, the target audio wavelet coefficients, and the preset mathematical model of the fused audio signal, solve the mathematical model to obtain the target high sampling rate audio.
[0030] It should be noted that in step S400, the target audio time-domain signal and the target audio wavelet coefficients are weighted and combined to obtain the target high sampling rate audio. This invention reconstructs high sampling rate audio by combining the time-domain information and frequency-domain information (wavelet coefficients) of the audio signal to obtain a better high sampling rate audio signal, thereby further improving the overall generation quality of the audio.
[0031] In the specific embodiments disclosed in this application, step S200 includes steps S210 and S220.
[0032] S210. Obtain low-sampling-rate audio and a preset audio silence signal removal calculation formula, calculate low-sampling-rate speech audio, and the low-sampling-rate speech audio is the audio data of the sampled-rate audio with silence segments removed.
[0033] It should be noted that the low-sampling-rate audio obtained in this step includes both silent and audio portions. The main goal of audio processing is to increase the sampling rate of the audio portion. Removing the silent portion in this step can significantly improve the efficiency of subsequent audio processing.
[0034] Specifically, step S210 includes steps S211, S212 and S213.
[0035] S211. Based on the low sampling rate audio and the preset audio segmentation rules, divide the low sampling rate audio into at least one low sampling rate audio segment.
[0036] S212. Based on the low sampling rate audio segments and the preset audio judgment model, the low sampling rate audio segments are divided into speech segments and silence segments.
[0037] S213. Calculate the low-sampling-rate speech audio based on the speech segment and the preset audio combination method.
[0038] In this method, the audio signal is first divided into 10 segments, and the spectrum of each segment is further divided into 6 sub-bands (80Hz-250Hz, 250Hz-500Hz, 500Hz-1kHz, 1kHz-2kHz, 2kHz-3kHz, 3kHz-4kHz). The energy characteristics of each sub-band are calculated. Then, a Gaussian mixture model is used to calculate the probability that each sub-band is silence or speech, and its log-likelihood ratio is calculated. If the log-likelihood ratio of any of the 6 sub-bands exceeds a threshold, the segment is considered a silence segment and is removed. Finally, the remaining audio segments are combined together. The silence removal technique is common knowledge in the field and can be implemented using existing algorithms such as VAD silence removal, so it will not be elaborated upon in this application.
[0039] S220. Based on the low sampling rate audio, the target audio sampling rate, the low sampling rate speech audio, and the preset audio interpolation calculation formula, the initial high sampling rate audio is calculated. The length of the initial high sampling rate audio is equal to the length of the low sampling rate audio, and the sampling rate of the initial high sampling rate audio is equal to the target audio sampling rate.
[0040] In this method, low-sampling-rate audio is processed by interpolation, and the number of sampling points of the audio signal is increased according to the target audio sampling rate to facilitate calculations in subsequent steps.
[0041] Specifically, step S220 includes steps S221, S222, and S223.
[0042] S221. Extract the duration of the low-sampling-rate audio based on the low-sampling-rate audio.
[0043] S222. Calculate the high-sampling-rate speech audio based on the low-sampling-rate speech audio and the preset audio interpolation calculation formula.
[0044] S223. Calculate the initial high-sampling-rate audio based on the duration of the high-sampling-rate speech audio, the low-sampling-rate audio, and the preset audio extension model.
[0045] In this method, two polynomial cubic interpolation functions are used to interpolate the low-sampling-rate audio signal to increase the number of sampling points in the audio signal. This ensures that the length of the source audio and the target high-sampling-rate audio are consistent within the same time frame (e.g., 16,000 sampling points per second), resulting in an initial high-sampling-rate audio (note that this high-sampling-rate audio suffers from over-smoothing and cannot be directly used as the final result). This facilitates subsequent audio fusion calculations. This audio interpolation technique is common knowledge in the field and can be implemented using existing algorithms such as the Bicubic interpolation method, and will not be elaborated upon in this application.
[0046] In the specific embodiments disclosed in this application, step S300 includes steps S310, S320 and S330.
[0047] S310. Calculate the target audio time-domain signal based on the initial high-sampling-rate audio and the preset audio time-domain reconstruction mathematical model.
[0048] S320. Perform discrete wavelet transform on the target audio time-domain signal to obtain the initial wavelet coefficients.
[0049] S330. Reconstruct the mathematical model based on the initial wavelet coefficients and the preset audio wavelet coefficients, and calculate the target audio wavelet coefficients.
[0050] In this method, the initial high-sampling-rate audio time-domain signal is extracted, and the detailed information of the audio time-domain signal is supplemented by an algorithm. Then, the target audio time-domain signal is transformed and the target audio wavelet coefficients are calculated. The techniques for supplementing the details of the audio time-domain signal and converting the initial wavelet coefficients into the target audio wavelet coefficients are common knowledge in the field and can be implemented using existing algorithms such as AudioUNet and MS-Net. Therefore, they will not be elaborated further in this application.
[0051] In the specific embodiments disclosed in this application, step S400 includes steps S410, S420, and S430.
[0052] S410. Perform discrete wavelet transform on the target audio time-domain signal to obtain the initial wavelet coefficients.
[0053] S420. Calculate the final wavelet coefficients of the target audio based on the initial wavelet coefficients, the target audio wavelet coefficients, and the preset weights.
[0054] S430. Perform discrete wavelet inverse transform on the final wavelet coefficients of the target audio to calculate the target high sampling rate audio.
[0055] It should be noted that in this application, a discrete wavelet transform is performed on the target audio time-domain signal, and then the result is combined with the target audio wavelet coefficients by applying weights to obtain the final wavelet coefficients of the target audio. Finally, a discrete wavelet inverse transform is performed. The specific calculation formula for obtaining the final target high-sampling-rate audio is as follows:
[0056]
[0057]
[0058]
[0059]
[0060] Where D represents the initial high-sampling-rate audio, xi represents the i-th low-sampling-rate audio input, yi represents the actual high-sampling-rate audio corresponding to xi, and n is the number of audio files. For the target audio time domain signal, Let represent the target audio wavelet coefficients, θ and w represent the parameters to be learned, fθ represents the mapping relationship learned by AuduNet and MS-Net, M represents the final wavelet coefficients of the target audio, ⊙ represents the element-wise dot product, DWT represents the discrete wavelet transform, and I DWT represents the inverse discrete wavelet transform.
[0061] The loss function for the mathematical model of fused audio signals is as follows:
[0062]
[0063] in, To obtain the target high-sample-rate audio from the true high-sample-rate audio and the reconstructed high-sample-rate audio. The L1 norm of the difference, ||θ||1 is the L1 norm of θ, and λ is the regularization parameter of θ.
[0064] Then, backpropagation is used to update the model parameters until the loss function converges, at which point the difference between the generated target high-sampling-rate audio signal and the real high-sampling-rate audio signal is minimized. This application further optimizes the high-sampling-rate audio obtained by high-frequency reconstruction from low-sampling-rate audio, thereby greatly enhancing the overall quality of the generated audio.
[0065] Example 2:
[0066] like Figure 2 As shown, this embodiment provides an apparatus for improving audio generation quality, the apparatus including...
[0067] Data acquisition module 1 is used to acquire low sampling rate audio, target audio sampling rate, and audio processing model. The audio processing method includes an audio preprocessing mathematical model and an audio signal reconstruction mathematical model.
[0068] Audio processing module 2 is used to calculate the initial high sampling rate audio based on the low sampling rate audio, the target audio sampling rate, and the audio preprocessing mathematical model;
[0069] Audio analysis module 3 is used to reconstruct a mathematical model based on the initial high-sampling-rate audio and audio signal, and calculate the target audio time-domain signal and target audio wavelet coefficients.
[0070] Audio reconstruction module 4 is used to solve the mathematical model based on the target audio time domain signal, the target audio wavelet coefficients and the preset fused audio signal mathematical model to obtain the target high sampling rate audio.
[0071] In some specific embodiments, the audio processing module 2 includes:
[0072] The first calculation unit 21 is used to acquire low sampling rate audio and a preset audio silence signal removal calculation formula, and calculate low sampling rate speech audio. The low sampling rate speech audio is the audio data of the sampling rate audio with silence segments deleted.
[0073] The second calculation unit 22 is used to calculate the initial high sampling rate audio based on the low sampling rate audio, the target audio sampling rate, the low sampling rate speech audio, and the preset audio interpolation calculation formula. The length of the initial high sampling rate audio is equal to the length of the low sampling rate audio, and the sampling rate of the initial high sampling rate audio is equal to the target audio sampling rate.
[0074] In some specific embodiments, the first computing unit 21 includes:
[0075] The first segmentation unit 211 is used to divide the low-sampling-rate audio into at least one low-sampling-rate audio segment according to the low-sampling-rate audio and the preset audio segmentation rules.
[0076] The first classification unit 212 divides low-sampling-rate audio segments into speech segments and silence segments based on low-sampling-rate audio segments and a preset audio judgment model.
[0077] The third calculation unit 213 is used to calculate low-sampling-rate speech audio based on speech segments and a preset audio combination method.
[0078] In some specific embodiments, the second computing unit 22 includes:
[0079] The first extraction unit 221 is used to extract the duration of low-sampling-rate audio based on the low-sampling-rate audio.
[0080] The fourth calculation unit 222 is used to calculate the high sampling rate speech audio based on the low sampling rate speech audio and the preset audio interpolation calculation formula;
[0081] The fifth calculation unit 223 is used to calculate the initial high-sampling-rate audio based on the high-sampling-rate speech audio, the duration of the low-sampling-rate audio, and a preset audio extension model.
[0082] In some specific embodiments, the audio analysis module 3 includes:
[0083] The sixth calculation unit 31 is used to calculate the target audio time domain signal based on the initial high sampling rate audio and the preset audio time domain reconstruction mathematical model.
[0084] The seventh calculation unit 32 is used to perform discrete wavelet transform on the target audio time-domain signal to obtain the initial wavelet coefficients.
[0085] The eighth calculation unit 33 is used to reconstruct the mathematical model based on the initial wavelet coefficients and the preset audio wavelet coefficients, and calculate the target audio wavelet coefficients.
[0086] In some specific embodiments, the audio reconstruction module 4 includes:
[0087] The ninth calculation unit 41 is used to perform discrete wavelet transform on the target audio time-domain signal to obtain the initial wavelet coefficients;
[0088] The tenth calculation unit 42 is used to calculate the final wavelet coefficients of the target audio based on the initial wavelet coefficients, the target audio wavelet coefficients, and the preset weights.
[0089] The eleventh calculation unit 43 is used to perform discrete wavelet inverse transform on the final wavelet coefficients of the target audio to calculate the target high sampling rate audio.
[0090] It should be noted that the specific manner in which each module performs its operation in the apparatus described in the above embodiments has been described in detail in the embodiments of the method, and will not be elaborated here.
[0091] Example 3:
[0092] Corresponding to the above method embodiments, this embodiment also provides a device for improving audio generation quality. The device for improving audio generation quality described below and the method for improving audio generation quality described above can be referred to in correspondence.
[0093] Figure 3 This is a block diagram illustrating a device 800 for improving audio generation quality according to an exemplary embodiment. Figure 3 As shown, the device 800 for improving audio generation quality may include a processor 801 and a memory 802. The device 800 may also include one or more of a multimedia component 803, an I / O interface 804, and a communication component 805.
[0094] The processor 801 controls the overall operation of the audio quality improvement device 800 to complete all or part of the steps in the above-described method for improving audio quality. The memory 802 stores various types of data to support the operation of the audio quality improvement device 800. This data may include, for example, instructions for any application or method operating on the audio quality improvement device 800, and application-related data such as contact data, sent and received messages, images, audio, video, etc. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 803 may include a screen and an audio component. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in the memory 802 or transmitted via the communication component 805. The audio component also includes at least one speaker for outputting audio signals. I / O interface 804 provides an interface between processor 801 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons can be virtual or physical buttons. Communication component 805 is used for wired or wireless communication between the audio quality enhancement device 800 and other devices. Wireless communication includes, for example, Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof; therefore, the corresponding communication component 805 may include a Wi-Fi module, a Bluetooth module, or an NFC module.
[0095] In an exemplary embodiment, the device 800 for improving audio generation quality may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method for improving audio generation quality.
[0096] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the method for improving audio generation quality described above. For example, the computer-readable storage medium may be the memory 802 including the program instructions described above, which may be executed by the processor 801 of the device 800 for improving audio generation quality to perform the method for improving audio generation quality described above.
[0097] Example 4:
[0098] Corresponding to the above method embodiments, this embodiment also provides a readable storage medium. The readable storage medium described below corresponds to and can be referred to in relation to the method for improving audio generation quality described above.
[0099] A readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the method for improving audio generation quality described in the above method embodiments.
[0100] Specifically, the readable storage medium can be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or any other readable storage medium capable of storing program code.
[0101] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0102] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for improving audio generation quality, characterized in that, include: The low-sampling-rate audio, the target audio sampling rate, and the audio processing mathematical model are obtained. The audio processing mathematical model includes an audio preprocessing mathematical model and an audio signal reconstruction mathematical model. The initial high-sampling-rate audio is calculated based on the low-sampling-rate audio, the target audio sampling rate, and the audio preprocessing mathematical model. Based on the initial high-sampling-rate audio and the audio signal reconstruction mathematical model, the target audio time-domain signal and the target audio wavelet coefficients are calculated. Based on the target audio time-domain signal, the target audio wavelet coefficients, and a preset mathematical model of the fused audio signal, the mathematical model is solved to obtain the target high-sampling-rate audio. The step of obtaining the target high-sampling-rate audio by solving the mathematical model based on the target audio time-domain signal, the target audio wavelet coefficients, and a preset fused audio signal mathematical model includes: The initial wavelet coefficients are obtained by performing a discrete wavelet transform on the target audio time-domain signal. The final wavelet coefficients of the target audio are calculated based on the initial wavelet coefficients, the target audio wavelet coefficients, and the preset weights. The target high-sampling-rate audio is obtained by performing an inverse discrete wavelet transform on the final wavelet coefficients of the target audio.
2. The method for improving audio generation quality according to claim 1, characterized in that... The step of calculating the initial high-sampling-rate audio based on the low-sampling-rate audio, the target audio sampling rate, and the audio preprocessing mathematical model includes: Based on the low sampling rate audio and the preset audio silence signal removal calculation formula, a low sampling rate speech audio is obtained, wherein the low sampling rate speech audio is the audio data of the sampling rate audio with silence segments removed; An initial high-sampling-rate audio is calculated based on the low-sampling-rate audio, the target audio sampling rate, the low-sampling-rate speech audio, and a preset audio interpolation formula. The length of the initial high-sampling-rate audio is equal to the length of the low-sampling-rate audio, and the sampling rate of the initial high-sampling-rate audio is equal to the target audio sampling rate.
3. The method for improving audio generation quality according to claim 2, characterized in that... The step of calculating the low-sampling-rate speech audio based on the low-sampling-rate audio and the preset audio silence signal removal calculation formula includes: Based on the low sampling rate audio and the preset audio segmentation rules, the low sampling rate audio is divided into at least one low sampling rate audio segment; Based on the low-sampling-rate audio segments and the preset audio judgment model, the low-sampling-rate audio segments are divided into speech segments and silence segments; The low-sampling-rate audio is calculated based on the speech segment and the preset audio combination method.
4. The method for improving audio generation quality according to claim 2, characterized in that... The step of calculating the initial high-sampling-rate audio based on the low-sampling-rate audio, the target audio sampling rate, the low-sampling-rate speech audio, and a preset audio interpolation formula includes: Based on the low-sampling-rate audio, the duration of the low-sampling-rate audio is extracted; The high-sampling-rate speech audio is calculated based on the low-sampling-rate speech audio and the preset audio interpolation calculation formula; The initial high-sampling-rate audio is calculated based on the high-sampling-rate speech audio, the duration of the low-sampling-rate audio, and the preset audio expansion model.
5. The method for improving audio generation quality according to claim 1, characterized in that... The step of reconstructing the mathematical model based on the initial high-sampling-rate audio and audio signal to calculate the target audio time-domain signal and target audio wavelet coefficients includes: The target audio time-domain signal is calculated based on the initial high-sampling-rate audio and the preset audio time-domain reconstruction mathematical model. The initial wavelet coefficients are obtained by performing a discrete wavelet transform on the target audio time-domain signal. The mathematical model is reconstructed based on the initial wavelet coefficients and the preset audio wavelet coefficients, and the target audio wavelet coefficients are calculated.
6. An apparatus for improving the quality of audio generation, characterized in that, include: The data acquisition module is used to acquire low-sampling-rate audio, target audio sampling rate, and audio processing mathematical model, wherein the audio processing mathematical model includes an audio preprocessing mathematical model and an audio signal reconstruction mathematical model; An audio processing module is used to calculate an initial high-sampling-rate audio based on the low-sampling-rate audio, the target audio sampling rate, and the audio preprocessing mathematical model. The audio analysis module is used to reconstruct a mathematical model based on the initial high-sampling-rate audio and the audio signal, and calculate the target audio time-domain signal and the target audio wavelet coefficients. The audio reconstruction module is used to solve the mathematical model based on the target audio time-domain signal, the target audio wavelet coefficients, and a preset fused audio signal mathematical model to obtain the target high sampling rate audio. The audio reconstruction module includes: The ninth calculation unit is used to perform discrete wavelet transform on the target audio time-domain signal to obtain initial wavelet coefficients; The tenth calculation unit is used to calculate the final wavelet coefficients of the target audio based on the initial wavelet coefficients, the target audio wavelet coefficients, and the preset weights. The eleventh calculation unit is used to perform discrete wavelet inverse transform on the final wavelet coefficients of the target audio to calculate the target high sampling rate audio.
7. The apparatus for improving audio generation quality according to claim 6, characterized in that, The audio processing module includes: The first calculation unit is used to obtain the low sampling rate audio and a preset audio silence signal removal calculation formula, and calculate the low sampling rate speech audio, wherein the low sampling rate speech audio is the audio data of the sampling rate audio with the silence segment deleted; The second calculation unit is used to calculate an initial high sampling rate audio based on the low sampling rate audio, the target audio sampling rate, the low sampling rate speech audio, and a preset audio interpolation calculation formula. The length of the initial high sampling rate audio is equal to the length of the low sampling rate audio, and the sampling rate of the initial high sampling rate audio is equal to the target audio sampling rate.
8. The apparatus for improving audio generation quality according to claim 7, characterized in that, The first computing unit includes: The first segmentation unit is used to divide the low-sampling-rate audio into at least one low-sampling-rate audio segment according to the low-sampling-rate audio and a preset audio segmentation rule. The first classification unit divides the low-sampling-rate audio segments into speech segments and silence segments based on the low-sampling-rate audio segments and the preset audio judgment model. The third calculation unit is used to calculate the low-sampling-rate speech audio based on the speech segment and a preset audio combination method.
9. The apparatus for improving audio generation quality according to claim 7, characterized in that, The second computing unit includes: The first extraction unit is used to extract the duration of the low sampling rate audio based on the low sampling rate audio. The fourth calculation unit is used to calculate the high sampling rate speech audio based on the low sampling rate speech audio and the preset audio interpolation calculation formula; The fifth calculation unit is used to calculate the initial high-sampling-rate audio based on the high-sampling-rate speech audio, the duration of the low-sampling-rate audio, and a preset audio extension model.
10. The apparatus for improving audio generation quality according to claim 6, characterized in that, The audio analysis module includes: The sixth calculation unit is used to calculate the target audio time-domain signal based on the initial high sampling rate audio and the preset audio time-domain reconstruction mathematical model; The seventh calculation unit is used to perform discrete wavelet transform on the target audio time-domain signal to obtain initial wavelet coefficients; The eighth calculation unit is used to reconstruct a mathematical model based on the initial wavelet coefficients and the preset audio wavelet coefficients, and calculate the target audio wavelet coefficients.
11. A device for improving audio generation quality, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the method for improving audio generation quality as described in any one of claims 1 to 5 when executing the computer program.
12. A readable storage medium, characterized in that: The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method for improving audio generation quality as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Method and system for dually recognizing semantics and vocal prints based on waveform time-frequency domain analysis
CN110349593A
High-frequency optimization method and device for audios and medium
CN112562703A