Methods, devices, equipment and media for optimizing human voice quality
By employing a multi-level optimization method for Fourier transform and harmonic audio data, combined with joint coding of reference audio data, the problem of poor human voice quality optimization in existing technologies has been solved, achieving higher speech clarity and timbre preservation.
Patent Information
- Application Number
- CN202510686734.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-05-27
AI Technical Summary
Existing voice quality optimization methods are prone to introducing distortion when processing complex speech content and have difficulty in effectively recovering speech feature information, resulting in poor optimization results.
The original audio spectrum data is generated through Fourier transform, the target harmonic frequency is determined and the harmonic audio data is extracted, and joint semantic coding is performed in combination with reference audio data to form joint audio coding features. Based on the harmonic audio coding features, multi-level optimization is performed, and finally semantic decoding is performed to form optimized audio data.
While maintaining the original timbre, the sound quality has been significantly improved, solving the problem of poor human voice quality optimization in existing technologies and achieving higher speech clarity and timbre preservation.
Smart Images

Figure CN120375838B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio data processing technology, and more specifically, to a method, apparatus, device, and medium for optimizing human voice quality. Background Technology
[0002] With the continuous development of audio technology, voice quality optimization has received widespread attention in the field of audio data processing technology. The goal of voice quality optimization is to improve the speech quality in raw audio data through technical means, making it clearer, more natural, and more expressive; alternatively, it can be used to optimize singing, making it more pleasant to listen to. However, existing voice quality optimization methods still have many technical bottlenecks, resulting in optimization effects that are difficult to meet practical needs.
[0003] In existing technologies, voice quality optimization is mainly achieved through the following methods:
[0004] Frequency domain processing method: By analyzing the frequency domain of the audio signal, speech features are extracted and enhanced or repaired. However, this method is prone to introducing distortion when processing complex speech content, especially in the high-frequency part, which may lead to inaccurate speech pitch and affect the overall sound quality.
[0005] Time-domain processing methods: By analyzing and processing the time series of audio signals, the clarity and fluency of speech can be improved. However, this method has limited effectiveness in dealing with nonlinear distortion and is difficult to effectively recover the characteristic information of speech.
[0006] In other words, all existing voice quality optimization technologies suffer from poor performance. Summary of the Invention
[0007] In view of this, the purpose of this application is to provide a method, apparatus, device and medium for optimizing human voice quality, so as to improve the problem of poor human voice quality optimization in the prior art.
[0008] To achieve the above objectives, this application adopts the following technical solution:
[0009] A method for optimizing human voice quality includes:
[0010] Perform a Fourier transform on the raw audio data to generate the raw audio spectrum data;
[0011] The target harmonic frequency is determined from the original audio spectrum data, and harmonic audio data is extracted from the original audio data based on the target harmonic frequency.
[0012] Joint semantic coding is performed on the reference audio data and the original audio data to form joint audio coding features, wherein the reference audio data and the original audio data have the same sound content, and the sound quality of the reference audio data is higher than that of the original audio data;
[0013] The harmonic audio data is semantically encoded to form harmonic audio coding features;
[0014] Based on the harmonic audio coding features, the joint audio coding features are optimized to form optimized audio coding features;
[0015] The optimized audio coding features are semantically decoded to form optimized audio data.
[0016] In a preferred embodiment of this application, in the above-described method for optimizing voice quality, the step of optimizing the joint audio coding features based on the harmonic audio coding features to form optimized audio coding features includes:
[0017] Based on the harmonic audio coding features, the joint audio coding features are optimized at multiple levels to form multiple optimized features. The granularity of processing the coding features differs between every two levels of optimization.
[0018] By integrating the multiple levels of optimization features, an optimized audio coding feature is formed.
[0019] In a preferred embodiment of this application, in the above-described method for optimizing voice quality, the step of optimizing the joint audio coding features at multiple levels based on the harmonic audio coding features to form multiple levels of optimized features includes:
[0020] For each level of optimization, based on the granularity corresponding to the optimization at that level, the harmonic audio coding features and the joint audio coding features are segmented respectively to form at least one segmented harmonic coding feature and at least one segmented joint coding feature corresponding to the optimization at that level. The number of segmented harmonic coding features formed based on the maximum granularity segmentation is 1, and the number of segmented joint coding features formed based on the maximum granularity segmentation is 1.
[0021] For each of the segmented harmonic coding features, the segmented joint coding features corresponding to the segmented harmonic coding features are optimized based on the segmented harmonic coding features to form the optimized joint coding features corresponding to the segmented harmonic coding features.
[0022] Based on the position coding feature corresponding to each of the segmented harmonic coding features, position fusion is performed on the optimized joint coding feature corresponding to each of the segmented harmonic coding features to form the position fused coding feature corresponding to each of the segmented harmonic coding features;
[0023] The positional fusion coding features corresponding to each segmented harmonic coding feature of the same level of optimization are aggregated to form the level optimization feature corresponding to that level of optimization.
[0024] In a preferred embodiment of this application, in the above-described voice quality optimization method, the step of optimizing the segmented joint coding feature corresponding to each segmented harmonic coding feature to form an optimized joint coding feature corresponding to that segmented harmonic coding feature includes:
[0025] In the i-th optimization, based on the segmented harmonic coding feature, attention coding is performed on the output coding feature of the (i-1)-th optimization, and the output coding feature of the (i-1)-th optimization is connected to the result of attention coding to form the output coding feature of the i-th optimization, wherein the output coding feature of the 0th optimization is the segmented joint coding feature corresponding to the segmented harmonic coding feature, and i is a positive odd number;
[0026] In the j-th optimization, the segmented harmonic coding features are subjected to linear transformation and nonlinear activation processing to form an optimized mapping parameter distribution. Then, the optimized mapping parameter distribution and the output coding features of the (j-1)-th optimization are multiplied bitwise to form the output coding features of the j-th optimization, where j is an even number greater than or equal to 2.
[0027] The output coding feature of the last optimization is determined as the optimized joint coding feature corresponding to the segmented harmonic coding feature, wherein the last optimization is an odd-numbered optimization.
[0028] In a preferred embodiment of this application, the step of fusing the multiple levels of optimization features to form optimized audio coding features in the above-described human voice quality optimization method includes:
[0029] The hierarchical optimization feature with the largest granularity among the multiple hierarchical optimization features is determined as the first hierarchical optimization feature;
[0030] For each hierarchical optimization feature whose granularity is not the largest among the multiple hierarchical optimization features, the mean feature distance between at least two segmented harmonic coding features and at least two segmented joint coding features corresponding to that hierarchical optimization feature is determined. The hierarchical optimization feature is formed based on the at least two segmented harmonic coding features and the at least two segmented joint coding features, and the at least two segmented harmonic coding features and the at least two segmented joint coding features are formed by segmenting the harmonic audio coding features and the joint audio coding features according to the corresponding granularity.
[0031] In each level of optimization feature where the granularity is not the largest, the level of optimization feature with the smallest mean feature distance is determined and used as the second level of optimization feature;
[0032] Based on the first-level optimization features and the second-level optimization features, optimized audio coding features are determined.
[0033] In a preferred embodiment of this application, the step of performing joint semantic coding on the reference audio data and the original audio data to form joint audio coding features in the above-mentioned human voice quality optimization method includes:
[0034] The reference audio data is subjected to Fourier transform to form reference audio spectrum data, and the reference audio spectrum data is subjected to convolution processing to form reference audio convolution vector;
[0035] The original audio spectrum data corresponding to the original audio data is convolved to form the original audio convolution vector;
[0036] The reference audio convolution vector and the original audio convolution vector are subjected to semantic mining to form joint audio coding features.
[0037] In a preferred embodiment of this application, the step of performing semantic mining on the reference audio convolutional vector and the original audio convolutional vector to form joint audio coding features in the above-mentioned human voice quality optimization method includes:
[0038] The target input vector and the reference audio convolution vector are concatenated to form a concatenated vector. In the first association semantic mining stage, the target input vector is the original audio convolution vector. In each subsequent association semantic mining stage, the target input vector is the target output vector of the previous association semantic mining stage.
[0039] The concatenated vector is downsampled to form the target output vector for the current semantic mining stage;
[0040] The target output vector of the last associated semantic mining stage is determined as the joint audio coding feature.
[0041] This application also provides a human voice quality optimization device, comprising:
[0042] The Fourier transform module is used to perform Fourier transform on the raw audio data to form raw audio spectrum data;
[0043] The harmonic extraction module is used to determine the target harmonic frequency from the original audio spectrum data, and extract harmonic audio data from the original audio data based on the target harmonic frequency;
[0044] A joint semantic coding module is used to perform joint semantic coding on the reference audio data and the original audio data to form joint audio coding features, wherein the reference audio data and the original audio data have the same sound content, and the sound quality of the reference audio data is higher than that of the original audio data;
[0045] The harmonic semantic coding module is used to perform semantic coding on the harmonic audio data to form harmonic audio coding features;
[0046] The coding feature optimization module is used to optimize the joint audio coding features based on the harmonic audio coding features to form optimized audio coding features;
[0047] The semantic decoding module is used to perform semantic decoding on the optimized audio coding features to form optimized audio data.
[0048] Based on the above, this application also provides an electronic device, including:
[0049] Memory, used to store computer programs;
[0050] A processor connected to the memory is used to execute the computer program stored in the memory to implement the above-described human voice quality optimization method.
[0051] Based on the above, this application also provides a computer-readable storage medium storing a computer program that, when executed, performs the various steps of the above-described human voice quality optimization method.
[0052] The method, apparatus, device, and medium for optimizing human voice quality provided in this application firstly perform Fourier transform on the original audio data to form original audio spectrum data; secondly, determine the target harmonic frequency from the original audio spectrum data, and extract harmonic audio data from the original audio data based on the target harmonic frequency; then, perform joint semantic coding on the reference audio data and the original audio data to form joint audio coding features; subsequently, perform semantic coding on the harmonic audio data to form harmonic audio coding features; further, optimize the joint audio coding features based on the harmonic audio coding features to form optimized audio coding features; finally, perform semantic decoding on the optimized audio coding features to form optimized audio data. Based on the above, on the one hand, since joint semantic encoding is performed on the reference audio data and the original audio data, and the sound quality of the reference audio data is higher than that of the original audio data, the resulting joint audio encoding features have a better representation ability of semantic information in the dimension of sound quality. Furthermore, and importantly, harmonic audio data has a significant impact on the representation of timbre. Therefore, further optimization through the semantic features corresponding to the harmonic audio data can make the resulting optimized audio encoding features more effective in maintaining the original timbre dimension. Based on this, the optimized audio data formed by semantic decoding can take into account both sound quality and timbre, thereby improving the problem of poor human voice quality optimization in the existing technology. Attached Figure Description
[0053] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings.
[0054] Figure 1 A structural block diagram of an electronic device provided in an embodiment of this application.
[0055] Figure 2 This is a flowchart illustrating the voice quality optimization method provided in this application embodiment.
[0056] Figure 3 This is a schematic diagram illustrating the correspondence between segmented harmonic coding features and segmented joint coding features provided in the embodiments of this application.
[0057] Figure 4 This is a schematic diagram illustrating the optimization of the segmentation joint coding features provided in the embodiments of this application.
[0058] Figure 5 A block diagram of a voice quality optimization device provided in an embodiment of this application. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0060] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0061] like Figure 1 As shown in the illustration, this application provides an electronic device. The electronic device may include a memory, a processor, and a voice quality optimization device.
[0062] Specifically, the memory and the processor are electrically connected directly or indirectly to enable data transmission or interaction. For example, the memory and the processor can be electrically connected via one or more communication buses or signal lines. The voice quality optimization device includes at least one software functional module stored in the memory in the form of software or firmware. The processor is used to execute executable computer programs stored in the memory, such as the software functional modules and computer programs included in the voice quality optimization device, to implement the voice quality optimization method provided in the embodiments of this application.
[0063] Optionally, the memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0064] Optionally, the processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a system on chip (SoC), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0065] Understandable. Figure 1 The structure shown is for illustrative purposes only; the electronic device may also include components that are more advanced than those shown. Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown may include, for example, a communication unit for exchanging information with other devices (such as audio acquisition devices).
[0066] Combination Figure 2 This application also provides a method for optimizing human voice quality applicable to the aforementioned electronic device. The method steps defined in the process related to the human voice quality optimization method can be implemented by the electronic device. The following will describe... Figure 2 The specific process shown will be explained in detail.
[0067] Step S110: Perform Fourier transform on the original audio data to form the original audio spectrum data.
[0068] In this embodiment, the electronic device can perform Fourier transform on the raw audio data (the specific processing procedure can refer to relevant prior art) to form raw audio spectrum data (i.e., the corresponding spectrum diagram). The raw audio data can be a single audio frame or multiple audio frames. Furthermore, in a specific application scenario, the voice quality optimization method can be used to optimize singing; therefore, the raw audio data can be the user's singing voice.
[0069] Step S120: Determine the target harmonic frequency from the original audio spectrum data, and extract harmonic audio data from the original audio data based on the target harmonic frequency.
[0070] In this embodiment, after forming the original audio spectrum data, the electronic device can determine the target harmonic frequency from the original audio spectrum data and extract harmonic audio data from the original audio data based on the target harmonic frequency. For example, in one implementation, the fundamental frequency is typically the most significant peak in the original audio spectrum data, i.e., the frequency component with the largest amplitude. That is, the fundamental frequency in the original audio spectrum data is first determined, and then the target harmonic spectrum is determined based on this fundamental frequency. For example:
[0071] f k =k*f0 (k=2,3,4,...);
[0072] Where f0 is the fundamental frequency and k is the harmonic order.
[0073] It should be noted that the target harmonic frequency can include the harmonic frequency corresponding to each of multiple harmonic orders, or it can include the harmonic frequency corresponding to only one harmonic order. Then, based on the determined target harmonic frequency, the original audio data can be filtered using a corresponding bandpass filter (BPF) or notch filter to obtain the corresponding harmonic audio data. For example, if the target harmonic frequencies are 2f0 and 3f0, two bandpass filters can be designed to extract these two harmonic components respectively. If only all harmonic components need to be retained, this can be achieved using multiple bandpass filters or a single multi-band filter.
[0074] Step S130: Perform joint semantic encoding on the reference audio data and the original audio data to form joint audio coding features.
[0075] In this embodiment, the electronic device can perform joint semantic encoding on the reference audio data and the original audio data to form joint audio encoding features. The reference audio data and the original audio data have the same sound content (e.g., the same lyrics in the same song), and the sound quality of the reference audio data is higher than that of the original audio data; for example, the reference audio data may be the original vocals of the song. This allows the joint audio encoding features to contain semantic information from the high-quality reference audio data, thus providing a better representation of sound quality through the joint audio encoding features.
[0076] Step S140: Semantic encoding is performed on the harmonic audio data to form harmonic audio coding features.
[0077] In this embodiment of the application, after the harmonic audio data is generated, the electronic device can perform semantic encoding on the harmonic audio data to form harmonic audio encoding features. That is, corresponding semantic information is extracted from the harmonic audio data, and harmonic semantic information plays an important role in the representation of timbre, so that the harmonic audio encoding features can effectively represent the timbre of the original audio data (corresponding to the person).
[0078] Step S150: Based on the harmonic audio coding features, optimize the joint audio coding features to form optimized audio coding features.
[0079] In this embodiment, after forming the harmonic audio coding feature and the joint audio coding feature, the electronic device can optimize the joint audio coding feature based on the harmonic audio coding feature to form an optimized audio coding feature. That is, since the optimized audio coding feature mainly focuses on optimizing the sound quality, it also carries semantic information about timbre from the reference audio data. Therefore, direct decoding would result in a significant change in timbre. Thus, optimization using harmonic audio coding features that characterize timbre can preserve the original timbre semantic information to the greatest extent possible.
[0080] Step S160: Semantically decode the optimized audio coding features to form optimized audio data.
[0081] In this embodiment, after forming the optimized audio coding features, the electronic device can perform semantic decoding on the optimized audio coding features to form optimized audio data. That is, the optimized audio coding features can be mapped to the audio time domain to form optimized audio data. Thus, the optimized audio data has higher sound quality and the same timbre compared to the original audio data, thereby achieving reliable optimization of human voice quality.
[0082] Based on the above, on the one hand, since joint semantic encoding is performed on the reference audio data and the original audio data, and the sound quality of the reference audio data is higher than that of the original audio data, the resulting joint audio encoding features have a better representation ability of semantic information in the dimension of sound quality. Furthermore, and importantly, harmonic audio data has a significant impact on the representation of timbre. Therefore, further optimization through the semantic features corresponding to the harmonic audio data can make the resulting optimized audio encoding features more effective in maintaining the original timbre dimension. Based on this, the optimized audio data formed by semantic decoding can take into account both sound quality and timbre, thereby improving the problem of poor human voice quality optimization in the existing technology.
[0083] Firstly, regarding step S130, it should be noted that the specific method of jointly semantically encoding the reference audio data and the original audio data is not limited and can be selected according to actual needs.
[0084] For example, in one specific implementation, the reference audio data and the original audio data can be embedded in the temporal and spatial domains respectively (e.g., through a recurrent neural network (RNN) or a long short-term memory network (LSTM)) to obtain corresponding embedding features. Then, the two embedding features can be concatenated, summed, and subjected to attention to achieve fusion, thereby forming a joint audio coding feature carrying the semantic information of the two audio data.
[0085] For example, in another specific implementation, considering that temporal semantic mining directly operates on audio data in the temporal domain, which is easily affected by noise, the above step S130 may further include steps S131, S132 and S133, the specific contents of each step are as follows.
[0086] Step S131: Perform Fourier transform on the reference audio data to form reference audio spectrum data, and perform convolution processing on the reference audio spectrum data to form a reference audio convolution vector.
[0087] In this embodiment, the reference audio data can be Fourier transformed to form reference audio spectrum data, and then convolved to form a reference audio convolution vector. That is, the time-domain audio signal can be converted into a corresponding spectrogram, and then convolution processing can be performed on the spectrogram to achieve semantic mining. It should be noted that convolution operations can focus on local regions of the spectrogram, which is very useful for capturing local features of useful signals. Furthermore, since noise in the spectrogram usually manifests as random, unstructured peaks, it is relatively easy to identify and separate in the spectrogram. Then, through effective learning of the convolution kernel during the convolution operation, these noise peaks can be reliably filtered out, thereby improving the accuracy of audio signal processing.
[0088] Step S132: Perform convolution processing on the original audio spectrum data corresponding to the original audio data to form the original audio convolution vector.
[0089] In this embodiment, the original audio spectrum data corresponding to the original audio data can be convolved to form an original audio convolution vector, that is, the original audio convolution vector is formed by convolving the original audio spectrum data formed in step S110. Alternatively, in other embodiments, the original audio data can be subjected to Fourier transform again.
[0090] Step S133: Perform semantic mining on the reference audio convolution vector and the original audio convolution vector to form joint audio coding features.
[0091] In this embodiment of the application, after forming the reference audio convolution vector and the original audio convolution vector, semantic mining is performed on the reference audio convolution vector and the original audio convolution vector to form joint audio coding features. That is, the reference audio convolution vector and the original audio convolution vector are fused so that the formed joint audio coding features can represent the semantic information in both the reference audio convolution vector and the original audio convolution vector.
[0092] It is understood that in step S133 above, the specific method for performing semantic mining on the reference audio convolution vector and the original audio convolution vector is not limited. For example, in a specific implementation, in order to better capture the semantic correlation between the original audio and the reference audio, step S133 above may include the following:
[0093] First, the target input vector and the reference audio convolution vector can be concatenated to form a concatenated vector. In the first association semantic mining stage, the target input vector is the original audio convolution vector. In each subsequent association semantic mining stage, the target input vector is the target output vector of the previous association semantic mining stage. In this way, the fusion of the reference audio convolution vector and the original audio convolution vector can be achieved in multiple stages.
[0094] Secondly, the concatenated vector can be downsampled (for example, by mean pooling or max pooling) to form the target output vector for the current association semantic mining stage;
[0095] Then, the target output vector of the last associated semantic mining stage can be determined as the joint audio coding feature.
[0096] Secondly, it should be noted that the specific method of semantic encoding of the harmonic audio data is not limited and can be selected according to actual needs.
[0097] For example, in one specific implementation, the harmonic audio data can be embedded in the time-space domain (e.g., through a recurrent neural network (RNN) or a long short-term memory network (LSTM)) to obtain the corresponding harmonic audio coding features.
[0098] For example, in another specific implementation, the harmonic audio data can be Fourier transformed to form corresponding spectrum data. Then, the spectrum data can be convolved to obtain the corresponding harmonic audio coding features.
[0099] Thirdly, it should be noted that the specific method for optimizing the joint audio coding features in step S150 is not limited and can be selected according to actual needs.
[0100] For example, in one specific implementation, the joint audio coding features can be subjected to cross-attention processing based on the harmonic audio coding features to form optimized audio coding features.
[0101] For example, in another specific implementation, considering that timbre plays a relatively important role, in order to fully integrate the semantic information in the harmonic audio coding features into the joint audio coding features, the above step S150 may further include steps S151 and S152, the specific contents of each step are as follows.
[0102] Step S151: Based on the harmonic audio coding features, the joint audio coding features are optimized at multiple levels to form multiple levels of optimized features.
[0103] In this embodiment, the joint audio coding features can be optimized at multiple levels based on the harmonic audio coding features, forming multiple levels of optimized features. The granularity of processing the coding features differs between each two levels of optimization; for example, the first level of optimization corresponds to a first granularity, and the second level of optimization corresponds to a second granularity.
[0104] Step S152: Integrate the multiple levels of optimization features to form optimized audio coding features.
[0105] In this embodiment, after forming the multiple levels of optimization features, these features can be fused to form optimized audio coding features. Based on this, by fusing multiple levels of optimization features, the optimization results from different levels can be combined, capturing more dimensions of information and improving the comprehensiveness and accuracy of the coding.
[0106] It is understood that in step S151 above, the specific method of optimizing the joint audio coding features at multiple levels is not limited. In order to optimize the audio features step by step from macro to micro and ensure the precision and comprehensiveness of the optimization, step S151 above may further include steps S151a, S151b, S151c and S151d, and the specific contents of each step are as follows.
[0107] Step S151a: For each level of optimization, based on the granularity corresponding to the optimization at that level, the harmonic audio coding features and the joint audio coding features are segmented respectively to form at least one segmented harmonic coding feature and at least one segmented joint coding feature corresponding to the optimization at that level.
[0108] In this embodiment, for each optimization level, based on the granularity corresponding to that level, the harmonic audio coding features and the joint audio coding features are segmented to form at least one segmented harmonic coding feature and at least one segmented joint coding feature corresponding to that level. Specifically, the number of segmented harmonic coding features formed based on the maximum granularity segmentation is 1, and the number of segmented joint coding features formed based on the maximum granularity segmentation is 1. That is, in an optimization level, the harmonic audio coding features can be directly used as segmented harmonic coding features, and the joint audio coding features can be directly used as segmented joint coding features.
[0109] Step S151b: For each of the segmented harmonic coding features, optimize the segmented joint coding feature corresponding to the segmented harmonic coding feature based on the segmented harmonic coding feature to form the optimized joint coding feature corresponding to the segmented harmonic coding feature.
[0110] In this embodiment, after forming the segmented harmonic coding feature and the segmented joint coding feature, for each segmented harmonic coding feature, the corresponding segmented joint coding feature is optimized based on that feature to form an optimized joint coding feature. This involves fusing the semantic information from the segmented harmonic coding feature into the joint coding feature. Furthermore, it should be noted that the correspondence between the segmented harmonic coding feature and the joint coding feature is as follows: they are formed based on the same granularity of segmentation, and their positions in the harmonic audio coding feature and the joint audio coding feature are also the same. For example... Figure 3 As shown, there is a correspondence between segmented harmonic coding feature 11 and segmented joint coding feature 21, a correspondence between segmented harmonic coding feature 12 and segmented joint coding feature 22, a correspondence between segmented harmonic coding feature 13 and segmented joint coding feature 23, and a correspondence between segmented harmonic coding feature 14 and segmented joint coding feature 24.
[0111] Step S151c: Based on the position coding feature corresponding to each segmented harmonic coding feature, perform position fusion on the optimized joint coding feature corresponding to each segmented harmonic coding feature to form a position fused coding feature corresponding to each segmented harmonic coding feature.
[0112] In this embodiment, after forming the optimized joint coding features, the optimized joint coding features corresponding to each segmented harmonic coding feature can be positionally fused based on the positional coding features corresponding to each segmented harmonic coding feature, forming a positionally fused coding feature corresponding to each segmented harmonic coding feature. By introducing positional coding features and performing positional fusion on the optimized joint features, the relevant information between different semantic features in the audio signal can be better preserved, avoiding the loss of important semantic information during the optimization process. Furthermore, the positional coding features can be formed by embedding the parameters of the center position of the corresponding segmented harmonic coding feature, specifically the coordinates of the center position in the harmonic audio coding feature. Moreover, the corresponding optimized joint coding features and positional coding features can be added or concatenated to achieve fusion.
[0113] Step S151d: Aggregate the positional fusion coding features corresponding to each segmented harmonic coding feature of the same level of optimization to form the level optimization feature corresponding to that level of optimization.
[0114] In the embodiments of this application, the positional fusion coding features corresponding to each segmented harmonic coding feature corresponding to the optimization at the same level can be aggregated (such as convolution processing after splicing, or addition operation, etc.) to form the hierarchical optimization feature corresponding to that level of optimization.
[0115] It is understood that in step S151b above, the specific method for optimizing the segmented joint coding features corresponding to the segmented harmonic coding features is not limited. For example, in a specific implementation, in order to achieve full fusion of the segmented harmonic coding features and the segmented joint coding features, and to avoid problems such as large computational resource overhead or difficulty in effectively guaranteeing reliability during the full fusion process, step S151b above may further include the following (in conjunction with...). Figure 4 ):
[0116] First, in the i-th optimization, based on the segmented harmonic coding features, attention coding (i.e., cross-attention processing) is performed on the output coding features of the (i-1)-th optimization, and the output coding features of the (i-1)-th optimization are connected to the result of attention coding (e.g., the output coding features of the (i-1)-th optimization and the result of attention coding are added together) to form the output coding features of the i-th optimization. Here, the output coding features of the 0th optimization are the segmented joint coding features corresponding to the segmented harmonic coding features, and i is a positive odd number, i.e., the optimizations of the 1st, 3rd, 5th, etc., are attention codings.
[0117] Secondly, in the j-th optimization, the segmented harmonic coding features are subjected to linear transformation (i.e., multiplied with a weight matrix and then added with a bias matrix; it should be noted that the parameters of the linear transformation can be the same in each optimization, i.e., to achieve the sharing of the optimization mapping parameter distribution, or they can be different to improve accuracy) and nonlinear activation processing (such as through the sigmoid function) to form the optimization mapping parameter distribution. Then, the optimization mapping parameter distribution and the output coding features of the (j-1)-th optimization are multiplied bitwise to form the output coding features of the j-th optimization, where j is an even number greater than or equal to 2, i.e., the optimizations of the 2nd, 4th, 6th, etc., are bitwise multiplication operations.
[0118] Then, the output coding feature of the last optimization can be determined as the optimized joint coding feature corresponding to the segmented harmonic coding feature, wherein the last optimization belongs to an odd-numbered optimization.
[0119] It's important to note that cross-attention can simultaneously consider the global relationship between two feature vectors, capturing their complex dependencies. This global perspective helps retain more semantic information during feature fusion. When the relationship between two feature vectors is complex (e.g., non-linear), cross-attention can better model this relationship, thus improving the fusion effect. However, cross-attention requires calculating the similarity of all element pairs between the two feature vectors, resulting in high computational complexity, especially with high feature dimensions. Linear transformations, non-linear activation processing, and bitwise multiplication are relatively simple, with fewer parameters and lower computational resource consumption. However, non-linear mappings that rely on a single feature vector may not fully capture the complex relationship between the two feature vectors, thus reducing the comprehensiveness of the fusion. Therefore, by alternating between the two fusion methods, their advantages and disadvantages can be complemented.
[0120] It is understood that the specific method of fusing the multiple hierarchical optimization features in step S152 above is not limited. For example, in a specific implementation, in order to balance the richness of semantic information and the consumption of computing resources, step S152 above may further include:
[0121] First, the hierarchical optimization feature with the largest granularity among the multiple hierarchical optimization features can be determined as the first hierarchical optimization feature;
[0122] Secondly, for each hierarchical optimization feature whose granularity is not the largest among the multiple hierarchical optimization features, the mean feature distance between at least two segmented harmonic coding features and at least two segmented joint coding features corresponding to that hierarchical optimization feature is determined (i.e., the mean feature distance between the corresponding coding features). The hierarchical optimization feature is formed based on the at least two segmented harmonic coding features and the at least two segmented joint coding features, and the at least two segmented harmonic coding features and the at least two segmented joint coding features are formed by segmenting the harmonic audio coding features and the joint audio coding features according to the corresponding granularity, as described above.
[0123] Then, among the optimized features at each level that is not the largest in granularity, the optimized feature with the smallest mean feature distance can be determined as the second optimized feature. In this way, the granularity with the highest segmentation accuracy can be determined.
[0124] Finally, optimized audio coding features can be determined based on the first-level optimized features and the second-level optimized features. For example, the first-level optimized features and the second-level optimized features can be concatenated and then convolutional to obtain optimized audio coding features.
[0125] Fourthly, regarding step S160, it should be noted that the specific method for semantic decoding of the optimized audio coding features is not limited and can be selected according to actual needs.
[0126] For example, in one specific implementation, the semantic decoding process may include:
[0127] First, the optimized audio coding features are mapped to a high-dimensional feature space through a fully connected layer to obtain high-dimensional decoding features. Second, the size of the high-dimensional decoding features is gradually increased using a deconvolution layer with transposed convolution. Then, the size of the output coding features of the deconvolution layer is further increased using a bilinear interpolation layer with an upsampling layer. Next, the size is further increased using a transposed convolution layer with another deconvolution layer. Finally, the size is restored to the original audio spectrum data size using a bilinear interpolation layer with an upsampling layer, thus forming the corresponding optimized audio spectrum data. Finally, the optimized audio spectrum data can be subjected to an inverse Fourier transform to convert the data from the frequency domain to the time domain, thus obtaining the optimized audio data.
[0128] Combination Figure 5 This application also provides a voice quality optimization device applicable to the aforementioned electronic devices. The voice quality optimization device may include a Fourier transform module, a harmonic extraction module, a joint semantic coding module, a harmonic semantic coding module, a coding feature optimization module, and a semantic decoding module.
[0129] The Fourier transform module is used to perform Fourier transform on the original audio data to form original audio spectrum data. In this embodiment, the Fourier transform module can be used to perform... Figure 2 The relevant content regarding the Fourier transform module in step S110 can be found in the preceding description of step S110.
[0130] The harmonic extraction module is used to determine the target harmonic frequency from the original audio spectrum data, and extract harmonic audio data from the original audio data based on the target harmonic frequency. In this embodiment, the harmonic extraction module can be used to perform... Figure 2 The relevant content regarding the harmonic extraction module in step S120 shown can be found in the previous description of step S120.
[0131] The joint semantic coding module is used to perform joint semantic coding on the reference audio data and the original audio data to form joint audio coding features, wherein the reference audio data and the original audio data have the same sound content, and the sound quality of the reference audio data is higher than that of the original audio data. In this embodiment, the joint semantic coding module can be used to perform... Figure 2 The relevant content regarding the joint semantic coding module in step S130 shown can be found in the previous description of step S130.
[0132] The harmonic semantic encoding module is used to perform semantic encoding on the harmonic audio data to form harmonic audio encoding features. In this embodiment, the harmonic semantic encoding module can be used to perform... Figure 2 The relevant content regarding the harmonic semantic coding module in step S140 shown can be found in the previous description of step S140.
[0133] The encoding feature optimization module is used to optimize the joint audio encoding features based on the harmonic audio encoding features to form optimized audio encoding features. In this embodiment, the encoding feature optimization module can be used to perform... Figure 2 The relevant content regarding the coding feature optimization module in step S150 shown can be found in the previous description of step S150.
[0134] The semantic decoding module is used to perform semantic decoding on the optimized audio coding features to form optimized audio data. In this embodiment, the semantic decoding module can be used to execute... Figure 2 The relevant content regarding the semantic decoding module in step S160 shown can be found in the preceding description of step S160.
[0135] In this embodiment of the application, corresponding to the above-described method for optimizing human voice quality applied to the electronic device, a computer-readable storage medium is also provided, which stores a computer program that executes the various steps of the human voice quality optimization method when the computer program is run.
[0136] The steps executed by the aforementioned computer program during runtime will not be described in detail here, but can be found in the explanation of the human voice quality optimization method above.
[0137] In summary, the voice quality optimization method, apparatus, device, and medium provided in this application firstly perform Fourier transform on the original audio data to form original audio spectrum data; secondly, determine the target harmonic frequency from the original audio spectrum data, and extract harmonic audio data from the original audio data based on the target harmonic frequency; then, perform joint semantic coding on the reference audio data and the original audio data to form joint audio coding features; subsequently, perform semantic coding on the harmonic audio data to form harmonic audio coding features; further, optimize the joint audio coding features based on the harmonic audio coding features to form optimized audio coding features; finally, perform semantic decoding on the optimized audio coding features to form optimized audio data. Based on the above, on the one hand, since joint semantic encoding is performed on the reference audio data and the original audio data, and the sound quality of the reference audio data is higher than that of the original audio data, the resulting joint audio encoding features have a better representation ability of semantic information in the dimension of sound quality. Furthermore, and importantly, harmonic audio data has a significant impact on the representation of timbre. Therefore, further optimization through the semantic features corresponding to the harmonic audio data can make the resulting optimized audio encoding features more effective in maintaining the original timbre dimension. Based on this, the optimized audio data formed by semantic decoding can take into account both sound quality and timbre, thereby improving the problem of poor human voice quality optimization in the existing technology.
[0138] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0139] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0140] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, electronic device, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks. It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. In the absence of further restrictions, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0141] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for optimizing human voice quality, characterized in that, include: Perform a Fourier transform on the raw audio data to generate the raw audio spectrum data; The target harmonic frequency is determined from the original audio spectrum data, and harmonic audio data is extracted from the original audio data based on the target harmonic frequency. Joint semantic coding is performed on the reference audio data and the original audio data to form joint audio coding features, wherein the reference audio data and the original audio data have the same sound content, and the sound quality of the reference audio data is higher than that of the original audio data; The harmonic audio data is semantically encoded to form harmonic audio coding features; Based on the harmonic audio coding features, the joint audio coding features are optimized to form optimized audio coding features; The optimized audio coding features are semantically decoded to form optimized audio data.
2. The method for optimizing human voice quality according to claim 1, characterized in that, The step of optimizing the joint audio coding features based on the harmonic audio coding features to form optimized audio coding features includes: Based on the harmonic audio coding features, the joint audio coding features are optimized at multiple levels to form multiple optimized features. The granularity of processing the coding features differs between every two levels of optimization. By integrating the multiple levels of optimization features, an optimized audio coding feature is formed.
3. The method for optimizing human voice quality according to claim 2, characterized in that, The step of optimizing the joint audio coding features at multiple levels based on the harmonic audio coding features to form multiple levels of optimized features includes: For each level of optimization, based on the granularity corresponding to the optimization at that level, the harmonic audio coding features and the joint audio coding features are segmented respectively to form at least one segmented harmonic coding feature and at least one segmented joint coding feature corresponding to the optimization at that level. The number of segmented harmonic coding features formed based on the maximum granularity segmentation is 1, and the number of segmented joint coding features formed based on the maximum granularity segmentation is 1. For each of the segmented harmonic coding features, the segmented joint coding features corresponding to the segmented harmonic coding features are optimized based on the segmented harmonic coding features to form the optimized joint coding features corresponding to the segmented harmonic coding features. Based on the position coding feature corresponding to each of the segmented harmonic coding features, position fusion is performed on the optimized joint coding feature corresponding to each of the segmented harmonic coding features to form the position fused coding feature corresponding to each of the segmented harmonic coding features; The positional fusion coding features corresponding to each segmented harmonic coding feature of the same level of optimization are aggregated to form the level optimization feature corresponding to that level of optimization.
4. The method for optimizing human voice quality according to claim 3, characterized in that, The step of optimizing the segmented joint coding feature corresponding to each segmented harmonic coding feature to form the optimized joint coding feature corresponding to the segmented harmonic coding feature includes: In the i-th optimization, based on the segmented harmonic coding feature, attention coding is performed on the output coding feature of the (i-1)-th optimization, and the output coding feature of the (i-1)-th optimization is connected to the result of attention coding to form the output coding feature of the i-th optimization, wherein the output coding feature of the 0th optimization is the segmented joint coding feature corresponding to the segmented harmonic coding feature, and i is a positive odd number; In the j-th optimization, the segmented harmonic coding features are subjected to linear transformation and nonlinear activation processing to form an optimized mapping parameter distribution. Then, the optimized mapping parameter distribution and the output coding features of the (j-1)-th optimization are multiplied bitwise to form the output coding features of the j-th optimization, where j is an even number greater than or equal to 2. The output coding feature of the last optimization is determined as the optimized joint coding feature corresponding to the segmented harmonic coding feature, wherein the last optimization is an odd-numbered optimization.
5. The method for optimizing human voice quality according to claim 2, characterized in that, The step of fusing the multiple levels of optimization features to form optimized audio coding features includes: The hierarchical optimization feature with the largest granularity among the multiple hierarchical optimization features is determined as the first hierarchical optimization feature; For each hierarchical optimization feature whose granularity is not the largest among the multiple hierarchical optimization features, the mean feature distance between at least two segmented harmonic coding features and at least two segmented joint coding features corresponding to that hierarchical optimization feature is determined. The hierarchical optimization feature is formed based on the at least two segmented harmonic coding features and the at least two segmented joint coding features, and the at least two segmented harmonic coding features and the at least two segmented joint coding features are formed by segmenting the harmonic audio coding features and the joint audio coding features according to the corresponding granularity. In each level of optimization feature where the granularity is not the largest, the level of optimization feature with the smallest mean feature distance is determined and used as the second level of optimization feature; Based on the first-level optimization features and the second-level optimization features, optimized audio coding features are determined.
6. The method for optimizing human voice quality according to any one of claims 1-5, characterized in that, The step of jointly semantically encoding the reference audio data and the original audio data to form joint audio coding features includes: The reference audio data is subjected to Fourier transform to form reference audio spectrum data, and the reference audio spectrum data is subjected to convolution processing to form reference audio convolution vector; The original audio spectrum data corresponding to the original audio data is convolved to form the original audio convolution vector; The reference audio convolution vector and the original audio convolution vector are subjected to semantic mining to form joint audio coding features.
7. The method for optimizing human voice quality according to claim 6, characterized in that, The step of performing semantic mining on the reference audio convolution vector and the original audio convolution vector to form joint audio coding features includes: The target input vector and the reference audio convolution vector are concatenated to form a concatenated vector. In the first association semantic mining stage, the target input vector is the original audio convolution vector. In each subsequent association semantic mining stage, the target input vector is the target output vector of the previous association semantic mining stage. The concatenated vector is downsampled to form the target output vector for the current semantic mining stage; The target output vector of the last associated semantic mining stage is determined as the joint audio coding feature.
8. A human voice quality optimization device, characterized in that, include: The Fourier transform module is used to perform Fourier transform on the raw audio data to form raw audio spectrum data; The harmonic extraction module is used to determine the target harmonic frequency from the original audio spectrum data, and extract harmonic audio data from the original audio data based on the target harmonic frequency; A joint semantic coding module is used to perform joint semantic coding on reference audio data and the original audio data to form joint audio coding features, wherein the reference audio data and the original audio data have the same sound content, and the sound quality of the reference audio data is higher than that of the original audio data; The harmonic semantic coding module is used to perform semantic coding on the harmonic audio data to form harmonic audio coding features; The coding feature optimization module is used to optimize the joint audio coding features based on the harmonic audio coding features to form optimized audio coding features; The semantic decoding module is used to perform semantic decoding on the optimized audio coding features to form optimized audio data.
9. An electronic device, characterized in that, include: memory for storing computer programs; A processor connected to the memory is used to execute a computer program stored in the memory to implement the human voice quality optimization method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains a computer program that, when executed, performs the human voice quality optimization method according to any one of claims 1-7.
Citation Information
Patent Citations
Sound conversion optimization method and system
CN108847249A
Audio optimization method, system and equipment based on artificial intelligence
CN118841023A