Audio synthesis method and device, electronic equipment, medium and program product

By determining and adjusting the loudness range of human voice audio in audio synthesis, the problems of high professional threshold and low efficiency in existing technologies are solved, realizing convenient and efficient audio synthesis and sound quality improvement.

CN121506064APending Publication Date: 2026-02-10BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411087748.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In existing technologies, the mixing and synthesis of vocals and accompaniment requires manual operation by professional audio engineers, which has a high professional threshold and low efficiency. Furthermore, machine-synthesized vocals lack the details of real vocals, resulting in long synthesis times.

Method used

By acquiring the vocal and accompaniment audio from the reference audio, determining their loudness range, and adjusting the loudness of the second vocal audio based on the comparison results until it matches the first vocal audio, the two audios are then mixed to generate the target audio.

Benefits of technology

It simplifies the audio synthesis process, improves synthesis efficiency, and enhances the sound quality and consistency of the target audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506064A_ABST
    Figure CN121506064A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of music engineering, and discloses an audio synthesis method and device, electronic equipment, a medium and a program product. The invention provides an audio synthesis method. The audio synthesis method comprises the following steps: acquiring a first human voice audio and an accompaniment audio in a reference audio; determining a first loudness range based on the loudness of the first human voice audio; acquiring a second human voice audio corresponding to the reference audio; determining a second loudness range based on the loudness of the second voice audio; based on a first comparison result between the second loudness range and the first loudness range, adjusting the loudness of the second human voice audio to obtain a first target human voice audio; and mixing the first target human voice audio and the accompaniment audio to obtain a target audio. The synthesis difficulty of the music audio can be simplified, the synthesis time is shortened, and the music synthesis efficiency can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of music engineering technology, specifically to audio synthesis methods, apparatus, electronic devices, media, and program products. Background Technology

[0002] In related technologies, in order to mix and synthesize vocals and accompaniment, professional audio engineers are required to manually synthesize using digital signal processing (DSP) systems and software. This not only has a high professional threshold, but also results in low audio synthesis efficiency due to the long synthesis time. Summary of the Invention

[0003] In view of this, the present disclosure provides an audio synthesis method, apparatus, electronic device, medium, and program product to solve the problem of low audio synthesis efficiency.

[0004] In a first aspect, this disclosure provides an audio synthesis method, the method comprising:

[0005] Extract the first vocal audio and the accompaniment audio from the reference audio;

[0006] Determine the first loudness range based on the loudness of the first human voice audio.

[0007] Obtain the second human voice audio corresponding to the reference audio;

[0008] Determine the second loudness range based on the loudness of the second human voice audio.

[0009] Based on the first comparison result between the second loudness range and the first loudness range, the loudness of the second human voice audio is adjusted to obtain the first target human voice audio;

[0010] The target audio is obtained by mixing the first target vocal audio and the accompaniment audio.

[0011] Secondly, this disclosure provides an audio synthesis apparatus, the apparatus comprising:

[0012] The first acquisition module is used to acquire the first human voice audio and the accompaniment audio in the reference audio;

[0013] The first processing module is used to determine a first loudness range based on the loudness of the first human voice audio.

[0014] The second acquisition module is used to acquire the second human voice audio corresponding to the reference audio;

[0015] The second processing module is used to determine the second loudness range based on the loudness of the second human voice audio.

[0016] The first adjustment module is used to adjust the loudness of the second human voice audio based on a first comparison result between the second loudness range and the first loudness range, so as to obtain the first target human voice audio.

[0017] The synthesis module is used to mix the first target vocal audio and the accompaniment audio to obtain the target audio.

[0018] Thirdly, this disclosure provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the audio synthesis method of the first aspect or any corresponding embodiment described above.

[0019] Fourthly, this disclosure provides a computer-readable storage medium storing computer instructions for causing a computer to perform the audio synthesis method of the first aspect or any corresponding embodiment described above.

[0020] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the audio synthesis method of the first aspect or any corresponding embodiment thereof.

[0021] The audio synthesis method provided in this embodiment can adjust the loudness of the second human voice audio in a targeted manner by using the first loudness range of the first human voice audio in the reference audio as a benchmark through loudness range matching. This not only simplifies the audio synthesis process and makes the synthesis of the target audio more convenient and efficient, but also effectively improves the sound quality of the synthesized target audio, thereby effectively improving the audio synthesis efficiency. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the specific embodiments of this disclosure or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0023] Figure 1 This is a schematic flowchart of an audio synthesis method provided according to an embodiment of the present disclosure;

[0024] Figure 2 This is a schematic flowchart of another audio synthesis method provided according to an embodiment of the present disclosure;

[0025] Figure 3 This is a flowchart illustrating another audio synthesis method provided according to an embodiment of the present disclosure;

[0026] Figure 4 This is a flowchart illustrating another audio synthesis method provided according to an embodiment of the present disclosure;

[0027] Figure 5 This is a flowchart illustrating another audio synthesis method provided according to an embodiment of the present disclosure;

[0028] Figure 6 This is a flowchart illustrating yet another audio synthesis method provided according to an embodiment of the present disclosure;

[0029] Figure 7 This is a structural block diagram of an audio synthesis apparatus according to an embodiment of the present disclosure;

[0030] Figure 8 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0032] In related technologies, mixing and synthesizing vocals and accompaniment requires professional audio engineers to manually synthesize using digital signal processing (DSP) systems and software. This process is highly dependent on the skills and experience of audio engineers, and the professional threshold for synthesis is high.

[0033] Furthermore, machine-synthesized vocals lack the subtle details of real human voices, which means that audio engineers need to spend a lot of time making targeted adjustments during the synthesis process to achieve the ideal sound, resulting in low synthesis efficiency.

[0034] In view of this, the present disclosure provides an audio synthesis method that, after determining the first loudness range of the first human voice audio in the reference audio, adjusts the second loudness range of the second human voice audio corresponding to the reference audio based on the first loudness range, so that the loudness of the first target human voice audio can be aligned with the loudness of the first human voice audio. Then, the first target human voice audio is mixed with the accompaniment audio in the reference audio, which not only makes the audio synthesis process more convenient and efficient, but also effectively improves the sound quality of the synthesized target audio and improves the audio synthesis efficiency.

[0035] According to an embodiment of this disclosure, an audio synthesis method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0036] This embodiment provides an audio synthesis method that can be used in electronic devices such as mobile phones and tablets. Figure 1 This is a flowchart of an audio synthesis method according to an embodiment of the present disclosure, such as... Figure 1 As shown, the process includes the following steps:

[0037] Step S101: Obtain the first human voice audio and the accompaniment audio from the reference audio.

[0038] Reference audio refers to the standard used for audio synthesis, containing vocals and background music (BGM). This reference audio can be a complete audio file or a segment of a specified audio file, depending on the requirements. Specifically, the first vocal audio is the audio segment in the reference audio corresponding to the vocals, and the accompaniment audio is the audio segment in the reference audio corresponding to the background music. The vocals corresponding to the first vocal audio can be real human voices. In some optional examples, the first vocal audio and accompaniment audio can be obtained by performing Music Source Separation (MSS) on the reference audio.

[0039] Step S102: Determine the first loudness range based on the loudness of the first human voice audio.

[0040] To facilitate subsequent audio synthesis and establish a clear benchmark for human voice processing, the loudness range of the first human voice audio is defined based on its loudness. This provides a clear reference for processing the second human voice audio, enabling better matching of audio effects.

[0041] In some optional implementation scenarios, the process of determining the first loudness range can be as follows: First, the overall audio energy of the first human voice audio is detected, and then processed using a specified loudness calculation formula to obtain the average loudness of the first human voice audio. For example, the average loudness of the first human voice audio can be determined using a standardized loudness calculation formula, as shown in the following formula:

[0042]

[0043] Where L(t) refers to the instantaneous loudness at sampling time t, and T refers to the total duration of the first human voice audio. avgIt refers to the average loudness of the first human voice audio.

[0044] The first human voice audio is sampled and quantized, and then converted into a time-domain signal through inverse Fourier transform to clarify the signal amplitude changes of the first human voice audio at each sampling time.

[0045] To improve the accuracy of the first loudness range, the maximum and minimum loudness values ​​of the first human voice audio at each sampling time are determined, thus obtaining the sub-first loudness range DR corresponding to each sampling time. The maximum loudness value can be represented by L. peak =max(|x(t)|) represents the minimum loudness value, which can be represented by L. min =min(|x(t)|) represents the signal amplitude of the first human voice audio at the corresponding sampling time. The expression for the sub-first loudness range DR corresponding to each sampling time can be (L min ,L peak ).

[0046] By identifying the sub-first loudness range corresponding to each sampling time, the loudness range of the first human voice audio at different sampling times can be determined, thereby obtaining the overall dynamic loudness range of the first human voice audio, that is, obtaining the first loudness range of the first human voice audio.

[0047] Step S103: Obtain the second human voice audio corresponding to the reference audio.

[0048] The second human voice audio refers to the human voice audio to be used for audio synthesis. The human voice corresponding to the second human voice audio can be a real human voice or a machine-synthesized voice, which can be determined according to the requirements.

[0049] Step S104: Determine the second loudness range based on the loudness of the second human voice audio.

[0050] To facilitate targeted adjustments to the second human voice audio and improve audio quality, the loudness of the second human voice audio is processed to clarify the overall loudness range of the second human voice audio and determine the second loudness range.

[0051] In some optional examples, when processing the loudness of the second human voice audio, the same analysis method as that used for the loudness of the first human voice audio is employed. This ensures that the two audios are based on the same standard, making the loudness results of the two audios consistent and comparable, thereby reducing errors during subsequent adjustments.

[0052] Step S105: Based on the first comparison result between the second loudness range and the first loudness range, adjust the loudness of the second human voice audio to obtain the first target human voice audio.

[0053] By comparing the second loudness range with the first loudness range, the difference between the two loudness ranges can be clearly identified, thus obtaining the first comparison result.

[0054] Since both the first and second loudness ranges are dynamic ranges that change randomly over time, in order to make the loudness of the second human voice audio better match the loudness of the first human voice audio in the reference audio, the loudness of the second human voice audio is adjusted based on the first comparison result, so that the loudness of the obtained first target human voice audio is more in line with the loudness characteristics of the first human voice audio. This will help improve the overall consistency and coordination of the target audio during subsequent audio synthesis.

[0055] Step S106: Mix the first target human voice audio and the accompaniment audio to obtain the target audio.

[0056] The loudness-adjusted first target human voice audio is mixed with the accompaniment audio so that the human voice in the first target human voice audio can match the background music in the reference audio, thereby obtaining the target audio corresponding to the human voice in the second human voice audio.

[0057] The audio synthesis method provided in this embodiment can adjust the loudness of the second human voice audio in a targeted manner by using the first loudness range of the first human voice audio in the reference audio as a benchmark through loudness range matching. This not only simplifies the audio synthesis process and makes the synthesis of the target audio more convenient and efficient, but also effectively improves the sound quality of the synthesized target audio, thereby effectively improving the audio synthesis efficiency.

[0058] This embodiment provides an audio synthesis method that can be used in electronic devices such as mobile phones and tablets. Figure 2 This is a flowchart of an audio synthesis method according to an embodiment of the present disclosure, such as... Figure 2 As shown, the process includes the following steps:

[0059] Step S201: Obtain the first vocal audio and the accompaniment audio from the reference audio. See details below. Figure 1 Step S101 of the illustrated embodiment will not be described again here.

[0060] Step S202: Determine a first loudness range based on the loudness of the first human voice audio. For details, please refer to [link to relevant documentation]. Figure 1 Step S102 of the illustrated embodiment will not be described again here.

[0061] Step S203: Obtain the second human voice audio corresponding to the reference audio. For details, please refer to [link to relevant documentation]. Figure 1 Step S103 of the illustrated embodiment will not be described again here.

[0062] Step S204: Determine the second loudness range based on the loudness of the second human voice audio. For details, please refer to [link to relevant documentation]. Figure 1 Step S104 of the illustrated embodiment will not be described again here.

[0063] Step S205: Based on the first comparison result between the second loudness range and the first loudness range, adjust the loudness of the second human voice audio to obtain the first target human voice audio.

[0064] Specifically, step S205 includes:

[0065] Step S2051: Determine the maximum signal decibel value corresponding to the first human voice audio, and determine the signal compression threshold based on the maximum signal decibel value.

[0066] To make the loudness of the second human voice audio more uniform and to allow for reasonable adjustment of the loudness of the second human voice audio, the maximum signal decibel value corresponding to the first human voice audio is determined to clarify the upper limit of the loudness of the first human voice audio. Then, based on this maximum signal decibel value, the signal compression threshold is determined to avoid distortion or other adverse effects during subsequent adjustments, thus ensuring the integrity and accuracy of the second human voice audio signal.

[0067] For example, the conversion formula between decibel value and loudness is as follows:

[0068]

[0069] Among them, L p The sound pressure level is expressed in decibels (dB), p represents the sound pressure being measured in Pascals (Pa), and p0 is the reference sound pressure, which is usually taken as 2 × 10⁻⁶ in air. 5 Pa.

[0070] As shown by the above formula, there is a 20-fold relationship between decibel value and loudness. Therefore, to make compression more reasonable, when determining the signal compression threshold, the difference between the maximum signal decibel value and 20 can be used as the signal compression threshold. This threshold can then be used as a constraint for signal compression, ensuring its rationality.

[0071] Step S2052: Determine the signal compression ratio based on the ratio of the second loudness range to the first loudness range.

[0072] The signal compression ratio reflects the relative loudness range between the second and first human voice audio frequencies. The expression for the signal compression ratio (ratio) is:

[0073] By determining the signal compression ratio, the loudness of the second human voice audio can be kept relatively balanced with that of the first human voice audio to a certain extent during the compression process, thereby avoiding situations where the loudness is too high or too low, which helps to improve the overall quality of the second human voice audio.

[0074] Since the first loudness range is a set of sub-first loudness ranges corresponding to multiple sampling times, and the second loudness range is a set of sub-second loudness ranges corresponding to multiple sampling times, the signal compression ratio is also a set of sub-signal compression ratios corresponding to multiple sampling times, which can ensure the targeting and reliability of the adjustment.

[0075] Step S2053: In response to the fact that the decibel value of the current second human voice audio signal is greater than the signal compression threshold, the loudness of the current second human voice audio signal is adjusted based on the first difference between the decibel value of the current second human voice audio signal and the signal compression threshold and the signal compression ratio, so as to obtain the first target human voice audio.

[0076] The current second human voice audio signal is one of the audio signals in the second human voice audio. In adjusting the loudness of the second human voice audio, a segment-by-segment matching method is used to detect whether the decibel value of the current second human voice audio signal at the current sampling time is greater than the signal compression threshold. If the decibel value of the current second human voice audio signal is greater than the signal compression threshold, it indicates that the loudness of the second human voice audio is too high and needs to be compressed. Therefore, based on the first difference between the decibel value of the current second human voice audio signal and the signal compression threshold, and the signal compression ratio, the loudness of the current second human voice audio signal is appropriately compressed to enhance the audibility of the second human voice audio, thereby ultimately obtaining the first target human voice audio.

[0077] In some optional implementation scenarios, the formula for signal conditioning can be as follows:

[0078]

[0079] Where y(t) is the adjusted second human voice audio signal at the current sampling time, x(t) is the input second human voice audio signal at the current sampling time, and ratio is the signal compression ratio at the current sampling time. dB (t) is the decibel value of the current second human voice audio signal at the current sampling time, and threshold is the signal compression threshold at the current sampling time.

[0080] By compressing the current second human voice audio signal using the above signal adjustment formula, the loudness adjustment method can be made more reasonable. Without losing sound quality, the loudness of the adjusted first target human voice audio can be more coordinated and more audible.

[0081] In some optional implementations, in response to the current second human voice audio signal having a decibel value less than or equal to a signal compression threshold, no loudness adjustment processing is performed on the current second human voice audio signal. That is, in response to the current second human voice audio signal having a decibel value less than or equal to a signal compression threshold, it is reasonable to characterize the loudness of the current second human voice audio signal. Therefore, no loudness adjustment processing is performed on the current second human voice audio signal to avoid sound quality loss or signal distortion due to unnecessary adjustments, thereby helping to save computing resources and improve processing efficiency.

[0082] In some optional implementation scenarios, the formula for adjusting the loudness of the second human voice audio signal can be as follows:

[0083]

[0084] This ensures the timely adjustment of audio signals.

[0085] Step S206: Mix the first target vocal audio and the accompaniment audio to obtain the target audio. For details, please refer to [link to details]. Figure 1 Step S106 of the illustrated embodiment will not be described again here.

[0086] The audio synthesis method provided in this embodiment, in the process of adjusting the loudness of the second human voice audio, makes targeted adjustments by determining the signal compression threshold and signal compression ratio, so that the loudness of the second human voice audio can be kept relatively balanced with the loudness of the first human voice audio to a certain extent, avoiding the situation of excessive or insufficient volume, thereby effectively improving the audio quality and audibility, and helping to improve the user's listening experience.

[0087] In some optional implementation scenarios, extended processing can be applied to audio signals with excessively low loudness in the second human voice audio signal to enhance the dynamics of the music.

[0088] This embodiment provides an audio synthesis method that can be used in electronic devices such as mobile phones and tablets. Figure 3 This is a flowchart of an audio synthesis method according to an embodiment of the present disclosure, such as... Figure 3 As shown, the process includes the following steps:

[0089] Step S301: Obtain the first vocal audio and the accompaniment audio from the reference audio.

[0090] Step S302: Determine the first loudness range based on the loudness of the first human voice audio.

[0091] Step S303: Obtain the second human voice audio corresponding to the reference audio.

[0092] Step S304: Determine the second loudness range based on the loudness of the second human voice audio.

[0093] Step S305: Based on the first comparison result between the second loudness range and the first loudness range, adjust the loudness of the second human voice audio to obtain the first target human voice audio.

[0094] Step S306: Mix the first target human voice audio and the accompaniment audio to obtain the target audio.

[0095] Specifically, step S306 includes:

[0096] Step S3061: Based on the spectral distribution of the first human voice audio, determine the first spectral distribution corresponding to the first frequency band.

[0097] To enhance the naturalness of the second human voice audio, the first human voice audio is subjected to spectrum-based processing to determine the spectrum distribution of the first human voice audio, and then to clarify the first spectrum distribution of the first human voice audio in the first frequency range, so as to reflect the energy distribution of the first human voice audio in the first frequency range through the first spectrum distribution.

[0098] The first frequency band can be understood as the frequency band that requires targeted adjustment. For example, the first frequency band can be a low-frequency band (approximately 20Hz-250Hz), a mid-frequency band (approximately 250Hz-2kHz), or a high-frequency band (approximately 2kHz-5kHz and above). The low-frequency band mainly includes the fundamental frequency and low-frequency harmonics of the sound, used to reflect the fullness and power of the sound. The mid-frequency band mainly includes the formants of the sound, used to form the sound's characteristics and intelligibility. The high-frequency band mainly includes the sharpness and clarity of the sound, used to reflect the clarity and detail of the sound. Preferably, the first frequency band can be the high-frequency band, which helps to enrich the sound details of the second vocal audio when adjusting it subsequently.

[0099] In some optional implementation scenarios, the overall spectral distribution of the first human voice audio can be determined by performing a Fourier transform on the audio, and then the first spectral distribution can be determined based on the frequency range corresponding to the first frequency band. The relevant formulas for performing the Fourier transform are as follows:

[0100] S(f)=∫x(t)e -j2πft dt;

[0101] Where x(t) is the audio signal of the first human voice in the current time period, and S(f) is the spectral amplitude value at frequency f.

[0102] Step S3062: Based on the spectral distribution of the first target human voice audio, determine the second spectral distribution corresponding to the first frequency band.

[0103] The same analysis and processing are performed on the spectral distribution of the first target human voice audio to determine the second spectral distribution corresponding to the first frequency band in the overall spectral distribution of the first target human voice audio.

[0104] Step S3063: Based on the first spectral distribution, adjust the second spectral distribution to obtain the second target human voice audio, which is then used to mix with the accompaniment audio to obtain the target audio.

[0105] To make the target audio obtained after mixing more balanced in frequency, the second spectral distribution of the first target human voice audio is adjusted using the first spectral distribution of the first human voice audio within the same first frequency band. This makes the frequency of the adjusted second target human voice audio more matched with that of the first human voice audio, thereby enhancing the integration of the second target human voice audio with the accompaniment audio. This makes the resulting target audio sound more natural and harmonious, and improves the overall quality of the target audio.

[0106] In some optional implementations, step S3063 above includes:

[0107] Step a1: Determine the initial value of the first spectrum of the first spectrum distribution.

[0108] To make the frequencies of the second target human voice audio more prominent in the first frequency band, the first spectral starting point of the first spectral distribution is determined and used as a reference. In the subsequent enhancement adjustment, the expressiveness and characteristics of the human voice can be enhanced.

[0109] Step a2: In response to the current second spectrum value being greater than the first spectrum starting value, determine the second difference between the current second spectrum value and the first spectrum starting value.

[0110] The current second spectrum value is one of the spectrum values ​​in the second spectrum distribution. During the adjustment of the second spectrum distribution, if the current second spectrum value is greater than the first spectrum starting value, it indicates that the current second spectrum value needs to be specifically enhanced. Therefore, a second difference between the current second spectrum value and the first spectrum starting value is determined to clarify the difference between the two.

[0111] Step s3: Based on the second difference and the specified enhancement factor, adjust the current second spectrum value to obtain the third spectrum distribution.

[0112] To enhance the contrast of the spectral distribution and make the second target vocal audio clearer and more distinguishable when mixed with the accompaniment audio, the current second spectral value is specifically adjusted using the second difference and a specified enhancement factor. This results in richer frequency layers in the second target vocal audio within the first frequency band, thereby improving the overall quality of the second vocal audio. The third spectral distribution refers to the overall spectral distribution of the second target vocal audio after adjusting the second spectral distribution. The above process can be represented by the following formula:

[0113] Y(f) = X(f) + α(X(f) - T(f)); where f refers to the second spectral value greater than the initial value of the first spectral distribution, X(f) refers to the first spectral distribution, T(f) refers to the second spectral distribution, and α refers to the enhancement factor. The enhancement factor can be determined based on the frequency domain energy corresponding to the first spectral distribution, thus ensuring the rationality of spectral enhancement. For example, since the first frequency band is a high-frequency band, the value of the enhancement factor can be between 30% and 50% of the frequency domain energy corresponding to the first spectral distribution. That is, if the frequency domain energy of the first spectral distribution is 100, then the value of the enhancement factor can be between 30 (30% of 100) and 50 (50% of 100). In this way, when adjusting the second spectral distribution based on the enhancement factor, the frequency domain energy of the first spectral distribution can be referenced for targeted adjustment, thereby helping to maintain the relative balance and rationality of the audio during the adjustment process.

[0114] In some optional examples, the current second spectrum value is not adjusted in response to the current second spectrum value being less than or equal to the initial value of the first spectrum, thereby helping to improve adjustment efficiency.

[0115] Step a4: Based on the third spectral distribution, the second target human voice audio is obtained.

[0116] In some optional examples, step a4 above includes:

[0117] Step a31: Determine the target harmonic energy based on the first harmonic spectrum distribution of the first human voice audio.

[0118] Step a32: Determine the second harmonic spectrum distribution in the third spectrum distribution;

[0119] Step a33: Determine the harmonic modulation order based on the target harmonic energy;

[0120] Step a34: Based on the harmonic adjustment order and the specified adjustment factor, adjust the second harmonic spectrum distribution to obtain the second target human voice audio.

[0121] Specifically, to make the second target vocal audio fuller and more vivid, the harmonic spectrum of the first vocal audio is processed. Based on its first harmonic spectrum distribution, the energy distribution of each harmonic component in the first vocal audio can be determined, thus determining the target harmonic energy. The target harmonic energy is set based on the desired audio effect and the characteristics of the first vocal audio, and is the harmonic energy level expected to be achieved in the second target vocal audio.

[0122] To further refine and optimize the harmonic components, the harmonic modulation order is determined based on the target harmonic energy. The harmonic modulation order determines the precision and range of adjustment to the second harmonic spectral distribution.

[0123] To ensure the rationality of harmonic modulation, a specified modulation factor is determined based on the frequency domain energy corresponding to the first spectral distribution to control the intensity and amplitude of modulation. Then, based on the harmonic modulation order and the specified modulation factor, the second harmonic spectral distribution can be specifically adjusted using the following formula:

[0124] H(f) = βS(nf), where n is the order of the harmonic, n = 2, 3, 4 (the specific value can be determined according to actual needs), β is the specified adjustment factor, S(f) is the second harmonic spectrum distribution, and H(f) is the adjusted second harmonic spectrum distribution.

[0125] Preferably, the value range of the specified adjustment factor can be between 0% and 3% of the frequency domain energy corresponding to the first spectral distribution, which helps to make the harmonic characteristics of the second target human voice audio more in line with expectations, thereby optimizing the timbre and sound quality of the audio and making it sound more natural and comfortable.

[0126] By adjusting the second harmonic spectrum distribution in the above manner, the harmonic consistency between the second target human voice audio and the first human voice audio can be enhanced, thereby helping to improve the overall harmony and coherence of the audio.

[0127] In some optional examples, step a4 above also includes:

[0128] Step a44: Based on the spectral phase of the first spectral distribution, obtain the first phase distribution;

[0129] Step a45: Based on the spectral phase of the third spectral distribution, obtain the second phase distribution corresponding to the first frequency band interval and the third phase distribution corresponding to the second frequency band interval. The frequency values ​​in the second frequency band interval are less than the frequency values ​​in the first frequency band interval.

[0130] Step a46: Based on the first phase distribution and the third phase distribution, adjust the second phase distribution to obtain the second target human voice audio.

[0131] Specifically, in order to improve the transparency and detail of the audio, the spectral phase of the first spectral distribution is analyzed in a targeted manner to clarify the phase distribution of the first spectral distribution and thus obtain the first phase distribution.

[0132] Since the third spectral distribution refers to the overall spectral distribution of the second target human voice audio after adjusting the second spectral distribution, a targeted analysis of the spectral phase of the third spectral distribution is performed to reasonably improve the sound quality of the second human voice audio. This results in the second phase distribution corresponding to the first frequency band and the third phase distribution corresponding to the second frequency band. The frequency values ​​within the second frequency band are lower than those within the first frequency band.

[0133] The third phase distribution can clearly define the phase distribution benchmark of the second human voice audio. Then, when adjusting the second phase distribution based on the first and third phase distributions, over-adjustment can be avoided, thus obtaining the second target human voice audio that meets expectations.

[0134] For example, the process of adjusting the second phase distribution can be expressed by the following formula:

[0135] φ new (f)=φ orig (f)+δφ(f);

[0136] Where, φ orig (f) refers to the second phase distribution, φ new (f) refers to the adjusted second phase distribution, where δ is determined based on the first and third phase distributions. For example, the first phase distribution is the phase of the first human voice audio in the high-frequency range, and the third phase distribution is the phase of the second human voice audio in the mid-frequency range. Setting δ between 5% of the first phase distribution and 10% of the third phase distribution helps improve the sound quality, clarity, and balance of the second human voice audio.

[0137] The audio synthesis method provided in this embodiment can effectively improve the quality and blending of the target audio by analyzing and adjusting the spectrum of human voice audio, thereby helping to ensure the fullness and naturalness of the target audio in the spectrum.

[0138] This embodiment provides an audio synthesis method that can be used in electronic devices such as mobile phones and tablets. Figure 4 This is a flowchart of an audio synthesis method according to an embodiment of the present disclosure, such as... Figure 4 As shown, the process includes the following steps:

[0139] Step S401: Obtain the first vocal audio and the accompaniment audio from the reference audio.

[0140] Step S402: Determine the first loudness range based on the loudness of the first human voice audio.

[0141] Step S403: Obtain the second human voice audio corresponding to the reference audio.

[0142] Step S404: Determine the second loudness range based on the loudness of the second human voice audio.

[0143] Step S405: Based on the first comparison result between the second loudness range and the first loudness range, adjust the loudness of the second human voice audio to obtain the first target human voice audio.

[0144] Step S406: Mix the first target human voice audio and the accompaniment audio to obtain the target audio.

[0145] Step S407: Based on the spectral distribution of the reference audio, obtain the fourth spectral distribution and determine the first reference spectral energy of the main frequency band corresponding to the reference audio.

[0146] To ensure that the spectrum of the target audio and the spectrum of the reference audio are consistent and harmonious, the distribution of frequency components is understood based on the overall spectrum distribution of the reference audio, thus obtaining a fourth spectrum distribution.

[0147] The main frequency band refers to the frequency range in audio where energy is relatively concentrated. By determining the first reference spectrum energy of the main frequency band corresponding to the reference audio, it is helpful to make targeted comparisons and adjustments later.

[0148] Step S408: Based on the spectral distribution of the target audio, a fifth spectral distribution is obtained, and the second reference spectral energy of the main frequency band corresponding to the target audio is determined.

[0149] Similarly, based on the overall spectral distribution of the target audio, the distribution of its frequency components is understood, thus obtaining the fifth spectral distribution. The fifth spectral distribution reflects the energy distribution of the target audio at various frequencies.

[0150] To improve the targeting of the adjustment, the energy of the second reference spectrum corresponding to the main frequency band of the target audio is determined to clarify the spectrum distribution of the main frequency band in the target audio.

[0151] Step S409: Based on the ratio of the fourth spectral distribution to the fifth spectral distribution, adjust the spectral energy distribution of the target audio to obtain the first intermediate audio.

[0152] By comparing the fourth and fifth spectral distributions, we can understand their energy differences across frequencies. Furthermore, based on their ratio, we can clarify the relative energy relationships between the reference and target audio at various frequencies. Based on this ratio, we adjust the spectral energy distribution of the target audio so that the resulting first intermediate audio's spectral energy distribution is closer to the spectral characteristics of the reference audio than the original target audio. This helps improve the sound quality of the target audio, making its frequency energy distribution more reasonable.

[0153] For example, the process of adjusting the spectral energy distribution of the target audio based on the ratio of the fourth spectral distribution to the fifth spectral distribution can be represented by the following formula:

[0154]

[0155] Where X(f) is the spectrum of the target audio, S orig (f) is the fifth spectral distribution, S ref Y(f) is the fourth spectral distribution, and Y(f) is the spectrum of the first intermediate audio frequency.

[0156] Step S4010: Based on the comparison result of the first reference spectral energy and the second reference spectral energy, adjust the spectral energy distribution of the first intermediate audio to obtain the first target audio.

[0157] Based on the comparison results, the spectral energy distribution of the first intermediate audio is further adjusted. If the first reference spectral energy is greater than the second reference spectral energy, the energy of the first intermediate audio in the main frequency band is increased; conversely, if the first reference spectral energy is less than the second reference spectral energy, the energy of the first intermediate audio in the main frequency band is decreased, so that the obtained first target audio is closer to the reference audio in terms of spectral energy distribution.

[0158] In some optional implementation scenarios, the process of adjusting the spectral energy distribution of the first intermediate audio can be represented by the following formula:

[0159] The gain factor is a gain coefficient determined based on the difference between the energy of the first reference spectrum and the energy of the second reference spectrum.

[0160] The audio synthesis method provided in this embodiment can make the final synthesized first target audio closer to the reference audio by performing spectrum processing and dynamic range optimization on the second person's audio, thereby effectively improving the quality and consistency of the synthesized audio.

[0161] This embodiment provides an audio synthesis method that can be used in electronic devices such as mobile phones and tablets. Figure 5This is a flowchart of an audio synthesis method according to an embodiment of the present disclosure, such as... Figure 5 As shown, the process includes the following steps:

[0162] Step S501: Obtain the first vocal audio and the accompaniment audio from the reference audio.

[0163] Step S502: Determine the first loudness range based on the loudness of the first human voice audio.

[0164] Step S503: Obtain the second human voice audio corresponding to the reference audio.

[0165] Step S504: Determine the second loudness range based on the loudness of the second human voice audio.

[0166] Step S505: Based on the first comparison result between the second loudness range and the first loudness range, adjust the loudness of the second human voice audio to obtain the first target human voice audio.

[0167] Step S506: Mix the first target human voice audio and the accompaniment audio to obtain the target audio.

[0168] Step S507: Perform channel separation processing on the reference audio to obtain the first audio corresponding to the left channel and the second audio corresponding to the right channel.

[0169] To ensure that the stereo characteristics of the target audio are consistent with those of the reference audio, the reference audio is subjected to channel separation processing to obtain the first audio corresponding to the left channel and the second audio corresponding to the right channel. Subsequent channel-specific processing helps to improve the accuracy of analysis and adjustment.

[0170] Step S508: Determine the first channel energy of the first audio and the second channel energy of the second audio.

[0171] The expression for the energy of the first channel is as follows: E left =∑ n |S left [n]| 2 S left [n] refers to the first audio frequency. The expression for the energy of the second channel is as follows: E right =∑ n |S right [n] 2 S left [n] refers to the second audio frequency.

[0172] Step S509: Based on the energy of the first channel, the energy of the second channel, and the third comparison result of the energy of the first channel and the energy of the second channel, adjust the target audio to obtain the second target audio.

[0173] Based on the third comparison result of the energy of the first channel and the energy of the second channel, the energy difference between the left and right channels in the reference audio can be clearly identified. Subsequently, when the energy of the first channel and the energy of the second channel are used to adjust the target audio, the channel energy can be balanced, thereby obtaining a second target audio that is more in line with the reference audio.

[0174] In some optional implementations, step S509 above includes:

[0175] Step c1: Perform channel separation processing on the target audio to obtain the third audio corresponding to the left channel and the fourth audio corresponding to the right channel;

[0176] Step c2: Based on the third comparison result and the ratio of the energy of the first channel to the energy of the second channel, determine the first target gain corresponding to the left channel;

[0177] Step c3: Based on the third comparison result and the ratio of the energy of the second channel to the energy of the first channel, determine the second target gain corresponding to the right channel;

[0178] Step c4: Adjust the third audio through the first target gain and the fourth audio through the second target gain to obtain the second target audio.

[0179] Specifically, the target audio is subjected to channel separation processing to obtain the third audio corresponding to the left channel and the fourth audio corresponding to the right channel.

[0180] Based on the third comparison result, the energy strength relationship between the left and right channels in the reference audio can be clearly identified. Therefore, when determining the first target gain for the left channel and the second target gain for the right channel, it can be determined whether the gain adjustment should be increased or decreased. For example, if the third comparison result indicates that the energy of the first channel is greater than that of the second channel, then the gain of the right channel needs to be increased, and the gain of the left channel decreased. If the third comparison result indicates that the energy of the first channel is less than that of the second channel, then the gain of the left channel needs to be increased, and the gain of the right channel decreased.

[0181] Based on the ratio of the energy of the first channel to the energy of the second channel, the gain value of the first target gain corresponding to the left channel can be obtained. Then, based on the third comparison result, the sign of this gain value ("+" or "-") can be determined, thus obtaining the first target gain. Similarly, based on the ratio of the energy of the second channel to the energy of the first channel, the gain value of the second target gain corresponding to the right channel can be obtained. Then, based on the third comparison result, the sign of this gain value ("+" or "-") can be determined, thus obtaining the second target gain.

[0182] The third audio is adjusted by the first target gain and the fourth audio is adjusted by the second target gain to balance the output of the left and right channels of the target audio, thereby obtaining the second target audio.

[0183] In some optional examples, step c4 above includes:

[0184] Step c41: Adjust the third audio value using the first target gain to obtain the fifth audio value;

[0185] Step c42: Adjust the fourth audio frequency using the second target gain to obtain the sixth audio frequency;

[0186] Step c43: Determine the cross-correlation relationship between the first audio and the second audio.

[0187] Step c44: Based on the cross-correlation relationship, the time difference between the fifth and sixth audio frequencies is adjusted to obtain the second target audio.

[0188] Specifically, in order to reduce the time delay or phase difference between the left and right channels, the third audio is adjusted by the first target gain to obtain the fifth audio; the fourth audio is adjusted by the second target gain to obtain the sixth audio, so as to obtain the fifth audio corresponding to the left channel and the sixth audio corresponding to the right channel after the target audio is adjusted.

[0189] The process of determining the fifth audio element: _Y left [n] = X left [n]×gain left , where X left [n] represents the third audio frequency, gain left The first target gain. The process of determining the sixth audio frequency: _Y right [n] = X right [n]×gain right , where X right [n] represents the fourth audio frequency, gain right This is the gain for the second objective.

[0190] To mimic the stereo width of the reference audio, the cross-correlation relationship between the first and second audio is determined, and then the time difference between the fifth and sixth audio is adjusted. The most suitable delay value can be found to obtain the second target audio, which is closer to the stereo width of the reference audio.

[0191] The formula used to adjust the time difference between the fifth and sixth audio frequencies can be as follows:

[0192] τ opt It refers to the optimal sample delay, that is, the maximum delay value in the cross-correlation function R[k].

[0193] R[k]= n S left [n]×S ri g ht [n+k], where k is the maximum delay value to be determined.

[0194] The audio synthesis method provided in this embodiment can ensure that the stereo characteristics of the synthesized second target audio are consistent with the stereo characteristics of the reference audio, while effectively improving the overall spatial sense and listening quality, thereby effectively improving the quality of audio synthesis.

[0195] As one or more specific application embodiments of this disclosure, the process of performing audio synthesis processing on the second human voice audio can be as follows: Figure 6 As shown. After obtaining the reference audio, the reference audio is subjected to source separation processing to obtain the first vocal audio and the accompaniment audio. The first loudness range, stereo sound, and the first spectral distribution corresponding to the first frequency band of the first vocal audio are determined respectively.

[0196] After obtaining the second vocal audio, a second loudness range is determined. Based on a first comparison between the second and first loudness ranges, the loudness of the second vocal audio is dynamically adjusted to obtain a first target vocal audio, which is then mixed with the accompaniment audio to obtain the target audio. Channel separation processing is performed on the reference audio to obtain a first audio corresponding to the left channel and a second audio corresponding to the right channel. The first channel energy of the first audio and the second channel energy of the second audio are determined. Channel separation processing is performed on the target audio to obtain a third audio corresponding to the left channel and a fourth audio corresponding to the right channel. Based on the first channel energy, the second channel energy, and a third comparison between the first and second channel energy, the third and fourth audios are specifically adjusted to obtain the second target audio. Finally, based on the spectral distribution of the reference audio, the spectral distribution of the second target audio is adjusted to obtain the first target audio.

[0197] Furthermore, in order to optimize the first target audio, the spectral distribution in the first target audio corresponding to the first frequency band is adjusted based on the first spectral distribution in the first human voice audio corresponding to the first frequency band, thereby obtaining the final synthesized audio.

[0198] The aforementioned music synthesis method enables automated audio synthesis, improving synthesis efficiency. Furthermore, precise spectral processing and dynamic range optimization effectively enhance the quality and auditory experience of the final synthesized audio, making it more closely match the audio characteristics of the reference frequency and fulfilling synthesis expectations.

[0199] This embodiment also provides an audio synthesis apparatus for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0200] This embodiment provides an audio synthesis device, such as... Figure 7 As shown, it includes:

[0201] The first acquisition module 701 is used to acquire the first human voice audio and the accompaniment audio in the reference audio;

[0202] The first processing module 702 is used to determine a first loudness range based on the loudness of the first human voice audio.

[0203] The second acquisition module 703 is used to acquire the second human voice audio corresponding to the reference audio;

[0204] The second processing module 704 is used to determine a second loudness range based on the loudness of the second human voice audio.

[0205] The first adjustment module 705 is used to adjust the loudness of the second human voice audio based on a first comparison result between the second loudness range and the first loudness range, so as to obtain the first target human voice audio.

[0206] Synthesis module 706 is used to mix the first target human voice audio and the accompaniment audio to obtain the target audio.

[0207] In some alternative implementations, the first adjustment module 705 includes:

[0208] The first processing unit is used to determine the maximum signal decibel value corresponding to the first human voice audio, and to determine the signal compression threshold based on the maximum signal decibel value;

[0209] The second processing unit is used to determine the signal compression ratio based on the ratio of the second loudness range to the first loudness range;

[0210] The first adjustment unit is used to adjust the loudness of the current second human voice audio signal in response to the fact that the decibel value of the current second human voice audio signal is greater than the signal compression threshold, based on the first difference between the decibel value of the current second human voice audio signal and the signal compression threshold and the signal compression ratio, so as to obtain the first target human voice audio.

[0211] In some alternative implementations, the first adjustment module 705 further includes:

[0212] The first execution unit is configured to not perform loudness adjustment processing on the current second human voice audio signal in response to the current second human voice audio signal being less than or equal to the signal compression threshold.

[0213] In some alternative implementations, the synthesis module 706 includes:

[0214] The first unit is used to determine the first spectral distribution corresponding to the first frequency band based on the spectral distribution of the first human voice audio.

[0215] The second unit is used to determine a second spectral distribution corresponding to the first frequency band based on the spectral distribution of the first target human voice audio.

[0216] The second adjustment unit is used to adjust the second spectrum distribution based on the first spectrum distribution to obtain the second target human voice audio, which is then mixed with the accompaniment audio to obtain the target audio.

[0217] In some alternative implementations, the second adjustment unit includes:

[0218] The first determining unit is used to determine the first spectrum initial value of the first spectrum distribution;

[0219] The second execution unit is configured to determine a second difference between the current second spectrum value and the first spectrum starting value in response to the current second spectrum value being greater than the first spectrum starting value.

[0220] The third execution unit is used to adjust the current second spectrum value based on the second difference and the specified enhancement factor to obtain the third spectrum distribution;

[0221] The third processing unit is used to obtain the second target human voice audio based on the third spectral distribution.

[0222] In some optional implementations, the second adjustment unit further includes:

[0223] The fourth execution unit is used to not adjust the current second spectrum value in response to the current second spectrum value being less than or equal to the first spectrum starting value.

[0224] In some optional implementations, the third processing unit includes:

[0225] The fifth execution unit is used to determine the target harmonic energy based on the first harmonic spectrum distribution of the first human voice audio.

[0226] The sixth execution unit is used to determine the second harmonic spectrum distribution in the third spectrum distribution;

[0227] The seventh execution unit is used to determine the harmonic adjustment order based on the target harmonic energy, and adjust the second harmonic spectrum distribution based on the harmonic adjustment order and the specified adjustment factor to obtain the second target human voice audio.

[0228] In some optional implementations, the third processing unit further includes:

[0229] The eighth execution unit is used to obtain the first phase distribution based on the spectral phase of the first spectral distribution;

[0230] The ninth execution unit is used to obtain a second phase distribution corresponding to the first frequency band interval and a third phase distribution corresponding to the second frequency band interval based on the spectral phase of the third spectral distribution, wherein the frequency value in the second frequency band interval is less than the frequency value in the first frequency band interval.

[0231] The tenth execution unit is used to adjust the second phase distribution based on the first phase distribution and the third phase distribution to obtain the second target human voice audio.

[0232] In some alternative embodiments, the apparatus further includes:

[0233] The third processing module is used to obtain a fourth spectral distribution based on the spectral distribution of the reference audio, and to determine the first reference spectral energy of the main frequency band corresponding to the reference audio.

[0234] The fourth processing module is used to obtain the fifth spectral distribution based on the spectral distribution of the target audio, and to determine the second reference spectral energy of the main frequency band corresponding to the target audio.

[0235] The second adjustment module is used to adjust the spectral energy distribution of the target audio based on the ratio of the fourth spectral distribution to the fifth spectral distribution to obtain the first intermediate audio.

[0236] The third adjustment module is used to adjust the spectral energy distribution of the first intermediate audio based on the comparison result of the first reference spectral energy and the second reference spectral energy, so as to obtain the first target audio.

[0237] In some alternative embodiments, the apparatus further includes:

[0238] The separation processing module is used to perform channel separation processing on the reference audio to obtain the first audio corresponding to the left channel and the second audio corresponding to the right channel;

[0239] The fifth processing module is used to determine the energy of the first channel of the first audio and the energy of the second channel of the second audio.

[0240] The fourth adjustment module is used to adjust the target audio based on the energy of the first channel, the energy of the second channel, and the third comparison result of the energy of the first channel and the energy of the second channel to obtain the second target audio.

[0241] In some alternative implementations, the fourth adjustment module includes:

[0242] The fifth processing unit is used to perform channel separation processing on the target audio to obtain the third audio corresponding to the left channel and the fourth audio corresponding to the right channel;

[0243] The sixth processing unit is used to determine the first target gain corresponding to the left channel based on the third comparison result and the ratio of the energy of the first channel to the energy of the second channel.

[0244] The seventh processing unit is used to determine the second target gain corresponding to the right channel based on the third comparison result and the ratio of the energy of the second channel to the energy of the first channel.

[0245] The eighth processing unit is used to adjust the third audio through the first target gain and the fourth audio through the second target gain to obtain the second target audio.

[0246] In some optional implementations, the eighth processing unit includes:

[0247] The third adjustment unit is used to adjust the third audio through the first target gain to obtain the fifth audio;

[0248] The fourth adjustment unit is used to adjust the fourth audio frequency through the second target gain to obtain the sixth audio frequency.

[0249] The relationship determination unit is used to determine the cross-correlation relationship between the first audio and the second audio.

[0250] The fifth adjustment unit is used to adjust the time difference between the fifth and sixth audio frequencies based on the cross-correlation relationship to obtain the second target audio.

[0251] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0252] In this embodiment, the audio synthesis device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0253] This disclosure also provides an electronic device having the above-described features. Figure 7 The audio synthesis device shown.

[0254] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of an electronic device provided in an optional embodiment of this disclosure, such as... Figure 8As shown, the electronic device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules as needed. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 8 Take a processor 10 as an example.

[0255] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GPA), or any combination thereof.

[0256] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0257] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0258] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0259] The electronic device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 8 Taking the example of a connection between China and Israel via a bus.

[0260] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touch screen.

[0261] This disclosure also provides a computer-readable storage medium in which the methods described in this disclosure can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium after being downloaded over a network. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium may be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium may also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code that, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0262] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0263] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0264] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0265] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0266] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0267] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. An audio synthesis method, characterized in that, The method includes: Retrieve the first vocal audio and the accompaniment audio from the reference audio; Based on the loudness of the first human voice audio, a first loudness range is determined; Obtain the second human voice audio corresponding to the reference audio; Based on the loudness of the second human voice audio, a second loudness range is determined; Based on the first comparison result between the second loudness range and the first loudness range, the loudness of the second human voice audio is adjusted to obtain the first target human voice audio; The target audio is obtained by mixing the first target human voice audio and the accompaniment audio.

2. The method according to claim 1, characterized in that, The step of adjusting the loudness of the second human voice audio based on a first comparison result between the second loudness range and the first loudness range to obtain the first target human voice audio includes: Determine the maximum signal decibel value corresponding to the first human voice audio, and determine the signal compression threshold based on the maximum signal decibel value; The signal compression ratio is determined based on the ratio of the second loudness range to the first loudness range; In response to the fact that the decibel value of the current second human voice audio signal is greater than the signal compression threshold, the loudness of the current second human voice audio signal is adjusted based on the first difference between the decibel value of the current second human voice audio signal and the signal compression threshold and the signal compression ratio, to obtain a first target human voice audio, wherein the current second human voice audio signal is one of the audio signals in the second human voice audio.

3. The method according to claim 2, characterized in that, The step of adjusting the loudness of the second human voice audio based on a first comparison result between the second loudness range and the first loudness range to obtain the first target human voice audio further includes: In response to the fact that the decibel value of the current second human voice audio signal is less than or equal to the signal compression threshold, no loudness adjustment processing is performed on the current second human voice audio signal.

4. The method according to claim 1, characterized in that, The process of mixing the first target human voice audio and the accompaniment audio to obtain the target audio includes: Based on the spectral distribution of the first human voice audio, a first spectral distribution corresponding to the first frequency band is determined; Based on the spectral distribution of the first target human voice audio, a second spectral distribution corresponding to the first frequency band is determined; Based on the first spectral distribution, the second spectral distribution is adjusted to obtain a second target human voice audio, which is then mixed with the accompaniment audio to obtain the target audio.

5. The method according to claim 4, characterized in that, The step of adjusting the second spectral distribution based on the first spectral distribution to obtain the second target human voice audio includes: Determine the first initial value of the first spectrum distribution; In response to the fact that the current second spectrum value is greater than the first spectrum starting value, a second difference between the current second spectrum value and the first spectrum starting value is determined, wherein the current second spectrum value is one of the spectrum values ​​in the second spectrum distribution; Based on the second difference and the specified enhancement factor, the current second spectrum value is adjusted to obtain a third spectrum distribution; Based on the third spectral distribution, the second target human voice audio is obtained.

6. The method according to claim 5, characterized in that, The step of adjusting the second spectral distribution based on the first spectral distribution to obtain the second target human voice audio further includes: In response to the current second spectrum value being less than or equal to the first spectrum starting value, the current second spectrum value is not adjusted.

7. The method according to claim 5 or 6, characterized in that, The process of obtaining the second target human voice audio based on the third spectral distribution includes: The target harmonic energy is determined based on the first harmonic spectrum distribution of the first human voice audio. Determine the second harmonic spectrum distribution in the third spectrum distribution; Based on the target harmonic energy, determine the harmonic modulation order; Based on the harmonic adjustment order and the specified adjustment factor, the second harmonic spectrum distribution is adjusted to obtain the second target human voice audio.

8. The method according to claim 7, characterized in that, The process of obtaining the second target human voice audio based on the third spectral distribution further includes: Based on the spectral phase of the first spectral distribution, a first phase distribution is obtained; Based on the spectral phase of the third spectral distribution, a second phase distribution corresponding to the first frequency band interval and a third phase distribution corresponding to the second frequency band interval are obtained, wherein the frequency value in the second frequency band interval is less than the frequency value in the first frequency band interval. Based on the first phase distribution and the third phase distribution, the second phase distribution is adjusted to obtain the second target human voice audio.

9. The method according to claim 1, characterized in that, The method further includes: Based on the spectral distribution of the reference audio, a fourth spectral distribution is obtained, and the first reference spectral energy of the main frequency band corresponding to the reference audio is determined. Based on the spectral distribution of the target audio, a fifth spectral distribution is obtained, and the second reference spectral energy of the main frequency band corresponding to the target audio is determined; Based on the ratio of the fourth spectral distribution to the fifth spectral distribution, the spectral energy distribution of the target audio is adjusted to obtain the first intermediate audio. Based on the comparison results of the first reference spectral energy and the second reference spectral energy, the spectral energy distribution of the first intermediate audio is adjusted to obtain the first target audio.

10. The method according to claim 1, characterized in that, The method further includes: The reference audio is subjected to channel separation processing to obtain a first audio corresponding to the left channel and a second audio corresponding to the right channel; Determine the first channel energy of the first audio and the second channel energy of the second audio; Based on the energy of the first channel, the energy of the second channel, and a third comparison result of the energy of the first channel and the energy of the second channel, the target audio is adjusted to obtain the second target audio.

11. The method according to claim 10, characterized in that, The step of adjusting the target audio based on the first channel energy, the second channel energy, and a third comparison result between the first channel energy and the second channel energy to obtain the second target audio includes: The target audio is subjected to channel separation processing to obtain a third audio corresponding to the left channel and a fourth audio corresponding to the right channel; Based on the third comparison result and the ratio of the energy of the first channel to the energy of the second channel, the first target gain corresponding to the left channel is determined; Based on the third comparison result and the ratio of the energy of the second channel to the energy of the first channel, the second target gain corresponding to the right channel is determined. The third audio is adjusted by the first target gain, and the fourth audio is adjusted by the second target gain to obtain the second target audio.

12. The method according to claim 11, characterized in that, The step of adjusting the third audio through the first target gain and adjusting the fourth audio through the second target gain to obtain the second target audio includes: The third audio is obtained by adjusting the first target gain; The fourth audio signal is obtained by adjusting the second target gain; Determine the cross-correlation relationship between the first audio and the second audio; Based on the cross-correlation relationship, the time difference between the fifth and sixth audio frequencies is adjusted to obtain the second target audio.

13. An audio synthesis device, characterized in that, The device includes: The first acquisition module is used to acquire the first human voice audio and the accompaniment audio in the reference audio; The first processing module is used to determine a first loudness range based on the loudness of the first human voice audio. The second acquisition module is used to acquire a second human voice audio corresponding to the reference audio; The second processing module is used to determine a second loudness range based on the loudness of the second human voice audio. The first adjustment module is used to adjust the loudness of the second human voice audio based on a first comparison result between the second loudness range and the first loudness range, so as to obtain the first target human voice audio. A synthesis module is used to mix the first target human voice audio and the accompaniment audio to obtain the target audio.

14. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the audio synthesis method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the audio synthesis method according to any one of claims 1 to 12.

16. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the audio synthesis method according to any one of claims 1 to 12.