Human voice reproduction method and human voice reproduction system

By combining the fast synthesis module and the refined synthesis module, the problem of time-consuming voice replication in the in-vehicle intelligent voice system is solved, and fast and highly similar voice replacement is achieved, improving the user experience.

CN116343742BActive Publication Date: 2025-09-23CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310179438.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-27
Publication Date
2025-09-23
Estimated Expiration
2043-02-27

AI Technical Summary

Technical Problem

The existing voice replication technology of in-vehicle intelligent voice systems takes a long time, and it is difficult to quickly set the user's preferred voice as the voice of the voice assistant, resulting in a poor user experience.

Method used

By combining the fast synthesis module and the refined synthesis module, and using the human voice database and cloud server for rapid matching and training, replica audio is generated to achieve fast synthesis and refined synthesis, improving the user experience in terms of replication speed and similarity respectively.

Benefits of technology

It achieves the rapid replacement of car announcement sounds in a short period of time, and ensures a high similarity between the announcement sounds and the user's preferred voice, thus improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116343742B_ABST
    Figure CN116343742B_ABST
Patent Text Reader

Abstract

The present invention relates to a human voice replication method, comprising the following steps: receiving human voice data to be replicated; calling first human voice data with the highest matching degree with the human voice data to be replicated from a human voice database; performing sound replication training based on the human voice data to be replicated to obtain second human voice data; if the sound replication training is not completed, synthesizing replication audio based on the first human voice data, and if the sound replication training is completed, synthesizing replication audio based on the second human voice data; and outputting replication audio. The present invention also proposes a human voice replication system. The present invention has the following characteristics: performing fast synthesis and refined synthesis, being able to use the rapidly synthesized replication audio to quickly replace the car's announcement sound, thereby improving the user experience in terms of replication speed, being able to use the refined synthesized replication audio to ensure a high similarity between the car's announcement sound and the user's preferred human voice, thereby improving the user experience in terms of replication similarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an automobile, and in particular to a human voice reproduction method and a human voice reproduction system. Background Art

[0002] In-vehicle voice interaction can greatly help users control their vehicles. However, in current in-vehicle intelligent voice systems, the voice assistant's voice, timbre and other characteristics are the same in the same car model, or there are only a few options to choose from, which poses a technical problem that makes it difficult to meet the diverse and personalized needs of users.

[0003] The existing voice replication technology can replicate the collected voice samples through a voice replication training model, but the voice replication training takes a long time. When the user wants to set his or her favorite voice as the voice of the voice assistant, after the user completes the voice sample collection, he or she needs to wait for a long time (about 20 minutes). Only after the voice replication training is completed can the replicated voice be set as the voice of the voice assistant. There is a technical problem that it is difficult to quickly set the user's favorite voice as the voice of the voice assistant, and the user experience is poor. Summary of the Invention

[0004] The purpose of the present invention is to provide a human voice replication method and a human voice replication system to alleviate or eliminate at least one of the above-mentioned technical problems.

[0005] The human voice reproduction method of the present invention comprises the following steps:

[0006] Receive the human voice data to be replicated;

[0007] Retrieving first vocal data having the highest matching degree with the vocal data to be replicated from a vocal database;

[0008] Performing voice replication training based on the to-be-replicated human voice data to obtain second human voice data;

[0009] If the sound replication training is not completed, synthesizing the replication audio based on the first vocal data; if the sound replication training is completed, synthesizing the replication audio based on the second vocal data;

[0010] The reproduced audio is output.

[0011] Optionally, the calling of the first vocal data having the highest matching degree with the vocal data to be replicated from the vocal database includes the following steps:

[0012] Analyze and extract the vocal features of the to-be-replicated vocal data, and retrieve the first vocal data having the highest matching degree with the vocal features from a vocal database.

[0013] Optionally, the analyzing and extracting the vocal features of the vocal data to be replicated includes the following steps: intercepting a portion of the vocal data to be replicated as the vocal data to be analyzed, and analyzing and extracting the vocal features of the vocal data to be analyzed.

[0014] Optionally, the following steps are also included:

[0015] Receive the text to be broadcast;

[0016] Determine whether the voice reproduction training is completed, and if so, synthesize the reproduction audio based on the second human voice data and the to-be-broadcast text; otherwise, synthesize the reproduction audio based on the first human voice data and the to-be-broadcast text;

[0017] The reproduced audio is output.

[0018] A human voice replication system described in the present invention includes a cloud server, which is used to: receive human voice data to be replicated; call first human voice data with the highest matching degree with the human voice data to be replicated from a human voice database; perform sound replication training based on the human voice data to be replicated to obtain second human voice data; if the sound replication training is not completed, synthesize replicated audio based on the first human voice data; if the sound replication training is completed, synthesize replicated audio based on the second human voice data; and output the replicated audio.

[0019] Optionally, the cloud server includes a fast synthesis module and a refined synthesis module; the fast synthesis module is used to call the first vocal data with the highest matching degree with the vocal data to be replicated from the vocal database, and synthesize the replicated audio based on the first vocal data; the refined synthesis module is used to perform sound replication training based on the vocal data to be replicated, obtain second vocal data, and synthesize the replicated audio based on the first vocal data.

[0020] Optional equipment also includes microphone, car head unit and speaker;

[0021] The microphone is used to obtain the human voice data to be reproduced input by the user;

[0022] The vehicle computer is used to send the to-be-replicated human voice data to the cloud server, receive the reproduced audio output by the cloud server, and control the speaker to broadcast the reproduced audio;

[0023] The speaker is used to broadcast the reproduced audio.

[0024] The present invention has the following characteristics: it can perform fast synthesis and refined synthesis, and can use the fast synthesized replica audio to quickly replace the car's broadcast sound, thereby improving the user's experience in terms of replication speed; it can use the refined synthesized replica audio to ensure a high similarity between the car's broadcast sound and the user's favorite human voice, thereby improving the user's experience in terms of replication similarity. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 Flowchart of the human voice recording process described in the specific implementation manner;

[0026] Figure 2 is a flowchart of the vocal synthesis process described in the specific implementation manner;

[0027] Figure 3 Schematic diagram of the human voice replication system described in the specific implementation manner. DETAILED DESCRIPTION

[0028] The following describes the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art will readily appreciate the other advantages and benefits of the present invention from the disclosure herein. The present invention may also be implemented or applied through various other specific embodiments, and the various details in this specification may be modified or altered based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are intended only to illustrate the present invention and are not intended to limit the scope of protection of the present invention.

[0029] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.

[0030] The present invention proposes a method for replicating human voice, comprising the following steps: receiving human voice data to be replicated; calling the first human voice data with the highest matching degree with the human voice data to be replicated from a human voice database; performing sound replication training based on the human voice data to be replicated to obtain second human voice data; if the sound replication training is not completed, synthesizing the replicated audio based on the first human voice data; if the sound replication training is completed, synthesizing the replicated audio based on the second human voice data; and outputting the replicated audio. The above-mentioned technical solution is adopted to perform fast synthesis and refined synthesis, and the fast synthesized replicated audio can be used to quickly replace the car's announcement sound, thereby improving the user experience in terms of replication speed, and the refined synthesized replicated audio can be used to ensure a high similarity between the car's announcement sound and the user's preferred human voice, thereby improving the user experience in terms of replication similarity.

[0031] In some embodiments, retrieving the first vocal data that best matches the vocal data to be replicated from the vocal database includes the following steps: analyzing and extracting vocal features of the vocal data to be replicated, and retrieving the first vocal data that best matches the vocal features from the vocal database. As a specific example, the vocal features are voiceprints, and the first vocal data that best matches the voiceprint of the vocal data to be replicated is found and retrieved from the vocal database.

[0032] In some embodiments, analyzing and extracting vocal features from the to-be-replicated vocal data includes the following steps: intercepting a portion of the to-be-replicated vocal data as the to-be-analyzed vocal data, and analyzing and extracting vocal features from the to-be-analyzed vocal data. The above technical solution can improve the speed of analysis and extraction, and thus the speed of vocal replication. As a specific example, during the vocal recording process, three sentences are recorded, and the rapid synthesis module only analyzes and extracts the vocal features of one of the three sentences.

[0033] In some embodiments, the voice replication method further comprises the following steps: receiving a text to be broadcast; determining whether voice replication training is complete; if so, synthesizing a replicated audio based on the second voice data and the text to be broadcast; otherwise, synthesizing a replicated audio based on the first voice data and the text to be broadcast; and outputting the replicated audio. As a specific example, the first voice data and the second voice data include loudness data, pitch data, and timbre data.

[0034] As a specific embodiment, the voice replication method is applied to a voice replication system, which includes a cloud server, a microphone, a car computer and a speaker. The cloud server is provided with a fast synthesis module and a refined synthesis module; the voice replication method includes a voice recording process and a voice synthesis process.

[0035] like Figure 1 As shown in FIG, the voice recording process includes the following steps:

[0036] S101: Entering the voice recording interface of the vehicle terminal through the vehicle terminal screen;

[0037] S102: The screen of the vehicle computer displays the voice recording text, the user reads the voice recording text aloud, the reading audio is picked up by the microphone, and then the reading audio is sent to the vehicle computer;

[0038] S103: The vehicle computer recognizes the reading audio and compares and analyzes the recognition result with the voice recording text to determine whether the reading audio meets the requirements. If not, the user is guided to re-record the reading audio until the reading audio meets the requirements.

[0039] S104: The vehicle computer sends the voice data to be replicated to the cloud server via the network. The cloud server distributes the voice data to be replicated to the fast synthesis module and the refined synthesis module. The voice data to be replicated typically includes voice recording text and reading audio.

[0040] S105: The rapid synthesis module intercepts a portion of the vocal data to be replicated as the vocal data to be analyzed, analyzes and extracts the vocal features of the vocal data to be analyzed, and calls the first vocal data with the highest matching degree with the vocal features from the vocal database pre-stored in the cloud server; the refined synthesis module analyzes and extracts the vocal features of the vocal data to be replicated, performs training based on the sound replication training model and the extracted vocal features, and obtains the second vocal data.

[0041] like Figure 2 As shown in Figure 2, the vocal synthesis process includes the following steps:

[0042] S201: The vehicle computer sends the message to be broadcast to the cloud server via the network;

[0043] S202: The cloud server first determines whether the refined synthesis module has completed training. If the refined synthesis module has not completed training, the cloud server sends the to-be-broadcast copy to the fast synthesis module. The fast synthesis module synthesizes a replica audio based on the first human voice data and the to-be-broadcast copy. If the refined synthesis module has completed training, the cloud server sends the to-be-broadcast copy to the refined synthesis module. The refined synthesis module synthesizes a replica audio based on the second human voice data and the to-be-broadcast copy.

[0044] S203: The cloud server sends the reproduced audio to the vehicle computer, and broadcasts the reproduced audio through the speaker.

[0045] In the specific implementation, the vehicle computer is provided with a first receiving module, a second receiving module, a first output module, a second output module, a speech recognition module and a speech analysis module. The first receiving module can receive the reading audio picked up by the microphone, the second receiving module can receive the reproduced audio output by the cloud server, the first output module can output the human voice data to be reproduced to the cloud server, the second output module can output the reproduced audio to the speaker for broadcast, the speech recognition module can recognize the reading audio, and the speech analysis module can compare and analyze the recognition result of the reading audio with the human voice input text to determine whether the reading audio meets the requirements.

[0046] In practice, a voice-replicating app can be developed on the vehicle's operating system. Users can operate the app by inputting commands through the vehicle's screen. After opening the voice-replicating app and selecting voice recording, the vehicle's screen will enter the voice recording interface, where a voice recording text will be displayed. The vehicle's screen and voice assistant will guide the user through voice recording via text and voice prompts, reminding them to record in a quiet environment. The user then reads aloud based on the displayed voice recording text. The audio will be picked up by the car's microphone and fed into the vehicle. Users can also have family members, friends, etc. record their voices and set their voices as the voice of the voice assistant.

[0047] In specific implementation, after the car computer receives the reading audio, it needs to recognize the reading audio and compare and analyze the recognition result with the voice recording copy. If the recognition result of the user's reading audio is inconsistent with the voice recording copy, it will reduce the accuracy of the voice reproduction system's analysis of the human voice characteristics. Therefore, after the user's recorded reading audio is sent to the car computer, the car computer will test the quality of the reading audio to determine whether it meets the requirements. The test indicators include: sound clarity and consistency with the voice recording copy. If the test results do not meet the target requirements, re-entry is required. At this time, the car computer's screen and voice assistant will guide the user to re-enter through text and voice broadcasts, reminding the user to record in a quiet environment. At the same time, the car computer's screen displays the voice recording copy that needs to be re-entered.

[0048] In practice, the car's audio system is equipped with a voice assistant. During daily use, the car's microphone transmits the user's voice commands to the system, where voice recognition technology converts these commands into command code and sends it to the command execution mechanism. Simultaneously, the system sends the text the voice assistant needs to speak to a cloud server's fast synthesis module or refined synthesis module. The cloud server synthesizes the reproduced audio and sends it to the car system, where it is played through the speakers. Once the refined synthesis module completes the entire voice reproduction process, the car's audio system uses the reproduced audio synthesized by the refined synthesis module for subsequent use.

[0049] A human voice replication system of the present invention includes a cloud server, which is used to: receive human voice data to be replicated; call first human voice data with the highest matching degree with the human voice data to be replicated from a human voice database; perform sound replication training based on the human voice data to be replicated to obtain second human voice data; if the sound replication training is not completed, synthesize replicated audio based on the first human voice data; if the sound replication training is completed, synthesize replicated audio based on the second human voice data; and output the replicated audio.

[0050] In some embodiments, the cloud server includes a fast synthesis module and a refined synthesis module; the fast synthesis module is used to retrieve the first vocal data that has the highest match with the vocal data to be replicated from the vocal database, and synthesize the replicated audio based on the first vocal data; the refined synthesis module is used to perform sound replication training based on the vocal data to be replicated, obtain second vocal data, and synthesize the replicated audio based on the first vocal data. In specific implementations, the cloud server also includes a receiving module, a judging module, and an output module. The receiving module is capable of receiving the vocal data to be replicated and the text to be broadcast, the output module is capable of outputting the replicated audio, and the judging module is capable of judging whether the refined synthesis module has completed training.

[0051] In some embodiments, the voice replication system also includes a microphone, a vehicle computer and a speaker; the microphone is used to obtain the voice data to be replicated input by the user; the vehicle computer is used to send the voice data to be replicated to the cloud server, receive the replicated audio output by the cloud server, and control the speaker to play the replicated audio; the speaker is used to play the replicated audio; the vehicle computer includes a screen, which can display the voice input text and input operation instructions.

[0052] In some embodiments, the vehicle computer is an onboard control device that can receive external control commands and control in-vehicle functional devices, including in-vehicle multimedia devices and in-vehicle display screens. The cloud server can be a cloud software platform that can analyze and process the acquired information during operation.

[0053] Using the aforementioned voice replication system, users can record their voices into the car's computer via the in-car microphone. The car then sets the voice assistant's voice to a voice similar to the user's, and the voice assistant uses a voice similar to the user's recorded voice for daily announcements. Furthermore, by setting up a fast synthesis module and a refined synthesis module in the cloud server, the fast synthesis module can quickly complete voice replication, allowing users to experience the voice immediately after completing voice recording, which can enhance the user experience. After the refined synthesis module completes training, it can obtain a replica audio that is closer to the user's recorded voice, which also enhances the user experience.

[0054] The above embodiments are merely preferred embodiments for fully illustrating the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the scope of protection of the present invention. In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example" or "some examples" etc. mean that the specific features, structures, materials or characteristics of the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification.

Claims

1. A method for reproducing human voice, characterized in that: The following steps are involved: Receive the human voice data to be replicated; Retrieving first vocal data having the highest matching degree with the vocal data to be replicated from a vocal database; Performing voice replication training based on the to-be-replicated human voice data to obtain second human voice data; If the sound replication training is not completed, synthesizing the replication audio based on the first vocal data; if the sound replication training is completed, synthesizing the replication audio based on the second vocal data; The reproduced audio is output.

2. The method for reproducing human voice according to claim 1, wherein: The step of retrieving the first vocal data having the highest matching degree with the vocal data to be replicated from the vocal database comprises the following steps: Analyze and extract the vocal features of the to-be-replicated vocal data, and retrieve the first vocal data having the highest matching degree with the vocal features from a vocal database.

3. The method for reproducing human voice according to claim 2, wherein: The analyzing and extracting the vocal features of the vocal data to be replicated includes the following steps: intercepting a portion of the vocal data to be replicated as the vocal data to be analyzed, and analyzing and extracting the vocal features of the vocal data to be analyzed.

4. The method for reproducing human voice according to claim 1, wherein: The following steps are also included: Receive the report to be broadcast; Determine whether the voice reproduction training is completed, and if so, synthesize the reproduction audio based on the second human voice data and the to-be-broadcast text; otherwise, synthesize the reproduction audio based on the first human voice data and the to-be-broadcast text; The reproduced audio is output.

5. A human voice reproduction system, characterized in that: The system includes a cloud server configured to: receive vocal data to be replicated; retrieve first vocal data having the highest matching degree with the vocal data to be replicated from a vocal database; perform sound replication training based on the vocal data to be replicated to obtain second vocal data; synthesize replicated audio based on the first vocal data if the sound replication training is not completed, and synthesize replicated audio based on the second vocal data if the sound replication training is completed; The reproduced audio is output.

6. The human voice reproduction system according to claim 5, characterized in that: The cloud server includes a fast synthesis module and a refined synthesis module; the fast synthesis module is used to call the first vocal data with the highest matching degree with the vocal data to be replicated from the vocal database, and synthesize the replicated audio based on the first vocal data; the refined synthesis module is used to perform sound replication training based on the vocal data to be replicated, obtain second vocal data, and synthesize the replicated audio based on the first vocal data.

7. The human voice reproduction system according to claim 5, characterized in that: It also includes microphones, car computers and speakers; The microphone is used to obtain the human voice data to be reproduced input by the user; The vehicle computer is used to send the to-be-replicated human voice data to the cloud server, receive the reproduced audio output by the cloud server, and control the speaker to broadcast the reproduced audio; The speaker is used to broadcast the reproduced audio.

Citation Information

Patent Citations

  • Voice endpoint detection method and system based on deep learning

    CN110706694A

  • Audio recognition method and device, electronic equipment and computer readable storage medium

    CN114512117A