Speech synthesis method, electronic device, and computer-readable storage medium

By extracting and fusing voiceprint feature information and using a unified speech synthesis model to generate personalized speech, the problem of long recording time and high cost of personalized speech in electronic devices is solved, realizing fast and low-cost personalized speech generation and improving user experience.

CN116665635BActive Publication Date: 2026-07-31HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2022-02-18
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In the existing technology, the voice interaction function of electronic devices lacks personalized voice settings, resulting in an insufficient user experience, and personalized voice recording requires a lot of time and cost.

Method used

By extracting voiceprint feature information from users' original speech data, personalized target speech data is generated using a unified speech synthesis model, reducing recording time and cost, and supporting the adjustment and fusion of multiple voiceprint feature information.

Benefits of technology

It enables rapid recording and efficient generation of personalized voice messages, improving user experience and reducing collection time and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116665635B_ABST
    Figure CN116665635B_ABST
Patent Text Reader

Abstract

This application relates to the field of terminal technology, and particularly to speech synthesis methods, electronic devices, and computer-readable storage media. In this method, a first electronic device extracts first voiceprint feature information, including at least one of timbre, prosody, and style, from at least one sentence of first raw speech data input by a first user. Based on a speech synthesis model, it directly generates first target speech data according to the first voiceprint feature information and the content of a first target text. This eliminates the need for the first user to input large or long amounts of speech data, and also eliminates the need for precise matching between the user's input speech data and the specified text content, thus reducing the time and cost of speech data acquisition. Furthermore, it eliminates the need to train a speech synthesis model corresponding to the first user based on their speech data, reducing model training time and consequently reducing the recording time for personalized speech. Personalized speech recording can be as short as a few seconds, improving recording efficiency and enhancing user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of terminal technology, and in particular relates to speech synthesis methods, electronic devices and computer-readable storage media. Background Technology

[0002] With the rapid development of electronic technology, more and more electronic devices are equipped with voice interaction functions, enabling users to interact with the devices via voice or for the devices to play content through voice. Currently, during voice interaction between users and electronic devices, the devices typically use one or more default voice settings, such as the system default voice or a celebrity's voice. How to intelligently set personalized voices for voice interaction functions is a problem that urgently needs to be solved. Summary of the Invention

[0003] This application provides a speech synthesis method, an electronic device, and a computer-readable storage medium, which can realize intelligent and efficient personalized voice settings and improve the user experience of voice interaction.

[0004] In a first aspect, embodiments of this application provide a speech synthesis method applied to a first electronic device, the method including:

[0005] The first electronic device acquires the first user's first raw voice data;

[0006] The first electronic device acquires the first voiceprint feature information corresponding to the first original voice data, and the first voiceprint feature information includes at least one of the timbre, rhythm and style related to the first user.

[0007] The first electronic device generates first target speech data corresponding to the first target text content based on the first voiceprint feature information and the first target text content using a speech synthesis model.

[0008] In the aforementioned speech synthesis method, the first electronic device can extract first voiceprint feature information, including at least one of timbre, prosody, and style, from at least one sentence of first raw speech data input by the first user. Based on the first voiceprint feature information and the first target text content, it can directly generate personalized first target speech data using a speech synthesis model. This eliminates the need for the first user to input large or long amounts of speech data, and also eliminates the need for precise matching between the user's input speech data and the specified text content, thus reducing speech data acquisition time and costs. Furthermore, it eliminates the need to train a speech synthesis model corresponding to the first user based on their speech data, reducing model training time and consequently reducing the recording time for personalized speech. Personalized speech recordings can be as short as a few seconds, improving recording efficiency and enhancing user experience.

[0009] Understandably, the first voiceprint feature information can be a multi-dimensional vector (e.g., a 256-dimensional vector), and it can be one or more of the following: timbre, prosody, and style, which do not contain speech content. For example, the first voiceprint feature information can be a multi-dimensional vector representing timbre, or a multi-dimensional vector representing style, or a multi-dimensional vector representing both timbre and prosody, or a multi-dimensional vector representing timbre and style, or a multi-dimensional vector representing timbre, prosody, and style, and so on.

[0010] For example, the first electronic device may generate first target speech data corresponding to the first target text content based on the first voiceprint feature information and the first target text content using a speech synthesis model, which may include:

[0011] The first electronic device inputs the first voiceprint feature information and the first target text content into the speech synthesis model to generate the first target speech data.

[0012] In the speech synthesis method provided by this implementation, the speech synthesis model can directly process the input first voiceprint feature information and the first target text content to generate the first target speech data corresponding to the first target text content. That is, the first voiceprint feature information corresponding to the first original speech data is extracted directly from the first original speech data, not determined based on the model parameters of the speech synthesis model. Therefore, for all users, a unified speech synthesis model can be used to generate target speech data, without needing to train personalized speech synthesis models for each user separately. This reduces adaptive training of the model and shortens the recording time for personalized speech, allowing recordings to be as short as a few seconds.

[0013] In one possible implementation, after the first electronic device acquires the first voiceprint feature information corresponding to the first original voice data, the method may further include:

[0014] The first electronic device selects the second voiceprint feature information and adjusts the first voiceprint feature information based on the second voiceprint feature information to obtain the third voiceprint feature information; wherein, the third voiceprint feature information is a fusion of the first voiceprint feature information and the second voiceprint feature information;

[0015] The first electronic device generates first target speech data corresponding to the first target text content based on the first voiceprint feature information and the first target text content, using a speech synthesis model. Specifically, this may include:

[0016] The first electronic device generates first target speech data corresponding to the first target text content based on the speech synthesis model, according to the third voiceprint feature information and the first target text content.

[0017] In the speech synthesis method provided by this implementation, the first electronic device can use one or more second voiceprint feature information to adjust the first voiceprint feature information to obtain third voiceprint feature information. This third voiceprint feature information not only possesses at least one of the characteristics such as timbre, rhythm, and style corresponding to the first voiceprint feature information, but also possesses at least one of the characteristics such as timbre, rhythm, and style corresponding to the second voiceprint feature information. This can enhance the sound characteristics of the first electronic device in voice interaction or content playback, greatly increase the enjoyment of the first electronic device in voice interaction or content playback, and improve the user experience.

[0018] The second voiceprint feature information can be the voiceprint feature information corresponding to the voice input already present in the first electronic device, the voiceprint feature information corresponding to voice input subsequently downloaded by the first user, the voiceprint feature information corresponding to personalized voice input previously recorded by the first user, and so on. In other words, the second voiceprint feature information is the voiceprint feature information corresponding to the existing voice input in the first electronic device.

[0019] In one example, the first electronic device selects second voiceprint feature information and adjusts the first voiceprint feature information based on the second voiceprint feature information to obtain third voiceprint feature information, which may specifically include:

[0020] The first electronic device displays the identifier of the first voiceprint feature information and the identifier of the fourth voiceprint feature information;

[0021] In response to the adjustment operation of the first user, the first electronic device determines the second voiceprint feature information and the second weight of the second voiceprint feature information from the fourth voiceprint feature information;

[0022] The first electronic device determines the first weight of the first voiceprint feature information according to the second weight, and generates the third voiceprint feature information according to the first voiceprint feature information, the first weight, the second voiceprint feature information and the second weight.

[0023] In the speech synthesis method provided by this implementation, the second voiceprint feature information can be selected by the first user, so that the first user can adjust the first voiceprint feature information according to actual needs by selecting the second voiceprint feature information, thereby improving the user experience.

[0024] Optionally, the first user can directly select the second voiceprint feature information by adjusting the weight of the fourth voiceprint feature information.

[0025] For example, the first electronic device can display the identifiers of the first and fourth voiceprint feature information, along with corresponding edit buttons, on a display interface. When the first user clicks the edit button corresponding to the first voiceprint feature information, the first electronic device can display the voice management interface corresponding to the first voiceprint feature information. The voice management interface can display the identifier of the fourth voiceprint feature information and the weight bars corresponding to each fourth voiceprint feature information. Initially, the weight bars corresponding to each fourth voiceprint feature information can all be 0. When the first user wants to select a certain fourth voiceprint feature information as the second voiceprint feature information, the first user can adjust the weight bar corresponding to that fourth voiceprint feature information to set the second weight of the fourth voiceprint feature information to be greater than 0.

[0026] Based on the adjustment operation of the first user, the first electronic device can determine the second voiceprint feature information and the second weight of the second voiceprint feature information from the fourth voiceprint feature information, and determine the first weight of the first voiceprint feature information based on the second weight of the second voiceprint feature information (for example, the first weight = 1 - the sum of all the second weights). Subsequently, the first electronic device can generate the third voiceprint feature information based on the first voiceprint feature information, the first weight, the second voiceprint feature information and the second weight. For example, the third voiceprint feature information can be obtained by weighted summing of the first voiceprint feature information and the second voiceprint feature information using the first weight and the second weight.

[0027] In another example, the first electronic device selects second voiceprint feature information and adjusts the first voiceprint feature information based on the second voiceprint feature information to obtain third voiceprint feature information, which may include:

[0028] The first electronic device displays the identifier of the first voiceprint feature information and the identifier of the fourth voiceprint feature information;

[0029] In response to the first user's selection operation of the fourth voiceprint feature information, the first electronic device determines the second voiceprint feature information;

[0030] When the second voiceprint feature information is one, the first electronic device displays an adjustment bar, one end of which is the first voiceprint feature information and the other end is the second voiceprint feature information;

[0031] In response to the adjustment operation of the first user on the adjustment bar, the first electronic device determines the first weight of the first voiceprint feature information and the second weight of the second voiceprint feature information, and generates the third voiceprint feature information based on the first voiceprint feature information, the first weight, the second voiceprint feature information and the second weight.

[0032] In the speech synthesis method provided by this implementation, the first user can also first select the second voiceprint feature information, and then adjust the first weight of the second voiceprint feature information and the second weight of the second voiceprint feature information.

[0033] For example, the first electronic device can display identifiers for the first and fourth voiceprint feature information, along with corresponding editing buttons, on a display interface. When the first user clicks the editing button corresponding to the first voiceprint feature information, the first electronic device can display a voice management interface for that first voiceprint feature information. This voice management interface can display identifiers for the fourth voiceprint feature information and corresponding selection boxes for each fourth voiceprint feature information. The first user can then select one or more second voiceprint feature information pieces through the selection boxes to adjust the first voiceprint feature information.

[0034] Specifically, when a first user selects a second voiceprint feature to adjust the first voiceprint feature, the first electronic device can display an adjustment bar on the display interface. One end of the adjustment bar displays the first voiceprint feature, and the other end displays the second voiceprint feature. The first user can slide the adjustment bar to adjust the second weight of the second voiceprint feature and the first weight of the first voiceprint feature.

[0035] When a first user selects multiple second voiceprint feature information, the first electronic device can display the identifier of each second voiceprint feature information and the corresponding weight bar on the display interface. The first user can set the second weight of each second voiceprint feature information through the weight bar.

[0036] It should be understood that the third voiceprint feature information can be stored as a new voiceprint feature information in the first electronic device; that is, the third voiceprint feature information and the first voiceprint feature information can be two voiceprint feature information that coexist. When the first user chooses to use the third voiceprint feature information for voice interaction or content playback, the first electronic device can synthesize the third voiceprint feature information with the target text content to obtain the target speech data corresponding to the target text content, and then perform voice interaction or content playback.

[0037] Alternatively, the third voiceprint feature information can be directly used as the adjusted voiceprint feature information of the first voiceprint feature information. That is, when the first user chooses to use the first voiceprint feature information for voice interaction or content playback, the first electronic device can synthesize the adjusted third voiceprint feature information with the first target text content to obtain the first target speech data corresponding to the first target text content, and then use it for voice interaction or content playback.

[0038] Optionally, the first original voice data is the voice data input during the voice interaction; or, the first original voice data is the voice data in the uploaded audio file.

[0039] Optionally, the first original speech data is speech data with a duration longer than a preset duration and containing arbitrary content; or, the first original speech data is speech data containing specified text content.

[0040] "Any content" means that the text content contained in the first original voice data can be any text content, that is, the text content contained in the first original voice data can be determined by the first user, and is not a specific text content. The preset duration can be set according to the actual scenario, for example, the preset duration can be set to any value such as 3 seconds.

[0041] In the speech synthesis method provided by this implementation, the first electronic device can record personalized speech based on a short first original speech data containing any text content. This eliminates the need for the first user to input a large or long amount of speech data, and also eliminates the need for the first user's input speech data to be precisely matched with the specified text content. This reduces the time and cost of voice data acquisition, allowing personalized speech recording to be as short as a few seconds, thereby improving the efficiency of personalized speech recording and enhancing the user experience.

[0042] In one possible implementation, the first electronic device acquiring the first user's raw voice data may include:

[0043] The first electronic device receives a voice command input by the first user, the voice command including an instruction to record the user's voice intent;

[0044] The first electronic device identifies the voice command, which includes the user's voice intent, as the first raw voice data.

[0045] In the speech synthesis method provided by this implementation, when online recording of personalized speech is initiated based on the voice command of the first user (e.g., a voice command containing an implicit recording intent), the first electronic device can directly determine the voice command containing the recording intent (e.g., "Xiaoyi Xiaoyi, learn my speech") as the first original voice data input by the first user in the personalized speech recording. The first user does not need to input any other voice data, which simplifies the input process of the first original voice data, allows the user to quickly experience the speech synthesis features without interaction, and improves the user experience.

[0046] In one example, the first electronic device may also determine the voice command containing the recording intent, along with one or more sentences of voice data A following the voice command, as the first raw voice data. This allows for the extraction of first voiceprint feature information through a larger or longer amount of voice data, which can improve the effect of personalized voice recording and enhance the user experience. Here, voice data A may be another voice command received by the first electronic device after receiving the voice command containing the recording intent; or, voice data A may be the user-input voice data received by the first electronic device after actively guiding the user to engage in voice interaction following the receipt of the voice command containing the recording intent.

[0047] In another possible implementation, the first electronic device acquiring the first user's first raw voice data may include:

[0048] The first electronic device acquires a voice command input by the first user, the voice command including an instruction to record the user's voice intent, and outputs interactive information based on the voice command, the interactive information being used to prompt the first user to input the first original voice data;

[0049] The first electronic device acquires the first raw voice data input by the first user.

[0050] In the speech synthesis method provided by this implementation, after offline recording of personalized speech is initiated based on the voice command of the first user (e.g., a voice command containing an explicit recording intent), the first electronic device can output interactive information, enabling the first user to interact with the first electronic device via voice based on the interactive information. During the voice interaction, the first electronic device can acquire the first raw voice data input by the first user. The interactive information is used to instruct the first user to input the first raw voice data in one or more of the following ways: (1) freely say one or more sentences; (2) upload an audio file; (3) read aloud specified text content.

[0051] Optionally, when voice input is inconvenient in the current environment, such as in a noisy environment or when the microphone of the first electronic device is unusable, the first user can upload an existing audio file to record personalized voice. This can expand the recording scenarios for personalized voice, making it easier for users to record personalized voice in various application scenarios, improving the recording effect of personalized voice, and enhancing the user experience.

[0052] Optionally, when the first user needs to record other users (e.g., user A), but user A is unable to input voice, for example, when user A is not at the current recording location, the first electronic device can obtain user A's audio file, such as obtaining the audio file of user A sent by other electronic devices (e.g., the electronic device to which user A belongs), and record user A's personalized voice by uploading user A's audio file. This allows for quick and convenient recording of personalized voices for other users, improving the user experience.

[0053] In one example, the first electronic device may also determine both the voice command containing the recording intent and the voice data input by the first user as the first original voice data input by the first user, or determine both the voice command containing the recording intent and the voice data in the audio file uploaded by the first user as the first original voice data input by the first user, so as to increase the duration or quantity of the first original voice data and improve the recording effect of personalized voice.

[0054] For example, when the voice data or audio file entered by the first user is relatively small, such as when the first user only enters a short voice data sentence or the audio file only includes a short voice data sentence, the first electronic device can determine the voice command with recording intent input by the first user and the entered voice data, or both the voice command with recording intent input by the first user and the voice data in the audio file, as the first original voice data.

[0055] The shortness of the voice data can be determined by whether its duration is less than a preset threshold. Specifically, if the duration of the voice data is less than the preset threshold, the first electronic device can determine that the voice data is short. The preset threshold can be set by technicians according to the actual scenario; for example, it can be set to any value such as 3 seconds or 5 seconds.

[0056] In one example, after the first electronic device acquires the first voiceprint feature information corresponding to the first original voice data, the method may further include:

[0057] The first electronic device automatically switches the interactive voice of the first electronic device to the voice corresponding to the first voiceprint feature information.

[0058] In the speech synthesis method provided by this implementation, after obtaining the first voiceprint feature information, the first electronic device can directly and automatically set the personalized voice corresponding to the first voiceprint feature information as the interactive voice or broadcast voice of the first electronic device, enabling the first electronic device to perform voice interaction or content broadcasting through the personalized voice corresponding to the first voiceprint feature information. For example, when starting online recording of personalized voice based on the voice command of the first user (e.g., a voice command containing implicit recording intent), after obtaining the first voiceprint feature information, the first electronic device can directly and automatically set the personalized voice corresponding to the first voiceprint feature information as the interactive voice or broadcast voice of the first electronic device, allowing the user to quickly experience the speech synthesis features without interaction and improving the user experience.

[0059] In another example, after the first electronic device acquires the first voiceprint feature information corresponding to the first raw voice data, the method may further include:

[0060] The first electronic device receives a switching instruction, which is used to instruct the switching of the interactive voice of the first electronic device;

[0061] The first electronic device switches its interactive voice to the voice corresponding to the first voiceprint feature information according to the switching instruction.

[0062] In the speech synthesis method provided by this implementation, the first electronic device can also switch the interactive voice based on the switching command of the first user, so as to switch to the personalized voice desired by the first user when the first user needs to switch the interactive voice, thereby improving the user experience.

[0063] The switching command can be contained in the voice data input by the first user, meaning the first user can directly input the switching command via voice. Alternatively, the switching command can be generated based on a preset operation by the first user on the display interface. For example, the first electronic device can display an interactive voice switching button on the display interface, and the first user can switch the interactive voice based on the switching button. That is, when the first electronic device detects that the switching button has been triggered, it can generate a switching command to switch the interactive voice.

[0064] In one possible implementation, the method may further include:

[0065] The first electronic device acquires the second user's second raw voice data;

[0066] The first electronic device acquires the fifth voiceprint feature information corresponding to the second original voice data, wherein the fifth voiceprint feature information includes at least one of the timbre, rhythm and style related to the second user;

[0067] The first electronic device inputs the fifth voiceprint feature information and the second target text content into the speech synthesis model to generate the second target speech data.

[0068] For example, the first electronic device acquiring the second user's second raw voice data may include:

[0069] The first electronic device receives an audio file sent by the second electronic device and determines the voice data in the audio file as the second original voice data. The second original voice data contained in the audio file is voice data with a duration greater than a preset duration and containing arbitrary content.

[0070] It should be understood that the first electronic device generates second target speech data from the second user's second original speech data for voice interaction, and this is the same speech synthesis model used by the first electronic device to generate first target speech data from the first user's first original speech data for voice interaction. That is, when the first electronic device uses different users' voices for voice interaction or content playback, it can generate target speech data corresponding to different users using the same speech synthesis model, without needing to train different speech synthesis models for each user. This saves time on adaptive model training, reduces the recording time for personalized speech, allowing personalized speech recordings to be as short as a few seconds, improving recording efficiency and enhancing user experience.

[0071] For example, before the first electronic device generates first target speech data corresponding to the first target text content based on the first voiceprint feature information and the first target text content using a speech synthesis model, the method may further include:

[0072] The first electronic device acquires the interactive voice data input by the first user;

[0073] The first electronic device determines the first target text content to be interacted with by the first electronic device based on the interactive voice data.

[0074] In one possible implementation, the speech synthesis model is used to obtain text feature information corresponding to the first target text content, map the text feature information corresponding to the first target text content and the first voiceprint feature information to the acoustic feature information corresponding to the first target text content, and convert the acoustic feature information corresponding to the first target text content into the first target speech data.

[0075] In training the speech synthesis model, different voiceprint feature information is used to adjust the mapping relationship learned by the speech synthesis model, so that the speech synthesis model learns the mapping relationship between voiceprint feature information, text feature information and acoustic feature information.

[0076] It should be understood that a speech synthesis model may include a text front-end processing module, a duration alignment module, an acoustic module, and a vocoder. The text front-end processing module extracts features from the first target text content to obtain text feature information corresponding to the first target text content. The duration alignment module expands the text feature information to frame-level text feature information based on the text feature information and the first voiceprint feature information, i.e., establishing a duration-based correspondence between the text feature information and the acoustic feature information. The acoustic module outputs the corresponding acoustic feature information based on the first voiceprint feature information and the frame-level text feature information. The vocoder synthesizes the first target speech data corresponding to the first target text content based on the acoustic feature information.

[0077] The acoustic module can include a first LSTM network and a second LSTM network. The acoustic module is primarily used to map textual feature information to acoustic feature information. Since textual feature information itself does not contain the speaker's voice characteristics, during training, voiceprint feature information from different speakers can be incorporated as a reference. This allows the acoustic module to adjust the mapping relationship between textual and acoustic feature information based on the voiceprint feature information of different speakers, enabling it to learn the influence of different speakers' voiceprint feature information on this mapping relationship.

[0078] This means that both voiceprint and text feature information can be input into the first LSTM network simultaneously for network training. Therefore, with a large amount of training data containing sufficiently rich voiceprint and text feature information, the first LSTM network can gradually learn the mapping relationship between (text feature information and voiceprint feature information) and acoustic feature information. This eliminates the need for the acoustic module to learn the mapping relationship between text feature information and acoustic feature information for each speaker, meaning that the acoustic module does not need to be retrained for each speaker.

[0079] Secondly, embodiments of this application provide a speech synthesis device applied to a first electronic device, the device comprising:

[0080] The voice acquisition module is used to acquire the first user's raw voice data.

[0081] The voiceprint extraction module is used to obtain the first voiceprint feature information corresponding to the first original speech data, wherein the first voiceprint feature information includes at least one of the timbre, rhythm and style related to the first user.

[0082] The speech synthesis module is used to generate first target speech data corresponding to the first target text content based on the first voiceprint feature information and the first target text content, using a speech synthesis model.

[0083] For example, the speech synthesis module is specifically used to input the first voiceprint feature information and the first target text content into the speech synthesis model to generate the first target speech data.

[0084] In one possible implementation, the device may further include:

[0085] A voiceprint adjustment module is used to select second voiceprint feature information and adjust the first voiceprint feature information based on the second voiceprint feature information to obtain third voiceprint feature information; wherein, the third voiceprint feature information is a fusion of the first voiceprint feature information and the second voiceprint feature information.

[0086] The speech synthesis module can also be used to generate first target speech data corresponding to the first target text content based on the speech synthesis model, according to the third voiceprint feature information and the first target text content.

[0087] In one example, the voiceprint adjustment module is specifically used to display the identifier of the first voiceprint feature information and the identifier of the fourth voiceprint feature information; in response to the adjustment operation of the first user, to determine the second voiceprint feature information and the second weight of the second voiceprint feature information from the fourth voiceprint feature information; to determine the first weight of the first voiceprint feature information according to the second weight, and to generate the third voiceprint feature information according to the first voiceprint feature information, the first weight, the second voiceprint feature information and the second weight.

[0088] In another example, the voiceprint adjustment module is further configured to display the identifiers of the first voiceprint feature information and the fourth voiceprint feature information; determine the second voiceprint feature information in response to the first user's selection operation on the fourth voiceprint feature information; when there is only one second voiceprint feature information, display an adjustment bar, one end of which is the first voiceprint feature information and the other end of which is the second voiceprint feature information; and determine the first weight of the first voiceprint feature information and the second weight of the second voiceprint feature information in response to the first user's adjustment operation on the adjustment bar, and generate the third voiceprint feature information based on the first voiceprint feature information, the first weight, the second voiceprint feature information, and the second weight.

[0089] Optionally, the first original voice data is the voice data input during the voice interaction; or, the first original voice data is the voice data in the uploaded audio file.

[0090] Optionally, the first original speech data is speech data with a duration longer than a preset duration and containing arbitrary content; or, the first original speech data is speech data containing specified text content.

[0091] In one possible implementation, the voice acquisition module is specifically configured to receive a voice command input by the first user, the voice command including an instruction to record the user's voice intent; and to determine the voice command including the instruction to record the user's voice intent as the first raw voice data.

[0092] In another possible implementation, the voice acquisition module is further configured to acquire voice commands input by the first user, the voice commands including instructions to record the user's voice intent, and output interactive information based on the voice commands, the interactive information being used to prompt the first user to input the first raw voice data; and acquire the first raw voice data input by the first user.

[0093] In one example, the device may further include:

[0094] The voice switching module is used to automatically switch the interactive voice of the first electronic device to the voice corresponding to the first voiceprint feature information.

[0095] In another example, the device may further include: an instruction acquisition module for acquiring a switching instruction, the switching instruction being used to instruct switching the interactive voice of the first electronic device;

[0096] The voice switching module is also used to switch the interactive voice of the first electronic device to the voice corresponding to the first voiceprint feature information according to the switching instruction.

[0097] In one possible implementation, the voice acquisition module is further configured to acquire second raw voice data from the second user;

[0098] The voiceprint extraction module is further configured to obtain the fifth voiceprint feature information corresponding to the second original speech data, wherein the fifth voiceprint feature information includes at least one of the timbre, rhythm and style related to the second user;

[0099] The speech synthesis module is further configured to input the fifth voiceprint feature information and the second target text content into the speech synthesis model to generate the second target speech data.

[0100] For example, the voice acquisition module is further configured to receive an audio file sent by a second electronic device, and determine the voice data in the audio file as the second original voice data, wherein the second original voice data contained in the audio file is voice data with a duration greater than a preset duration and containing arbitrary content.

[0101] For example, the device may further include:

[0102] The text content determination module is used to acquire interactive voice data input by the first user; and to determine the first target text content to be interacted by the first electronic device based on the interactive voice data.

[0103] In one possible implementation, the speech synthesis model is used to obtain text feature information corresponding to the first target text content, map the text feature information corresponding to the first target text content and the first voiceprint feature information to the acoustic feature information corresponding to the first target text content, and convert the acoustic feature information corresponding to the first target text content into the first target speech data.

[0104] In training the speech synthesis model, different voiceprint feature information is used to adjust the mapping relationship learned by the speech synthesis model, so that the speech synthesis model learns the mapping relationship between voiceprint feature information, text feature information and acoustic feature information.

[0105] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the speech synthesis method described in any one of the first aspects above.

[0106] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a computer, causes the computer to implement the speech synthesis method described in any one of the first aspects above.

[0107] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the speech synthesis method described in any one of the first aspects.

[0108] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0109] Figure 1This is a schematic diagram of the structure of the electronic device to which the speech synthesis method provided in the embodiments of this application is applicable;

[0110] Figure 2 This is a schematic diagram of an application scenario for voice interaction and recording of raw voice data provided in the embodiments of this application. Figure 1 ;

[0111] Figure 3 This is a schematic diagram of an application scenario for voice interaction and recording of raw voice data provided in the embodiments of this application. Figure 2 ;

[0112] Figure 4 This is a schematic diagram of an application scenario for voice interaction and recording of raw voice data provided in the embodiments of this application. Figure 3 ;

[0113] Figure 5 This is a schematic diagram of an application scenario where raw voice data is entered through a settings interface, as provided in an embodiment of this application.

[0114] Figure 6 This is a schematic diagram of the structure of the speech synthesis model provided in the embodiments of this application;

[0115] Figure 7 This is a schematic diagram of the acoustic module provided in the embodiments of this application;

[0116] Figure 8 This is a schematic diagram of an application scenario for adjusting the first voiceprint feature information provided in an embodiment of this application;

[0117] Figure 9 This is a schematic flowchart of the speech synthesis method provided in the embodiments of this application;

[0118] Figure 10 This is a schematic diagram of the speech synthesis device provided in the embodiments of this application. Detailed Implementation

[0119] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0120] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0121] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0122] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0123] Furthermore, the term "multiple" mentioned in the embodiments of this application should be interpreted as two or more.

[0124] The steps involved in the speech synthesis method provided in this application are merely examples, and not all steps are mandatory, nor are all information or message contents required. They can be added or removed as needed during use. The same step or step or message with the same function in this application can be referenced and learned from each other in different embodiments.

[0125] The business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0126] With the rapid development of electronic technology, more and more electronic devices are equipped with voice interaction functions to facilitate user interaction with the devices via voice or to enable the devices to play content through voice. Among these functions, electronic devices can have personalized voice recording capabilities, allowing users to record their own personalized voice messages for voice interaction, thereby improving the user experience. However, typical personalized voice recording requires users to first record a large or long amount of audio text (e.g., at least 20 sentences or more). Furthermore, the recorded audio text must be noise-free audio data matching the specified text content, not arbitrary audio input by the user. This allows the server to train a specific voice model corresponding to that user, enabling the electronic device to output personalized voice messages based on that model. Since personalized voice messages for different users require separate training of corresponding voice models, and training each user's voice model typically takes a considerable amount of time, such as 20 to 30 minutes or more. This method results in high costs for voice data acquisition, low recording efficiency, and significantly reduces the user experience.

[0127] To address the aforementioned problems, embodiments of this application provide a speech synthesis method, an electronic device, and a computer-readable storage medium. In this method, the electronic device can acquire first raw speech data input by a first user. Subsequently, the electronic device can extract first voiceprint feature information from the first raw speech data. The first voiceprint feature information may include at least one of timbre, prosody, and style. Based on the first voiceprint feature information and first target text content, the electronic device can generate first target speech data corresponding to the first target text content using a speech synthesis model, so as to perform voice interaction or voice broadcasting through the first target speech data.

[0128] In this embodiment, the electronic device can extract first voiceprint feature information containing at least one of timbre, rhythm, and style from at least one sentence of first original speech data entered by the user. Based on the first voiceprint feature information and the first target text content, it can directly generate personalized first target speech data using a speech synthesis model. This eliminates the need for the user to enter a large or long amount of speech data, and also eliminates the need for the user-entered speech data to precisely match the specified text content, thus reducing the time and cost of speech data acquisition. Furthermore, it eliminates the need to train a corresponding speech synthesis model based on the user's speech data, reducing model training time and the recording time for personalized speech. This allows for recordings as short as a few seconds, improving recording efficiency, enhancing user experience, and demonstrating strong usability and practicality.

[0129] In this application embodiment, the electronic device can be a mobile phone, tablet computer, wearable device, in-vehicle device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA) or desktop computer, etc., that supports voice input (i.e., can collect voice data entered by the user) and voice broadcast. This application embodiment does not limit the specific type of electronic device.

[0130] The following first describes the electronic device involved in the embodiments of this application. Please refer to... Figure 1 , Figure 1 A schematic diagram of an electronic device 100 is shown.

[0131] Electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, antenna 1, antenna 2, a mobile communication module 140, a wireless communication module 150, an audio module 160, a speaker 160A, a receiver 160B, a microphone 160C, a headphone jack 160D, a sensor module 170, buttons 180, a display screen 190, etc. The sensor module 170 may include a pressure sensor 170A, a gyroscope sensor 170B, a magnetic sensor 170C, an accelerometer sensor 170D, a fingerprint sensor 170E, a touch sensor 170F, etc.

[0132] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0133] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0134] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0135] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0136] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0137] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0138] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0139] In some embodiments, the processor 110 can extract the first voiceprint feature information corresponding to the original speech data. In other embodiments, the processor 110 can load a speech synthesis model and, during voice interaction, determine the target text content to be interacted with, so as to input the first voiceprint feature information and the target text content into the speech synthesis model, thereby generating target speech data corresponding to the target text content through the speech synthesis model based on the first voiceprint feature information and the target text content, wherein the target speech data has the sound characteristics corresponding to the first voiceprint feature information.

[0140] In some embodiments, the processor 110 can also recognize the user's intent based on the user's input voice data, such as recognizing the user's intent to record personalized voice, or recognizing the user's intent to switch interactive voices on electronic devices, or recognizing the user's intent to adjust the first voiceprint feature information, and so on.

[0141] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0142] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0143] The wireless communication function of electronic device 100 can be implemented through antenna 1, antenna 2, mobile communication module 140, wireless communication module 150, modem processor, and baseband processor.

[0144] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.

[0145] The mobile communication module 140 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 140 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 140 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 140 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 140 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 140 and at least some modules of the processor 110 may be housed in the same device.

[0146] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 160A, receiver 160B, etc.) or displays images or videos through the display screen 190. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 140 or other functional modules.

[0147] The wireless communication module 150 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 150 can be one or more devices integrating at least one communication processing module. The wireless communication module 150 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 150 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0148] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 140, and antenna 2 is coupled to wireless communication module 150, enabling electronic device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).

[0149] Electronic device 100 implements display functions through a GPU, a display screen 190, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 190 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0150] The display screen 190 is used to display images, videos, etc. The display screen 190 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 190, where N is a positive integer greater than 1.

[0151] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0152] Internal memory 121 can be used to store computer executable program code, which includes instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of electronic device 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located in the processor.

[0153] Electronic device 100 can implement audio functions such as music playback and recording through audio module 160, speaker 160A, receiver 160B, microphone 160C, headphone jack 160D, and application processor.

[0154] The audio module 160 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 160 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 160 may be located in the processor 110, or some functional modules of the audio module 160 may be located in the processor 110.

[0155] The speaker 160A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 160A.

[0156] The receiver 160B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a telephone call or voice message, the receiver 160B can be brought close to the ear to listen to the voice.

[0157] Microphone 160C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 160C, inputting the sound signal into microphone 160C. Electronic device 100 may have at least one microphone 160C. In some embodiments, electronic device 100 may have two microphones 160C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 160C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.

[0158] The 160D headphone jack is used to connect wired headphones. The 160D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0159] Pressure sensor 170A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 170A can be disposed on display screen 190. There are many types of pressure sensors 170A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 170A, the capacitance between the electrodes changes. Electronic device 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 190, electronic device 100 detects the intensity of the touch operation based on pressure sensor 170A. Electronic device 100 can also calculate the touch position based on the detection signal from pressure sensor 170A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.

[0160] The gyroscope sensor 170B can be used to determine the motion attitude of the electronic device 100. In some embodiments, the gyroscope sensor 170B can determine the angular velocity of the electronic device 100 about three axes (i.e., the x, y, and z axes). The gyroscope sensor 170B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 170B detects the angle of the shake of the electronic device 100, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the electronic device 100 by moving in the opposite direction, thus achieving image stabilization. The gyroscope sensor 170B can also be used in navigation and motion-sensing game scenarios.

[0161] The magnetic sensor 170C includes a Hall sensor. The electronic device 100 can use the magnetic sensor 170C to detect the opening and closing of the flip cover. In some embodiments, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip cover based on the magnetic sensor 170C. Then, based on the detected opening and closing state of the cover or the flip cover, features such as automatic flip unlocking can be set.

[0162] The 170D accelerometer can detect the magnitude of acceleration of electronic device 100 in various directions (typically three axes). When electronic device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of electronic device and applied to applications such as screen orientation switching and pedometers.

[0163] The fingerprint sensor 170E is used to collect fingerprints. The electronic device 100 can utilize the characteristics of the collected fingerprints to achieve fingerprint unlocking, accessing application locks, taking photos with fingerprints, answering calls with fingerprints, etc.

[0164] Touch sensor 170F, also known as a "touch device," can be located on display screen 190. The touch sensor 170F and display screen 190 together form a touchscreen, also known as a "touchscreen." Touch sensor 170F detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 190. In other embodiments, touch sensor 170F may also be located on the surface of electronic device 100, in a different position than display screen 190.

[0165] Buttons 180 include a power button, volume buttons, etc. Buttons 180 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.

[0166] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. For example, the software system of electronic device 100 can adopt a layered architecture such as Android operating system (OS), Harmony OS, or iOS.

[0167] The speech synthesis method provided in this application embodiment will be described in detail below with reference to the accompanying drawings and specific application scenarios.

[0168] In this embodiment, when a user wants to record personalized voice so that the electronic device can perform voice interaction or content playback through personalized voice, the user can record raw voice data. After acquiring the raw voice data, the electronic device can extract first voiceprint feature information from the raw voice data. The first voiceprint feature information may include at least one of timbre, rhythm, and style. Subsequently, the electronic device can generate target voice data corresponding to the target text content based on the first voiceprint feature information and the target text content, using a speech synthesis model. The target voice data and the raw voice data are identical in at least one of timbre, rhythm, and style, so that the electronic device can perform voice interaction or content playback through personalized target voice data.

[0169] The original voice data can be voice data input by the user that is longer than a preset duration and contains arbitrary content, or it can be voice data from an audio file uploaded by the user that is longer than a preset duration and contains arbitrary content. "Arbitrary content" means that the text content included in the original voice data can be any text content. That is, this embodiment does not specifically limit the text content included in the original voice data entered by the user; users can input original voice data containing any text content to record personalized voice messages. The preset duration can be set according to the actual scenario; for example, the preset duration can be set to any value such as 3 seconds. It is understood that the original voice data can also be voice data read aloud by the user based on specified text content. That is, the original voice data can also be voice data longer than a preset duration and containing specified text content.

[0170] For example, timbre refers to the distinctive characteristics of a sound; for instance, timbre can refer to the different waveform characteristics of different sounds. Prosody is used to characterize features such as intonation, pitch, emphasis, pauses, and rhythm. Style is used to characterize features such as tone, emotion, and accent.

[0171] It should be noted that the speech synthesis method provided in this application can be applied to applications that can perform voice interaction or content playback, such as voice assistants (also known as smart assistants or intelligent assistants), map applications, browser applications, or e-book applications. A voice assistant is an intelligent application that can perform specific functions through intelligent dialogue and real-time question-and-answer interaction. For example, a voice assistant could be Xiao Ai, a voice assistant based on the Android operating system. Xiaoyi based on HarmonyOS or based on iOS And so on. The following will exemplify the speech synthesis method provided in the embodiments of this application by taking its application in a voice assistant as an example.

[0172] For example, when a user wants to record personalized voice messages, they can activate the voice assistant and interact with it via voice to initiate recording and input raw voice data. Alternatively, the user can open the voice assistant's settings interface to initiate recording and input raw voice data.

[0173] The following will explain the methods for recording personalized voice data: 1. Recording personalized voice data via voice interaction to input raw voice data; 2. Recording personalized voice data via the settings interface to input raw voice data.

[0174] 1. Start recording personalized voice data through voice interaction to input raw voice data.

[0175] In this embodiment of the application, when a user wants to record personalized voice, the user can start the voice assistant and interact with the voice assistant to start recording personalized voice through voice interaction, and can also input raw voice data through voice interaction.

[0176] For example, a user can wake up a voice assistant using voice data and initiate personalized voice recording by interacting with the voice assistant, thereby inputting raw voice data through voice interaction. That is, when interacting with the user through the voice assistant, the electronic device can recognize the voice data input by the user to determine the user's intention to record personalized voice, and thus initiate personalized voice recording based on the user's recording intention and obtain the raw voice data entered by the user.

[0177] The method for waking up the voice assistant can be specifically set according to the actual scenario, and this application embodiment does not impose any restrictions on it. For example, the voice assistant can be woken up by inputting voice data containing a first preset keyword, or by touching a specific button, or by touching a specific icon, and so on. The first preset keyword can be customized by the user or set by default by the electronic device. For example, the electronic device can set the first preset keyword to "Xiaoyi Xiaoyi" by default. For example, the user can customize the first preset keyword to "Hello Xiaoyi", etc.

[0178] It should be understood that recording intent can include explicit recording intent and implicit recording intent. Explicit recording intent refers to the intention to record personalized voice that is clearly expressed in the user's input voice data; that is, the electronic device can directly determine that the user wants to record personalized voice based on the user's input voice data. Implicit recording intent refers to the intention to record personalized voice that is implicitly contained in the user's input voice data; that is, the electronic device can infer that the user wants to record personalized voice based on the user's input voice data.

[0179] In one example, the electronic device can determine the user's recording intent based on keywords contained in the user's input voice data. When the user's input voice data contains a second preset keyword, the electronic device can determine that the user has an explicit recording intent. When the user's input voice data contains a third preset keyword, the electronic device can determine that the user has an implicit recording intent.

[0180] The second preset keyword may include one or more words that explicitly indicate the recording of personalized voice, such as recording, creating, or cloning. The third preset keyword may include one or more words that implicitly contain the recording of personalized voice, such as learning, imitating, copying, or emulating.

[0181] For example, during voice interaction with a voice assistant, when the electronic device detects voice data such as "record my voice", "create Xiaoming's voice", "clone my voice", or "clone my timbre", the electronic device can determine that the voice data contains a second preset keyword, thus confirming that the user has an explicit recording intention.

[0182] For example, during voice interaction with a voice assistant, when the electronic device detects voice data such as "learn my speech," "imitate my speech," "follow my example," or "follow my example," it can determine that the voice data contains a third preset keyword, thus confirming that the user has an implicit recording intention. In other words, upon detecting this voice data, the electronic device can infer that the user wants to record personalized voice messages.

[0183] It should be noted that when a user wants to record personalized voice messages, they can first enter voice data containing the first preset keyword to activate the voice assistant, and then enter voice data containing the second or third preset keyword to start recording the personalized voice message. Alternatively, a user can enter voice data containing both the first and second preset keywords, or both the first and third preset keywords, to activate the voice assistant and start recording the personalized voice message.

[0184] For example, a user can first type "Xiaoyi Xiaoyi" to activate the voice assistant. Then, the user can type "I want to create my personalized voice" to start recording personalized voice. After the electronic device receives "I want to create my personalized voice," it can determine that the user has an explicit recording intention. At this point, the electronic device can start recording personalized voice and obtain the raw voice data entered by the user.

[0185] For example, a user can type "Hey Celia, say this" to wake up the voice assistant and start recording personalized voice. After the electronic device receives "Hey Celia, say this," it can determine that the user has an implicit intention to record. At this point, the electronic device can wake up the voice assistant and start recording personalized voice to obtain the user's original voice data.

[0186] In another example, electronic devices can directly input user-supplied voice data into a speech recognition model to directly determine the user's recording intent, that is, to directly determine whether the user has an explicit or implicit intention to record personalized voice.

[0187] In another example, electronic devices can directly input user-supplied voice data into a speech recognition model. The speech recognition model can first identify whether the user has the intention to record personalized voice. If the user has the intention to record personalized voice, the model can then determine whether the user's recording intention is explicit or implicit based on the user's input voice data, and so on.

[0188] It should be noted that the above-described method for determining the user's recording intent is merely an illustrative explanation and should not be construed as a limitation on the embodiments of this application. Electronic devices may also determine the user's recording intent in other ways; that is, the embodiments of this application do not impose specific limitations on the method for determining the recording intent.

[0189] In one example, when the recording intent is explicit, the electronic device can acquire the user-entered raw voice data through offline recording. When the recording intent is implicit, the electronic device can acquire the user-entered raw voice data through online recording.

[0190] The following will explain the specific application scenarios: (I) Obtaining raw voice data through offline recording; (II) Obtaining raw voice data through online recording.

[0191] (i) Obtaining raw audio data through offline recording

[0192] In offline recording, electronic devices can interact with users via voice assistants and acquire raw voice data recorded by the user during this interaction. Offline recording refers to recording where the user is aware of the raw voice data input process. For example, in offline recording, the electronic device can output interactive information to prompt the user to input raw voice data. The user can then input raw voice data based on this interactive information.

[0193] For example, after initiating offline recording of personalized voice based on a user's voice command (e.g., a voice command containing an explicit recording intent), the electronic device can output interactive information through a voice assistant, enabling the user to interact with the voice assistant based on the interactive information. During the voice interaction, the electronic device can acquire the user's raw voice data. The interactive information instructs the user to input the raw voice data in one or more of the following ways: (1) freely say one or more sentences; (2) upload an audio file; (3) read aloud specified text content.

[0194] Optionally, when voice input is inconvenient in the current environment, such as in a noisy environment or when the microphone of the electronic device is unusable, users can upload existing audio files to record personalized voice messages. This expands the recording scenarios for personalized voice messages, making it easier for users to record personalized voice messages in various application scenarios, improving the recording effect of personalized voice messages, and enhancing the user experience.

[0195] Optionally, when a user needs to record another user's (e.g., user A) voice input, but user A is unable to do so, for example, when user A is not at the current recording location, the electronic device can obtain user A's audio file. For example, it can obtain the audio file of user A sent by another electronic device (such as the electronic device to which user A belongs), and record user A's personalized voice by uploading user A's audio file. This allows for quick and convenient recording of personalized voices for other users, improving the user experience.

[0196] Optionally, the electronic device can use a voice assistant to read interactive information aloud, and / or can display interactive information as text on a display interface. The following will illustrate this with examples of the electronic device reading interactive information aloud via a voice assistant and displaying interactive information as text on a display interface.

[0197] Please see Figure 2 , Figure 2 This application illustrates an application scenario where raw voice data is recorded via voice interaction, as provided in an embodiment of this application. Figure 1 .

[0198] In this application scenario, when the electronic device receives the user's input, "Hey Celia, I want to create my own voice," it can wake up the voice assistant and start recording personalized voice messages. After waking up the voice assistant, the electronic device can use the voice assistant to read out interactive information A. For example, interactive information A could be, "Creating your voice requires recording a few sentences. Please say what you want to say, or upload your audio file, or read aloud specified text content." Figure 2 As shown in (a), the electronic device can display card 200 on the display interface. Card 200 can display the input method in interactive information A, namely, "Say one or more sentences freely" 210, "Upload audio file" 220, and "Read the following specified text content" 230. Among them, "Upload audio file" 220 includes an upload button 221. "Read the following specified text content" 230 includes specified text content for the user to read aloud (hereinafter, the specified text content corresponding to interactive information A will be referred to as the first specified text content) 231. For example, the first specified text content 231 could be "The world is beautiful, we should face every day with a smile."

[0199] In one example, the electronic device can also directly open the voice recording interface corresponding to the voice assistant, which can display... Figure 2 The input method shown in (a) includes an upload button 221 and a first specified text content 231. The following explanation will use the example of directly displaying interactive information, specified text content, and an upload button on the display interface.

[0200] For example, the first specified text content 231 corresponding to interactive information A can be switched by the user as needed. For instance, a refresh button corresponding to the first specified text content 231 can also be displayed in card 200. Figure 2 (Not shown in the image). When a user wants to switch the first specified text content 231, the user can touch the refresh button. After the electronic device detects that the refresh button has been touched, it can update the first specified text content 231 displayed on the display interface, thereby obtaining the new first specified text content and displaying it on the display interface. Similarly, other displayed specified text contents can also be switched by the user as needed.

[0201] Optionally, after the voice assistant broadcasts interactive information A, the electronic device can start recording audio through the microphone, that is, it can start acquiring the user's voice input data through the microphone. In other words, after the voice assistant broadcasts interactive information A, the user can directly speak what they want to say or directly read aloud the first specified text content 231 in card 200 to input voice data. At this time, the electronic device can acquire the user's voice input data through the microphone.

[0202] Optionally, after the voice assistant broadcasts interactive information A, the electronic device can obtain the user's selection operation on the display interface to determine whether to activate the microphone for sound recording based on the selection operation. For example, when the user clicks "Say one or more sentences" 210 or "Read the following specified text" 230, the electronic device can activate the microphone to record sound and obtain the user's input voice data. For example, when the user clicks the upload button 221 in "Upload audio file" 220, the electronic device may not activate the microphone for sound recording.

[0203] In one possible implementation, when a user wants to input raw voice data by freely speaking one or more sentences, the user can directly say what they want to say, or they can click... Figure 2 (a) shows “Say one or more sentences freely”210, followed by saying what you want to say.

[0204] For example, the electronic device can acquire the first sentence entered by the user (e.g., "I'll think about it tomorrow"). The electronic device can determine whether the duration of the first sentence is greater than a preset duration (e.g., 3 seconds). When it is determined that the duration of the first sentence is greater than 3 seconds, for example, when the duration of the first sentence is 5 seconds, such as... Figure 2 As shown in (b), the electronic device can continue to broadcast interactive information B via voice assistant. For example, interactive information B could be "We have received your first sentence. You can continue to say what you want to say to continue recording and improve the effect, or you can say to end recording and exit voice acquisition."

[0205] When a user wants to continue recording voice data, they can continue to say what they want to say directly. For example, a user can continue to say, "Everything is the best arrangement."

[0206] The electronic device can retrieve the second sentence entered by the user (i.e., "Everything is the best arrangement"). Then, as... Figure 2 As shown in (b), the electronic device can continue to broadcast interactive information C via voice assistant. For example, interactive information C could be "We have received your second sentence. You can continue to say what you want to say to continue recording and improve the effect, or you can say 'End recording and exit voice acquisition'."

[0207] When a user wants to exit voice recording, they can type "End Recording". When the electronic device detects "End Recording", it can terminate the voice interaction with the voice assistant and determine the user's first and second sentences ("I'll think about it tomorrow") as the raw voice data. Personalized voice recordings can then be created based on shorter raw voice data (e.g., 10s or 20s). Once completed, the voice assistant can announce the completion status. For example... Figure 2 As shown in (b), the electronic device can announce "Your voice production is complete" via a voice assistant.

[0208] For example, the electronic device can acquire the first sentence entered by the user (e.g., "Okay"). The electronic device can determine whether the duration of the first sentence is greater than a preset duration (e.g., 3 seconds). When it is determined that the duration of the first sentence is less than or equal to 3 seconds, for example, when the duration of the first sentence is 1 second, such as... Figure 2 As shown in (c), the electronic device can continue to broadcast interactive information B via voice assistant. For example, interactive information B could be "We have received your first sentence. The currently recorded voice is too short. Please continue to say what you want to say to continue recording and improve the effect."

[0209] Users can continue to say what they want based on the interaction information B. For example, a user can continue to say, "Everything is the best arrangement."

[0210] The electronic device can acquire the second sentence entered by the user (e.g., "Everything is the best arrangement"). At this point, the electronic device can further determine whether the sum of the durations of the first and second sentences is greater than 3 seconds. For example, if the duration of the second sentence is 7 seconds, meaning the sum of the durations of the first and second sentences is 8 seconds, then... Figure 2 As shown in (c), the electronic device can continue to broadcast interactive information C through the voice assistant. For example, interactive information C can be "We have received your second sentence. You can continue to say what you want to say to continue recording and improve the effect, or you can say to end the recording and exit the voice acquisition."

[0211] When a user wants to exit voice recording, they can type "End Recording". When the electronic device detects "End Recording", it can terminate the voice interaction with the voice assistant and determine the user's first and second sentences ("Okay") as the raw voice data. Personalized voice recordings can then be created based on shorter raw voice data (e.g., 8 seconds). Once completed, the voice assistant can announce the completion status. For example, ... Figure 2 As shown in (c), the electronic device can announce "Your voice production is complete" via a voice assistant.

[0212] Please see Figure 3 , Figure 3 This application illustrates an application scenario where raw voice data is recorded via voice interaction, as provided in an embodiment of this application. Figure 2 .

[0213] In another possible implementation, in displaying as Figure 2 After the voice recording interface shown in (a), when a user wants to record raw voice data by uploading an audio file, such as in a noisy environment, when the microphone of the electronic device cannot record voice, or when the user wants to record personalized voice of another user who is not in the current recording location, such as... Figure 3 As shown in (a), the user can touch the upload button 221 in "Upload Audio File" 220. Figure 3As shown in (b), after detecting that the upload button 221 has been touched, the electronic device can pop up an upload window 222, which displays a browse button 224 and a file input box 223. The user can select an audio file to upload by clicking the browse button 224. Alternatively, after detecting that the upload button 221 has been touched, the electronic device can jump to the upload interface, which displays the browse button 224 and the file input box 223, allowing the user to upload audio files through the upload interface. The electronic device obtains the audio file uploaded by the user and can determine the voice data in the audio file as the user's original voice data. The audio file contains voice data of arbitrary content, and the duration of the voice data contained in the audio file is longer than a preset duration (such as 3 seconds). For example, the duration of the voice data contained in the audio file can be any duration such as 5 seconds, 10 seconds, or 30 seconds, so that the user does not need to record a lot or a long amount of voice data, that is, personalized voice recording can be performed with a shorter duration of voice data, thus improving the user experience.

[0214] Understandably, when a user wants to input raw voice data by reading a specified text, the user can directly read the first specified text content 231 displayed on card 200. The electronic device can acquire the first specified text content 231 read by the user and can identify the first specified text content 231 read by the user as the raw voice data input by the user.

[0215] In this embodiment of the application, when recording voice data via voice input, the user can also combine speaking one or more sentences freely with reading a specified text content to record voice data. For example, the user can first speak one or more sentences freely, and then read a specified text content. The electronic device can determine both the sentence or more sentences spoken by the user and the specified text content read by the user as the original voice data recorded by the user.

[0216] Please see Figure 4 , Figure 4 This application illustrates an application scenario where raw voice data is recorded via voice interaction, as provided in an embodiment of this application. Figure 3 .

[0217] In this application scenario, when the electronic device receives the user's input, "Hey Celia, I want to create my own voice," it can wake up the voice assistant and start recording personalized voice messages. After waking up the voice assistant, the electronic device can use the voice assistant to read out interactive information D, such as, "Creating your voice requires recording a few sentences. Please say what you want to say, or upload your audio file, or read aloud specified text content." Figure 4As shown in (a), the electronic device can display card 400 on the display interface. Card 400 can display the input method in interactive information D, such as "voice input" 410 and "upload audio file" 420.

[0218] The "Upload Audio File" section 420 includes an upload button 421. The "Voice Input" section 410 may include "You can freely say one or more sentences, or read the following specified text content to perform voice input," and may include a second specified text content 411 for the user to read aloud. For example, the second specified text content 411 could be "The world is beautiful, we should face every day with a smile."

[0219] When a user wants to input voice data, they can directly say what they want to say or read aloud a second specified text. For example, a user can directly say, "I will think about it tomorrow."

[0220] The electronic device can capture the first sentence entered by the user (i.e., "I'll think about it tomorrow"). Subsequently, as... Figure 4 As shown in (b), the electronic device can continue to broadcast interactive information E via voice assistant. For example, interactive information E could be "We have received your first sentence. You can continue to say what you want to say, or read the following specified text content to continue recording and improve the effect, or you can say 'End recording and exit voice acquisition'." At this time, the electronic device can display card 430 on the display interface, and card 430 can display the third specified text content 431 for the user to read aloud.

[0221] Optionally, the third specified text content 431 may be the same as or different from the second specified text content 411. For example, when the second specified text content 411 is not read aloud, that is, when the voice data acquired by the electronic device does not contain the second specified text content 411, the third specified text content 431 may be the same as the second specified text content 411, that is, the electronic device may continue to display "The world is beautiful, we should face every day with a smile" on the card 430. For example, when the second specified text content 411 has been read aloud, that is, when the voice data acquired by the electronic device contains the second specified text content 411, the third specified text content 431 may be different from the second specified text content 411, for example, the third specified text content 431 may be "Tomorrow will be a new day, let's cheer each other on".

[0222] It should be understood that the above-described determination of the third specified text content 431 based on the reading of the second specified text content 411 is merely an illustrative explanation and should not be construed as a limitation on the embodiments of this application. In the embodiments of this application, the electronic device may also display different specified text content for the user to read aloud each time interactive information is output. That is, regardless of whether the second specified text content 411 corresponding to interactive information D is read aloud, the third specified text content 431 corresponding to interactive information E may be different from the second specified text content 411 corresponding to interactive information D.

[0223] This application scenario is illustrated by taking the third specified text content 431 as "Tomorrow will be a new day, let's cheer each other on" as an example.

[0224] When a user wants to continue recording voice data, they can either speak directly or read aloud the third specified text content 431. For example, if a user reads aloud the third specified text content 431, that is, "Tomorrow will be a new day, let's cheer each other on."

[0225] The electronic device can capture the second sentence entered by the user (i.e., "Tomorrow will be a new day, let's cheer each other on"). For example... Figure 4 As shown in (c), the electronic device can continue to broadcast interactive information F via voice assistant. For example, interactive information F could be "We have received your second sentence. You can continue to say what you want to say, or read the following specified text content to continue recording and improve the effect, or you can say 'End recording and exit voice acquisition'." Similarly, the electronic device can also display card 440 on the display interface, which can display a fourth specified text content 441 for the user to read aloud.

[0226] Optionally, the fourth specified text content 441 may be the same as the second specified text content 411, or the same as the third specified text content 431, or the fourth specified text content 441 may be different from both the second specified text content 411 and the third specified text content 431.

[0227] For example, when the second specified text content 411 is not read aloud, that is, when the voice data acquired by the electronic device does not contain the second specified text content 411, the fourth specified text content 441 can be the same as the second specified text content 411, that is, the electronic device can continue to display "The world is beautiful, we should face every day with a smile" in the card 440.

[0228] For example, when the third specified text content 431 is not read aloud, that is, when the voice data acquired by the electronic device does not contain the third specified text content 431, the fourth specified text content 441 can be the same as the third specified text content 431. That is, the electronic device can continue to display "Tomorrow will be a new day, let's cheer each other on" on card 440.

[0229] For example, when both the second specified text content 411 and the third specified text content 431 have been read aloud, that is, when the voice data acquired by the electronic device contains the second specified text content 411 and the third specified text content 431, the fourth specified text content 441 may be different from both the second specified text content 411 and the third specified text content 431.

[0230] It should be noted that the statement that the voice data acquired by the electronic device does not contain the second specified text content 411 and the third specified text content 431 can mean that none of the voice data acquired by the electronic device contains the second specified text content 411 and none of the voice data acquired by the electronic device contains the third specified text content 431. Similarly, the statement that the voice data acquired by the electronic device contains the second specified text content 411 and the third specified text content 431 can mean that a certain voice data acquired by the electronic device contains both the second specified text content 411 and the third specified text content 431, or it can mean that voice data A acquired by the electronic device contains the second specified text content 411 and voice data B acquired by the electronic device contains the third specified text content 431.

[0231] This application scenario is illustrated by taking the fourth specified text content 441 as an example, which is "The world is beautiful, we should face every day with a smile".

[0232] When a user wants to exit voice recording, they can type "End Recording". Upon detecting "End Recording", the electronic device can terminate the voice interaction with the voice assistant and identify the user's first and second sentences ("I'll think about it tomorrow") as the raw voice data. This raw data is then used to create a personalized voice message. Once completed, the voice assistant can announce the completion status. For example,... Figure 4 As shown in (d), electronic devices can announce "Your voice production is complete" via a voice assistant.

[0233] In one example, the electronic device may also identify both the user-inputted voice commands and the user-recorded voice data as the user-recorded original voice data, or it may identify both the user-inputted voice commands and the voice data in the user-uploaded audio file as the user-recorded original voice data.

[0234] For example, when the user-entered voice data or audio file contains limited voice data—such as a single short voice sentence or audio file containing only a short voice sentence—the electronic device may determine the user-inputted voice command and the entered voice data, or both the user-inputted voice command and the voice data in the audio file, as the original voice data. It should be understood that the voice command may be an instruction to wake up the voice assistant and / or an instruction to initiate personalized voice recording. For example, a voice command might be "Hey Celia." Or, for example, a voice command might be "Hey Celia, I want to create my voice."

[0235] The shortness of voice data can be determined by whether its duration is less than a preset threshold. Specifically, if the duration of the voice data is less than the preset threshold, the electronic device can determine that the voice data is short. The preset threshold can be set by technicians according to the actual scenario; for example, it can be set to any value such as 3 seconds or 5 seconds.

[0236] (ii) Obtaining raw voice data through online recording

[0237] For example, when the user's recording intent is implicit, the electronic device can acquire raw voice data through online recording. Online recording refers to directly identifying the user's input voice commands as raw voice data, eliminating the need for the user to input additional raw voice data, thus making the recording process seamless for the user. In other words, during online recording, the electronic device can directly acquire the user's input voice commands and identify them as the user's recorded raw voice data. These voice commands can be commands to wake up a voice assistant and / or commands to initiate personalized voice recording.

[0238] In other words, when online recording of personalized voice is initiated based on the user's voice command (such as a voice command containing an implicit recording intent), the electronic device can directly identify the voice command containing the recording intent (such as "Hey Celia, learn my voice") as the user's original voice data in the personalized voice recording process. This eliminates the need for the user to input any additional voice data, simplifies the process of inputting original voice data, allows users to quickly experience the features of voice synthesis without interaction, and improves the user experience.

[0239] For example, when a user first enters "Xiaoyi Xiaoyi" to wake up the voice assistant, and then enters "Learn to speak" to start recording personalized voice, the electronic device can determine that the user wants to record personalized voice. At this time, the electronic device can directly identify "Xiaoyi Xiaoyi, learn to speak" or "Learn to speak" as the user's original voice data, without needing to obtain the user's original voice data through voice interaction with the voice assistant.

[0240] For example, when a user enters "Xiaoyi Xiaoyi, imitate my speech" to start recording personalized voice, the electronic device can determine that the user wants to record personalized voice. At this time, the electronic device can directly identify "Xiaoyi Xiaoyi, imitate my speech" as the user's original voice data, without the need for the voice assistant to interact with the user to obtain the user's original voice data.

[0241] In one example, the electronic device can also determine the original voice data as a voice command containing the recording intent, along with one or more sentences of voice data A following the voice command. This allows for the extraction of first voiceprint feature information through a larger or longer amount of voice data, which can improve the effect of personalized voice recording and enhance the user experience. Here, voice data A can be another voice command received by the electronic device after receiving the voice command containing the recording intent; or, voice data A can be the voice data received by the electronic device after actively guiding the user to engage in voice interaction following the receipt of the voice command containing the recording intent.

[0242] For example, when an electronic device receives the user's voice command "Xiaoyi Xiaoyi, imitate my speech" and the subsequent voice command "Imitate my speech, you are such a smart person", the electronic device can identify the voice commands "Xiaoyi Xiaoyi, imitate my speech" and "Imitate my speech, you are such a smart person" together as raw voice data to improve the effect of personalized voice recording and enhance the user experience.

[0243] For example, when an electronic device receives the user's voice command "Hey Celia, imitate me," it can proactively output "Would you like to say something?" to guide the user to input more voice data. For instance, based on the electronic device's guidance, the user could input "The weather is so nice today." At this point, the electronic device can combine the voice command "Hey Celia, imitate me" with the user's input "The weather is so nice today" to form the raw voice data, thereby improving the personalized voice recording effect and enhancing the user experience.

[0244] 2. Start personalized voice recording through the settings interface to input raw voice data.

[0245] In this embodiment, when a user wants to record personalized voice, the user can open the settings interface corresponding to the voice assistant and start recording personalized voice through the settings interface. For example, the user can touch a specific button on the settings interface to start recording personalized voice. After starting the recording of personalized voice, the electronic device can display the input method on the display interface for the user to select to input the original voice data. The input method can be: (1) speaking one or more sentences freely; (2) uploading an audio file; (3) reading aloud specified text content.

[0246] When a user selects (1) to freely speak one or more sentences to input raw voice data, the electronic device can jump to the recording interface, where a record button can be displayed, allowing the user to input voice data by touching the record button. When a user selects (2) to upload an audio file to input raw voice data, the electronic device can jump to the upload interface or pop up an upload window, allowing the user to upload an audio file through the upload interface or upload window, thereby inputting raw voice data. When a user selects (3) to read aloud specified text content to input raw voice data, the electronic device can jump to the recording interface, where a record button and specified text content for the user to read aloud can be displayed, allowing the user to input voice data by reading aloud the specified text content after touching the record button.

[0247] Please see Figure 5 , Figure 5 This illustration shows an application scenario diagram of inputting raw voice data through a settings interface, as provided in an embodiment of this application.

[0248] like Figure 5 As shown in (a), when a user wants to record personalized voice, the user can open the corresponding settings interface 500 of the voice assistant. The settings interface 500 can display voice settings, as well as wake-up settings, AAA settings, BBB settings, CCC settings, DDD settings, and EEE settings.

[0249] Users can click on the voice settings item to enter the voice settings interface 510, where they can customize their voice. For example... Figure 5 As shown in (b), the voice settings interface 510 may display an "Add Sound" button 511, allowing users to start recording personalized voice messages by touching the "Add Sound" button 511. The voice settings interface 510 may also include the official voice of the voice assistant and a corresponding description. For example, the official voice may include a child's voice and a description such as "Innocent, playful, and incredibly cute."

[0250] like Figure 5As shown in (c), when the electronic device detects that the sound addition button 511 has been touched, it can enter the personalized voice recording interface 520. The recording interface 520 may display a message such as "Voice recording requires recording a few sentences of yours. Please select any of the following recording methods to record your voice data." Simultaneously, the recording interface 520 may also display selection options for the recording method, namely, a first option 521, a second option 522, and a third option 523. The first option 521 allows you to freely speak one or more sentences; the second option 522 allows you to upload an audio file; and the third option 523 allows you to read aloud specified text content.

[0251] When a user wants to input data by speaking one or more sentences freely, they can tap the first option, 521. For example... Figure 5 As shown in (d), after the electronic device detects that the first selection item 521 has been touched, it can display a recording interface 530, which may display a record button 531. The user can press and hold the record button 531 to input voice data; that is, the user can press and hold the record button 531 and freely speak one or more sentences of any content. After speaking, the user can release the record button 531 to end the voice input. The electronic device can acquire the voice data during the period when the record button 531 is pressed and hold, and determine the acquired voice data as the user's original voice data. It should be understood that the recording interface 530 may also include a play button 532 and a save button 533. The user can save the recorded voice data using the save button 533 and play the recorded voice data using the play button 532.

[0252] When a user wants to record audio by uploading a file, they can tap the second option 522. At this point, the electronic device can redirect to the upload interface or display an upload window, allowing the user to upload the audio file. The electronic device can then identify the voice data from the uploaded audio file as the user's original recorded voice data.

[0253] When a user wants to input text by reading it aloud, they can touch the third option 523. The electronic device will then switch to the recording interface, displaying the specified text for the user to read aloud. The recording interface also displays a record button, a play button, and a save button. The user can press and hold the record button and read the specified text displayed on the screen to input their voice. After recording, the user can save the recorded voice data using the save button or play it using the play button.

[0254] For example, when displaying specified text content for the user to read aloud in the recording interface, a refresh button can also be displayed on the recording interface. When the user wants to switch the specified text content, the user can touch the refresh button. After the electronic device detects that the refresh button has been touched, it can update the specified text content displayed in the recording interface, that is, obtain the new specified text content, and display the new specified text content on the recording interface.

[0255] The process of extracting the first voiceprint feature information corresponding to the original speech data will be explained in detail below.

[0256] In this embodiment of the application, after acquiring the original voice data entered by the user, the electronic device can extract the first voiceprint feature information from the original voice data. It should be understood that the first voiceprint feature information can be a multi-dimensional vector (e.g., a 256-dimensional vector), and the first voiceprint feature information can be one or more of timbre, rhythm, and style, etc., that do not contain voice content.

[0257] For example, the first voiceprint feature information can be a multi-dimensional vector representing timbre, or a multi-dimensional vector representing timbre and rhythm, or a multi-dimensional vector representing timbre and style, or a multi-dimensional vector representing timbre, rhythm and style, and so on.

[0258] For example, an electronic device can acquire acoustic feature information corresponding to the original speech data and input this acoustic feature information into a voiceprint feature extraction model to directly obtain the first voiceprint feature information corresponding to the original speech data. The voiceprint feature extraction model can consist of a three-layer Long Short-Term Memory (LSTM) network. Each LSTM layer can have 768 nodes. The voiceprint feature extraction model can use the output of the last hidden state of the last LSTM layer as the first voiceprint feature information corresponding to the original speech data.

[0259] In one example, the acoustic feature information corresponding to the original speech data can be Mel-Frequency Cepstral Coefficients (MFCCs), for example, an 80-dimensional MFCC. The frame length of the MFCC can be 50 milliseconds, and the frame shift can be 12.5 milliseconds. The electronic device can use any method to obtain the acoustic feature information corresponding to the original speech data; this application embodiment does not impose specific limitations on this.

[0260] It should be understood that a voiceprint feature extraction model can be trained using a large amount of training speech data from different speakers. For example, training speech data from different speakers can be obtained using a uniform sampling rate (e.g., 16kHz). After obtaining the training speech data, operations such as data cleaning, noise reduction, reverberation reduction, and volume normalization can be performed on each training speech data set. Additionally, standard channel labels can be assigned to each training speech data set. Training speech data from different sources will have different standard channel labels.

[0261] For example, unsupervised data clustering can be performed on the training speech data using a variational autoencoder (VAE) to obtain standard speaker labels corresponding to each training speech data. Both the standard speaker labels and the standard channel labels can be encoded as 256-dimensional one-hot vectors.

[0262] During training, the electronic device first acquires the acoustic feature information corresponding to each training speech data point. This acoustic feature information is then input into the voiceprint feature extraction model for processing, resulting in training voiceprint feature information output by the model. Subsequently, the electronic device inputs this training voiceprint feature information into a speaker discriminator and a channel discriminator for processing, obtaining the speaker label predicted by the speaker discriminator and the channel label predicted by the channel discriminator. Then, the electronic device adjusts the parameters of the voiceprint feature extraction model based on the speaker label predicted by the speaker discriminator and the standard speaker label corresponding to the training speech data, as well as the channel label predicted by the channel discriminator and the standard channel label corresponding to the training speech data. The model is then trained again based on the acoustic feature information corresponding to each training speech data point until the error between the speaker label predicted by the speaker discriminator and the standard speaker label is less than a first preset threshold, and the error between the channel label predicted by the channel discriminator and the corresponding standard channel label is greater than a second preset threshold. This completes the training of the voiceprint feature extraction model.

[0263] In other words, the embodiments of this application can perform adversarial training through a channel discriminator to eliminate the influence of different channel data on the voiceprint feature extraction model, so that the voiceprint feature information extracted by the voiceprint feature extraction model is independent of the channel, thereby improving the accuracy of the voiceprint feature extraction model.

[0264] It should be noted that the first and second preset thresholds can be set according to the specific scenario. Both the speaker discriminator and the channel discriminator can include deep neural networks (DNNs) and softmax layers, and the speaker discriminator and the channel discriminator are trained simultaneously.

[0265] In one possible implementation, after obtaining the first voiceprint feature information, the electronic device can directly set the personalized voice corresponding to the first voiceprint feature information as the default voice for the voice assistant, enabling the voice assistant to perform voice interaction or content playback using the personalized voice corresponding to the first voiceprint feature information. For example, when performing voice interaction or content playback through the voice assistant, the electronic device can acquire the target text content to be interacted with and synthesize the target text content and the first voiceprint feature information to obtain target voice data corresponding to the target text content. This allows the voice assistant to perform voice interaction or content playback using the target voice data, enabling the voice assistant to play the target text content using the personalized voice corresponding to the first voiceprint feature information.

[0266] For example, after an electronic device initiates online recording of personalized voice based on the user's input "Hey Celia, imitate my voice," and obtains the first voiceprint feature information corresponding to "Hey Celia, imitate my voice," the electronic device can acquire the target text content to be interacted with (e.g., "This is your beautiful voice!"), and synthesize the target text content and the first voiceprint feature information to obtain the target voice data corresponding to the target text content. Subsequently, the electronic device can use a voice assistant to broadcast the target voice data to be imitated; that is, the voice broadcast by the voice assistant, "This is your beautiful voice!", has the same sound characteristics as the user's input, "Hey Celia, imitate my voice!"

[0267] In another possible implementation, after obtaining the first voiceprint feature information, the electronic device can switch the voice assistant's broadcast voice to the personalized voice corresponding to the first voiceprint feature information when it receives a switching command input by the user. Subsequently, when the electronic device performs voice interaction or content playback through the voice assistant, it can obtain the target text content to be interacted with and synthesize the target text content and the first voiceprint feature information to obtain the target voice data corresponding to the target text content, so that the voice assistant can perform voice interaction or content playback through the target voice data.

[0268] In one example, after obtaining the first voiceprint feature information, the electronic device can proactively ask the user whether to switch the voice assistant's broadcast voice to the personalized voice corresponding to the first voiceprint feature information. For example, after obtaining the first voiceprint feature information, the electronic device can provide a voice assistant broadcast switching prompt, and / or display a switching prompt on the display interface. The switching prompt is used to prompt the user to switch the voice assistant's broadcast voice to the personalized voice corresponding to the first voiceprint feature information. When the user wants to switch the voice assistant's broadcast voice to the personalized voice corresponding to the first voiceprint feature information, the user can input a switching command via voice. When the electronic device detects the switching command, it can switch the voice assistant's broadcast voice to the personalized voice corresponding to the first voiceprint feature information.

[0269] For example, after obtaining the first voiceprint feature information, the electronic device can use a voice assistant to announce, "Your personalized voice has been recorded. You can say 'Use new voice' to switch to the latest voice." When the user wants to use the recorded personalized voice, they can input "Use new voice." Upon detecting "Use new voice," the electronic device can determine that a switching command has been detected. At this point, the electronic device can switch the voice assistant's announcement to the personalized voice corresponding to the first voiceprint feature information, allowing the voice assistant to perform voice interaction or content playback using the personalized voice corresponding to the first voiceprint feature information.

[0270] In another example, during voice interaction via a voice assistant, the electronic device can recognize the user's input voice data to determine if it contains a switching command. When a switching command is detected, for example, when the user inputs "play it in Xiaoming's voice," the electronic device can determine that the input contains a switching command. At this point, the electronic device can search for the target voice the user wants to switch to; for example, the target voice could be Xiaoming's voice. That is, the electronic device can search for the first voiceprint feature information corresponding to Xiaoming's original voice data. Once the target voice is found, the electronic device can set it as the voice assistant's playback voice. Subsequent voice interactions or content playback via the voice assistant can then utilize the target voice.

[0271] It should be understood that when the target voice is not found, the electronic device can display a prompt window. This window can inform the user that the target voice they wish to switch to was not found and can ask the user if they want to record the target voice. When the user confirms the recording of the target voice, that is, when the voice data acquired by the electronic device contains the second preset keyword, the electronic device can acquire the original voice data through offline recording and obtain the first voiceprint feature information corresponding to the original voice data to obtain the target voice.

[0272] It should be noted that after obtaining the first voiceprint feature information corresponding to the original voice data, the electronic device can associate the first voiceprint feature information with the user corresponding to the original voice data, and can also associate the first voiceprint feature information with the user account corresponding to the electronic device. Therefore, when it is necessary to use the user's voice for voice interaction or content playback on other electronic devices (electronic devices logged into with the same user account as the aforementioned electronic devices), the other electronic devices can directly obtain the first voiceprint feature information corresponding to the user based on the user account, and can synthesize target voice data based on the first voiceprint feature information and target text content for voice interaction or content playback, without requiring the user to input the original voice data on other electronic devices, and without requiring other electronic devices to extract voiceprint feature information, thus realizing cross-device use of voiceprint feature information and improving user experience.

[0273] The following is a detailed explanation of the process by which an electronic device synthesizes the target text content and the first voiceprint feature information to obtain target speech data.

[0274] In this embodiment of the application, the electronic device can input the first voiceprint feature information and the target text content into the speech synthesis model, so as to synthesize the first voiceprint feature information and the target text content through the speech synthesis model to obtain the target speech data corresponding to the target text content.

[0275] The speech synthesis model can be trained using massive amounts of training data. During training, voiceprint features from different speakers can be incorporated as a reference to adjust the mapping between text and acoustic features. This allows the speech synthesis model to learn the influence of different speakers' voiceprint features on this mapping. Therefore, with a large amount of training data containing sufficiently rich voiceprint and text features, the trained speech synthesis model can learn the mapping between text and acoustic features. This eliminates the need for retraining the speech synthesis model for each speaker, meaning it doesn't require retraining for each individual speaker.

[0276] In practice, the speech synthesis model can directly obtain acoustic feature information with the speaker's vocal characteristics based on text feature information, any speaker's voiceprint feature information, and the mapping relationship between (text feature information and voiceprint feature information) and acoustic feature information. Therefore, target speech data with the speaker's vocal characteristics can be synthesized based on this acoustic feature information. Thus, in practice, only a small amount of speech data containing arbitrary content needs to be recorded by the user for voiceprint feature extraction, without requiring the user to record large or long amounts of speech data. Furthermore, there is no need for the user-recorded speech data to precisely match the specified text. Additionally, there is no need for adaptive training of the speech synthesis model based on the user's recorded speech data to obtain a personalized model for the user. This reduces speech data recording time, lowers the requirements for speech data recording, reduces the cost of speech data acquisition, saves time on adaptive training, and allows personalized speech recordings to be as short as a few seconds. This improves the efficiency of personalized speech recording, enhances the user experience, and saves resources such as speech detection, speech recognition, and GPU retraining, thus reducing the power consumption of electronic devices.

[0277] In one implementation, please refer to Figure 6 , Figure 6 A schematic diagram of a speech synthesis model provided in an embodiment of this application is shown.

[0278] like Figure 6 As shown, the speech synthesis model 600 may include a text front-end processing module (Encoder) 601, a duration alignment module (Alignment) 602, an acoustic module (Decoder) 603, and a vocoder 604. The text front-end processing module 601 is used to extract features from the target text content to obtain text feature information corresponding to the target text content. The duration alignment module 602 is used to expand the text feature information to frame-level text feature information based on the text feature information and the first voiceprint feature information, that is, to establish a duration-based correspondence between the text feature information and the acoustic feature information. The acoustic module 603 is used to output acoustic feature information based on the first voiceprint feature information and the frame-level text feature information. The vocoder 604 is used to synthesize target speech data corresponding to the target text content based on the acoustic feature information.

[0279] For example, the text front-end processing module 601 may include a preprocessing unit and an LSTM network. The preprocessing unit in the text front-end processing module 601 can process the target text content, such as performing linear transformations. The LSTM network in the text front-end processing module 601 can obtain the text feature information corresponding to the preprocessed target text content, that is, it can integrate the linguistic information (such as word segmentation, parts of speech, prosody, etc.) contained in the phoneme sequence corresponding to the target text content to obtain the text feature information.

[0280] The duration alignment module 602 may include an LSTM network and a Gaussian sampling unit. The LSTM network in the duration alignment module 602 can predict the mean and variance of the duration corresponding to each phoneme in the text feature information based on the first voiceprint feature information. The Gaussian sampling unit in the duration alignment module 602 can sample the mean and variance to obtain the duration corresponding to each phoneme, thus indicating the mapping relationship between phonemes and speech frames. In other words, the duration alignment module 602 can expand the text feature information into frame-level text feature information based on the first voiceprint feature information to establish a duration-based correspondence between frame-level text feature information and acoustic feature information.

[0281] Please see Figure 7 , Figure 7 A schematic diagram of an acoustic module provided in an embodiment of this application is shown.

[0282] like Figure 7 As shown, the acoustic module 603 may include a first LSTM network and a second LSTM network. The acoustic module 603 is primarily used to map text feature information to acoustic feature information. Since the text feature information itself does not contain the speaker's voice characteristics, during training, voiceprint feature information from different speakers can be incorporated as a reference. This allows the mapping relationship between text feature information and acoustic feature information to be adjusted based on the voiceprint feature information of different speakers, enabling the acoustic module 603 to learn the influence of the voiceprint feature information of different speakers on this mapping relationship.

[0283] That is, voiceprint feature information and text feature information can be input into the first LSTM network at the same time for network training. Therefore, under the condition of a large amount of training data with sufficiently rich voiceprint feature information and text feature information, the first LSTM network can gradually learn the mapping relationship between (text feature information and voiceprint feature information) and acoustic feature information, so that the acoustic module 603 does not need to learn the mapping relationship between text feature information and acoustic feature information for each speaker, that is, it does not need to retrain the acoustic module 603 for each speaker.

[0284] In addition, textual feature information can be input into a second LSTM network so that the second LSTM network can further utilize linguistic information, thereby ensuring that the mapping relationship from textual feature information to acoustic feature information is stable and reliable enough.

[0285] For example, the acoustic module 603 can adopt an autoregressive structure. For the text feature information of each frame, after processing by the acoustic module 603 to obtain the acoustic feature information corresponding to that frame, the acoustic feature information corresponding to that frame can be input into the first LSTM network to participate in the prediction of the acoustic feature information corresponding to the text feature information of the next frame.

[0286] For example, the acoustic module 603 may also include a preprocessing unit, which can preprocess the text feature information, the voiceprint feature information and the acoustic feature information of the previous frame respectively, so as to remove redundant information in the text feature information and dilate the voiceprint feature information, so that the first LSTM network can better map (text feature information and voiceprint feature information) with acoustic feature information, thereby improving the processing efficiency of the acoustic module 603.

[0287] Therefore, when using it, when predicting the acoustic feature information corresponding to the text feature information of each frame, the acoustic module 603 can obtain the acoustic feature information of the current speech frame based on the frame-level text feature information, the first voiceprint feature information and the acoustic feature information of the previous frame output by the duration alignment module 602, thereby obtaining the acoustic feature information corresponding to the target text content based on the acoustic feature information of each speech frame.

[0288] The vocoder 604 can employ a high-fidelity generative adversarial network (HIFI-GAN), which can synthesize target speech data corresponding to the target text content by using the acoustic feature information corresponding to the target text content through HIFI-GAN.

[0289] It should be understood that in typical personalized voice recording, a specific speech synthesis model for that user is trained based on a large or long amount of voice text entered by the user. The electronic device can then output the user's personalized voice based on this specific speech synthesis model. In other words, in typical personalized voice recording, once the speech synthesis model for a particular user is trained, its corresponding voiceprint features are determined. If adjustments to these voiceprint features are needed, the user must enter a large or long amount of new voice text for retraining the model, resulting in a newly trained speech synthesis model with new voiceprint features. In other words, in typical personalized voice recording, the voiceprint features of the speech synthesis model are determined based on the model parameters, making it impossible for the user to visually adjust these features.

[0290] In this embodiment, the first voiceprint feature information corresponding to the first original speech data is directly extracted first, and then the extracted first voiceprint feature information is input into a unified speech synthesis model to generate target speech data corresponding to the first target text content. That is, in this embodiment, the first voiceprint feature information corresponding to the first original speech data is directly extracted from the first original speech data, not determined based on model parameters. Therefore, after extracting the first voiceprint feature information, it can be visualized, allowing the user to visually adjust it. The first electronic device can then input the adjusted first voiceprint feature information into the speech synthesis model to generate the target speech data.

[0291] In one possible implementation, after obtaining the first voiceprint feature information, the electronic device can adjust the first voiceprint feature information. For example, it can use one or more second voiceprint feature information to adjust the first voiceprint feature information to obtain mixed voiceprint feature information (or it can be called third voiceprint feature information). This mixed voiceprint feature information not only possesses at least one of the characteristics such as timbre, rhythm, and style corresponding to the first voiceprint feature information, but also possesses at least one of the characteristics such as timbre, rhythm, and style corresponding to the second voiceprint feature information. This can enhance the voice characteristics of the voice assistant in voice interaction or content playback, increase the fun of voice interaction or content playback, and improve the user experience.

[0292] The second voiceprint feature information can be the voiceprint feature information corresponding to the voice assistant's built-in voice, the voiceprint feature information corresponding to voice downloaded by the user later, the voiceprint feature information corresponding to personalized voice recorded by the user before, and so on. In other words, the second voiceprint feature information is the voiceprint feature information corresponding to the voice already in the voice assistant.

[0293] In this embodiment, the electronic device can adjust the first voiceprint feature information based on one or more second voiceprint feature information selected by the user to obtain mixed voiceprint feature information. It should be understood that the user can access the voice management interface corresponding to the voice assistant to select one or more second voiceprint feature information to adjust the first voiceprint feature information. The voice management interface may display identifiers for the first voiceprint feature information and identifiers for existing voice features in the electronic device.

[0294] For example, when interacting with a voice assistant, a user can input voice data to access the voice management interface of the voice assistant. Alternatively, a user can open the settings interface of the voice assistant and access the voice management interface by touching a specific button in the settings interface. For instance, even when not interacting with the voice assistant, a user can open the settings interface of the voice assistant and then access the voice management interface by touching a specific button in the settings interface.

[0295] For example, after obtaining the first voiceprint feature information or switching the voice assistant's broadcast voice to the personalized voice corresponding to the first voiceprint feature information, the electronic device can broadcast prompts via the voice assistant and / or display prompts on the display interface. The prompts are used to suggest that the user can adjust the first voiceprint feature information. When the user wants to adjust the first voiceprint feature information, the user can enter the corresponding voice data based on the prompts to access the voice management interface, where they can select one or more second voiceprint feature information to adjust the first voiceprint feature information.

[0296] For example, when interacting with a voice assistant, the electronic device can acquire the user's input voice data to determine whether the user intends to adjust the first voiceprint feature information. When it is determined that the user intends to adjust the first voiceprint feature information, for example, when the voice data contains "adjust my voice", the electronic device can open the voice management interface corresponding to the voice assistant, so that the user can select one or more second voiceprint feature information to adjust the first voiceprint feature information through the voice management interface.

[0297] It should be understood that when adjusting the first voiceprint feature information based on one or more second voiceprint feature information selected by the user, the weights of the first voiceprint feature information and / or the second voiceprint feature information can be adjusted. The electronic device can obtain the first weight corresponding to the first voiceprint feature information and the second weight corresponding to each second voiceprint feature information, and can perform a weighted summation of the first voiceprint feature information and the second voiceprint feature information based on the first weight and each second weight to obtain the mixed voiceprint feature information.

[0298] For example, the first voiceprint feature information is adjusted using the second voiceprint feature information A and the second voiceprint feature information B, and the second weight corresponding to the second voiceprint feature information A is Q. a The second weight corresponding to the second voiceprint feature information B is Q. b When the first weight corresponding to the first voiceprint feature information is Q1, the electronic device can obtain the mixed voiceprint feature information as (the second voiceprint feature information A*Q). a +Second voiceprint feature information B*Q b+First voiceprint feature information *Q1).

[0299] For example, the first weight corresponding to the first voiceprint feature information and the second weight corresponding to each of the second voiceprint feature information can be set by default by the electronic device or by user-defined settings. Optionally, after the user selects one or more second voiceprint feature information, the electronic device can set the first weight corresponding to the first voiceprint feature information and the second weight corresponding to each of the second voiceprint feature information by default. For example, the electronic device can set the first weight corresponding to the first voiceprint feature information and the second weight corresponding to each of the second voiceprint feature information to be the same by default.

[0300] Optionally, when selecting one or more second voiceprint feature information, the user can customize the second weight corresponding to each second voiceprint feature information. The electronic device can determine the first weight corresponding to the first voiceprint feature information based on the second weight corresponding to each second voiceprint feature information. For example, the user can customize the second weight corresponding to each second voiceprint feature information through a visual method such as a slider.

[0301] Please see Figure 8 , Figure 8 This illustration shows an application scenario diagram of adjusting voiceprint feature information provided in an embodiment of this application.

[0302] After obtaining the first voiceprint feature information or switching the personalized voice corresponding to the first voiceprint feature information to the voice assistant's broadcast voice, the electronic device can adjust the prompts through the voice assistant's voice broadcast. For example, the electronic device can broadcast "You have switched to your new voice. You can also say 'View My Voice' to open the voice management interface. In the voice management interface, you can rename the voice or adjust the voice."

[0303] like Figure 8 As shown in (a), when the user inputs "View my voice", the electronic device can jump to the voice management interface 800 corresponding to the voice assistant. The voice management interface 800 can display the identifiers of the voiceprint feature information corresponding to each existing voice of the voice assistant (such as AAA, Celebrity B, Celebrity C, and Child's Voice) and the corresponding editing buttons. AAA can be the identifier corresponding to the first voiceprint feature information, while Celebrity B, Celebrity C, and Child's Voice are the identifiers corresponding to the second voiceprint feature information already available in the voice assistant. When the user wants to adjust the first voiceprint feature information, the user can click the editing button corresponding to the first voiceprint feature information, i.e., click the editing button corresponding to AAA.

[0304] like Figure 8As shown in (b), after the electronic device detects that the edit button corresponding to AAA has been touched, it can jump to the details interface 810 corresponding to AAA. In the details interface 810, the identifiers of existing second voiceprint feature information in the voice assistant (i.e., voiceprint feature information other than AAA) can be displayed. The user can select one or more second voiceprint feature information and set the second weight corresponding to the selected second voiceprint feature information via a slider, that is, set the second weight corresponding to each selected second voiceprint feature information. In other words, the user can select second voiceprint feature information by setting the weight via a slider, so as to adjust the first voiceprint feature information based on the selected second voiceprint feature information. For example, the second weight corresponding to celebrity B can be set to 0.4, the second weight corresponding to celebrity C to 0, and the second weight corresponding to the child's voice to 0.2 by using a slider. At this time, the electronic device can determine that the second voiceprint feature information selected by the user is celebrity B and the child's voice, and the weight of celebrity B is 0.4 and the weight of the child's voice is 0.2. Therefore, the electronic device can determine that the first weight corresponding to AAA is (1-0.4-0.2)=0.4, so that the adjusted mixed voiceprint feature information can be obtained as (the first voiceprint feature information corresponding to AAA*0.4+the second voiceprint feature information corresponding to celebrity B*0.4+the second voiceprint feature information corresponding to the child's voice*0.2).

[0305] Or, such as Figure 8 As shown in (c), after the electronic device detects that the edit button corresponding to AAA has been touched, it can jump to the details interface 820 corresponding to the first voiceprint feature information. The details interface 820 can display the identifiers of the existing second voiceprint feature information in the voice assistant, the selection boxes 821 corresponding to each second voiceprint feature information, and the OK button 822. The user can click the selection box 821 to select the second voiceprint feature information (e.g., selecting celebrity B), and can confirm the selection by clicking the OK button 822.

[0306] like Figure 8 As shown in (d), after the electronic device detects that the OK button 822 has been clicked, it can obtain the second voiceprint feature information selected by the user (i.e., celebrity B) and display an adjustment bar on the display interface. It should be understood that one end of the adjustment bar represents celebrity B, and the other end represents AAA. The user can slide the adjustment bar to adjust the second weight corresponding to celebrity B and the first weight corresponding to AAA.

[0307] At this point, the electronic device can determine the adjustment point (i.e. Figure 8 The first distance between the black dot (in the diagram) and AAA, and the second distance between the adjustment point and celebrity B, are used to determine the first weight corresponding to AAA and the second weight corresponding to celebrity B. The smaller the distance, the greater the weight; the greater the distance, the smaller the weight.

[0308] Assuming the first distance between the adjustment point and AAA is Ra, and the second distance between the adjustment point and celebrity B is Rb, the electronic device can determine the first weight corresponding to AAA as Rb / (Ra+Rb) and the second weight corresponding to celebrity B as Ra / (Ra+Rb).

[0309] Subsequently, the electronic device can adjust the first voiceprint feature information using the first voiceprint feature information, the first weight, each of the second voiceprint feature information and each of the second weight to obtain the mixed voiceprint feature information.

[0310] It should be understood that the mixed voiceprint feature information can be stored as a new voiceprint feature information in the electronic device; that is, the mixed voiceprint feature information and the first voiceprint feature information can be two voiceprint feature information that coexist. When a user chooses to use the mixed voiceprint feature information for voice interaction or content playback, the electronic device can synthesize the mixed voiceprint feature information with the target text content to obtain the target speech data corresponding to the target text content, so that the voice assistant can perform voice interaction or content playback based on the target speech data.

[0311] Alternatively, the mixed voiceprint feature information can be directly used as the adjusted voiceprint feature information of the first voiceprint feature information. When the user chooses to use the first voiceprint feature information for voice interaction or content playback, the electronic device can synthesize the adjusted mixed voiceprint feature information with the target text content to obtain the target speech data corresponding to the target text content, so that the voice assistant can perform voice interaction or content playback based on the target speech data.

[0312] Based on the above embodiments, the speech synthesis method provided in this application will be exemplarily described below using the example of recording personalized speech through voice interaction. The content of the above embodiments is applicable to this embodiment. Please refer to... Figure 9 , Figure 9 A schematic flowchart of a speech synthesis method provided in an embodiment of this application is shown. This method can be applied to a first electronic device. Figure 9 As shown, the method may include:

[0313] S901, The first electronic device acquires the first user's first voice data;

[0314] In one example, a first user can interact with a first electronic device via voice. For instance, the first user can input the voice command "Hey Celia, I want to create my voice," and the first electronic device can identify the first user's voice command "Hey Celia, I want to create my voice" as first voice data.

[0315] S902, The first electronic device determines the recording intention of the first user based on the first voice data and obtains the first original voice data recorded by the first user;

[0316] For example, the first electronic device can determine the recording intent of the first user based on the voice command input by the first user, that is, determine whether the first user intends to record personalized voice. The specific method for determining the recording intent of the first user can refer to the description of determining the recording intent in the aforementioned "I. Starting the recording of personalized voice through voice interaction to input raw voice data", which will not be repeated here.

[0317] Optionally, in one example, the first electronic device may further determine whether the first user's recording intent is an implicit recording intent or an explicit recording intent.

[0318] For example, when the recording intent is explicit, the first electronic device can initiate offline recording to acquire the first raw voice data entered by the first user. For instance, it can initiate recording as follows: Figure 2 or Figure 3 or Figure 4 The offline recording shown is used to obtain the first original voice data entered by the first user. That is, in the offline recording, the first user can enter the first original voice data by (1) freely saying one or more sentences; (2) uploading audio files; (3) reading aloud specified text content.

[0319] For example, when voice input is inconvenient in the current environment, such as in a noisy environment or when the microphone of the first electronic device is unusable, the first user can upload an existing audio file to record personalized voice. For example, when the first user needs to record another user (e.g., user A), but user A is unable to input voice, such as when user A is not at the current recording location, the first electronic device can obtain user A's audio file, for example, obtain the audio file of user A sent by a second electronic device (such as the electronic device to which user A belongs), and record personalized voice corresponding to user A by uploading user A's audio file.

[0320] It should be understood that the first raw audio data can be audio data with a duration greater than the preset duration and containing arbitrary content. Alternatively, the first raw audio data can be audio data with a duration greater than the preset duration and containing specified text content.

[0321] The preset duration can be set according to the specific scenario. For example, the preset duration can be set to any value such as 3 seconds, depending on the actual scenario. "Any content" means that the text content contained in the first raw voice data can be any text content that is not specified, that is, the electronic device or system does not limit the text content recorded.

[0322] In one example, after determining the recording intention of the first user, the electronic device can output interactive information to prompt the first user to input first raw voice data. At this time, the first electronic device can obtain the first raw voice data input by the first user; for example, the voice data received after outputting the interactive information can be identified as the first raw voice data input by the first user.

[0323] For example, when the recording intent is implicit, the first electronic device can initiate online recording to acquire the first raw voice data entered by the first user. That is, the first electronic device can directly identify the voice command input by the first user (e.g., "Xiaoyi Xiaoyi, learn to speak") as the first raw voice data entered by the first user, simplifying the process of entering the first raw voice data, allowing users to quickly experience the features of speech synthesis without interaction, and improving the user experience.

[0324] For example, the first electronic device can also, after determining the recording intent, use the user's voice content received after the recording intent as the first raw voice data. For instance, after receiving the voice command "Hey Celia, say it like me," and determining the recording intent, the voice assistant can initiate an interaction, and then use the voice data received from the interaction as the first raw voice data. User: "Hey Celia, say it like me." Voice Assistant: "Okay, let's chat for a bit. Do you like…?" Then the user's subsequent responses can be used as the first raw voice data. The interaction can be one round or multiple rounds. The form of the interaction is also not limited; for example, it can be an interaction initiated by the user, or other voice commands input by the user. For example, User: "Hey Celia, say it like me"; Voice Assistant: "Okay."; User: "How's the weather today?" (This voice is used as the first raw voice data).

[0325] S903. The first electronic device obtains the first voiceprint feature information corresponding to the first original speech data through the voiceprint feature extraction model.

[0326] The first voiceprint feature information can be a multi-dimensional vector (e.g., a 256-dimensional vector), and it can be one or more of the following: timbre, prosody, and style, which do not contain speech content. For example, the first voiceprint feature information can be a multi-dimensional vector representing timbre, or a multi-dimensional vector representing timbre and prosody, or a multi-dimensional vector representing timbre and style, or a multi-dimensional vector representing timbre, prosody, and style, and so on.

[0327] It should be understood that the process by which the first electronic device obtains the first voiceprint feature information corresponding to the original speech data through the voiceprint feature extraction model can be referred to in the aforementioned extraction process of the first voiceprint feature information, and will not be repeated here.

[0328] S904. The first electronic device acquires the second voice data of the first user, and sets the voice corresponding to the first voiceprint feature information as the interactive voice of the first electronic device according to the second voice data.

[0329] For example, a first electronic device can receive second voice data input by a first user. The second voice data may include a switching instruction, which is used to switch the interactive voice of the first electronic device to the voice corresponding to the first voiceprint feature information. In this case, the first electronic device can switch its interactive voice to a personalized voice corresponding to the first voiceprint feature information based on the second voice data.

[0330] For example, after obtaining the first user's first voiceprint feature information, if the first electronic device receives the first user's voice input "use new voice" as the second voice data, the first electronic device can switch the interactive voice of the first electronic device to the newly recorded voice of the first user, that is, switch the interactive voice of the first electronic device to the voice corresponding to the first voiceprint feature information.

[0331] In another example, the switching instruction can also be generated based on the first user's preset operation on the display interface. For example, the first electronic device can display a switching button for interactive voice on the display interface. The first user can switch the interactive voice of the first electronic device to the first user's voice based on the switching button. At this time, the first electronic device can generate a switching instruction based on the first user's switching operation, so as to switch the interactive voice of the first electronic device to the first user's voice according to the switching instruction.

[0332] In another example, after the first electronic device acquires the first voiceprint feature information corresponding to the first original voice data, it can automatically switch the interactive voice of the first electronic device to the voice corresponding to the first voiceprint feature information. For example, after the first electronic device starts online recording of personalized voice based on the first user's input "Hey Celia, imitate my voice," and obtains the first voiceprint feature information corresponding to "Hey Celia, imitate my voice," the first electronic device can automatically switch the interactive voice of the first electronic device to the voice corresponding to the first voiceprint feature information.

[0333] S905, The first electronic device acquires the third voice data of the first user;

[0334] Specifically, after setting the voice corresponding to the first voiceprint feature information as the interactive voice of the first electronic device, the first electronic device can receive third voice data from the first user. For example, it can receive voice data from the first user asking "How's the weather today?" and output interactive content based on the third voice data.

[0335] S906. The first electronic device determines the first target text content to be interacted with based on the third voice data;

[0336] It should be understood that after the first electronic device recognizes the first user's intention to inquire about the weather based on the third voice data, it obtains relevant weather content such as "Today's weather is sunny, with a temperature of 18 to 24 degrees" from the "weather" related skills, and can determine the obtained weather content as the first target text content to be interacted with.

[0337] S907. The first electronic device generates first target speech data corresponding to the first target text content based on the first target text content and the first voiceprint feature information, using a speech synthesis model.

[0338] For example, the first electronic device inputs the first voiceprint feature information and the first target text content into the speech synthesis model to obtain the first target speech data output by the speech synthesis model.

[0339] For example, the speech synthesis model is used to obtain text feature information corresponding to the first target text content, map the text feature information and the first voiceprint feature information corresponding to the first target text content to the acoustic feature information corresponding to the first target text content, and can convert the acoustic feature information corresponding to the first target text content into the first target speech data corresponding to the first target text content.

[0340] In training the speech synthesis model, different voiceprint feature information is used to adjust the mapping relationship learned by the speech synthesis model, so that the speech synthesis model learns the mapping relationship between voiceprint feature information, text feature information and acoustic feature information.

[0341] In one possible implementation, the structure of the speech synthesis model can be as follows: Figure 6 and Figure 7 As shown. It should be understood that the specific details of the speech synthesis model and the generation of the first target speech data through the speech synthesis model can be found in the aforementioned description of the speech synthesis model, and will not be repeated here.

[0342] S908, the first electronic device performs voice interaction through the first target voice data.

[0343] For example, after obtaining the first target speech data output by the speech synthesis model, the first electronic device can output the first target speech data. For instance, it can read out the weather content "Today's weather is sunny, and the temperature is 18 to 24 degrees Celsius" through a voice containing the first voiceprint feature information.

[0344] For example, when the first voiceprint feature information contains timbre, the timbre of the first target speech data is the same as the timbre in the first voiceprint feature information. That is, the first electronic device can read the weather content "Today's weather is sunny, and the temperature is 18 to 24 degrees" by using the timbre of the first original speech data.

[0345] For example, when the first voiceprint feature information contains prosody, the prosody of the first target speech data is the same as the prosody in the first voiceprint feature information. That is, the first electronic device can read the weather content "Today's weather is sunny, and the temperature is 18 to 24 degrees" through the prosody of the first original speech data.

[0346] For example, when the first voiceprint feature information includes timbre and style, the timbre of the first target speech data is the same as the timbre in the first voiceprint feature information, and the style of the first target speech data is the same as the style in the first voiceprint feature information. That is, the first electronic device can read the weather content "Today the weather is sunny, and the temperature is 18 to 24 degrees" by using the timbre and style of the first original speech data.

[0347] For example, when the first voiceprint feature information includes timbre, rhythm, and style, the timbre of the first target speech data is the same as the timbre in the first voiceprint feature information, and the rhythm of the first target speech data is also the same as the rhythm in the first voiceprint feature information. At the same time, the style of the first target speech data is the same as the style in the first voiceprint feature information. That is, the first electronic device can read the weather content "Today the weather is sunny, and the temperature is 18 to 24 degrees" by using the timbre, rhythm, and style of the first original speech data.

[0348] In one example, after the first electronic device acquires the first voiceprint feature information corresponding to the original voice data, the first electronic device can select the second voiceprint feature information and adjust the first voiceprint feature information based on the second voiceprint feature information to obtain the third voiceprint feature information; wherein, the third voiceprint feature information is a fusion of the first voiceprint feature information and the second voiceprint feature information.

[0349] For example, the second voiceprint feature information can be selected by the first user. Optionally, the first electronic device can display an identifier for the first voiceprint feature information and an identifier for the fourth voiceprint feature information, so that the first user can select one or more second voiceprint feature information from the fourth voiceprint feature information according to the identifier of the fourth voiceprint feature information to adjust the first voiceprint feature information.

[0350] In one possible implementation, the first user can directly select the second voiceprint feature information by adjusting the weight of the fourth voiceprint feature information.

[0351] For example, the first electronic device can display the identifiers of the first and fourth voiceprint feature information, along with corresponding edit buttons, on a display interface. When the first user clicks the edit button corresponding to the first voiceprint feature information, the first electronic device can display the voice management interface corresponding to the first voiceprint feature information. The voice management interface can display the identifier of the fourth voiceprint feature information and the weight bars corresponding to each fourth voiceprint feature information. Initially, the weight bars corresponding to each fourth voiceprint feature information can all be 0. When the first user wants to select a certain fourth voiceprint feature information as the second voiceprint feature information, the first user can adjust the weight bar corresponding to that fourth voiceprint feature information to set the second weight of the fourth voiceprint feature information to be greater than 0.

[0352] Based on the adjustment operation of the first user, the first electronic device can determine the second voiceprint feature information and the second weight of the second voiceprint feature information from the fourth voiceprint feature information, and can determine the first weight of the first voiceprint feature information based on the second weight of the second voiceprint feature information (for example, the first weight = 1 - the sum of all the second weights). Subsequently, the first electronic device can generate the third voiceprint feature information based on the first voiceprint feature information, the first weight, the second voiceprint feature information and the second weight. For example, the third voiceprint feature information can be obtained by weighted summing of the first voiceprint feature information and the second voiceprint feature information using the first weight and the second weight.

[0353] In another possible implementation, the first user may first select the second voiceprint feature information, and then adjust the first weight of the second voiceprint feature information and the second weight of the second voiceprint feature information.

[0354] For example, the first electronic device can display identifiers for the first and fourth voiceprint feature information, along with corresponding editing buttons, on a display interface. When the first user clicks the editing button corresponding to the first voiceprint feature information, the first electronic device can display a voice management interface for that first voiceprint feature information. This voice management interface can display identifiers for the fourth voiceprint feature information and corresponding selection boxes for each fourth voiceprint feature information. The first user can then select one or more second voiceprint feature information pieces through the selection boxes to adjust the first voiceprint feature information.

[0355] Specifically, when a first user selects a second voiceprint feature to adjust the first voiceprint feature, the first electronic device can display an adjustment bar on the display interface. One end of the adjustment bar displays the first voiceprint feature, and the other end displays the second voiceprint feature. The first user can slide the adjustment bar to adjust the second weight of the second voiceprint feature and the first weight of the first voiceprint feature.

[0356] When a first user selects multiple second voiceprint feature information, the first electronic device can display the identifier of each second voiceprint feature information and the corresponding weight bar on the display interface. The first user can set the second weight of each second voiceprint feature information through the weight bar.

[0357] It should be noted that selecting the second voiceprint feature information to adjust the specific content of the first voiceprint feature information can also be done as follows: Figure 8 As shown.

[0358] In one example, after obtaining the third voiceprint feature information, the first electronic device can determine the third voiceprint feature information as the adjusted voiceprint feature information of the first voiceprint feature information, and can directly generate the first target speech data corresponding to the first target text content based on the speech synthesis model according to the third voiceprint feature information and the first target text content. That is, if the speech corresponding to the first voiceprint feature information is the interactive speech of the first electronic device, after obtaining the third voiceprint feature information, the first electronic device can directly determine the speech corresponding to the third voiceprint feature information as the interactive speech of the first electronic device, so as to output the first target speech data through the speech corresponding to the third voiceprint feature information.

[0359] In another example, after obtaining the third voiceprint feature information, the first electronic device can identify the third voiceprint feature information as the new voiceprint feature information; that is, the third voiceprint feature information and the first voiceprint feature information can be two voiceprint feature information that coexist. Subsequently, when a switching command to switch the third voiceprint feature information to the interactive voice of the first electronic device is detected, the first electronic device can switch its interactive voice to the voice corresponding to the third voiceprint feature information, so as to perform voice interaction through the voice corresponding to the third voiceprint feature information.

[0360] For example, the method may also include:

[0361] A first electronic device acquires second raw speech data from a second user and records the second user's speech based on the second raw speech data to perform voice interaction. For example, the first electronic device can acquire fifth voiceprint feature information corresponding to the second raw speech data. The fifth voiceprint feature information includes at least one of the second user's timbre, rhythm, and style. Subsequently, the first electronic device can input the fifth voiceprint feature information and the second target text content into a speech synthesis model to generate second target speech data.

[0362] In one example, a first electronic device can receive an audio file sent by a second electronic device and identify the speech data in the audio file as second raw speech data. The second raw speech data contained in the audio file is speech data with a duration greater than a preset duration and containing arbitrary content.

[0363] In another example, the first electronic device can acquire second raw voice data recorded by the second user, such as acquiring one or more sentences spoken freely by the second user, or acquiring voice data of the second user reading a specified text. The duration of the second raw voice data recorded by the second user is longer than a preset duration.

[0364] It should be understood that the process of recording the second user's voice and conducting voice interaction through the second user's voice can be as described in S901 to S908 above, and will not be repeated here.

[0365] It should be noted that the voice interaction using the second user's second original voice data is the same voice synthesis model used for the voice interaction using the first user's first original voice data. In other words, the first electronic device can generate target voice data for different users using the same voice synthesis model for voice interaction, without needing to train different voice synthesis models for each user. This saves time on adaptive model training, reduces the recording time for personalized voice, and allows personalized voice recording to be as short as a few seconds, thereby effectively improving the recording efficiency of personalized voice and enhancing the user experience.

[0366] In this embodiment, when a user wants to generate personalized voice, the user can record one or more sentences of raw voice data. The electronic device can acquire the raw voice data recorded by the user and extract first voiceprint feature information from the raw voice data. The first voiceprint feature information may include at least one of timbre, rhythm, and style. Subsequently, the electronic device can perform voice interaction or voice broadcasting based on the first voiceprint feature information and the target text content, according to the target voice data corresponding to the target text content of the speech synthesis model.

[0367] In this embodiment, the electronic device can extract first voiceprint feature information containing at least one of timbre, rhythm, and style from one or more sentences of raw speech data entered by the user. It can then directly generate personalized target speech data based on the first voiceprint feature information and target text content using a speech synthesis model. This eliminates the need for users to input large or long amounts of speech data, and also eliminates the need for precise matching between the user-input speech data and the specified text content, thus reducing speech data acquisition time and cost. Furthermore, it eliminates the need for adaptive training of the speech synthesis model corresponding to the user's speech data, reducing model training time and personalized speech recording time. Personalized speech recordings can be as short as a few seconds, effectively improving recording efficiency and enhancing user experience.

[0368] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0369] Corresponding to the speech synthesis method described in the above embodiments, this application also provides a speech synthesis device, the various modules of which can correspondingly implement the various steps of the speech synthesis method.

[0370] Based on the above embodiments, for example, please refer to Figure 10 , Figure 10 A schematic diagram of a speech synthesis device provided in an embodiment of this application is shown. The content of the above embodiments is applicable to this embodiment and will not be repeated here. Figure 10 As shown, the device may include:

[0371] The voice acquisition module 1001 is used to acquire the first raw voice data of the first user. The details of acquiring the first raw voice data can be found in the aforementioned section "I. Starting Personalized Voice Recording via Voice Interaction to Input Raw Voice Data," or in the aforementioned section "II. Starting Personalized Voice Recording via the Settings Interface to Input Raw Voice Data," and will not be repeated here.

[0372] The voiceprint extraction module 1002 is used to obtain first voiceprint feature information corresponding to the first original speech data. The first voiceprint feature information includes at least one of timbre, rhythm, and style related to the first user. The extraction of the first voiceprint feature information can be referred to the aforementioned process of extracting first voiceprint feature information from the original speech data using a voiceprint feature extraction model, and will not be repeated here.

[0373] The speech synthesis module 1003 is used to generate first target speech data corresponding to the first target text content based on the first voiceprint feature information and the first target text content, according to the speech synthesis model.

[0374] In one example, the voice acquisition module 1001 can also be used to acquire the second raw voice data of the second user.

[0375] The voiceprint extraction module 1002 can also be used to extract voiceprint feature information corresponding to the second original speech data. The voiceprint feature information corresponding to the second original speech data includes at least one of the timbre, rhythm and style related to the second user.

[0376] The speech synthesis module 1003 can also be used to generate second target speech data corresponding to the second target text content based on the voiceprint feature information corresponding to the second original speech data and the second target text content, using a speech synthesis model.

[0377] The voiceprint extraction module 1002 may include a voiceprint feature extraction model. The voiceprint extraction module 1002 can obtain first voiceprint feature information corresponding to the first original speech data, or extract voiceprint feature information corresponding to the second original speech data, through the voiceprint feature extraction model. It should be understood that the process by which the voiceprint extraction module 1002 obtains second voiceprint feature information corresponding to the second original speech data through the voiceprint feature extraction model is similar to the process of obtaining the first voiceprint feature information, and will not be described in detail here.

[0378] The speech synthesis module 1003 may include a speech synthesis model to input first voiceprint feature information and first target text content into the speech synthesis model to generate first target speech data corresponding to the first target text content, or to input voiceprint feature information corresponding to second original speech data and second target text content into the speech synthesis model to generate second target speech data corresponding to the second target text content.

[0379] It should be understood that the structure of a speech synthesis model can be as follows: Figure 6 and Figure 7 As shown, it will not be elaborated further here.

[0380] In one example, the device may also include:

[0381] The intent recognition module is used to determine the user's intent based on the user's input voice data. For example, it can determine the user's intent to record personalized voice messages based on the user's input voice data. It can also determine the user's intent to switch between interactive voice commands on electronic devices based on the user's input voice data.

[0382] In another example, the device may also include:

[0383] The voice switching module is used to switch the interactive voice of electronic devices according to the switching command input by the user.

[0384] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0385] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0386] This application also provides an electronic device, which includes at least one memory, at least one processor, and a computer program stored in the at least one memory and executable on the at least one processor. When the processor executes the computer program, it causes the electronic device to implement the steps in any of the above-described method embodiments. Exemplarily, the structure of the electronic device can be as follows: Figure 1 As shown.

[0387] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a computer, causes the computer to perform the steps in any of the above-described method embodiments.

[0388] This application provides a computer program product that, when run on an electronic device, causes the electronic device to perform the steps in any of the above method embodiments.

[0389] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable storage medium can include at least: any entity or device capable of carrying computer program code to a device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable storage media cannot be electrical carrier signals or telecommunication signals.

[0390] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0391] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0392] In the embodiments provided in this application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0393] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0394] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A speech synthesis method characterized by, Applied to a first electronic device, the method includes: The first electronic device acquires the first user's first raw voice data; The first electronic device acquires the first voiceprint feature information corresponding to the first original voice data, and the first voiceprint feature information includes at least one of the timbre, rhythm and style related to the first user. The first electronic device selects the second voiceprint feature information and adjusts the first voiceprint feature information based on the second voiceprint feature information to obtain the third voiceprint feature information; wherein, the third voiceprint feature information is a fusion of the first voiceprint feature information and the second voiceprint feature information; The first electronic device generates first target speech data corresponding to the first target text content based on the third voiceprint feature information and the first target text content, using a speech synthesis model.

2. The method of claim 1, wherein, The first electronic device generates first target speech data corresponding to the first target text content based on the third voiceprint feature information and the first target text content, using a speech synthesis model, including: The first electronic device inputs the third voiceprint feature information and the first target text content into the speech synthesis model to generate the first target speech data.

3. The method of claim 1, wherein, The first electronic device selects the second voiceprint feature information and adjusts the first voiceprint feature information based on the second voiceprint feature information to obtain the third voiceprint feature information, specifically including: The first electronic device displays the identifier of the first voiceprint feature information and the identifier of the fourth voiceprint feature information; In response to the adjustment operation of the first user, the first electronic device determines the second voiceprint feature information and the second weight of the second voiceprint feature information from the fourth voiceprint feature information; The first electronic device determines the first weight of the first voiceprint feature information according to the second weight, and generates the third voiceprint feature information according to the first voiceprint feature information, the first weight, the second voiceprint feature information and the second weight.

4. The method of claim 1, wherein, The first electronic device selects the second voiceprint feature information and adjusts the first voiceprint feature information based on the second voiceprint feature information to obtain the third voiceprint feature information, including: The first electronic device displays the identifier of the first voiceprint feature information and the identifier of the fourth voiceprint feature information; In response to the first user's selection operation of the fourth voiceprint feature information, the first electronic device determines the second voiceprint feature information; When the second voiceprint feature information is one, the first electronic device displays an adjustment bar, one end of which is the first voiceprint feature information and the other end is the second voiceprint feature information; In response to the adjustment operation of the first user on the adjustment bar, the first electronic device determines the first weight of the first voiceprint feature information and the second weight of the second voiceprint feature information, and generates the third voiceprint feature information based on the first voiceprint feature information, the first weight, the second voiceprint feature information and the second weight.

5. The method of claim 1, wherein, The first original voice data is the voice data input during the voice interaction; or, the first original voice data is the voice data in the uploaded audio file.

6. The method of claim 1, wherein, The first original voice data is voice data with a duration greater than a preset duration and containing arbitrary content; or, the first original voice data is voice data containing specified text content.

7. The method of claim 1, wherein, The first electronic device acquires the first user's first raw voice data, including: The first electronic device receives a voice command input by the first user, the voice command including an instruction to record the user's voice intent; The first electronic device identifies the voice command, which includes the user's voice intent, as the first raw voice data.

8. The method of claim 1, wherein, The first electronic device acquires the first user's first raw voice data, including: The first electronic device acquires a voice command input by the first user, the voice command including an instruction to record the user's voice intent, and outputs interactive information based on the voice command, the interactive information being used to prompt the first user to input the first original voice data; The first electronic device acquires the first raw voice data input by the first user.

9. The method of claim 1, wherein, After the first electronic device acquires the first voiceprint feature information corresponding to the first original voice data, the method further includes: The first electronic device automatically switches the interactive voice of the first electronic device to the voice corresponding to the first voiceprint feature information.

10. The method of claim 1, wherein, After the first electronic device acquires the first voiceprint feature information corresponding to the first original voice data, the method further includes: The first electronic device receives a switching instruction, which is used to instruct the switching of the interactive voice of the first electronic device; The first electronic device switches its interactive voice to the voice corresponding to the first voiceprint feature information according to the switching instruction.

11. The method according to any one of claims 1 to 10, characterized in that, The method further includes: The first electronic device acquires the second user's second raw voice data; The first electronic device acquires the fifth voiceprint feature information corresponding to the second original voice data, wherein the fifth voiceprint feature information includes at least one of the timbre, rhythm and style related to the second user; The first electronic device inputs the fifth voiceprint feature information and the second target text content into the speech synthesis model to generate the second target speech data.

12. The method of claim 11, wherein, The first electronic device acquires the second user's second raw voice data, including: The first electronic device receives an audio file sent by the second electronic device and determines the voice data in the audio file as the second original voice data. The second original voice data contained in the audio file is voice data with a duration greater than a preset duration and containing arbitrary content.

13. The method according to any one of claims 1 to 10, characterized in that, Before the first electronic device generates the first target speech data corresponding to the first target text content based on the third voiceprint feature information and the first target text content using a speech synthesis model, the method further includes: The first electronic device acquires the interactive voice data input by the first user; The first electronic device determines the first target text content to be interacted with by the first electronic device based on the interactive voice data.

14. The method according to any one of claims 1 to 10, characterized in that, The speech synthesis model is used to obtain text feature information corresponding to the first target text content, map the text feature information corresponding to the first target text content and the third voiceprint feature information to the acoustic feature information corresponding to the first target text content, and convert the acoustic feature information corresponding to the first target text content into the first target speech data. In training the speech synthesis model, different voiceprint feature information is used to adjust the mapping relationship learned by the speech synthesis model, so that the speech synthesis model learns the mapping relationship between voiceprint feature information, text feature information and acoustic feature information.

15. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it causes the electronic device to implement the speech synthesis method as described in any one of claims 1 to 14.

16. A computer-readable storage medium, the computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the computer, it causes the computer to implement the speech synthesis method as described in any one of claims 1 to 14.