Voice prompting method and device, electronic equipment and storage medium

By acquiring the acoustic feature information of the target user or their authorized friend, and combining it with the prompt text to generate personalized voice prompts, the problem of monotonous voice prompts is solved, and the user experience is improved.

CN116631370BActive Publication Date: 2026-05-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-14
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

The voice prompts in existing technologies are too monotonous and cannot meet the personalized needs of different users.

Method used

By acquiring the acoustic feature information of the target user or their authorized friend, and combining it with the prompt text of the target node, speech synthesis is performed to generate prompt voice that simulates the user or authorized friend speaking.

Benefits of technology

Personalized voice prompts have been implemented, enhancing the user experience and making the voice prompts more in line with the user's familiar voice environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116631370B_ABST
    Figure CN116631370B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of voice synthesis, and discloses a voice prompting method and device, electronic equipment and a storage medium, the method comprising the following steps: if a target process goes to a target node requiring voice prompting, acquiring prompt text corresponding to the target node; according to a target user identifier of a target user participating in the target process, acquiring acoustic characteristic information of an associated user, the associated user being the target user or an authorized friend of the target user; performing voice synthesis according to the acoustic characteristic information of the associated user and the prompt text to obtain prompt voice simulating the speech of the associated user; and playing the prompt voice; the scheme can flexibly simulate the voice of the target user or the authorized friend of the target user to perform voice prompting. The scheme can be applied to various scenes such as cloud technology, artificial intelligence, intelligent transportation and auxiliary driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech synthesis technology, and more specifically, to a speech prompting method, apparatus, electronic device, and storage medium. Background Technology

[0002] In related technologies, the audio corresponding to voice prompts is uniformly generated according to preset acoustic feature information. For different users, the voice prompts sound the same, making the voice prompts too monotonous. Summary of the Invention

[0003] In view of the above problems, embodiments of this application propose a voice prompt method, apparatus, electronic device, and storage medium to improve the above problems.

[0004] According to one aspect of the embodiments of this application, a voice prompt method is provided, comprising: if a target process jumps to a target node that requires voice prompts, obtaining prompt text corresponding to the target node; obtaining acoustic feature information of an associated user based on the target user identifier of a target user participating in the target process, wherein the associated user is the target user or an authorized friend of the target user; performing speech synthesis based on the acoustic feature information of the associated user and the prompt text to obtain a prompt voice simulating the associated user speaking; and playing the prompt voice.

[0005] According to one aspect of the embodiments of this application, a voice prompt device is provided, comprising: a prompt text acquisition module, configured to acquire a prompt text corresponding to the target node if the target process jumps to a target node that requires a voice prompt; an acoustic feature information acquisition module, configured to acquire acoustic feature information of an associated user based on a target user identifier of a target user participating in the target process, wherein the associated user is the target user or an authorized friend of the target user; a speech synthesis module, configured to perform speech synthesis based on the acoustic feature information of the associated user and the prompt text to obtain a prompt voice simulating the speech of the associated user; and a playback module, configured to play the prompt voice.

[0006] In some embodiments, the voice prompt device further includes: an acoustic parameter adjustment information acquisition module, used to acquire acoustic parameter adjustment information corresponding to the target node; and an adjustment module, used to adjust the acoustic feature information according to the acoustic parameter adjustment information, wherein the adjusted acoustic feature information is used for speech synthesis for the target node.

[0007] In some embodiments, the speech synthesis module includes: a text analysis unit for performing text analysis on the prompt text to obtain a phoneme sequence; and a speech synthesis unit for performing speech synthesis by a vocoder based on the phoneme sequence and the acoustic feature information of the associated user to obtain a prompt speech simulating the speech of the associated user.

[0008] In some embodiments, the voice prompt device further includes: an ambient audio acquisition module for acquiring ambient audio of the environment where the target user is located; an ambient volume determination module for determining the ambient volume of the ambient audio; and a first playback parameter determination module for determining playback parameters of the prompt voice based on the ambient volume. In this embodiment, the playback module is further configured to play the prompt voice according to the playback parameters.

[0009] In other embodiments, the prompt voice is played by the terminal, and the voice prompt device further includes: a distance determination module for determining the distance between the target user and the terminal; and a second playback parameter determination module for determining the playback parameters of the prompt voice based on the distance; in this embodiment, the playback module is further configured to play the prompt voice according to the playback parameters.

[0010] In some embodiments, the target node is the first node in the target process that requires voice prompts after the target user identifier is obtained; in this embodiment, the acoustic feature information acquisition module includes: a query request sending unit, used to send a query request to the server, the query request being used to request the server to query the acoustic feature information of the associated user; the query request carries the target user identifier; and a receiving unit, used to receive the acoustic feature information of the associated user returned by the server; the server obtains the acoustic feature information of the associated user from the acoustic information based on the target user identifier.

[0011] In some embodiments, the voice prompt device further includes: a first storage module, configured to associate and store the user identifier of the associated user with the acoustic feature information of the associated user in a cache; a second acoustic feature information acquisition module, configured to acquire the acoustic feature information of the associated user from the cache when the target process reaches the first node that requires voice prompts; a second speech synthesis module, configured to perform speech synthesis based on the acoustic feature information acquired from the cache and the prompt text corresponding to the first node to obtain a first prompt voice simulating the associated user speaking; and a second playback module, configured to play the first prompt voice.

[0012] In some embodiments, the voice prompt device further includes: a voice acquisition module, configured to acquire the voice of the associated user in response to a voice authorization operation; a voice parsing module, configured to parse the voice of the associated user to obtain the acoustic feature information of the associated user; and a second storage module, configured to associate the acoustic feature information of the associated user with the user identifier of the associated user and store it in the acoustic information database.

[0013] In some embodiments, the target process is to transfer to the target node before the target user identifier is obtained; the voice prompt device further includes: a first default prompt voice acquisition module, used to acquire a first default prompt voice for the target node, wherein the first default prompt voice is obtained by speech synthesis based on default acoustic feature information and the prompt text corresponding to the target node; and a third playback module, used to play the first default prompt voice.

[0014] In other embodiments, the associated user is an authorized friend of the target user. The acoustic feature information acquisition module includes: an authorized friend set determination unit, used to determine the authorized friend set of the target user based on the target user identifier, the authorized friend set including an authorized friend identifier, the authorized friend identifier referring to the user identifier of a friend who authorized the target user to use voice; an authorized friend identifier acquisition unit, used to acquire an authorized friend identifier from the authorized friend set of the target user; and an acoustic feature information acquisition unit, used to acquire the acoustic feature information of the corresponding authorized friend based on the acquired authorized friend identifier.

[0015] In some embodiments, the acoustic feature information of the associated user carries validity information; the voice prompt device further includes: a validity verification module, configured to verify the validity of the acoustic feature information based on the validity information; a fourth playback module, configured to, if the acoustic feature information is determined to be invalid, play a first default prompt voice for the target node, the first default prompt voice being synthesized from default acoustic feature information and the prompt text; if the acoustic feature information is determined to be valid, then proceed to the speech synthesis module.

[0016] In some embodiments, the effectiveness information includes at least one of effectiveness node information and effectiveness time information; the validity verification module is further configured to: determine that the acoustic feature information is invalid if it is determined that the effectiveness node indicated by the effectiveness node information does not include the target node, or if it is determined that the current time is not within the effectiveness time period indicated by the effectiveness time information.

[0017] According to one aspect of the embodiments of this application, an electronic device is provided, including: a processor; a memory, the memory storing computer-readable instructions, which, when executed by the processor, implement the voice prompt method as described above.

[0018] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, on which computer-readable instructions are stored, which, when executed by a processor, implement the voice prompt method as described above.

[0019] According to one aspect of the embodiments of this application, a computer program product is provided, including computer instructions that, when executed by a processor, implement the voice prompt method as described above.

[0020] In this application, when the target process reaches a target node requiring voice prompts, a voice prompt simulating the target user's speech is generated using the acoustic feature information of the target user participating in the target process and the corresponding prompt text for that target node; alternatively, a voice prompt simulating the authorized friend of the target user participating in the target process is generated using the acoustic feature information of that authorized friend and the corresponding prompt text for that target node. The resulting voice prompt is then played. This achieves flexible voice prompts simulating either the speech of the target user participating in the target process or the speech of their authorized friend, effectively solving the problem of monotonous voice prompts in related technologies. Moreover, for the target user, whether the voice prompts simulate the speech of the target user or their authorized friend, the voice is familiar, creating a familiar voice environment and enhancing the user experience. Attached Figure Description

[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0022] Figure 1 This is a schematic diagram of a voice prompt system according to an embodiment of this application.

[0023] Figure 2 This is a flowchart illustrating a voice prompt method according to an embodiment of this application.

[0024] Figure 3 A is a schematic diagram of a first page according to an embodiment of this application.

[0025] Figure 3 B is a schematic diagram of a voice acquisition prompt page according to an embodiment of this application.

[0026] Figure 4 A is a schematic diagram of a second page according to an embodiment of this application.

[0027] Figure 4 B is a schematic diagram of a friend selection page according to an embodiment of this application.

[0028] Figure 5 This is illustrated according to an embodiment of the present application. Figure 2 The flowchart of step 220 in the corresponding embodiment.

[0029] Figure 6 This is a flowchart illustrating face payment according to a specific embodiment of this application.

[0030] Figure 7 This is a schematic diagram illustrating the fields included in the acoustic feature information according to an embodiment of this application.

[0031] Figure 8 This diagram illustrates the nodes in the facial recognition payment process that require voice prompts.

[0032] Figure 9 The process of providing voice prompts at a target node is shown.

[0033] Figure 10 This is a flowchart illustrating the use of voice by an authorized friend and the synthesis of voice based on the voice feature information of the authorized friend, according to an embodiment of this application.

[0034] Figure 11 The set of authorized friends for user T1 and user T2 is shown.

[0035] Figure 12 This is a block diagram of a voice prompt device according to an embodiment of this application.

[0036] Figure 13 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation

[0037] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0038] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0039] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0040] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily need to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0041] It should be noted that "multiple" in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0042] Figure 1 This is a schematic diagram of a voice prompt system according to an embodiment of this application, such as... Figure 1 As shown, the voice prompt system includes a terminal 110 and a server 120. The terminal 110 can communicate with the server 120 via a wired or wireless network. The terminal 110 can be a smartphone, tablet, laptop, desktop computer, smart speaker, in-vehicle terminal, payment terminal, smart voice interaction device, smart home appliance, aircraft, self-service terminal, voice broadcast terminal, or other electronic device that can interact with the user; no specific limitations are specified here.

[0043] The server 120 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0044] Terminal 110 can run an application (which can be simply referred to as an application). Users can interact with the terminal and the server based on the application. The application can be an instant messaging application, a shopping application, a navigation application, etc. The running application includes one or more processes, which include one or more nodes that require voice prompts. When the current node of the process is a node that requires voice prompts, voice prompts can be given according to the scheme of this application.

[0045] Specifically, terminal 110 sends a query request to the server based on the target user identifier of the target user participating in the target process. This query request carries the target user identifier. Then, server 120 obtains the acoustic feature information of the associated user (i.e., the target user or the target user's authorized friend) based on the query request and returns the obtained acoustic feature information of the associated user to terminal 110. Server 120 deploys an acoustic information database, and the server queries the acoustic information database to obtain the acoustic feature information of the associated user.

[0046] After receiving the acoustic feature information of the associated user, terminal 110 can perform speech synthesis based on the prompt text corresponding to the target node in the target process and the acoustic feature information of the associated user to obtain a prompt voice simulating the associated user's speech, and then play the prompt voice. It is worth mentioning that, if the processing capability of terminal 110 is sufficient, an acoustic information database can also be deployed on terminal 110, so that terminal 110 can execute the method of this application.

[0047] The embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving, which require voice prompts.

[0048] The implementation details of the technical solutions in the embodiments of this application are described in detail below:

[0049] Figure 2 This is a flowchart illustrating a voice prompt method according to an embodiment of this application. The method can be executed by a computer device with processing capabilities, such as… Figure 1 This method can also be implemented through interaction between the terminal and the server, such as in mid-terminals. (See reference...) Figure 2As shown, the method includes at least steps 210 to 240, which are described in detail below:

[0050] Step 210: If the target process jumps to the target node that requires voice prompts, obtain the prompt text corresponding to the target node.

[0051] The applications running on the terminal include processes that implement one or more functions. For example, in a payment application, the process could be a payment process or a transfer process; in a navigation application, the process could be a navigation process for the currently selected route; in a shopping application, the process could be a return application process or an order placement process; in an instant messaging application, the process could be a friend addition process; and in an enterprise's instant messaging application, the process could be a business processing process, a leave application process, an application approval process, etc., without specific limitations.

[0052] In this application, the currently running process in the application is called the target process. It is understood that the target process includes nodes that require voice prompts. The node in the target process that requires voice prompts is called the target node. It is understood that the target process includes multiple nodes, and at least includes nodes that require voice prompts; however, it may also include nodes that do not require voice prompts.

[0053] The prompt text corresponding to the target node indicates the content that needs to be prompted at the target node. Different nodes in the target process may require different voice prompts. Therefore, the prompt text corresponding to each node in the target process that needs voice prompts can be pre-set, and the prompt text can be associated with the node identifier of the corresponding node and stored to obtain the prompt text set of the target process.

[0054] Based on this, after determining the node identifier of the target node where the target process is located, the prompt text corresponding to the target node is obtained from the prompt text set corresponding to the target process according to the node identifier of the target node.

[0055] In some embodiments, after a process start event is detected, the process requested to be started by the process start event is taken as the target process in this application, and the flow status of the target process is detected, thereby determining the current node of the target process in real time. If the current node of the target process is a node that needs to be given a voice prompt, then the node is taken as the target node in this application, and then the prompt text corresponding to the target node is obtained from the prompt text set corresponding to the target process.

[0056] In some embodiments, to avoid empty access to the prompt text set corresponding to the target process, a first marker can be added to the node identifier of the node in the target process that requires voice prompts. Thus, when the node identifier of the node currently in which the target process is located is detected to carry the first marker, the prompt text set corresponding to the target process is accessed and the prompt text corresponding to the target node is obtained from it; otherwise, if the node identifier of the node currently in which the target process is located does not carry the first marker, it is not necessary to access the prompt text set corresponding to the target process.

[0057] Step 220: Based on the target user identifier of the target user participating in the target process, obtain the acoustic feature information of the associated user, where the associated user is the target user or an authorized friend of the target user.

[0058] The target user refers to a user participating in the target process. In some embodiments, the user participating in the target process can be a user who needs to handle the task corresponding to the target node in the target process, the user corresponding to the client that triggers the start of the target process, or a logged-in user in the application where the target process is located; no specific limitation is made here. It is understood that when the target process requires the assistance of multiple users, the user handling the task corresponding to different nodes in the target process may be different. In this case, the user corresponding to the current target node of the target process can be determined as the target user. When the target user is a logged-in user in the application where the target process is located, the target user is the same user at different nodes in the target process.

[0059] In some embodiments, an acoustic information database can be pre-built, and the user identifier of each user can be associated with the user's acoustic feature information and stored in the acoustic information database. Thus, in step 220, after obtaining the target user identifier of the target user, the acoustic feature information of the target user can be obtained from the acoustic information database.

[0060] Authorized friends of a target user refer to users who have authorized the target user to use their voice and are friends with the target user. For example, if user A and user B are friends, and user A authorizes user B to use user A's voice, then user B is user A's authorized friend.

[0061] In some embodiments, an authorized friend set can be constructed for each user, which includes the authorized friend identifiers of the user's authorized friends. Based on this, if the associated user is an authorized friend of the target user, in this embodiment, step 220 includes: determining the authorized friend set of the target user based on the target user identifier, the authorized friend set including the authorized friend identifier, which refers to the user identifier of a friend who authorized the target user to use voice; obtaining an authorized friend identifier from the target user's authorized friend set; and obtaining the acoustic feature information of the corresponding authorized friend from an acoustic information database based on the obtained authorized friend identifier.

[0062] In some embodiments, an authorized friend identifier may be randomly selected from the authorized friend set of the target user. In other embodiments, the authorized friend identifiers in the authorized friend set of the target user may be sorted to obtain an authorized friend ranking. When an authorized friend identifier needs to be obtained, the authorized friend identifier that is first in the authorized friend ranking is obtained each time. Of course, after obtaining an authorized friend identifier from the authorized friend set each time, the obtained authorized friend identifier is moved back to the last position in the authorized friend ranking.

[0063] In some embodiments, authorized friends can be sorted by their intimacy level or interaction frequency from highest to lowest, without specific limitations. Intimacy can be determined by the tags added by the user to their friends; for example, the intimacy level between a user and a friend marked as a star is higher than the intimacy level between the user and any friend not marked as a star.

[0064] In some embodiments, an authorization control for authorizing the use of a user's own voice can be set in the application's page. In this embodiment, before step 220, the method further includes: in response to a voice authorization operation, acquiring the voice of an associated user; parsing the voice of the associated user to obtain the acoustic feature information of the associated user; and associating the acoustic feature information of the associated user with the user identifier of the associated user and storing it in an acoustic information database.

[0065] When the associated user is the target user, the target user can trigger the voice authorization operation on the page; when the associated user is an authorized friend of the target user, the authorized friend can trigger the voice authorization operation on the page. The voice authorization operation can be an action that triggers an authorization control, such as clicking the authorization control or touching the authorization control.

[0066] Let's take user T1 as an example to explain the voice authorization process. Figure 3 A shows a schematic diagram of the first page, as follows: Figure 3As shown in Figure A, the first page 310 contains an authorization control 311. When a user triggers the authorization control 311 (e.g., a click operation (single or double click, touch operation, swipe operation, etc.), it constitutes a voice authorization operation. After detecting the user's triggering of the authorization control 311, the process proceeds to... Figure 3 Page 320, shown in B, is the voice acquisition prompt page.

[0067] The voice acquisition prompt page 320 includes a prompt text display area 321 and a voice acquisition control 322, wherein the prompt text display area 321 is used to display the prompt text. Figure 3 The prompt text in B is "Today is New Year's Day", and the user is prompted to read the prompt text in the prompt text display area 321.

[0068] When the user triggers the voice acquisition control 322, voice acquisition is initiated, thereby capturing the user T1's voice while the user T1 reads the prompt text. The user T1's voice can then be analyzed to obtain the user T1's acoustic feature information, which is then associated with the user T1's user identifier and stored in the acoustic information database.

[0069] For any user, it can be done according to... Figure 3 A and Figure 3 The process shown in B is used to perform voice authorization, so as to associate the user's acoustic feature information and the user's user identifier and store them in the acoustic information database.

[0070] In some embodiments, the application running on the terminal also includes a friend authorization control. Figure 4 A shows a schematic diagram of the second page in the application, such as Figure 4 As shown in Figure A, a friend authorization control 331 is provided on the second page 330. If a triggering operation on the friend authorization control 331 is detected, access can be made... Figure 4 As shown in B, the friend selection page 340 provides user identifiers for user T1's friends. Therefore, the user indicated by the selected user identifier in friend selection page 330 is the authorized friend authorized to use user T1's voice. Subsequently, the terminal can upload the user identifiers of the friends selected in friend selection page 330 to the server, allowing the server to construct the authorized friend set for each user based on the reported user identifiers of the authorized friends.

[0071] Of course, in other embodiments, the friend authorization control 312 and the authorization control 311 can be displayed on the same page, and can be set according to actual needs.

[0072] In some embodiments, it is possible to enter Figure 3After capturing the user's voice (as shown on page 320 in section B), a prompt message appears asking the user to authorize a friend to use the voice recording. If the user allows the authorized friend to use the captured voice recording, the user will be redirected to... Figure 4 The friend selection page 340 shown in B is for users to select friends on this friend selection page 340.

[0073] In some embodiments, based on the target user's authorization to use voice and the target user's friend's authorization to the target user to use the friend's voice, acoustic feature information can be randomly selected from the target user's acoustic feature information and the acoustic feature information of the target user's authorized friend for subsequent speech synthesis.

[0074] In some embodiments, a first process set that prioritizes the acoustic feature information of the target user and a second process set that prioritizes the acoustic feature information of the target user's authorized friends can be set. Thus, if it is determined that the target process belongs to the first process set, the acoustic feature information of the target user is obtained in step 220; otherwise, if it is determined that the target process belongs to the second process set, the acoustic feature information of an authorized friend of the target user is obtained in step 220.

[0075] Of course, if the target user's authorized friend set is empty, then in step 220, the target user's acoustic feature information is obtained; similarly, if the target user has not authorized the use of their own voice, then in step 220, the acoustic feature information of one of the target user's authorized friends is obtained.

[0076] In some embodiments, to avoid empty queries, a second tag can be added to the user identifier of a user authorized to use their own voice or a user with authorized friends. Thus, after obtaining the target user identifier of the target user, if it is determined that the target user identifier has been marked with the second tag, the step of obtaining the acoustic feature information of the associated user is executed; otherwise, if it is determined that the target user identifier has not been marked with the second tag, the first default prompt voice for the target node is obtained and played. The first default prompt voice is obtained by speech synthesis based on the default acoustic feature information and the prompt text corresponding to the target node. That is, in this case, it is not necessary to execute the step of obtaining the acoustic feature information of the associated user and the subsequent steps.

[0077] Acoustic feature information is used to indicate the acoustic features of the corresponding user. Specifically, the acoustic feature information parameters are used to indicate at least the timbre of the corresponding user. Furthermore, the acoustic feature information parameters may also include at least one of the following: speech rate, volume, frequency (or pitch) of the corresponding user.

[0078] Step 230: Speech synthesis is performed based on the acoustic feature information of the associated user and the prompt text to obtain a prompt voice that simulates the associated user speaking.

[0079] Speech synthesis, also known as text-to-speech (TTS) technology, primarily converts text into audible sound information. In this application, speech synthesis refers to converting prompt text into speech that simulates the user's speech.

[0080] In some embodiments, step 230 includes: performing text analysis on the prompt text to obtain a phoneme sequence; and using a vocoder to perform speech synthesis based on the phoneme sequence and the acoustic feature information of the associated user to obtain a prompt voice simulating the associated user speaking.

[0081] Since a user's voice is composed of different phonemes, such as the initials and finals of each character in Chinese, these initials and finals are called phonemes. Combining phonemes yields the pronunciation of each character. Text analysis of the prompt text refers to performing linguistic analysis on the prompt text to determine the phonemes corresponding to each character in the prompt text. The phonemes corresponding to each character in the prompt text are arranged according to the order of the characters in the prompt text, thus obtaining a phoneme sequence.

[0082] Specifically, a phoneme dictionary can be constructed, containing the phonemes corresponding to each character. The phonemes for each character in the prompt text can be retrieved from this dictionary. It is understandable that different languages ​​have different pronunciation rules; therefore, for each language, a phoneme dictionary corresponding to its pronunciation rules needs to be constructed.

[0083] A vocoder can be a neural network vocoder built using a neural network, a linear predictive vocoder, etc., without being specifically limited here. A neural network vocoder can be built using recurrent neural networks, feedforward neural networks, convolutional neural networks, fully connected neural networks, long short-time recurrent neural networks, transformer networks, etc.

[0084] In some embodiments, the vocoder can generate a Mel spectrum corresponding to the prompt text based on acoustic feature information and phoneme sequence, and then decode the Mel spectrum to obtain a prompt voice that simulates the user's speech.

[0085] Specifically, the vocoder can first determine the phoneme duration information based on the phoneme sequence and the acoustic feature information of the associated user. This phoneme duration information can indicate the duration of each phoneme in the phoneme sequence, the interval duration between adjacent phonemes, and whether the phoneme is stressed, etc. Then, based on the phoneme duration information, the phoneme sequence, and the timbre information of the associated user, a Mel spectrum corresponding to the prompt text is generated. Then, the Mel spectrum is decoded to obtain the prompt voice simulating the associated user speaking.

[0086] In some embodiments, the acoustic feature information of the associated user may only include timbre information used to indicate the associated user, while some acoustic features, such as volume and speech rate, may be preset. Therefore, in step 230, speech synthesis can be performed by combining the timbre information of the associated user, the preset volume, the preset speech rate, and the prompt text. In this case, it can be ensured that the synthesized prompt speech for different associated users differs in timbre, but the volume and speech rate remain consistent with the preset volume and preset speech rate.

[0087] In some embodiments, before step 230, the method further includes: obtaining acoustic parameter adjustment information corresponding to the target node; adjusting acoustic feature information according to the acoustic parameter adjustment information, and using the adjusted acoustic feature information for speech synthesis; in this embodiment, in step 230, speech synthesis is performed according to the adjusted acoustic feature information of the associated user and the prompt text.

[0088] In the target workflow, the importance of the content to be prompted varies at each node. For some important nodes in the target workflow, to ensure effective user reminders, acoustic parameter adjustment information can be pre-set for these nodes. This acoustic parameter adjustment information can be used to indicate increasing volume, slowing down speech, etc. Then, based on this acoustic parameter adjustment information, the corresponding acoustic parameters in the associated user's acoustic feature information are adjusted. It is understood that, because it is necessary to ensure that the obtained prompt voice simulates the associated user's speech, this acoustic parameter adjustment information does not involve timbre adjustment.

[0089] In some embodiments, for each node, the node identifier corresponding to the node is associated with and stored with the corresponding acoustic feature adjustment information, so that the acoustic feature adjustment information corresponding to the target node can be obtained based on the target node identifier of the target node.

[0090] The acoustic parameter adjustment information indicates the acoustic parameters that need to be adjusted and the amount of adjustment. Based on this information, the corresponding acoustic parameters are then adjusted. It is understandable that the acoustic parameters requiring adjustment may differ at different nodes, as may the amount of adjustment.

[0091] Of course, if the acoustic parameter adjustment information corresponding to the target node indicates that no adjustment is needed, then speech synthesis is performed directly based on the acquired acoustic feature information and the acquired prompt text corresponding to the target node.

[0092] Step 240: Play the prompt voice.

[0093] By playing the prompt voice, users can at least receive the corresponding prompt content based on the voice prompt.

[0094] In some embodiments, before step 240, the method further includes: acquiring ambient audio of the target user's environment; determining the ambient volume of the ambient audio; and determining playback parameters for the prompt voice based on the ambient volume. In this embodiment, step 240 includes: playing the prompt voice according to the playback parameters.

[0095] Ambient volume refers to the volume of ambient audio. It's understandable that when the ambient volume in the target user's environment is high, it can easily interfere with the user, potentially causing them to be unable to clearly hear the prompts. Therefore, in such cases, to reduce the impact of ambient audio on the target user, the playback parameters of the prompts can be adjusted.

[0096] Specifically, an audio acquisition device is installed on the terminal near the target user to acquire ambient audio. After obtaining the prompt voice corresponding to the target node through steps 210-230, the volume of the prompt voice can be determined, and then the playback parameters are determined based on the volume of the prompt voice and the ambient volume of the ambient audio.

[0097] In some embodiments, a first functional relationship between playback parameters and the volume of the prompt voice and the ambient volume can be preset, thereby allowing the playback parameters to be determined based on this first functional relationship.

[0098] Playback parameters may include playback rate (also known as playback speed parameter) and playback volume. In some embodiments, if it is determined that the ambient volume is lower than a first threshold and the ambient volume is lower than the volume of the prompt voice, the playback parameters can be determined to be preset playback parameters, such as a playback speed parameter of 1 and a playback volume equal to the volume of the prompt voice. If the ambient volume is higher than a second threshold and the second threshold is higher than the first threshold, the playback rate can be reduced (e.g., played with a playback speed parameter less than 1) and / or the playback volume of the prompt voice can be increased.

[0099] In other embodiments, the terminal plays a prompt voice message. Before step 240, the method further includes: determining the distance between the target user and the terminal; and determining the playback parameters of the prompt voice message based on the distance. In this embodiment, step 240 includes: playing the prompt voice message according to the playback parameters.

[0100] It is understandable that sound attenuates during its propagation through the environment. Therefore, if the target user is far from the terminal, they may not be able to clearly hear the prompt voice. Thus, in this embodiment, the playback parameters of the prompt voice are flexibly determined based on the distance between the target user and the terminal to reduce the probability of the user being unable to hear the prompt voice due to distance.

[0101] In some embodiments, the terminal may be equipped with a distance sensor to detect distance and obtain the distance between the target user and the terminal. This distance sensor may be an infrared ranging sensor, a radar ranging sensor, etc., and is not specifically limited thereto.

[0102] In other embodiments, if the target user enters the target process by triggering facial recognition, and the target user's facial image is obtained through the image acquisition device on the terminal, the distance between the target user and the terminal can be calculated using the acquired facial image.

[0103] In some embodiments, a second functional relationship can be set between playback parameters, the volume of the prompt voice, and the distance between the target user and the terminal, to determine the playback parameters of the prompt voice based on this second functional relationship. Determining the playback parameters based on the distance between the target user and the terminal allows for increasing the playback volume of the prompt voice when the target user is far from the terminal, and decreasing the playback volume or setting the playback volume of the prompt audio to the same level as the prompt voice volume when the target user is close to the terminal.

[0104] In other embodiments, the terminal plays a prompt voice message. Before step 240, the method further includes: acquiring ambient audio of the target user's environment; determining the distance between the target user and the terminal; and determining playback parameters for the prompt voice message based on the ambient volume of the ambient audio and the distance between the target user and the terminal. In this embodiment, step 240 includes: playing the prompt voice message according to the determined playback parameters.

[0105] In this embodiment, the playback parameters of the prompt voice are determined by combining the ambient volume and the distance between the target user and the terminal, thereby further ensuring that the target user can accurately hear the played prompt voice. Specifically, a third functional relationship between the playback parameters and the ambient volume, the distance between the target user and the terminal, and the volume of the prompt voice can be preset, and the playback parameters of the prompt voice can be determined based on this third functional relationship.

[0106] In this application, when the target process reaches a target node requiring voice prompts, a voice prompt simulating the target user's speech is generated using the acoustic feature information of the target user participating in the target process and the corresponding prompt text for that target node; alternatively, a voice prompt simulating the authorized friend of the target user participating in the target process is generated using the acoustic feature information of that authorized friend and the corresponding prompt text for that target node. The resulting voice prompt is then played. This achieves flexible voice prompts simulating either the speech of the target user participating in the target process or the speech of their authorized friend, effectively solving the problem of monotonous voice prompts in related technologies. Moreover, for the target user, whether the voice prompts simulate the speech of the target user or their authorized friend, the voice is familiar, creating a familiar voice environment and improving the user experience.

[0107] In some embodiments, the target node is the first node in the target process that requires voice prompts after the target user identifier is obtained; such as Figure 5 As shown, step 220 includes:

[0108] Step 510: Send a query request to the server. The query request is used to request the server to query the acoustic feature information of the associated user. The query request carries the target user identifier.

[0109] Step 520: Receive the acoustic feature information of the associated user returned by the server; the server queries the acoustic information to obtain the acoustic feature information of the associated user based on the target user identifier.

[0110] In this scenario, the acoustic information database is deployed on the server. Since the target node is the first node in the target process that needs to provide voice prompts after obtaining the target user's identifier, it needs to request the server to query and obtain the acoustic feature information of the associated user.

[0111] In some embodiments, if the associated user is an authorized friend of the target user, the server, upon receiving the query request, obtains the set of authorized friends of the target user based on the target user identifier, obtains an authorized friend identifier from the authorized friend set, and then obtains the acoustic feature information corresponding to the authorized friend identifier from the acoustic information database as the acoustic feature information of the associated user.

[0112] Of course, if the associated user is the target user, the server will retrieve the acoustic feature information corresponding to the target user identifier from the acoustic information database and send it to the terminal.

[0113] In some embodiments, after step 520, the method further includes: associating the user identifier of the associated user with the acoustic feature information of the associated user and storing it in a cache; when the target process moves to the target node and a first node that needs to provide a voice prompt is reached, retrieving the acoustic feature information of the associated user from the cache; performing speech synthesis based on the acoustic feature information retrieved from the cache and the prompt text corresponding to the first node to obtain a first prompt voice simulating the associated user speaking; and playing the first prompt voice.

[0114] In other words, if the node in the target process that requires voice prompts is not the first node in the target process that requires voice prompts after obtaining the target user's identifier, then there is no need to query the acoustic feature information from the server again. Instead, the acoustic feature information of the associated user can be obtained directly from the cache, thereby improving the utilization of network resources.

[0115] In some embodiments, if the target process is detected to have ended, the acoustic feature information of the associated user in the cache can be deleted, and the acoustic feature information can be requested from the server again after the target process is started next time.

[0116] Understandably, in Figure 5 In the corresponding embodiment, after a target process is started, during the process of the target process flowing between nodes, the user corresponding to the target process at each node is the same user. In other words, the target process is a process that only requires the participation of one user.

[0117] Understandably, if the terminal has sufficient processing and storage capabilities, the acoustic information database can also be deployed locally on the terminal. Thus, the terminal can obtain the acoustic feature information of associated users from the local acoustic information database as needed.

[0118] In some embodiments, the target process is to transfer to the target node before obtaining the target user identifier. The method further includes: obtaining a first default prompt voice for the target node, wherein the first default prompt voice is obtained by speech synthesis based on default acoustic feature information and the prompt text corresponding to the target node; and playing the first default prompt voice.

[0119] Because generating a simulated user-related voice prompt requires first obtaining the acoustic feature information of the associated user based on the target user identifier, and in this embodiment, the target process proceeds to the target node before obtaining the target user identifier, it is impossible to generate a simulated user-related voice prompt in this situation. Therefore, in this case, a first default prompt voice is synthesized based on the default acoustic feature information and the prompt text, and then played. Consequently, when the target process proceeds to the target node, a significant amount of time is spent waiting to obtain the target user identifier and subsequently the acoustic feature information of the associated user, resulting in an inability to provide timely and effective voice prompts to the user.

[0120] In some embodiments, the acoustic feature information of the associated user carries validity information. In this embodiment, the method further includes: verifying the validity of the acoustic feature information based on the validity information; if it is determined that the acoustic feature information is invalid, playing a first default prompt voice for the target node, the first default prompt voice being obtained by speech synthesis based on the default acoustic feature information and the prompt text; if it is determined that the acoustic feature information is valid, then executing step 230.

[0121] In some embodiments, the effective information includes at least one of effective node information and effective time information; if it is determined that the effective node indicated by the effective node information does not include the target node, or if it is determined that the current time is not within the effective time period indicated by the effective time information, then the acoustic feature information is determined to be invalid.

[0122] Among them, the effective node information is used to indicate the node to which the corresponding acoustic feature information applies (i.e., the effective node). It can be understood that the effective node information can indicate one or more effective nodes.

[0123] The effective time information is used to indicate the time period during which the corresponding acoustic feature information is effective (i.e., the effective time period). In some embodiments, after the server obtains the acoustic feature information of the associated user from the acoustic information database, an effective time period is set for the acoustic feature information. For example, the duration required for all nodes in the target process to complete one cycle is used as the duration of the effective time period, and the time when the server obtains the acoustic feature information from the acoustic information database is used as the start time of the effective time period. In other embodiments, the effective time period can also be set according to actual needs.

[0124] If the effective information only includes effective node information, and it is determined that the effective node indicated by the effective node information includes the target node, then the obtained acoustic feature information is determined to be valid.

[0125] If the effective information only includes the effective time information, and it is determined that the current time is within the effective time period indicated by the effective time information, then the obtained acoustic feature information is determined to be valid.

[0126] When the effective information includes effective node information and effective time information, if it is determined that the effective node indicated by the effective node information includes the target node, and it is determined that the current time is within the effective time period indicated by the effective time information, then the obtained acoustic feature information is determined to be valid.

[0127] In some embodiments, the effective information may also be applicable user information, which is used to indicate the user identifier (referred to as applicable user identifier for ease of description) of the user to whom the corresponding acoustic feature information is applicable. In this case, if it is determined that the applicable user identifier indicated by the applicable user information does not include the target user identifier, the acoustic feature information is determined to be invalid.

[0128] In this embodiment, before performing speech synthesis, the validity of the acquired acoustic feature information of the associated user is verified to ensure the validity of the acoustic feature information used for speech synthesis. This avoids the situation where the acoustic feature information actually used for speech synthesis is not the acoustic feature information of the associated user due to the acquisition of incorrect acoustic feature information, which would lead to the obtained prompt voice not matching the actual required prompt voice.

[0129] The solution of this application will now be described with reference to a specific embodiment. In this embodiment, the target process is a face payment process.

[0130] Figure 6 This is a flowchart illustrating face payment according to a specific embodiment of this application, such as... Figure 6 As shown, after initiating the facial payment process, the payment terminal executes steps 610-630. Step 610 involves acquiring a facial image. Specifically, in response to the user's facial payment operation, the image acquisition device can be activated to acquire the target user's facial image. Step 620: Obtain the target user identifier of the target user; Step 630: Send a query request to the server based on the target user identifier; Subsequently, the server responds to the query request and executes steps 640 and 650. Step 640: Obtain the acoustic feature information of the target user; Specifically, the server deploys an acoustic information database, and the query request carries the target user identifier. Therefore, the server queries the acoustic information database based on the target user identifier to obtain the acoustic feature information of the target user; Step 650: Return the acoustic feature information of the target user; After receiving the acoustic feature information returned from the server, the payment terminal executes step 660: Speech synthesis and playback; Specifically, the payment terminal stores the prompt text corresponding to each node in the face payment process that requires voice prompts. Therefore, when the target process reaches the target node that requires voice prompts, the payment terminal performs speech synthesis based on the obtained acoustic feature information of the target user and the prompt text corresponding to the target node.

[0131] In this embodiment, step 620 includes: in response to a face payment operation, acquiring a face image of the target user; performing face recognition on the face image to obtain a recognition result; and obtaining the target user identifier of the target user based on the recognition result.

[0132] Figure 7 This is a schematic diagram illustrating the fields included in the acoustic feature information according to an embodiment of this application, such as... Figure 7 As shown, the acoustic feature information includes the Meta field, OpenId field, Volume field, Freq field, Speed ​​field, and Pitch field. The OpenId field represents the user identifier, the Volume field represents the volume, the Freq field represents the frequency, the Speed ​​field represents the speech rate, and the Pitch field represents the timbre. The Meta field represents the activation information. As described above, the activation information may include at least one of activation node information, activation time information, and applicable user information. In some embodiments, the applicable user information may indicate the user identifier of the user who can use the device; for example, the user can be all of the user's friends, starred friends, or a specified number of friends.

[0133] Figure 8 The diagram illustrates the nodes in the facial recognition payment process that require voice prompts, such as... Figure 8 The facial recognition payment process includes three nodes requiring voice prompts: the facial image acquisition node, the payment confirmation node, and the payment result notification node. When the application navigates to the facial recognition page, the facial recognition payment process has proceeded to the facial image acquisition node; when the application navigates to the payment confirmation page, the process has proceeded to the payment confirmation node; and when the application navigates to the payment result confirmation page, the process has proceeded to the payment result notification node.

[0134] At the face image acquisition node, a voice prompt for face recognition is broadcast; at the payment confirmation node, a voice prompt for payment confirmation is broadcast; and at the payment result prompt node, a voice prompt for payment result is broadcast. The payment result is used to indicate whether the payment was successful. Furthermore, the payment result may also include the amount paid, etc.

[0135] Figure 9 The process of providing voice prompts at a target node is illustrated, such as... Figure 9As shown, after obtaining the acoustic feature information of the target user, it is determined whether the obtained acoustic feature information is valid. If valid, the obtained acoustic feature information is sent to the speech generation component; if invalid, default acoustic feature information is sent to the speech generation component. Then, the speech synthesis component performs speech synthesis. When the target user's acoustic feature information is valid, the speech synthesis component synthesizes speech based on the target user's acoustic feature information and the prompt text corresponding to the target node to obtain the prompt speech. When the target user's acoustic feature information is invalid, the speech synthesis component synthesizes speech based on the default acoustic feature information and the prompt text corresponding to the target node to obtain the prompt speech (the prompt speech obtained in this case is the first default prompt speech mentioned above). The generated prompt speech is then applied to the speech playback component, which plays the prompt speech.

[0136] For both the payment confirmation and payment result notification nodes in the facial recognition payment process, you can follow... Figure 9 The process shown is used to provide voice prompts.

[0137] For the face image acquisition node in the face payment process, since the target user's identifier is obtained through face recognition based on the face image acquired at the face image acquisition node, the target user's identifier has not yet been obtained when the face payment process transitions to the face image acquisition node. Therefore, at the face image acquisition node, speech synthesis can be performed using default acoustic feature information and the corresponding prompt text. When the current node in the face payment process is the face image acquisition node, a second default prompt voice is obtained for the face image acquisition node. This second default prompt voice is obtained through speech synthesis based on default acoustic feature information and the corresponding prompt text; the second default prompt voice is then played. In other words, at the face image acquisition node, since the target user's identifier has not yet been obtained, default acoustic feature information can be used for speech synthesis.

[0138] Figure 10 This is a flowchart illustrating the use of voice by an authorized friend and the synthesis of voice based on the voice feature information of the authorized friend, according to an embodiment of this application. Figure 10 Steps 1011-1013 in the diagram illustrate a flowchart for authorizing friends to use voice. For example... Figure 10 As shown, the process includes: the payment terminal executes step 1011, where the first user authorizes a friend; then, the payment application background executes step 1012, synchronizing the authorized friend information; then, the application authorization service executes step 1013, synchronizing the authorized friend set of the friend.

[0139] Figure 11The authorized friend sets of user T1 and user T2 are shown. In this embodiment, user T1's user identifier is user identifier A, and user T2's user identifier is user identifier B. Figure 11 As shown, user T1's authorized friends include users with user IDs B, C, and D, and user T2's authorized friends include users with user IDs A, E, and F.

[0140] like Figure 10 As shown, the process of obtaining the voice of an authorized friend for voice prompts includes: the payment terminal executes step 1021, the second user initiates the face payment process; the payment application background executes step 1022, requesting the second user's authorized friend set; the application authorization service executes step 1023, reading the second user's authorized friend set; subsequently, the application authorization service executes step 1024, returning the second user's authorized friend set; the payment application background executes step 1025, requesting the acoustic feature information of the authorized friend; the acoustic feature information query service executes step 1026, returning the acoustic feature information of the authorized friend; the payment application background executes step 1027, sending the acoustic feature information of the authorized friend. Afterwards, the payment terminal can perform speech synthesis based on the acoustic feature information of the authorized friend and the prompt text corresponding to the target node to obtain a prompt voice simulating the authorized friend speaking, and then play the prompt voice.

[0141] In related technologies, the audio corresponding to voice prompts is uniformly generated according to preset acoustic feature information. Therefore, the voice prompts sound too monotonous, resulting in a low user experience.

[0142] In this embodiment, regarding the technology of analyzing the voices of various users to obtain their acoustic feature information, at nodes in the facial recognition payment process that require voice prompts, speech synthesis is performed using the acoustic feature information of the target user and the corresponding prompt text to obtain a prompt voice simulating the target user's speech; or speech synthesis is performed using the acoustic feature information of the target user's authorized friend and the corresponding prompt text to obtain a prompt voice simulating the target user's authorized friend's speech. Then, the obtained prompt voice is played, achieving flexible voice prompts that simulate the speech of the user participating in the facial recognition payment process, or the speech of the authorized friend of the user participating in the facial recognition payment process. This allows for flexible voice prompts using familiar voices for different users, bringing surprise and novelty to users, eliminating their resistance and unfamiliarity with facial recognition payment devices, and improving the user experience.

[0143] The following describes an apparatus embodiment of this application, which can be used to perform the methods described in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the method embodiments described in the above embodiments of this application.

[0144] Figure 12 This is a block diagram of a voice prompt device according to an embodiment of this application, such as... Figure 12 As shown, the voice prompt device includes: a prompt text acquisition module 1210, used to acquire the prompt text corresponding to the target node if the target process jumps to the target node that requires voice prompts; an acoustic feature information acquisition module 1220, used to acquire the acoustic feature information of the associated user based on the target user identifier of the target user participating in the target process, wherein the associated user is the target user or an authorized friend of the target user; a speech synthesis module 1230, used to synthesize speech based on the acoustic feature information of the associated user and the prompt text to obtain a prompt voice simulating the associated user speaking; and a playback module 1240, used to play the prompt voice.

[0145] In some embodiments, the voice prompt device further includes: an acoustic parameter adjustment information acquisition module, used to acquire acoustic parameter adjustment information corresponding to the target node; and an adjustment module, used to adjust the acoustic feature information according to the acoustic parameter adjustment information, wherein the adjusted acoustic feature information is used for speech synthesis for the target node.

[0146] In some embodiments, the speech synthesis module 1230 includes: a text analysis unit for performing text analysis on the prompt text to obtain a phoneme sequence; and a speech synthesis unit for performing speech synthesis by a vocoder based on the phoneme sequence and the acoustic feature information of the associated user to obtain a prompt speech simulating the associated user speaking.

[0147] In some embodiments, the voice prompt device further includes: an ambient audio acquisition module for acquiring ambient audio of the environment where the target user is located; an ambient volume determination module for determining the ambient volume of the ambient audio; and a first playback parameter determination module for determining playback parameters of the prompt voice based on the ambient volume. In this embodiment, the playback module 1240 is further configured to play the prompt voice according to the playback parameters.

[0148] In other embodiments, the terminal plays a prompt voice message. The voice prompt device further includes: a distance determination module for determining the distance between the target user and the terminal; and a second playback parameter determination module for determining the playback parameters of the prompt voice message based on the distance. In this embodiment, the playback module 1240 is further configured to play the prompt voice message according to the playback parameters.

[0149] In some embodiments, the target node is the first node in the target process that needs to provide voice prompts after the target user identifier is obtained; in this embodiment, the acoustic feature information acquisition module 1220 includes: a query request sending unit, used to send a query request to the server, the query request being used to request the server to query the acoustic feature information of the associated user; the query request carries the target user identifier; a receiving unit, used to receive the acoustic feature information of the associated user returned by the server; the server obtains the acoustic feature information of the associated user from the acoustic information based on the target user identifier.

[0150] In some embodiments, the voice prompt device further includes: a first storage module for storing the user identifier of the associated user and the acoustic feature information of the associated user in a cache; a second acoustic feature information acquisition module for acquiring the acoustic feature information of the associated user from the cache when a first node that requires a voice prompt is reached after the target process moves to the target node; a second speech synthesis module for performing speech synthesis based on the acoustic feature information acquired from the cache and the prompt text corresponding to the first node to obtain a first prompt voice simulating the associated user speaking; and a second playback module for playing the first prompt voice.

[0151] In some embodiments, the voice prompt device further includes: a voice acquisition module for acquiring the voice of an associated user in response to a voice authorization operation; a voice parsing module for parsing the voice of the associated user to obtain the acoustic feature information of the associated user; and a second storage module for storing the acoustic feature information of the associated user in an acoustic information database in association with the user identifier of the associated user.

[0152] In some embodiments, the target process is to transfer to the target node before obtaining the target user identifier; the voice prompt device further includes: a first default prompt voice acquisition module, used to acquire a first default prompt voice for the target node, the first default prompt voice being obtained by speech synthesis based on default acoustic feature information and the prompt text corresponding to the target node; and a third playback module, used to play the first default prompt voice.

[0153] In other embodiments, the associated user is the target user's authorized friend. The acoustic feature information acquisition module 1220 includes: an authorized friend set determination unit, used to determine the target user's authorized friend set based on the target user identifier, the authorized friend set including authorized friend identifiers, the authorized friend identifiers referring to the user identifiers of friends who authorize the target user to use voice; an authorized friend identifier acquisition unit, used to acquire an authorized friend identifier from the target user's authorized friend set; and an acoustic feature information acquisition unit, used to acquire the acoustic feature information of the corresponding authorized friend based on the acquired authorized friend identifier.

[0154] In some embodiments, the acoustic feature information of the associated user carries validity information; the voice prompt device further includes: a validity verification module, used to verify the validity of the acoustic feature information according to the validity information; a fourth playback module, used to play a first default prompt voice for the target node if it is determined that the acoustic feature information is invalid, the first default prompt voice is obtained by speech synthesis based on the default acoustic feature information and the prompt text; if it is determined that the acoustic feature information is valid, then proceed to the speech synthesis module 1230.

[0155] In some embodiments, the effectiveness information includes at least one of effectiveness node information and effectiveness time information; the validity verification module is further configured to: determine that the acoustic feature information is invalid if it is determined that the effectiveness node indicated by the effectiveness node information does not include the target node, or if it is determined that the current time is not within the effectiveness time period indicated by the effectiveness time information.

[0156] Figure 13 A schematic diagram of a computer system suitable for implementing an electronic device according to embodiments of this application is shown. The electronic device may be... Figure 1 Regarding the terminals and other components, it should be noted that... Figure 13 The computer system 1300 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0157] like Figure 13 As shown, the computer system 1300 includes a Central Processing Unit (CPU) 1301, which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, based on programs stored in Read-Only Memory (ROM) 1302 or programs loaded from storage portion 1308 into Random Access Memory (RAM) 1303. The RAM 1303 also stores various programs and data required for system operation. The CPU 1301, ROM 1302, and RAM 1303 are interconnected via a bus 1304. An Input / Output (I / O) interface 1305 is also connected to the bus 1304.

[0158] The following components are connected to I / O interface 1305: an input section 1306 including a keyboard, mouse, etc.; an output section 1307 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to I / O interface 1305 as needed. Removable media 1311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1310 as needed so that computer programs read from them can be installed into storage section 1308 as needed.

[0159] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1309, and / or installed from removable medium 1311. When the computer program is executed by central processing unit (CPU) 1301, it performs various functions defined in the system of this application.

[0160] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0161] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0162] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0163] In another aspect, this application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable storage medium carries computer-readable instructions that, when executed by a processor, implement the methods in any of the above embodiments.

[0164] According to one aspect of this application, an electronic device is also provided, comprising: a processor; and a memory storing computer-readable instructions that, when executed by the processor, implement the methods of any of the above embodiments.

[0165] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the methods of any of the above embodiments.

[0166] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0167] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the methods according to the embodiments of this application.

[0168] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0169] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A voice prompt method, characterized in that, include: If the target process reaches a target node that requires voice prompts, obtain the prompt text corresponding to the target node; Based on the target user identifier of the target user participating in the target process, the acoustic feature information of the associated user is obtained from the acoustic information database. The associated user is the target user or an authorized friend indicated by an authorized friend identifier in the target user's authorized friend set. The authorized friend of the target user is a user who authorizes the target user to use voice and has a friend relationship with the target user. If the target user's authorized friend set is empty, the target user is taken as the associated user. If the acoustic information database does not include the acoustic feature information of the target user, the authorized friend of the target user will be used as the associated user; if an operation triggered by the user on the authorization control on the first page is detected, a voice collection prompt page will be displayed. If the user's voice is captured on the voice capture prompt page, the acoustic feature information extracted from the user's voice and the user's user identifier will be associated and stored in the acoustic information database; After the user's voice is captured from the voice capture page, a prompt message is displayed to authorize a friend to use the voice. If the user allows the authorized friend to use the voice, the user enters the friend selection page and adds the user's user ID to the authorized friend set of the friend selected on the friend selection page. Based on the acoustic feature information of the associated user and the prompt text, speech synthesis is performed to obtain a prompt voice that simulates the speech of the associated user; Play the aforementioned prompt voice.

2. The method according to claim 1, characterized in that, Before performing speech synthesis based on the acoustic feature information of the associated user and the prompt text to obtain a prompt voice simulating the speech of the associated user, the method further includes: Obtain the acoustic parameter adjustment information corresponding to the target node; The acoustic feature information is adjusted according to the acoustic parameter adjustment information, and the adjusted acoustic feature information is used for speech synthesis for the target node.

3. The method according to claim 1 or 2, characterized in that, The step of synthesizing speech based on the acoustic feature information of the associated user and the prompt text to obtain a prompt voice simulating the speech of the associated user includes: The prompt text is analyzed to obtain a phoneme sequence; The vocoder performs speech synthesis based on the phoneme sequence and the acoustic feature information of the associated user to obtain a prompt voice that simulates the associated user speaking.

4. The method according to claim 1, characterized in that, Before playing the prompt voice, the method further includes: Obtain the ambient audio of the target user's environment; Determine the ambient volume of the ambient audio; The playback parameters of the prompt voice are determined based on the ambient volume. Playing the prompt voice includes: Play the prompt voice according to the playback parameters.

5. The method according to claim 1, characterized in that, The method further includes, prior to the terminal playing the prompt voice, the method comprising: Determine the distance between the target user and the terminal; The playback parameters of the prompt voice are determined based on the distance. Playing the prompt voice includes: Play the prompt voice according to the playback parameters.

6. The method according to claim 1, characterized in that, The target node is the first node in the target process that requires voice prompts after the target user identifier is obtained; The step of obtaining acoustic feature information of associated users from the acoustic information database based on the target user identifier of the target user participating in the target process includes: A query request is sent to the server, the query request being used to request the server to query the acoustic feature information of the associated user; the query request carries the target user identifier; The server receives the acoustic feature information of the associated user returned by the server; the server queries the acoustic information database to obtain the acoustic feature information of the associated user based on the target user identifier.

7. The method according to claim 6, characterized in that, After receiving the acoustic feature information of the associated user returned by the server, the method further includes: The user identifier of the associated user is associated with the acoustic feature information of the associated user and stored in the cache; When the target process reaches the target node, the first node that needs to provide voice prompts retrieves the acoustic feature information of the associated user from the cache; Based on the acoustic feature information obtained from the cache and the prompt text corresponding to the first node, speech synthesis is performed to obtain a first prompt voice simulating the speech of the associated user; Play the first prompt voice message.

8. The method according to claim 1, characterized in that, The target process involves transferring to the target node before obtaining the target user identifier; the method further includes: A first default prompt voice for the target node is obtained, wherein the first default prompt voice is obtained by speech synthesis based on default acoustic feature information and the prompt text corresponding to the target node; Play the first default prompt voice.

9. The method according to claim 1, characterized in that, The acoustic feature information of the associated user carries the activation information; The method further includes: Verify the validity of the acoustic feature information based on the activation information; If the acoustic feature information is determined to be invalid, a first default prompt voice for the target node is played. The first default prompt voice is obtained by speech synthesis based on the default acoustic feature information and the prompt text. If the acoustic feature information is determined to be valid, then the process of synthesizing speech based on the acoustic feature information of the associated user and the prompt text is performed to obtain a prompt voice that simulates the associated user speaking.

10. The method according to claim 9, characterized in that, The effective information includes at least one of effective node information and effective time information; The step of verifying the validity of the acoustic feature information based on the validity information includes: If it is determined that the effective node indicated by the effective node information does not include the target node, or if it is determined that the current time is not within the effective time period indicated by the effective time information, then the acoustic feature information is determined to be invalid.

11. A voice prompt device, characterized in that, include: The prompt text acquisition module is used to acquire the prompt text corresponding to the target node if the target process jumps to the target node that requires voice prompts. The acoustic feature information acquisition module is used to acquire acoustic feature information of associated users from an acoustic information database based on the target user identifier of the target user participating in the target process. The associated user is the target user or an authorized friend indicated by an authorized friend identifier in the target user's authorized friend set. The authorized friend of the target user is a user who authorizes the target user to use voice and has a friend relationship with the target user. If the target user's authorized friend set is empty, the target user is taken as the associated user. If the acoustic information database does not include the acoustic feature information of the target user, the authorized friend of the target user will be used as the associated user; if an operation triggered by the user on the authorization control on the first page is detected, a voice collection prompt page will be displayed. If the user's voice is captured on the voice capture prompt page, the acoustic feature information extracted from the user's voice and the user's user identifier will be associated and stored in the acoustic information database; After the user's voice is captured from the voice capture page, a prompt message is displayed to authorize a friend to use the voice. If the user allows the authorized friend to use the voice, the user enters the friend selection page and adds the user's user ID to the authorized friend set of the friend selected on the friend selection page. The speech synthesis module is used to synthesize speech based on the acoustic feature information of the associated user and the prompt text to obtain a prompt voice that simulates the speech of the associated user; The playback module is used to play the prompt voice.

12. The apparatus according to claim 11, characterized in that, The voice prompt device also includes: An acoustic parameter adjustment information acquisition module is used to acquire acoustic parameter adjustment information corresponding to the target node; An adjustment module is used to adjust the acoustic feature information according to the acoustic parameter adjustment information, and the adjusted acoustic feature information is used for speech synthesis for the target node.

13. The apparatus according to claim 11 or 12, characterized in that, The speech synthesis module includes: A text analysis unit is used to perform text analysis on the prompt text to obtain a phoneme sequence; The speech synthesis unit is used by the vocoder to synthesize speech based on the phoneme sequence and the acoustic feature information of the associated user, so as to obtain a prompt speech that simulates the speech of the associated user.

14. The apparatus according to claim 11, characterized in that, The voice prompt device also includes: An environmental audio acquisition module is used to acquire the environmental audio of the environment in which the target user is located; An ambient volume determination module is used to determine the ambient volume of the ambient audio. The first playback parameter determination module is used to determine the playback parameters of the prompt voice based on the ambient volume. The playback module is configured to play the prompt voice according to the playback parameters.

15. The apparatus according to claim 11, characterized in that, The prompt voice is played by the terminal, and the voice prompt device further includes: A distance determination module is used to determine the distance between the target user and the terminal; The second playback parameter determination module is used to determine the playback parameters of the prompt voice based on the distance; The playback module is configured to play the prompt voice according to the playback parameters.

16. The apparatus according to claim 11, characterized in that, The target node is the first node in the target process that requires voice prompts after the target user identifier is obtained; The acoustic feature information acquisition module includes: A query request sending unit is used to send a query request to the server, the query request being used to request the server to query the acoustic feature information of the associated user; the query request carries the target user identifier. The receiving unit is used to receive the acoustic feature information of the associated user returned by the server; the server obtains the acoustic feature information of the associated user from the acoustic information database according to the target user identifier.

17. The apparatus according to claim 16, characterized in that, The voice prompt device also includes: The first storage module is used to associate and store the user identifier of the associated user with the acoustic feature information of the associated user in a cache. The second acoustic feature information acquisition module is used to acquire the acoustic feature information of the associated user from the cache when the target process is transferred to the target node and a voice prompt is required at the first node. The second speech synthesis module is used to perform speech synthesis based on the acoustic feature information obtained from the cache and the prompt text corresponding to the first node to obtain a first prompt voice that simulates the speech of the associated user. The second playback module is used to play the first prompt voice.

18. The apparatus according to claim 11, characterized in that, The target process is to proceed to the target node before obtaining the target user identifier; The voice prompt device also includes: The first default prompt voice acquisition module is used to acquire a first default prompt voice for the target node. The first default prompt voice is obtained by speech synthesis based on default acoustic feature information and the prompt text corresponding to the target node. The third playback module is used to play the first default prompt voice.

19. The apparatus according to claim 11, characterized in that, The acoustic feature information of the associated user carries activation information; the voice prompt device further includes: The validity verification module is used to verify the validity of the acoustic feature information based on the validity information. The fourth playback module is used to play a first default prompt voice for the target node if it is determined that the acoustic feature information is invalid. The first default prompt voice is obtained by speech synthesis based on the default acoustic feature information and the prompt text. If it is determined that the acoustic feature information is valid, the module performs speech synthesis based on the acoustic feature information of the associated user and the prompt text to obtain a prompt voice simulating the associated user speaking.

20. The apparatus according to claim 19, characterized in that, The effective information includes at least one of effective node information and effective time information; the validity verification module is configured to: determine that the acoustic feature information is invalid if it is determined that the effective node indicated by the effective node information does not include the target node, or if it is determined that the current time is not within the effective time period indicated by the effective time information.

21. An electronic device, characterized in that, include: processor; A memory storing computer-readable instructions that, when executed by the processor, implement the method as described in any one of claims 1-10.

22. A computer-readable storage medium having stored thereon computer-readable instructions that, when executed by a processor, implement the method as described in any one of claims 1-10.

23. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the method of any one of claims 1-10.