Call-out voice large model verbal skill generation method, storage medium, electronic equipment and product

By obtaining the collection of user interaction content and historical dialogue records, and using large language models to generate accurate voice reply speech, it solves the problem that intelligent robot reply speech cannot cover all scenarios, and improves user interaction experience and timeliness.

CN119943047APending Publication Date: 2025-05-06LINGXI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510107990.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing intelligent robots cannot cover user questions in all scenarios in telemarketing scenarios, and are prone to answering non-questions, resulting in poor user interaction experience.

Method used

By obtaining a collection of interactive content generated by interaction with the user, including the problem-related text related to the user's current problem and the node-related text of the business process node's next node, and combining historical dialogue records, a large language model is used to generate accurate voice reply speech.

Benefits of technology

It realizes the generation of accurate outbound voice speech in different real scenarios, ensuring the timeliness of interaction and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943047A_ABST
    Figure CN119943047A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent customer service, and particularly provides an outbound voice large model verbal skill generation method, a storage medium, electronic equipment and a product, and the method can comprise the steps: obtaining an interaction content set generated by interaction with a user, the interactive content set comprises a question associated text related to a current question of a user and a node associated text of a next service node of a service process node where the user is located; and based on the interactive content set and the historical dialogue record, generating a voice reply verbal skill corresponding to the current question. According to some embodiments of the invention, the timeliness and accuracy of voice verbal skill reply can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent customer service technology, and more specifically, to a method, storage medium, electronic device, and product for generating a large model of outbound voice scripts. Background Art

[0002] With the continuous development of artificial intelligence technology, intelligent robots are widely used in telemarketing scenarios in various fields. In telemarketing scenarios in different fields, in order to improve the interactive experience between users and intelligent robots, the broadcast process of intelligent robots is generally set in advance according to certain rules. Then, the flow of process nodes is realized through natural language processing technology. At the same time, each process node is set with the intelligent robot's reply script. When the interaction with the user enters a certain process node, the recorded recording can be directly broadcast. However, this reply script cannot cover user questions in all scenarios, and it is easy to give irrelevant answers.

[0003] Therefore, how to provide a more accurate technical solution for the method of generating large-scale outbound voice model scripts has become a technical problem that needs to be solved urgently. Summary of the invention

[0004] The purpose of some embodiments of the present application is to provide a method, storage medium, electronic device and product for generating a large model of outbound voice scripts. Through the technical solutions of the embodiments of the present application, accurate outbound voice scripts can be generated for different real scenarios, and the timeliness of the interaction is guaranteed, thereby improving the user experience.

[0005] In a first aspect, some embodiments of the present application provide a method for generating a large model of outbound voice scripts, comprising: obtaining a set of interactive content generated by interaction with a user, wherein the interactive content set includes: question-associated text related to a current question of the user and node-associated text of a next business node of a business process node where the user is located; based on the interactive content set and historical conversation records, generating a voice reply script corresponding to the current question.

[0006] Some embodiments of the present application generate corresponding voice reply scripts by acquiring a set of interactive content associated with gear issues and business process nodes generated by the interaction between users and intelligent robots, and combining it with historical conversation records. This can achieve timely generation of voice reply scripts based on the current scenario with high accuracy, ensure the timeliness of the interaction, and improve the user experience.

[0007] In some embodiments, obtaining the interactive content set generated by the interaction with the user includes: obtaining the question-related text and obtaining the node-related text; and constructing the interactive content set based on the question-related text and the node-related text.

[0008] Some embodiments of the present application construct an interactive content set by acquiring question-related text and node-related text, thereby providing data support for the subsequent generation of voice response scripts with higher accuracy.

[0009] In some embodiments, obtaining the question-associated text includes: inputting the conversation content and integration requirements generated by the interaction with the user into a large language model, identifying and outputting the current question; matching the current question with a preset answer in a speech set to obtain the question-associated text.

[0010] Some embodiments of the present application use a large language model, conversation content, and integrated requirements to identify the user's current problem, and then match the problem-related text with a set of scripts, providing data support for the subsequent generation of higher-precision voice response scripts.

[0011] In some embodiments, obtaining the node-associated text includes: inputting the conversation content and process node content into the large language model, identifying and outputting the next business node; and searching the node-associated text corresponding to the next business node from the speech set.

[0012] Some embodiments of the present application identify the user's next business node through a large language model, conversation content, and process node content, and then match the node-associated text with a set of speech scripts, providing data support for the subsequent generation of high-precision voice response speech scripts.

[0013] In some embodiments, generating the voice reply speech corresponding to the current question based on the set of interactive content and the historical conversation records includes: using streaming speech synthesis technology to synchronously generate the voice reply speech while acquiring the reply text corresponding to the question-associated text, the node-associated text and the historical conversation records, wherein the reply text is the original text in the question-associated text and the node-associated text or a combination of multiple text fragments.

[0014] Some embodiments of the present application use streaming speech synthesis technology to simultaneously generate voice response words while acquiring the corresponding reply text, which can improve the efficiency of speech synthesis and solve the delay problem of word response.

[0015] In some embodiments, the synchronous generation of the voice reply speech includes: obtaining the most similar text corresponding to the reply text from the interactive content set; calculating the most similar text and the reply text to obtain at least one editing interval; and simultaneously editing the at least one editing interval to obtain the voice reply speech.

[0016] Some embodiments of the present application can obtain voice reply words by finding the most similar text of the reply text from the interactive content collection and then calculating and editing it. The streaming speech synthesis method solves the reply delay problem and ensures the efficiency of the words reply.

[0017] In some embodiments, the at least one editing interval is edited simultaneously to obtain the voice reply words, including: creating a masking matrix equal to the length of the reply text, wherein the elements of each editing interval in the masking matrix have different values ​​from the elements of the non-editing interval; after masking the most similar text with the masking matrix, performing speech synthesis processing on each editing interval to obtain a speech editing result; integrating the speech editing result with the audio of the non-editing interval to obtain the voice reply words, thereby ensuring the timeliness of the interaction between the user and the intelligent robot.

[0018] Some embodiments of the present application edit the edited interval and integrate it with the audio of the non-edited interval to obtain voice response words, thereby improving the efficiency of speech synthesis.

[0019] In a second aspect, some embodiments of the present application provide a device for generating speech scripts for a large model of outbound voice calls, comprising: a content acquisition module, used to acquire a set of interactive content generated by interaction with a user, wherein the interactive content set includes: question-associated text related to the user's current question and node-associated text of the next business node of the business process node where the user is located; a speech generation module, used to generate voice reply speech corresponding to the current question based on the interactive content set and historical conversation records.

[0020] In a third aspect, some embodiments of the present application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the method described in any embodiment of the first aspect.

[0021] In a fourth aspect, some embodiments of the present application provide an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor can implement a method as described in any embodiment of the first aspect when executing the program.

[0022] In a fifth aspect, some embodiments of the present application provide a computer program product, wherein the computer program product comprises a computer program, wherein the computer program, when executed by a processor, can implement the method described in any embodiment of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of some embodiments of the present application, the drawings required for use in some embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0024] Figure 1 A system diagram for generating outbound voice large model speech scripts provided for some embodiments of the present application;

[0025] Figure 2 One of the flow charts of the method for generating the large model of outbound voice scripts provided in some embodiments of the present application;

[0026] Figure 3 Flowchart 2 of a method for generating a large model of outbound voice speech provided in some embodiments of the present application;

[0027] Figure 4 A block diagram of the device composition for generating a large model of outbound voice speech provided in some embodiments of the present application;

[0028] Figure 5 A schematic diagram of an electronic device is provided for some embodiments of the present application. DETAILED DESCRIPTION

[0029] The technical solutions in some embodiments of the present application will be described below in conjunction with the drawings in some embodiments of the present application.

[0030] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0031] In the related art, in order to improve the interaction efficiency between users and intelligent robots, in outbound telemarketing scenarios, the existing technology usually limits the intelligent robot's reply scripts to a limited script set. The script set is sorted according to common call scenarios. However, the biggest disadvantage of using a limited script set is that the script set used for reply cannot cover all user questions in real scenarios, resulting in the phenomenon that the intelligent robot answers irrelevant questions.

[0032] It can be seen from the above-mentioned related technologies that the scenarios of intelligent robot reply dialogues in the existing technology are relatively limited and the accuracy is low.

[0033] In view of this, some embodiments of the present application provide a method for generating outbound voice large model speech, which can obtain the interactive content set generated by the interaction with the user through the large language model; then the interactive content set and the historical conversation record are combined, and the streaming speech synthesis technology is used to generate the voice reply speech to realize the voice broadcast. Some embodiments of the present application can solve the limitations of the reply speech, improve the accuracy of the reply, and solve the delay problem of the speech reply.

[0034] The following is combined with Figure 1 The overall composition structure of the system for generating large-scale model speech of outbound voice provided by some embodiments of the present application is exemplified.

[0035] like Figure 1 As shown, some embodiments of the present application provide a system for generating a large model of outbound voice scripts, and the system for generating a large model of outbound voice scripts includes: a mobile terminal 100 and an intelligent robot 200. When the intelligent robot 200 performs an outbound call task in a telemarketing scenario, it interacts with the user of the mobile terminal 100. During the interaction process, the intelligent robot 200 can obtain an interactive content set for interacting with the user, and the interactive content set contains text data related to the user's questions and the business process nodes in which it is located. The intelligent robot 200 can generate a voice reply script that can reply to the user's current question based on the interactive content set and historical conversation records; finally, it is played to the user through the mobile terminal 100 by broadcasting the voice reply script.

[0036] In some embodiments of the present application, the mobile terminal 100 may be various types of communication devices, such as a smart phone, a smart watch, or a tablet computer, etc. The embodiments of the present application are not specifically limited here.

[0037] The following is combined with Figure 2 The implementation process of generating outbound voice scripts performed by the intelligent robot 200 provided in some embodiments of the present application is exemplified.

[0038] Please see attached Figure 2 , Figure 2 A flow chart of a method for generating a large model of outbound voice speech provided in some embodiments of the present application, the method for generating a large model of outbound voice speech may include:

[0039] S210, obtaining an interactive content set generated by interaction with the user, wherein the interactive content set includes: a question-associated text related to a current question of the user and a node-associated text of a next business node of a business process node where the user is located.

[0040] For example, in some embodiments of the present application, during the process of interacting with the user through the mobile terminal 100, the intelligent robot 200 obtains a set of interactive content related to the dialogue record generated by the user interaction, so that accurate reply words can be generated later.

[0041] In some embodiments of the present application, S210 may include:

[0042] S211, obtaining the question-related text and obtaining the node-related text.

[0043] For example, in some embodiments of the present application, the intelligent robot 200 can analyze the user's current problem and determine the problem-related text; it can also analyze the business process node where the user is currently located, determine the next business node, and then obtain the node-related text.

[0044] In some embodiments of the present application, S211 may include: inputting the dialogue content and integration requirements generated by the interaction with the user into a large language model, identifying and outputting the current question; matching the current question with a preset answer in a speech set to obtain the question-associated text.

[0045] For example, in some embodiments of the present application, through the powerful reasoning ability of the large language model, the user's current problem can be summarized based on the conversation content generated between the user and the intelligent robot 200. For example, write a prompt input to the large language model, such as the prompt example:

[0046] The following '===' is the conversation record, which summarizes the user's questions based on the overall conversation.

[0047] ===

[0048] {conversation_history}

[0049] ===

[0050] Combine the conversation record and attention requirements to succinctly summarize the user's current problem in one phrase. The user's problem is:

[0051] The prompt is input into the large language model, and the large language model can output the summarized user's current question. The attention requirement is the integration requirement mentioned above, which is "summarize the user's question based on the overall conversation." It can be understood that the prompt can be written according to the actual situation, and the embodiment of the present application is not specifically limited here.

[0052] After obtaining the user's current question, a text similarity algorithm is used to match similar future questions from the speech set; then the preset answer corresponding to the user's question is used as the user's question background knowledge (as a specific example of question-related text). For example, if the user's current question identified by the large language model is "not needed", and "not needed temporarily" is matched in the speech set, then the answer corresponding to "not needed temporarily" is used as the user's question background knowledge for the user's current question.

[0053] Among them, the large language model can be chatgpt, llama, qwen, etc. The speech set is a pre-sorted speech that may be used for reply in real scenarios. The speech set includes: multiple user questions and answers corresponding to each user question (as a specific example of a preset answer), as well as multiple process nodes and the speech under each process node. The speech set can be flexibly expanded, and the embodiments of the present application are not specifically limited here.

[0054] In some embodiments of the present application, S211 may include: inputting the conversation content and process node content into the large language model, identifying and outputting the next business node; and searching the node-associated text corresponding to the next business node from the speech set.

[0055] For example, in some embodiments of the present application, through the powerful reasoning ability of the large language model, the next business node of the user can be inferred based on the conversation content generated between the user and the intelligent robot 200 and the process node in the current telemarketing scenario. For example, a prompt is written to be input into the large language model to determine the next process node (that is, the next business node) of the current conversation, such as the prompt example:

[0056] The operation process is as follows:

[0057] Node|Node Target

[0058] ---|---

[0059] A|Introduce yourself

[0060] B|Ask the user whether they have received service notifications

[0061] C|Guide users to scan the QR code

[0062] D|Confirm that the user has completed the process

[0063] E|Polite hang up;

[0064] The conversation was recorded as follows:

[0065] {conversation_history}

[0066] The operation flow node for the next sentence (select only one node in the operation flow table) is:

[0067] The prompt is input into the large language model, and the large language model can output the next process node that the user is in after judgment. The operation process is the process node content. From the above prompt, it can be seen that there are 5 process nodes, ABCDE, in this telemarketing scenario.

[0068] After obtaining the next process node that the user is next at, the speech corresponding to the next process node is found from the speech combination, and the speech is used as the process node background knowledge (as a specific example of node-related text). For example, if the next process node identified by the large language model is D, then the speech corresponding to process node D in the speech set is used as the process node background knowledge of the next process node.

[0069] S212: construct the interactive content set based on the question-related text and the node-related text.

[0070] For example, in some embodiments of the present application, a set consisting of user problem background knowledge and process node background knowledge is taken as an interactive content set, or referred to as a background knowledge set.

[0071] S220, generating a voice response script corresponding to the current question based on the interaction content set and historical conversation records.

[0072] For example, in some embodiments of the present application, the reasoning and generation capabilities of a large language model can be used to combine a collection of interactive content and historical conversation records with the user to generate voice response scripts.

[0073] In some embodiments of the present application, S220 may include: using streaming speech synthesis technology to synchronously generate the voice reply script while obtaining the reply text corresponding to the question-associated text, the node-associated text and the historical conversation record, wherein the reply text is the original text in the question-associated text, the node-associated text or a combination of multiple text fragments.

[0074] For example, in some embodiments of the present application, the interactive content collection is combined with the historical conversation record as input, and the large language model generates the reply text in a streaming manner (as a specific example of the reply text). The reply text can be the original text in the user's problem background knowledge and the process node background knowledge, or it can be a combination of multiple text fragments in the background knowledge, thereby avoiding the problem that the results generated by the large language model do not conform to the actual business scenario. In addition, in order to improve the user experience, there is no need to wait for the large language model to generate the results before converting the reply text into speech. Instead, the voice reply text is generated synchronously during the streaming generation of the reply text for broadcast. The following example of writing a prompt is written by combining the interactive content collection with the historical conversation record as input:

[0075] You are communicating with the user over the phone. You need to select or combine the most appropriate words to reply to the user based on the conversation record and background knowledge.

[0076] {base_info}

[0077] {conversation_history}{role_name}:

[0078] Among them, {base_info} is a set of background knowledge, {conversation_history} is a historical conversation record, and {role_name} is a robot role (for example, gpt or salesperson).

[0079] By inputting the prompt into the large language model, the reply text can be output, and the voice reply text is generated synchronously during the reply text output process.

[0080] In some embodiments of the present application, S220 may include: obtaining the most similar text corresponding to the reply text from the interactive content set; calculating the most similar text and the reply text to obtain at least one editing interval; and simultaneously editing the at least one editing interval to obtain the voice reply script.

[0081] For example, in some embodiments of the present application, in the outbound call system, the response time of the intelligent robot is required to be at the ms level. The traditional speech synthesis model cannot achieve effectiveness, and this part adopts voice editing technology. Since most of the reply text generated by the large language model comes from the background knowledge set, the background knowledge set has recorded the voice in advance. Therefore, it is necessary to find the text C that is most similar to the reply text from the background knowledge set (as a specific example of the most similar text). Voice editing of text C can obtain the voice reply text.

[0082] In some embodiments of the present application, S220 may include: creating a masking matrix equal to the length of the reply text, wherein the elements of each editing interval in the masking matrix have different values ​​from the elements of the non-editing interval; after masking the most similar text using the masking matrix, performing speech synthesis processing on each editing interval to obtain a speech editing result; integrating the speech editing result with the audio of the non-editing interval to obtain the voice reply words.

[0083] For example, in some embodiments of the present application, a vector matrix (as a specific example of a masking matrix, referred to as a matrix) equal to the length of the reply speech text is established, and this matrix is ​​used to mask a specific interval in the speech corresponding to the reply speech text. In this matrix, the element corresponding to the editing interval is set to 0, and the rest (that is, the non-editing interval) is set to 1. After the reply speech text is masked using the masking matrix, all the editing intervals marked as 0 are subjected to speech synthesis processing, while the non-editing intervals marked as 1 remain as they are. The speech of multiple processed intervals (as a specific example of the result of speech editing) is spliced ​​and integrated with the audio of the unprocessed interval (that is, the non-editing interval) to obtain the final edited voice reply speech. In other words, in the process of speech editing, only the part of the interval that needs to be edited needs to be synthesized, and the other parts that are the same as text C can directly use their corresponding audio. This method can effectively improve the speed of speech synthesis while retaining the timbre and rhythm of the original recording.

[0084] Specifically, the editing interval is obtained by the following method: the editing interval from the reply speech text to the text C is calculated through the reply speech text and the text C. For example, the reply speech text is "Please reply to agree to handle it", and the text C to be edited is "Please reply to agree to cancel it". There is an editing interval from the reply speech text to the text C, and one is that "handle" changes to "cancel". For the identified editing interval, the time points to the start and end in the reply speech text recording are clearly defined. These time points are used to define the boundaries of the editing interval. Afterwards, through the masking operation of the above-mentioned embodiment of the present application, the corresponding part of the editing interval is set to 0, and the other parts are set to 1. This technology can support the calculation of the edits of multiple intervals at the same time, so that the masked positions can be dispersed, which improves the efficiency of voice editing. After that, the part of the editing interval is voice synthesized and spliced ​​with the audio of "Please reply to agree" to obtain the voice reply speech.

[0085] The following is combined with Figure 3 The specific process of generating outbound voice scripts provided by some embodiments of the present application is exemplified.

[0086] Please see attached Figure 3 , Figure 3A flow chart of a method for generating a large model of outbound voice speech provided for some embodiments of the present application.

[0087] The above process is explained below as an example.

[0088] S310, input the dialogue content and integration requirements generated by the interaction with the user into the large language model, identify and output the current question.

[0089] S320, matching the current question with the preset answers in the speech set to obtain the background knowledge of the user's question.

[0090] S330: Input the conversation content and process node content into the large language model, identify and output the next business node of the user.

[0091] S340, searching the process node background knowledge corresponding to the next business node from the speech set.

[0092] S350, constructing a background knowledge set based on the user problem background knowledge and the process node background knowledge.

[0093] S360 uses streaming speech synthesis technology to obtain the reply text corresponding to the background knowledge set and historical conversation records, and simultaneously generates voice reply words.

[0094] It can be understood that the specific implementation process of S310 to S360 can refer to the method embodiment provided above, and in order to avoid repetition, the detailed description is appropriately omitted here.

[0095] Through some of the above embodiments of the present application, it can be seen that the embodiments of the present application are based on the powerful generation capability of the large language model, can be combined with historical dialogue records, and interact with users more naturally. Compared with the existing dialogue management technology, it avoids answering irrelevant questions, and the intelligent robot replies more accurately. Through voice editing technology, recordings can be streamed, which effectively improves the speed of intelligent robot broadcasting.

[0096] Please refer to Figure 4 , Figure 4 The block diagram of the device for generating a large model of outbound voice speech provided by some embodiments of the present application is shown. It should be understood that the device for generating a large model of outbound voice speech corresponds to the above method embodiment and can perform each step involved in the above method embodiment. The specific functions of the device for generating a large model of outbound voice speech can be found in the description above. To avoid repetition, the detailed description is appropriately omitted here.

[0097] Figure 4The device for generating a large model of outbound voice speech includes at least one software function module that can be stored in a memory in the form of software or firmware or solidified in the device for generating a large model of outbound voice speech, and the device for generating a large model of outbound voice speech includes: a content acquisition module 410, used to acquire a set of interactive content generated by interaction with a user, wherein the interactive content set includes: question-associated text related to the user's current question and node-associated text of the next business node of the business process node where the user is located; a speech generation module 420, used to generate a voice reply speech corresponding to the current question based on the interactive content set and historical conversation records.

[0098] In some embodiments of the present application, the content acquisition module 410 is used to acquire the question-related text and the node-related text; and construct the interactive content set based on the question-related text and the node-related text.

[0099] In some embodiments of the present application, the content acquisition module 410 is used to input the dialogue content and integration requirements generated by the interaction with the user into the large language model, identify and output the current question; match the current question with the preset answer in the speech set to obtain the question-associated text.

[0100] In some embodiments of the present application, the content acquisition module 410 is used to input the conversation content and process node content into the large language model, identify and output the next business node; and search the node-associated text corresponding to the next business node from the speech set.

[0101] In some embodiments of the present application, the speech generation module 420 is used to use streaming speech synthesis technology to synchronously generate the voice response speech while acquiring the reply text corresponding to the question-associated text, the node-associated text and the historical conversation record, wherein the reply text is the original text in the question-associated text, the node-associated text or a combination of multiple text fragments.

[0102] In some embodiments of the present application, the speech generation module 420 is used to obtain the most similar text corresponding to the reply text from the interactive content set; calculate the most similar text and the reply text to obtain at least one editing interval; and simultaneously edit the at least one editing interval to obtain the voice reply speech.

[0103] In some embodiments of the present application, the speech generation module 420 is used to create a masking matrix with the same length as the reply text, wherein the elements of each editing interval in the masking matrix have different values ​​from the elements of the non-editing interval; after masking the most similar text with the masking matrix, each editing interval is subjected to speech synthesis processing to obtain a speech editing result; the speech editing result is integrated with the audio of the non-editing interval to obtain the voice reply speech.

[0104] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method, and will not be described in detail here.

[0105] Some embodiments of the present application further provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the operations of the method corresponding to any of the above methods provided in the above embodiments.

[0106] Some embodiments of the present application further provide a computer program product, which includes a computer program, wherein when the computer program is executed by a processor, it can implement the operations corresponding to any of the above methods provided in the above embodiments.

[0107] like Figure 5 As shown, some embodiments of the present application provide an electronic device 500, which includes: a memory 510, a processor 520, and a computer program stored in the memory 510 and executable on the processor 520, wherein the processor 520 can implement a method as described in any of the above embodiments when reading the program from the memory 510 through a bus 530 and executing the program.

[0108] Processor 520 can process digital signals and can include various computing structures, such as complex instruction set computer structure, reduced instruction set computer structure, or a structure that implements a combination of multiple instruction sets. In some examples, processor 520 can be a microprocessor.

[0109] The memory 510 may be used to store instructions executed by the processor 520 or data related to the execution of instructions. These instructions and / or data may include codes for implementing some or all functions of one or more modules described in the embodiments of the present application. The processor 520 of the disclosed embodiment may be used to execute instructions in the memory 510 to implement the method shown above. The memory 510 includes a dynamic random access memory, a static random access memory, a flash memory, an optical memory, or other memory known to those skilled in the art.

[0110] The above description is only an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.

[0111] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0112] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

Claims

1. A method for generating a large model of outbound voice, characterized in that: include: Acquire an interactive content set generated by the interaction with the user, wherein the interactive content set includes: a question-related text related to the user's current question and a node-related text of a next business node of the business process node where the user is located; Based on the interactive content set and historical conversation records, a voice reply script corresponding to the current question is generated.

2. The method according to claim 1, characterized in that The acquiring of the interactive content set generated by the interaction with the user includes: Obtain the question-related text and the node-related text; The interactive content set is constructed based on the question-related text and the node-related text.

3. The method according to claim 2, characterized in that The step of obtaining the question-related text includes: Inputting the dialogue content and integration requirements generated by the interaction with the user into the large language model, identifying and outputting the current question; The current question is matched with the preset answers in the speech set to obtain the question-related text.

4. The method according to claim 3, characterized in that The obtaining of the node associated text includes: Inputting the conversation content and process node content into the large language model, identifying and outputting the next business node; The node-associated text corresponding to the next business node is searched from the speech set.

5. The method according to any one of claims 1 to 4, characterized in that The generating, based on the interactive content set and the historical conversation record, a voice reply script corresponding to the current question includes: Using streaming speech synthesis technology, while obtaining the reply text corresponding to the question-associated text, the node-associated text and the historical conversation record, the voice reply script is synchronously generated, wherein the reply text is the original text in the question-associated text and the node-associated text or a combination of multiple text fragments.

6. The method according to claim 5, characterized in that The synchronously generating the voice reply speech includes: Acquire the most similar text corresponding to the reply text from the interactive content set; Calculating the most similar text and the reply text to obtain at least one editing interval; The at least one editing interval is edited simultaneously to obtain the voice reply script.

7. The method according to claim 6, characterized in that The simultaneously editing the at least one editing interval to obtain the voice reply speech includes: Creating a mask matrix of the same length as the reply text, wherein the value of each element of the edit interval in the mask matrix is ​​different from the value of the element of the non-edit interval; After masking the most similar text with the masking matrix, performing speech synthesis processing on each editing interval to obtain a speech editing result; The voice editing result is integrated with the audio of the non-editing interval to obtain the voice reply script.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program executes the method according to any one of claims 1 to 7 when executed by a processor.

9. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the computer program executes the method according to any one of claims 1 to 7 when being run by the processor.

10. A computer program product, characterized in that The computer program product comprises a computer program, wherein the computer program executes the method according to any one of claims 1 to 7 when executed by a processor.