Telephone sales verbal skill generation method and device, equipment and medium

By decoding, recognizing and generating strategies to process input voice and product descriptions and generate target sales pitches, it solves the problem of low efficiency in telephone sales processing, achieves precise control of intent and emotions, and improves transaction efficiency and user experience.

CN120708604APending Publication Date: 2025-09-26PING AN TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511047890.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing methods for generating sales scripts for telephone sales are inefficient and unable to cope with customers’ different intentions, tones, and emotional changes, resulting in reduced willingness to communicate and lower closing rates.

Method used

By obtaining input speech and product description, the decoding strategy is used to identify the intent type, the recognition strategy is used to identify the emotional state, the contextual speech strategy is used to analyze and process the intermediate speech, and the target speech is generated based on the generation strategy.

Benefits of technology

It improves the closing efficiency and user experience of telephone sales, achieves precise control of intentions and emotions, and improves processing efficiency and compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708604A_ABST
    Figure CN120708604A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence, and provides a telemarketing verbal skill generation method, device and equipment and a medium, which are applied to the fields of finance and medical treatment, and the method comprises the following steps: obtaining an input voice and a product description; decoding the input voice based on the decoding strategy to obtain an intention type; performing recognition processing on the input voice based on a recognition strategy to obtain an emotional state; analyzing and processing the input voice, the intention type and the emotional state by using a contextual verbal skill strategy to obtain an intermediate verbal skill; and performing generation processing on the intermediate verbal skill and the product description based on the generation strategy to obtain a target verbal skill. By implementing the embodiment of the invention, the input voice and the product description are respectively decoded, identified, analyzed and generated to obtain the target verbal skill, the transaction efficiency and the user experience of telemarketing are improved, and the method has strong landing practicability and sustainable optimization capability, so that the processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a method, device, equipment and medium for generating telephone sales speech. Background Art

[0002] Currently, traditional telephone sales techniques in the fields of finance, healthcare, etc. often proceed according to preset processes. When faced with different intentions, tones, and emotional changes of customers, it is impossible to make natural responses. Customers can easily identify them as "robots", reducing communication willingness and transaction rates. In addition, it is difficult to deal with ambiguity, implicit rejection or emotional clues in customer expressions, which may cause misjudgment or interruption of dialogue, resulting in low processing efficiency.

[0003] Therefore, the existing telephone sales speech generation method has the problem of low processing efficiency. Summary of the Invention

[0004] The embodiments of the present invention provide a method, device, equipment and medium for generating telephone sales speech, aiming to solve the problem of low processing efficiency in the telephone sales speech generation method in the prior art.

[0005] In order to solve the above problems, in a first aspect, an embodiment of the present invention provides a method for generating telephone sales speech, which includes:

[0006] Get input voice and product description;

[0007] Decoding the input speech based on a decoding strategy to obtain an intent type;

[0008] Recognize the input speech based on the recognition strategy to obtain the emotional state;

[0009] Analyzing and processing the input speech, the intention type, and the emotional state using a contextual speech strategy to obtain an intermediate speech;

[0010] The intermediate speech and the product description are generated based on a generation strategy to obtain a target speech.

[0011] In a second aspect, an embodiment of the present application provides a device for generating telephone sales speech, comprising:

[0012] An acquisition unit, used to acquire input speech and product description;

[0013] A decoding unit, configured to decode the input speech based on a decoding strategy to obtain an intent type;

[0014] A recognition unit, configured to perform recognition processing on the input speech based on a recognition strategy to obtain an emotional state;

[0015] An analysis unit, configured to analyze and process the input speech, the intention type, and the emotional state using a contextual speech strategy to obtain an intermediate speech;

[0016] A generating unit is used to generate the intermediate speech and the product description based on a generating strategy to obtain a target speech.

[0017] In a third aspect, an embodiment of the present application provides a computer device, comprising a memory and a processor connected to the memory; the memory is used to store a computer program, and the processor is used to run the computer program stored in the memory to execute the method described in the first aspect above.

[0018] In a fourth aspect, an embodiment of the present application provides a storage medium, wherein the storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the method described in the first aspect is implemented.

[0019] The embodiments of the present invention provide a method, apparatus, device, and medium for generating telephone sales speech. The method comprises: obtaining input speech and product description; decoding the input speech based on a decoding strategy to obtain an intent type; recognizing the input speech based on a recognition strategy to obtain an emotional state; analyzing the input speech, the intent type, and the emotional state using a contextual speech strategy to obtain an intermediate speech; and generating the intermediate speech and the product description based on a generation strategy to obtain a target speech. Therefore, the embodiments of the present invention improve the closing efficiency and user experience of telephone sales by respectively decoding, recognizing, analyzing, and generating the input speech and product description to obtain the target speech. The method has strong practicality and sustainable optimization capabilities, thereby improving processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 A flowchart of a method for generating telephone sales speech provided by an embodiment of the present invention;

[0022] Figure 2 A schematic block diagram of a device for generating telephone sales speech provided by an embodiment of the present invention;

[0023] Figure 3 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0025] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0026] It should also be understood that the terms used in the present specification are only for the purpose of describing particular embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0027] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0028] See also Figure 1 , Figure 1 The flowchart of the method for generating telephone sales speech provided by the embodiment of the present invention is as follows. Figure 1 As shown, an embodiment of the present invention provides a method for generating telephone sales scripts, which includes the following steps S110-S150.

[0029] S110: Acquire input voice and product description.

[0030] In this embodiment, the input voice may be an audio file of a telephone sale; and the product description may be a text describing the product of the telephone sale.

[0031] The application scenarios of the embodiments of the present invention can be in the financial field and the medical field; for example, in the financial field, it can be used for business promotion, etc., and in the medical field, it can be used for customer service, etc.

[0032] Through the above embodiments, it can be seen that acquiring the input voice and product description and performing subsequent targeted processing according to the input voice and product description improves subsequent data accuracy and processing efficiency.

[0033] S120: Decode the input speech based on a decoding strategy to obtain an intent type.

[0034] In this embodiment, after the input speech is acquired, the input speech is decoded based on a decoding strategy to obtain the intent type.

[0035] In this embodiment, the decoding strategy-based decoding process of the input speech to obtain the intent type specifically includes: using a nested template to template the input speech to obtain a discourse template; inputting the discourse template into a preset model, and using the preset model to perform contextual semantic decoding on the discourse template to obtain the intent type. For example, the nested template may be "Please identify the true intent of the following customer speech: "{input speech}", and possible intentions include: expressing interest, raising objections, rejection, vague rejection, question, no clear intention, etc. The intent type may include explicit intent and implicit intent, for example, the explicit intent may be inquiry, rejection, hesitation, etc., and the implicit intent may be precaution, interest, price sensitivity, etc.; the intent type may also include reasons; and the preset model may be a large language model.

[0036] Through the above embodiment, it can be seen that the input speech is templated using nested templates to obtain a discourse template; the discourse template is input into a preset model, and the preset model is used to perform contextual semantic decoding on the discourse template to obtain the intent type. Therefore, by decoding the input speech based on the decoding strategy to obtain the intent type, the large language model's deep understanding of the semantics of the input speech is introduced, making the speech more tailored to the needs, and the communication more humane and persuasive, achieving precise control of the intent type, thereby improving processing efficiency.

[0037] S130: Perform recognition processing on the input speech based on a recognition strategy to obtain an emotional state.

[0038] In this embodiment, after the input voice is acquired, the input voice is recognized based on a recognition strategy to obtain the emotional state.

[0039] In one embodiment, the performing recognition processing on the input speech based on the recognition strategy to obtain the emotional state includes:

[0040] Performing emotion recognition on the input speech to obtain an emotion curve graph;

[0041] Performing emotion analysis on the emotion curve graph to obtain an emotional state.

[0042] In this embodiment, performing emotion recognition on the input speech to obtain the emotion graph specifically includes: using a large model and a sentiment analysis algorithm to recognize the input speech to construct the emotion graph; the large model and sentiment analysis algorithm may be BERT combined with a BiGRU structure or ChatGLM instruction fine-tuning; and the emotion graph is used to dynamically adjust the tone of the speech. Emotion analysis processing is performed on the emotion graph to obtain the emotional state. For example, the emotional state can be determined from the timeline of the emotion graph as apathetic, neutral, anxious, calm, etc., that is, the emotion graph may include time and emotion, and the emotional state may include multiple emotional states.

[0043] Through the above embodiment, it can be seen that emotion recognition is performed on the input speech to obtain an emotion curve graph; emotion analysis is performed on the emotion curve graph to obtain the emotional state. Therefore, the input speech is recognized and processed based on the recognition strategy to obtain the emotional state, thereby achieving a deep understanding of the emotion of the input speech, making the speech more tailored to the needs, and the communication more humane and persuasive. This achieves precise control of the emotional state, thereby improving processing efficiency.

[0044] S140. Analyze and process the input speech, the intention type, and the emotional state using a contextual speech strategy to obtain an intermediate speech.

[0045] In this embodiment, after determining the intention type and the emotional state, the input speech, the intention type and the emotional state are analyzed and processed using a contextual speech strategy to obtain an intermediate speech.

[0046] In one embodiment, the analyzing and processing the input speech, the intention type, and the emotional state using the contextual speech strategy to obtain the intermediate speech includes:

[0047] Performing dialogue analysis on the input speech to obtain a current dialogue stage;

[0048] Performing path analysis using the intention type, the emotional state, and the current dialogue stage to obtain a dialogue path;

[0049] Performing type analysis on the speech path according to a preset type library to obtain the speech type;

[0050] The current dialogue stage, the intention type, the emotional state and the speech type are fused to obtain the intermediate speech.

[0051] In one embodiment, performing conversation analysis on the input speech to obtain the current conversation stage includes:

[0052] Decompose the sales dialogue according to the decomposition strategy to obtain the dialogue stages;

[0053] The input speech and the dialogue stage are matched to obtain the current dialogue stage.

[0054] In this embodiment, the decomposition of the sales conversation according to the decomposition strategy to obtain the conversation stages specifically includes: decomposition of the sales conversation according to the timeline or conversation nodes to obtain a plurality of conversation stages; wherein the conversation stages may include opening, interest stimulation, objection handling, conversion promotion, and deal closure, etc. The matching of the input speech with the conversation stages to obtain the current conversation stage specifically includes searching and matching the input speech with the conversation stages to obtain the current conversation stage; wherein the current conversation stage may include opening, interest stimulation, objection handling, conversion promotion, or deal closure, etc.

[0055] The method of performing path analysis based on the intent type, the emotional state, and the current conversation stage to obtain the speech path specifically includes: performing path analysis based on the intent type, the emotional state, and the current conversation stage in a preset path library to obtain the speech path. For example, if the current conversation stage is the product introduction stage (opening), the intent type is hesitation or slight rejection, and the emotional state is slightly negative, then matching and analyzing the preset path library may yield a speech path that transitions to preferential explanation combined with emotional comfort.

[0056] The method of performing type analysis on the speech path according to a preset type library to obtain the speech type specifically includes: performing type analysis on the speech path from a preset type library to obtain the speech type; wherein the speech type may include respect type, enthusiasm type, professional type, etc.

[0057] The intermediate speech is obtained by fusing the current dialogue stage, the intention type, the emotional state and the speech type, that is, the current dialogue stage, the intention type, the emotional state and the speech type are adjusted and fused to obtain the intermediate speech; for example, the intermediate speech can be a speech generated for the customer who is currently in the "introduction product" stage and has a hesitant tone or slight rejection, with professional language.

[0058] Through the above embodiments, it can be seen that the sales dialogue is decomposed according to the decomposition strategy to obtain the dialogue stage; the input voice and the dialogue stage are matched to obtain the current dialogue stage; the intention type, the emotional state and the current dialogue stage are used to perform path analysis to obtain the speech path; the speech path is type analyzed according to the preset type library to obtain the speech type; the current dialogue stage, the intention type, the emotional state and the speech type are fused to obtain the intermediate speech. Therefore, by using the contextual speech strategy to analyze and process the input voice, the intention type and the emotional state to obtain the intermediate speech, the process control based on the dialogue stage and the speech path selection driven by intention and emotion are realized, breaking through the problem of traditional "rigid scripts", realizing dynamic response and flexible transformation, thereby improving processing efficiency.

[0059] S150: Generate the intermediate speech and the product description based on the generation strategy to obtain the target speech.

[0060] In this embodiment, after the intermediate speech is obtained, the intermediate speech and the product description may be generated based on a generation strategy to obtain the target speech.

[0061] In one embodiment, the generating process of the intermediate speech and the product description based on the generation strategy to obtain the target speech includes:

[0062] Performing characteristic analysis on the product description to obtain product characteristics;

[0063] Performing style analysis using the intermediate speech to obtain a speech style;

[0064] Generate a first speech based on the product characteristics and the speech style;

[0065] The first speech is reviewed and processed based on the review strategy to obtain the target speech.

[0066] In this embodiment, the product description is subjected to characteristic analysis to obtain the product characteristics, such as high cost-effectiveness; the intermediate speech is subjected to style analysis to obtain the speech style, such as "friendly and trustworthy" and "high-value communication"; a first speech is generated based on the product characteristics and the speech style; and the first speech is reviewed based on the review strategy to obtain the target speech. Therefore, through the analysis and review processes of the above embodiments, professionalism and compliance are maintained in natural conversations, "nonsense" is avoided, customer trust is enhanced, and processing efficiency is improved.

[0067] In one embodiment, the step of reviewing the first speech based on the review strategy to obtain the target speech includes:

[0068] Using a term alignment strategy to align the first speech to obtain a second speech;

[0069] Filtering the second speech according to the speech evaluation strategy to obtain a third speech;

[0070] Reviewing the third speech technique based on a preset static rule base and a preset review method to obtain a risk level;

[0071] When the risk level is a preset risk level, the third speech is rewritten based on the compliance generator to obtain the target speech.

[0072] In this embodiment, the use of the term alignment strategy to align the first speech to obtain the second speech specifically includes: normalizing the first speech according to the FinTerms embedding table and the Logit control mechanism to obtain the second speech to ensure the generation of terminology standards.

[0073] The filtering of the second speech according to the speech evaluation strategy to obtain the third speech specifically includes: inputting the second speech into a language style consistency scoring model, and using the language style consistency scoring model to perform tone or grammar evaluation on the second speech to filter out style-mismatched content to obtain the third speech.

[0074] The review and processing of the third speech based on the preset static rule base and the preset review method to obtain the risk level specifically includes: performing rule filtering processing on the third speech according to the preset static rule base to obtain the processed third speech; reviewing and processing the processed third speech using the preset review method to obtain the risk level; wherein, the preset static rule base includes banned terms (such as "100% money back guarantee"), over-commitment sentences (such as "absolute value preservation"), industry restricted words (such as "risk-free"), etc.; the preset review method includes real-time regular review combined with keyword matching and semantic-level compliance model to determine the risk level; the risk levels include no risk, warning risk and interruption risk.

[0075] When the risk level is the preset risk level, the compliance generator rewrites the third script to obtain the target script. Specifically, when the risk level is the preset risk level, i.e., when a violation is detected, the compliance version generator (using RLHF training or instruction fine-tuning) can be invoked to automatically rewrite the third script to obtain the target script. For example, the target script could be: You are a telesales expert currently in the "objection handling" stage. A customer expresses hesitation, "I'm not sure whether to buy." Considering the product's "high cost-performance," generate a script using professional language, a friendly and credible style, avoiding exaggeration, and meeting financial regulatory standards. The preset risk level can be either a warning risk or an interruption risk. Therefore, combining the richness of model generation with the compliance requirements of legal review, a script security mechanism consisting of "pre-generation restrictions + post-generation review + reversible rewriting" has been established. This full-process compliance review mechanism ensures that the output target script has legal risk prevention and control capabilities, helping to meet regulatory requirements in industries such as finance and insurance, and improving processing efficiency.

[0076] In one embodiment, after generating the intermediate speech and the product description based on the generation strategy to obtain the target speech, the method further includes:

[0077] Use the target words to conduct conversations with customers.

[0078] In this embodiment, after obtaining the target sales pitch, the target sales pitch is used to conduct a conversation with the customer, thereby significantly improving the closing efficiency of telephone sales and user experience.

[0079] In the process of generating the target sales pitch, this solution employs sales pitch chain tracking: recording the entire process from customer intent to sales pitch generation to conversion results, thereby building a "high-conversion sales pitch chain." Clustering and attribution analysis can also be employed: using sentence vector clustering and keyword attribution to identify common features, such as the high-frequency trigger word "installment payment possible" and key tone changes such as indifference, hesitation, and laughter / acceptance. Feedback model optimization can also be employed: using high-conversion sales pitches as training samples or expanded prompts to optimize the model's response strategy. The system's innovative significance lies in its "data self-growth capability," continuously enhancing the model's sales capabilities in actual business operations, achieving "closed-loop learning" for intelligent sales and improving processing efficiency.

[0080] In summary, the embodiment of the present invention obtains input speech and product description; decodes the input speech based on a decoding strategy to obtain the intent type; recognizes the input speech based on a recognition strategy to obtain the emotional state; analyzes the input speech, the intent type, and the emotional state using a contextual speech strategy to obtain intermediate speech; and generates the intermediate speech and the product description based on a generation strategy to obtain the target speech. Therefore, the embodiment of the present invention improves the closing efficiency and user experience of telephone sales by decoding, recognizing, analyzing, and generating the input speech and product description to obtain the target speech, and has strong practicality and sustainable optimization capabilities, thereby improving processing efficiency.

[0081] Figure 2 This is a schematic block diagram of a device for generating telephone sales speech provided by an embodiment of the present invention. Figure 2 As shown, the embodiment of the present invention provides a telephone sales speech generation device 700 for implementing the above method. Figure 2 The telephone sales speech generation device 700 includes:

[0082] An acquisition unit 701 is used to acquire input speech and product description;

[0083] A decoding unit 702 is configured to decode the input speech based on a decoding strategy to obtain an intent type;

[0084] The recognition unit 703 is configured to perform recognition processing on the input speech based on a recognition strategy to obtain an emotional state;

[0085] An analyzing unit 704 is configured to analyze and process the input speech, the intention type, and the emotional state using a contextual speech strategy to obtain an intermediate speech;

[0086] The generating unit 705 is configured to generate the intermediate speech and the product description based on a generating strategy to obtain a target speech.

[0087] In some embodiments, when executing the step of analyzing and processing the intention type and the emotion curve graph using the contextual speech strategy to obtain the intermediate speech, the analyzing unit 704 is specifically configured to:

[0088] Performing dialogue analysis on the input speech to obtain a current dialogue stage;

[0089] Performing path analysis using the intention type, the emotional state, and the current dialogue stage to obtain a dialogue path;

[0090] Performing type analysis on the speech path according to a preset type library to obtain the speech type;

[0091] The current dialogue stage, the intention type, the emotional state and the speech type are fused to obtain the intermediate speech.

[0092] In some embodiments, when executing the step of generating the intermediate speech and the product description based on the generation strategy to obtain the target speech, the generation unit 705 is specifically configured to:

[0093] Performing characteristic analysis on the product description to obtain product characteristics;

[0094] Performing style analysis using the intermediate speech to obtain a speech style;

[0095] Generate a first speech based on the product characteristics and the speech style;

[0096] The first speech is reviewed and processed based on the review strategy to obtain the target speech.

[0097] In some embodiments, when executing the step of reviewing the first speech based on the review strategy to obtain the target speech, the generating unit 705 is specifically configured to:

[0098] Using a term alignment strategy to align the first speech to obtain a second speech;

[0099] Filtering the second speech according to the speech evaluation strategy to obtain a third speech;

[0100] Reviewing the third speech technique based on a preset static rule base and a preset review method to obtain a risk level;

[0101] When the risk level is a preset risk level, the third speech is rewritten based on the compliance generator to obtain the target speech.

[0102] In some embodiments, when executing the step of performing recognition processing on the input speech based on the recognition strategy to obtain the emotional state, the recognition unit 703 is specifically configured to:

[0103] Performing emotion recognition on the input speech to obtain an emotion curve graph;

[0104] Performing emotion analysis on the emotion curve graph to obtain an emotional state.

[0105] In some embodiments, when executing the step of performing dialogue analysis on the input speech to obtain the current dialogue stage, the analyzing unit 704 is specifically configured to:

[0106] Decompose the sales dialogue according to the decomposition strategy to obtain the dialogue stages;

[0107] The input speech and the dialogue stage are matched to obtain the current dialogue stage.

[0108] In some embodiments, after executing the step of generating the intermediate speech and the product description based on the generation strategy to obtain the target speech, the generation unit 705 is further configured to:

[0109] Use the target words to conduct conversations with customers.

[0110] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned device can refer to the corresponding description in the aforementioned method embodiment. For the convenience and brevity of description, it will not be repeated here.

[0111] The above device can be implemented in the form of a computer program, which can be used in Figure 3 Runs on the computer equipment shown.

[0112] See also Figure 3 , Figure 3 8 is a schematic block diagram of an electronic device provided by an embodiment of the present invention. The electronic device 800 can be a terminal or a server, wherein the terminal can be an electronic device with communication functions. The server can be a standalone server or a server cluster consisting of multiple servers.

[0113] See Figure 3 The electronic device 800 includes a processor 802 , a memory, and a network interface 805 connected via a system bus 801 , wherein the memory may include a non-volatile storage medium 803 and an internal memory 804 .

[0114] The non-volatile storage medium 803 can store an operating system 8031 ​​and a computer program 8032. The computer program 8032 includes program instructions, which, when executed, can cause the processor 802 to execute a method for generating telemarketing scripts.

[0115] The processor 802 is used to provide computing and control capabilities to support the operation of the entire electronic device 800.

[0116] The internal memory 804 provides an environment for the operation of the computer program 8032 in the non-volatile storage medium 803. When the computer program 8032 is executed by the processor 802, the processor 802 can execute a method for generating telephone sales speech.

[0117] The network interface 805 is used to communicate with other devices over the network. Figure 3The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention, and does not constitute a limitation on the electronic device 800 to which the solution of the present invention is applied. The specific electronic device 800 may include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0118] The processor 802 is configured to execute a computer program 8032 stored in the memory to implement the following steps:

[0119] Get input voice and product description;

[0120] Decoding the input speech based on a decoding strategy to obtain an intent type;

[0121] Recognize the input speech based on the recognition strategy to obtain the emotional state;

[0122] Analyzing and processing the input speech, the intention type, and the emotional state using a contextual speech strategy to obtain an intermediate speech;

[0123] The intermediate speech and the product description are generated based on a generation strategy to obtain a target speech.

[0124] In some embodiments, when the processor 802 implements the step of analyzing and processing the intent type and the emotion curve graph using the contextual speech strategy to obtain the intermediate speech, it is specifically configured to:

[0125] Performing dialogue analysis on the input speech to obtain a current dialogue stage;

[0126] Performing path analysis using the intention type, the emotional state, and the current dialogue stage to obtain a dialogue path;

[0127] Performing type analysis on the speech path according to a preset type library to obtain the speech type;

[0128] The current dialogue stage, the intention type, the emotional state and the speech type are fused to obtain the intermediate speech.

[0129] In some embodiments, when implementing the step of generating the intermediate speech and the product description based on the generation strategy to obtain the target speech, the processor 802 is specifically configured to:

[0130] Performing characteristic analysis on the product description to obtain product characteristics;

[0131] Performing style analysis using the intermediate speech to obtain a speech style;

[0132] Generate a first speech based on the product characteristics and the speech style;

[0133] The first speech is reviewed and processed based on the review strategy to obtain the target speech.

[0134] In some embodiments, when the processor 802 implements the step of reviewing the first speech based on the review strategy to obtain the target speech, it is specifically configured to:

[0135] Using a term alignment strategy to align the first speech to obtain a second speech;

[0136] Filtering the second speech according to the speech evaluation strategy to obtain a third speech;

[0137] Reviewing the third speech technique based on a preset static rule base and a preset review method to obtain a risk level;

[0138] When the risk level is a preset risk level, the third speech is rewritten based on the compliance generator to obtain the target speech.

[0139] In some embodiments, when implementing the step of performing recognition processing on the input speech based on the recognition strategy to obtain the emotional state, the processor 802 is specifically configured to:

[0140] Performing emotion recognition on the input speech to obtain an emotion curve graph;

[0141] Performing emotion analysis on the emotion curve graph to obtain an emotional state.

[0142] In some embodiments, when implementing the step of performing conversation analysis on the input speech to obtain the current conversation stage, the processor 802 is specifically configured to:

[0143] Decompose the sales dialogue according to the decomposition strategy to obtain the dialogue stages;

[0144] The input speech and the dialogue stage are matched to obtain the current dialogue stage.

[0145] In some embodiments, after implementing the step of generating the target speech from the intermediate speech and the product description based on the generation strategy, the processor 802 is further configured to:

[0146] Use the target words to conduct conversations with customers.

[0147] It should be understood that in the embodiment of the present invention, the processor 802 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0148] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.

[0149] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor performs the following steps:

[0150] Get input voice and product description;

[0151] Decoding the input speech based on a decoding strategy to obtain an intent type;

[0152] Recognize the input speech based on the recognition strategy to obtain the emotional state;

[0153] Analyzing and processing the input speech, the intention type, and the emotional state using a contextual speech strategy to obtain an intermediate speech;

[0154] The intermediate speech and the product description are generated based on a generation strategy to obtain a target speech.

[0155] In one embodiment, when the processor executes the program instructions to implement the step of analyzing and processing the intention type and the emotion curve graph using the contextual speech strategy to obtain the intermediate speech, it is specifically configured to:

[0156] Performing dialogue analysis on the input speech to obtain a current dialogue stage;

[0157] Performing path analysis using the intention type, the emotional state, and the current dialogue stage to obtain a dialogue path;

[0158] Performing type analysis on the speech path according to a preset type library to obtain the speech type;

[0159] The current dialogue stage, the intention type, the emotional state and the speech type are fused to obtain the intermediate speech.

[0160] In one embodiment, when the processor executes the program instructions to implement the step of generating the intermediate speech and the product description based on the generation strategy to obtain the target speech, the processor is specifically configured to:

[0161] Performing characteristic analysis on the product description to obtain product characteristics;

[0162] Performing style analysis using the intermediate speech to obtain a speech style;

[0163] Generate a first speech based on the product characteristics and the speech style;

[0164] The first speech is reviewed and processed based on the review strategy to obtain the target speech.

[0165] In one embodiment, when the processor executes the program instructions to implement the step of reviewing the first speech based on the review strategy to obtain the target speech, it is specifically configured to:

[0166] Using a term alignment strategy to align the first speech to obtain a second speech;

[0167] Filtering the second speech according to the speech evaluation strategy to obtain a third speech;

[0168] Reviewing the third speech technique based on a preset static rule base and a preset review method to obtain a risk level;

[0169] When the risk level is a preset risk level, the third speech is rewritten based on the compliance generator to obtain the target speech.

[0170] In one embodiment, when the processor executes the program instructions to implement the step of performing recognition processing on the input speech based on the recognition strategy to obtain the emotional state, the processor is specifically configured to:

[0171] Performing emotion recognition on the input speech to obtain an emotion curve graph;

[0172] Performing emotion analysis on the emotion curve graph to obtain an emotional state.

[0173] In one embodiment, when the processor executes the program instructions to implement the step of performing conversation analysis on the input speech to obtain the current conversation stage, the processor is specifically configured to:

[0174] Decompose the sales dialogue according to the decomposition strategy to obtain the dialogue stages;

[0175] The input speech and the dialogue stage are matched to obtain the current dialogue stage.

[0176] In one embodiment, after executing the program instructions to implement the step of generating the target speech from the intermediate speech and the product description based on the generation strategy, the processor is further configured to:

[0177] Use the target words to conduct conversations with customers.

[0178] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.

[0179] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0180] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0181] The steps in the methods of the embodiments of the present invention may be adjusted in order, combined, or deleted as needed. The units in the devices of the embodiments of the present invention may be combined, divided, or deleted as needed. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.

[0182] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing an electronic device (such as a personal computer, terminal, or network device) to execute all or part of the steps of the method described in various embodiments of the present invention.

[0183] The above is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims. The software tools, models, or components that appear in the embodiments of the present invention are only examples and do not represent actual use.

Claims

1. A method for generating telephone sales speech, characterized in that: include: Get input voice and product description; Decoding the input speech based on a decoding strategy to obtain an intent type; Recognize the input speech based on the recognition strategy to obtain the emotional state; Analyzing and processing the input speech, the intention type, and the emotional state using a contextual speech strategy to obtain an intermediate speech; The intermediate speech and the product description are generated based on a generation strategy to obtain a target speech.

2. The method according to claim 1, characterized in that The process of analyzing and processing the input speech, the intention type, and the emotional state using the contextual speech strategy to obtain the intermediate speech includes: Performing dialogue analysis on the input speech to obtain a current dialogue stage; Performing path analysis using the intention type, the emotional state, and the current dialogue stage to obtain a dialogue path; Performing type analysis on the speech path according to a preset type library to obtain the speech type; The current dialogue stage, the intention type, the emotional state and the speech type are fused to obtain the intermediate speech.

3. The method according to claim 1, characterized in that The generating process of the intermediate speech and the product description based on the generation strategy to obtain the target speech includes: Performing characteristic analysis on the product description to obtain product characteristics; Performing style analysis using the intermediate speech to obtain a speech style; Generate a first speech based on the product characteristics and the speech style; The first speech is reviewed and processed based on the review strategy to obtain the target speech.

4. The method according to claim 3, characterized in that The step of reviewing the first speech based on the review strategy to obtain the target speech includes: Using a term alignment strategy to align the first speech to obtain a second speech; Filtering the second speech according to the speech evaluation strategy to obtain a third speech; Reviewing the third speech technique based on a preset static rule base and a preset review method to obtain a risk level; When the risk level is a preset risk level, the third speech is rewritten based on the compliance generator to obtain the target speech.

5. The method according to claim 1, wherein The step of performing recognition processing on the input speech based on the recognition strategy to obtain the emotional state includes: Performing emotion recognition on the input speech to obtain an emotion curve graph; Performing emotion analysis on the emotion curve graph to obtain an emotional state.

6. The method according to claim 2, characterized in that The performing dialogue analysis on the input speech to obtain the current dialogue stage includes: Decompose the sales dialogue according to the decomposition strategy to obtain the dialogue stages; The input speech and the dialogue stage are matched to obtain the current dialogue stage.

7. The method according to claim 1, characterized in that After the intermediate speech and the product description are generated based on the generation strategy to obtain the target speech, the method further includes: Use the target words to conduct conversations with customers.

8. A device for generating telephone sales speech, characterized in that: include: An acquisition unit, used to acquire input speech and product description; A decoding unit, configured to decode the input speech based on a decoding strategy to obtain an intent type; A recognition unit, configured to perform recognition processing on the input speech based on a recognition strategy to obtain an emotional state; An analysis unit, configured to analyze and process the input speech, the intention type, and the emotional state using a contextual speech strategy to obtain an intermediate speech; A generating unit is used to generate the intermediate speech and the product description based on a generating strategy to obtain a target speech.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the method according to any one of claims 1 to 7 can be implemented.

Citation Information

Patent Citations

  • Product recommendation method and device based on voice emotion analysis, equipment and medium

    CN109949071A

  • Multi-style polishing intelligent customer service reply method and system and storage medium thereof

    CN119357337A

  • Intelligent customer service system based on large language model

    CN119938850A