Outbound call script generation method and device, computer device and readable storage medium
By collecting interactive data in real time and using semantic recognition and sentiment analysis models, outbound call scripts are dynamically generated, solving the problems of insufficient understanding of user intent and untimely emotional response in existing technologies, and achieving efficient, flexible and stable communication in outbound call scripts.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN YISHIHUOLALA TECH CO LTD
- Filing Date
- 2026-03-30
- Publication Date
- 2026-07-10
AI Technical Summary
Existing outbound call script generation technologies struggle to understand user intent in real time, fail to respond adequately to changes in emotions, and lack dynamic adaptability. This results in mismatched script strategies, rigid communication, unstable conversion rates, limited cross-scenario generalization capabilities, and high maintenance costs.
By collecting interactive data in real time, identifying user intent using a preset semantic recognition model, judging emotional state by combining a preset scenario analysis model, dynamically generating script strategies, and generating target outbound call script data.
It improves the accuracy and flexibility of outbound call script generation, reduces the risk of misjudgment and communication interruption, and enhances the stability of outbound call interactions and business conversion efficiency.
Smart Images

Figure CN122372676A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent voice interaction, and in particular to a method, apparatus and computer equipment for generating outbound call scripts. Background Technology
[0002] In outbound calling applications such as intelligent customer service, marketing outreach, business follow-up, and risk alerts, outbound calling systems need to complete information transmission, demand confirmation, or task advancement with target users within a limited call time. During outbound call interactions, user expression is characterized by being conversational, fragmented, and uncertain, and the user's emotional state may fluctuate as the conversation progresses. Understanding the user's current expression during outbound calls and generating scripted content that aligns with the current stage of the conversation and communication goals is a crucial factor affecting outbound calling efficiency, service experience, and business conversion rates, and is also a key research direction in related technical fields.
[0003] Existing outbound call script generation and arrangement technologies largely rely on fixed scripts, template splicing, or rule / keyword-based classification strategies to select scripts. On the one hand, user needs are often not expressed as explicit instructions, easily leading to implicit requests, subtle refusals, and missing information. This makes semantic processing methods based solely on keywords or static rules unable to accurately extract user intent, prone to misjudgment or omissions. On the other hand, existing technologies lack the ability to perceive and utilize changes in user emotions, lacking a mechanism for co-modeling emotional states and semantic intent. This makes it difficult to adjust script strategies in a timely manner when user emotions fluctuate, the pace of the conversation changes, or intent shifts, resulting in script mismatch, communication stagnation, decreased user experience, and unstable conversion rates. Furthermore, the requirements for script style and implementation methods vary significantly across different business types, user profiles, and different outbound call stages. Existing solutions based primarily on static scripts have limited generalization capabilities, high maintenance costs, and difficulty in adapting to complex scenarios.
[0004] Therefore, there is an urgent need for an outbound call script generation method that can collect user interaction data in real time during outbound calls, obtain intent data through semantic recognition, obtain emotion data through scenario analysis, and then determine the script generation strategy and generate target outbound call script data based on the collaboration of intent and emotion. This method aims to overcome the problems of insufficient understanding of user intent, untimely response to emotion changes, lack of dynamic adaptation capability of script strategies, and insufficient cross-scenario generalization capability in existing technologies. This would improve the accuracy, flexibility, and consistency of outbound call script generation, reduce script maintenance costs, and improve the stability and business effectiveness of outbound call interactions. Summary of the Invention
[0005] Therefore, it is necessary to provide a method, device, computer equipment, and readable storage medium for generating outbound call scripts to address the aforementioned technical problems, in order to solve the problems of insufficient understanding of user intent, untimely response to emotional changes, lack of dynamic adaptation capability of script strategies, and insufficient cross-scenario generalization capability of existing voice interaction technologies.
[0006] A method for generating outbound call scripts, the method comprising: During outbound calls, real-time collection of interaction data from target users is conducted. The interaction data is semantically processed by a preset semantic recognition model to obtain the target user's intent data during the outbound call process; The interaction data is processed by a preset scenario analysis model to determine the emotion of the target user during the outbound call process. Based on the intent data and the emotion data, a script generation strategy is determined; Based on the aforementioned script generation strategy, target outbound call script data is generated.
[0007] Optionally, the interaction data includes voice information, text information, interaction context information, and historical interaction information. The real-time collection of the target user's interaction data during the outbound call process includes: During the outbound call process, the voice information of the target user is acquired, and the voice information is subjected to speech recognition to obtain text information; Obtain the interaction context information corresponding to the current outbound call round, wherein the interaction context information includes at least the outbound call text information and the corresponding outbound call script text information input by the target user in the historical rounds; Obtain historical interaction information associated with the target user, including historical login request information and historical outbound call record information; The voice information, text information, interaction context information, and historical interaction information are integrated to obtain the target user's interaction data.
[0008] Optionally, the step of semantically processing the interaction data using a preset semantic recognition model to obtain the target user's intent data during the outbound call process includes: Based on the interaction data, semantic prompt information is constructed; Based on the semantic prompt information, a request is sent to the preset semantic recognition model to obtain the intent data of the target user during the outbound call process returned by the preset semantic recognition model based on the semantic prompt information. The intent data includes the target user's explicit demand information and potential demand information.
[0009] Optionally, the step of performing emotion judgment processing on the interaction data through a preset scenario analysis model to obtain the target user's emotion data during the outbound call process includes: Determine the interaction context information and historical interaction information of the target user; Based on the interaction context information and historical interaction information, the voice feature data and historical text content data of the target user are determined respectively. Based on the target user's voice feature data and historical text content data, construct emotion feature prompt information; Based on the emotional feature prompt information, a request is sent to the preset scenario analysis model to obtain the emotional data of the target user during the outbound call process returned by the preset scenario analysis model based on the emotional feature prompt information. The emotional data includes current emotional state data and emotional change data.
[0010] Optionally, determining the script generation strategy based on the intent data and the emotion data includes: Based on the current emotional state data and emotional change data, the current dialogue emotional pattern is determined; Based on the current dialogue emotion pattern, determine the first text adjustment parameter of the target outbound call script data; Based on the explicit and potential demand information, the intent processing mode is determined. Based on the intent processing mode, the second text adjustment parameters of the target outbound call script data are determined; Based on the first text adjustment parameters and the second text adjustment parameters, the corresponding speech generation strategy is determined.
[0011] Optionally, generating target outbound call script data based on the script generation strategy includes: Based on the script generation strategy, generate the outbound call script text for the current round; During the multi-round outbound call interaction, the target user's emotional data is acquired in real time, and emotional change data is determined based on the emotional data of adjacent rounds; When the emotion change data meets the preset change conditions, the script generation strategy is adjusted based on the updated emotion data to obtain an updated script generation strategy. Based on the updated script generation strategy, the outbound script text for at least one subsequent round is updated and generated to obtain the target outbound script data and output it.
[0012] Optionally, after generating the target outbound call script data based on the script generation strategy, the method further includes: Obtain customer feedback data from the target user regarding the target outbound call script data; Based on the customer feedback data, the feedback evaluation results are determined; Based on the feedback evaluation results, the target user profile database, the preset scenario analysis model, and the preset semantic recognition model are updated.
[0013] An outbound call script generation device, the device comprising: The first data collection module is used to collect the target user's interaction data in real time during the outbound call process; The first processing module is used to perform semantic processing on the interaction data through a preset semantic recognition model to obtain the intent data of the target user during the outbound call process; The second processing module is used to perform emotion judgment processing on the interaction data through a preset scenario analysis model to obtain the emotional data of the target user during the outbound call process; The first determining module is used to determine the speech generation strategy based on the intent data and the emotion data; The first generation module is used to generate target outbound call script data based on the script generation strategy.
[0014] A computer device includes a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the processor implements the outbound call script generation method described above when executing the computer-readable instructions.
[0015] A readable storage medium storing computer-readable instructions, which, when executed by a processor, implement the above-described outbound call script generation method.
[0016] The aforementioned outbound call script generation method collects target user interaction data in real time during the outbound call process; performs semantic processing on the interaction data using a preset semantic recognition model to obtain the target user's intent data during the outbound call process; performs emotion judgment processing on the interaction data using a preset scenario analysis model to obtain the target user's emotion data during the outbound call process; determines a script generation strategy based on the intent data and the emotion data; and generates target outbound call script data based on the script generation strategy. Through these steps, the targeting and adaptability of outbound call script generation can be improved, reducing the risk of communication interruptions and conversion failures due to misjudgment of user needs or neglect of emotional changes, thereby improving the stability of outbound call interactions and business conversion efficiency. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating an outbound call script generation method in one embodiment of the present invention; Figure 2 This is a schematic diagram of an outbound call script generation process according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the outbound call script generation device in one embodiment of the present invention; Figure 4 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] In one embodiment, such as Figure 1 As shown, a method for generating outbound call scripts is provided, including the following steps: 101. During outbound calls, collect the target user's interaction data in real time.
[0021] In this embodiment of the invention, the outbound call script generation method can be applied to an outbound call script generation platform. The outbound call script generation platform has functions such as outbound call script generation data processing, outbound call script generation data sending and receiving, and outbound call script generation data memory storage. It can be built based on a server or server cluster. The server or server cluster can be an electronic device with outbound call script generation data processing capabilities.
[0022] The aforementioned target user refers to an individual user or group of users being called in an outbound call task. They are the direct communication and business recipients in the outbound interaction. They are typically identified by at least one of the following: phone number, account identifier, customer number, or order number. Alternatively, they can be identified by a combination of enterprise customer organization identifier and contact person identifier, used to establish a stable conversational relationship during the call. For example, in a follow-up scenario, the target user can correspond to the person who handled a completed transaction, and relevant historical records can be continuously linked during the outbound call to support more targeted communication.
[0023] The aforementioned interactive data refers to the set of information generated around the target user during outbound call interactions and collected and maintained in real time. It is used to reflect the input, output, and contextual relationships during the call, including but not limited to voice information, text information, interactive context information, and historical interactive information. Voice information is reflected in the call audio stream and its effective segments, text information is reflected in the text content obtained from speech recognition, interactive context information is reflected in the correlation between the user's input content in different rounds within the same call and the speech content output by the outbound call platform, and historical interactive information is reflected in the target user's past outbound call records, login records, or related business records. For example, interactive data is merged at the session dimension and continuously updated as the call progresses, thereby providing a complete and continuous data foundation for subsequent semantic understanding, emotion judgment, and speech strategy generation.
[0024] 102. By using a preset semantic recognition model to perform semantic processing on the interaction data, the intent data of the target user during the outbound call process can be obtained.
[0025] In this embodiment of the invention, the aforementioned preset semantic recognition model can refer to a semantic understanding model used in outbound calling scenarios for deep understanding and comprehensive inference of natural language expressions. Its core capability is reflected in the overall modeling of contextual semantics, implicit intentions, and complex language phenomena.
[0026] Generally, the aforementioned pre-defined semantic recognition model possesses cross-round contextual understanding capabilities. It can combine the current round's expression with historical interaction content to make an overall judgment on the user's true intention, rather than relying solely on single-sentence keyword matching. For example, when a user expresses a relatively vague statement such as "It's not appropriate now" or "Let's talk about it next time," the aforementioned pre-defined semantic recognition model can also combine the preceding and following context with historical outbound call records to determine whether it is closer to a postponement of communication or a clear refusal.
[0027] In one possible embodiment, the aforementioned preset semantic recognition model can be based on a large-scale pre-trained language model and optimized in a targeted manner using business data to improve the accuracy of understanding in outbound call contexts.
[0028] In another possible embodiment, the aforementioned outbound call script generation platform can perform understanding and mapping of the language content in the interaction data according to the aforementioned preset semantic recognition model, thereby realizing the transformation of natural language expressions into structured semantic results that can participate in strategy decision-making. Specifically, the aforementioned outbound call script generation platform standardizes and organizes the interaction data, and combines the interaction context and historical interaction information during the model invocation process to comprehensively judge semantic relationships such as omission, transition, and negation in the language.
[0029] Generally speaking, semantic processing not only focuses on the surface meaning of a single sentence, but also on the continuous changes in semantics in multi-turn dialogues. For example, when a user repeatedly expresses "I've been quite busy lately" or "I really can't spare the time right now," the aforementioned outbound call script generation platform uses cross-turn input to make the semantic processing result point to a temporary suspension of communication rather than a simple state description.
[0030] The aforementioned intent data may refer to the data set obtained by the aforementioned outbound call script generation platform after completing semantic processing, which describes the target user's communication goals, attitudes, and behavioral tendencies during the outbound call process, and is used to reflect the communication results that the target user hopes to achieve in the current outbound call stage.
[0031] Specifically, the aforementioned intent data can include, but is not limited to, two levels: explicit needs information and potential needs information. Explicit needs information corresponds to the user's directly expressed requests, while potential needs information corresponds to the implicit intentions inferred by the outbound call script generation platform based on semantic processing results, interaction context, and historical interaction information. For example, if a user expresses "I'll think about it," the outbound call script generation platform records the explicit need as "no decision for now" in the intent data, and simultaneously records the potential need as "needing additional information" or "reducing communication pressure."
[0032] 103. By using a preset scenario analysis model to process the interaction data for emotion judgment, the emotional data of the target user during the outbound call process can be obtained.
[0033] In this embodiment of the invention, the aforementioned preset scenario analysis model can refer to an analysis model used to comprehensively understand outbound call interaction scenarios and output emotion-related judgment results. Its focus is on identifying the user's emotional state, changing trends, and interactive context of the conversation during the call.
[0034] In this embodiment, after acquiring the interaction data, the aforementioned outbound call script generation platform is responsible for organizing the input content for scenario analysis and initiating an analysis request to a preset scenario analysis model, ensuring that the emotion judgment process maintains consistent execution logic across different outbound call tasks. Specifically, the preset scenario analysis model can combine multi-round information to judge the current interaction scenario, such as when expressions of refusal or emotional intensification occur consecutively in the same call, and can identify when the dialogue has entered a sensitive or confrontational stage.
[0035] In one possible embodiment, the aforementioned preset scenario analysis model can be built based on a large-scale pre-trained model and adapted using outbound call business data to make it more consistent with the emotional and scenario characteristics of the outbound call context.
[0036] In another possible embodiment, the aforementioned outbound call script generation platform can analyze and infer emotion-related clues in the interaction data through the preset scenario analysis model, thereby identifying the target user's emotional state and its changes during the outbound call process. Specifically, the aforementioned emotion judgment processing does not rely solely on single-dimensional information, but rather combines voice features with text content for comprehensive analysis. For example, when performing emotion judgment processing, the aforementioned outbound call script generation platform extracts voice features such as speech rate, tone, volume, pause duration, and interruption frequency from the interaction data, while simultaneously extracting negative expressions, complaining language, strong emotional words, and politeness cues from the text. This information is then provided to the preset scenario analysis model for comprehensive judgment. Specifically, when the speech rate significantly increases, the tone rises, and continuous negative expressions appear in the text, the emotion judgment processing tends to determine that the emotion is fluctuating or resistant.
[0037] It should be noted that the above emotion judgment processing can also incorporate interactive context information, so that the model can refer to the changes in previous rounds when judging the current emotion, and avoid misjudgment caused by a single abnormal expression.
[0038] The aforementioned emotional data can refer to the data set obtained by the aforementioned outbound call script generation platform after completing the emotional judgment and processing, which is used to characterize the emotional state and changing trend of the target user during the outbound call process, including but not limited to current emotional state data and emotional change data.
[0039] Generally, the aforementioned outbound call script generation platform updates the emotion data after each round of interaction and compares the current result with the previous round's record to continuously maintain information on emotion changes. If the emotion state is stable in the previous round but shows obvious resistance in the next round, the emotion change data will be recorded as an increase in resistance from stable, which is used to trigger adjustments to subsequent script strategies.
[0040] 104. Determine the script generation strategy based on intent data and sentiment data.
[0041] In this embodiment of the invention, the above-mentioned script generation strategy may refer to the decision result formed by the above-mentioned outbound call script generation platform after integrating intent data and emotion data, which is used to guide the generation and adjustment of outbound call scripts in terms of content organization, expression mode and pace.
[0042] It should be noted that the above-mentioned script generation strategy is not a fixed script text, but rather a set of constraints and control rules applied to the script generation process to ensure that the generated script content not only meets the current needs and orientations of the target user, but also matches their emotional state and the stage of the conversation.
[0043] Specifically, after acquiring intent data, the aforementioned outbound call script generation platform identifies the explicit and potential needs currently expressed by the target user. After acquiring emotion data, it determines whether the target user is in a stable, fluctuating, or confrontational emotional state. Based on the combination of these two factors, it forms a corresponding set of strategy parameters to control the tone intensity, information disclosure level, pace of progress, proportion of reassuring expressions, and priority of responding to objections in the script.
[0044] For example, when intent data points to a clear need and the emotional state remains stable, the script generation strategy tends towards a standard approach; when intent data shows vague rejection or uncertainty and emotional data shows a fluctuating trend, the script generation strategy tends towards a reassurance and confirmation approach to reduce communication confrontation and maintain dialogue continuity.
[0045] More specifically, when intent data indicates that the target user is not making a decision yet, while emotion data shows an increased speaking speed and tense tone, the aforementioned outbound call script generation platform reduces the intensity of the push through script generation strategies, increases empathy and explanatory expressions, and postpones the confirmation of key information; while when intent data indicates that the target user has a clear need and emotion data remains stable, the aforementioned outbound call script generation platform improves information focus through script generation strategies, prioritizing the output of core content that meets the need.
[0046] In another possible embodiment, the above-mentioned script generation strategy can also be represented in a parameterized form, with different parameters corresponding to the tone, expression, response order and termination condition, and allowing dynamic adjustment based on changes in intent data or emotion data during multi-round outbound calls, thereby achieving adaptive control of the script generation process.
[0047] 105. Generate target outbound call script data based on the script generation strategy.
[0048] In this embodiment of the invention, the aforementioned target outbound call script data may refer to the set of result data output by the outbound call script generation platform after the script generation strategy is determined, which can be directly used for outbound call interaction. This includes, but is not limited to, outbound call script text content, and may be extended to include structured script fragment information and controllable variable fields to support script reuse and flexible adjustment in different business scenarios.
[0049] Understandably, when generating target outbound call scripts, the aforementioned platform will prioritize meeting the core communication goals indicated in the intent data, while also adhering to the constraints of emotion data to avoid pushing too fast or expressing oneself too aggressively.
[0050] For example, when the script generation strategy indicates that the target user has a clear need and is in a stable emotional state, the target outbound script data generated by the aforementioned outbound script generation platform focuses on directly responding to the need and concentrating on outputting key information; when the script generation strategy indicates that the target user's intention is unclear and their emotions are fluctuating, the target outbound script data reflects the approach of first reassuring and confirming, and then gradually guiding the focus of communication, thereby reducing the pressure of the conversation.
[0051] During the multi-round interactive adjustment process of outbound calls, the aforementioned outbound call script generation platform continuously maintains emotion change data and updates the script generation strategy when it detects that the emotion change data meets preset change conditions. Subsequently, based on the updated script generation strategy, the platform regenerates or adjusts the target outbound call script data for at least one subsequent round, enabling the script content within the same call to dynamically adapt to changes in the user's emotions.
[0052] By following the above steps, the rigidity of static scripts in complex interactive scenarios is avoided, enabling the wording to maintain business consistency while possessing flexible adaptability.
[0053] In one possible embodiment, it can be based on, as follows: Figure 2 The diagram illustrates an outbound call script generation process. The outbound call script generation platform continuously collects interaction data from the target user during the outbound call process. The interaction data includes voice information, text information, interaction context information, and historical interaction information. This data serves as the input basis for overall analysis, and semantic processing and emotion judgment processing are performed respectively. On the one hand, it identifies the target user's needs and intentions during the current outbound call process. On the other hand, it combines voice features and historical expressions to determine the target user's emotional state and changing trends.
[0054] Under the combined constraints of semantic and sentiment analysis results, the aforementioned outbound call script generation platform determines a script generation strategy adapted to the current interaction scenario and generates the outbound call script text accordingly. During multiple rounds of outbound call interactions, the platform continuously monitors emotional changes. When emotional fluctuations or changes in the dialogue scenario are detected, the platform dynamically adjusts the script generation strategy, updating the script content for subsequent rounds to obtain and output the target outbound call script data. After the outbound call ends, the platform updates the user profile and analysis model based on user feedback, enabling continuous optimization of intent understanding and sentiment adaptation in subsequent outbound calls.
[0055] In this embodiment of the invention, during outbound calling, interaction data of the target user is collected in real time; semantic processing of the interaction data is performed using a preset semantic recognition model to obtain the target user's intent data during the outbound calling process; emotional judgment processing of the interaction data is performed using a preset scenario analysis model to obtain the target user's emotional data during the outbound calling process; based on the intent data and emotional data, a script generation strategy is determined; and based on the script generation strategy, target outbound calling script data is generated. Through these methods, the targeting and adaptability of outbound calling script generation can be improved, reducing the risk of communication interruptions and conversion failures due to misjudgment of user needs or neglect of emotional changes, thereby improving the stability of outbound calling interactions and business conversion efficiency.
[0056] It is understood that in the specific implementation of this application, data related to intent data, emotion data, target outbound call script data, and interaction data are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0057] Optionally, in the step of collecting the target user's interaction data in real time during the outbound call process, the following steps can also be taken: acquiring the target user's voice information and performing speech recognition on the voice information to obtain text information; acquiring the interaction context information corresponding to the current outbound call round, which includes at least the outbound call text information and corresponding outbound call script text information input by the target user in previous rounds; acquiring historical interaction information associated with the target user, which includes historical login request information and historical outbound call record information; and integrating the voice information, text information, interaction context information, and historical interaction information to obtain the target user's interaction data.
[0058] In this embodiment of the invention, the aforementioned interactive data may include, but is not limited to, voice information, text information, interactive context information, and historical interactive information. The interactive context information reflects the connection between different rounds within the same call and generally includes the outbound call text information input by the target user in historical rounds and the corresponding outbound call script text information. By using interactive context information, judgments can be avoided based solely on single-sentence expressions, thereby improving the ability to understand the continuity of the dialogue.
[0059] The aforementioned historical interaction information includes historical login request information and historical outbound call records, and can be expanded to include past communication results, rejection records, or business processing status, etc., to supplement long-term behavioral characteristics that cannot be directly reflected in the current call. Generally speaking, it can be used to help determine the user's acceptance or sensitivity to outbound calls.
[0060] After collecting voice information, text information, interaction context information, and historical interaction information, the aforementioned outbound call script generation platform integrates and processes information from different sources to form a unified set of interaction data. During the integration process, various types of information are aligned through conversation identifiers and time sequence, enabling voice content, recognized text, contextual relationships, and historical summaries within the same round to be referenced together, thereby providing a consistent data foundation for subsequent semantic processing and sentiment assessment.
[0061] The above integration methods can reduce misunderstandings caused by fragmented information and improve the overall accuracy of analysis.
[0062] Optionally, in the step of semantically processing the interaction data through a preset semantic recognition model to obtain the target user's intent data during the outbound call process, semantic prompt information can also be constructed based on the interaction data; a request can be sent to the preset semantic recognition model based on the semantic prompt information to obtain the target user's intent data during the outbound call process returned by the preset semantic recognition model according to the semantic prompt information, the intent data including the target user's explicit demand information and potential demand information.
[0063] In this embodiment of the invention, the aforementioned semantic prompt information may refer to the data content generated by the outbound call script generation platform based on interactive data, which is used to drive the semantic recognition model to understand and infer. It includes the target user's current expression, key dialogue information from previous rounds, and historical interaction points related to the current outbound call task. For example, when the target user only expresses "it is inconvenient now", the semantic prompt information may also include the topic of the previous round of communication, thereby helping the model to determine whether the expression is closer to postponing communication or explicitly refusing.
[0064] In this embodiment, after constructing the semantic prompt information, the outbound call script generation platform transmits the semantic prompt information to a preset semantic recognition model via an interface to trigger semantic understanding, thereby completing the request process. Specifically, this process can be uniformly managed by the outbound call script generation platform to ensure consistency between the model call sequence and the session. Generally, each request corresponds to the current call round, enabling the semantic recognition result to accurately map to the current interaction state.
[0065] In one possible embodiment, the aforementioned intent data can be output by the preset semantic recognition model after completing semantic understanding, and received and parsed by the outbound call script generation platform to describe the target user's communication goals and behavioral tendencies in the current outbound call context, serving as input to the script generation strategy. It is understood that the aforementioned intent data can be recorded in the session state on the platform side and can be updated as the call progresses.
[0066] The aforementioned explicit demand information can refer to the portion of the intent data that directly corresponds to the target user's clearly expressed needs, such as the user's explicit acceptance, rejection, need for explanation, or request for extension. Generally, this can be derived from the target user's direct verbal expression, has a high degree of certainty, and is used to guide the main direction of script generation.
[0067] The aforementioned potential demand information can refer to the implicit requests or attitudes obtained through semantic inference from the aforementioned intent data. It is usually obtained by combining the expression, context, and historical interaction information. It is used to supplement the parts not fully covered by the explicit demand information. For example, if the target user expresses "I'll think about it first", the explicit demand information is reflected as not making a decision for the time being, while the potential demand information may be reflected as needing to supplement information or reduce the intensity of communication.
[0068] In another possible embodiment, the aforementioned outbound call script generation platform can construct semantic prompt information by organizing, filtering, and structuring interactive data. It should be noted that the platform retains the core expressions of the target user in the current round during the construction process and combines necessary contextual summaries and historical summaries to ensure that the semantic information has a complete context when input.
[0069] Optionally, in the step of processing the interaction data for emotion judgment through a preset scenario analysis model to obtain the target user's emotion data during the outbound call process, the interaction context information and historical interaction information of the target user can also be determined; based on the interaction context information and historical interaction information, the voice feature data and historical text content data of the target user can be determined respectively; based on the voice feature data and historical text content data of the target user, emotion feature prompt information can be constructed; and based on the emotion feature prompt information, a request can be sent to the preset scenario analysis model to obtain the target user's emotion data during the outbound call process returned by the preset scenario analysis model based on the emotion feature prompt information.
[0070] In this embodiment of the invention, the outbound call script generation platform first determines the interaction context information corresponding to the current outbound call round and retrieves historical interaction information associated with the target user to provide a continuous contextual basis for emotion judgment. Based on this, the outbound call script generation platform extracts voice feature data from the interaction data and combines it with historical text content data to construct input content for emotion analysis.
[0071] The aforementioned voice feature data can refer to the feature information extracted from the target user's voice information to reflect the speaking state and emotional tendency, including but not limited to changes in speech rate, pitch, volume, pause duration, speech continuity, and interruption frequency, which are used to characterize the target user's external emotional expression during the call.
[0072] Generally, the aforementioned outbound call script generation platform extracts voice features simultaneously with voice acquisition and recognition, ensuring that the voice feature data is time-aligned with the corresponding text content. For example, when the speaking speed increases significantly and the tone rises, the voice feature data can reflect tension, impatience, or resistance.
[0073] The aforementioned historical text content data refers to text summaries compiled from historical interaction information to aid in emotion judgment. These summaries reflect the target user's language expression habits, attitudes, and common response patterns in past outbound calls or related interactions, serving as supplementary emotional background information that is difficult to directly reflect in a single call. For example, the aforementioned outbound call script generation platform will filter and compress historical text content, retaining only fragments relevant to the current outbound call topic or emotion judgment. For instance, if expressions of refusal or complaint appear repeatedly in the historical text content, it can serve as an important reference for the current emotion judgment.
[0074] The aforementioned emotional feature prompts can refer to the input content constructed by the aforementioned outbound call script generation platform based on voice feature data and historical text content data, used to drive the preset scenario analysis model for emotion analysis. It can generally be organized in a structured or semi-structured form, combining the voice feature summary of the current round with the historical text content summary.
[0075] The aforementioned emotional state data can refer to the result data returned by the preset scenario analysis model, which describes the target user's emotional performance in the current outbound call phase, such as stable, positive, negative, or resistant states, and may include an intensity level, to provide real-time emotional reference for the script generation strategy.
[0076] The aforementioned emotional change data can refer to the change information obtained by comparing the current emotional state data with the emotional state of historical rounds. It is used to reflect the evolution trend of the target user's emotions as the outbound call process progresses, such as from stable to resistant, or from resistant to mild, and is used to determine whether the script strategy needs to be adjusted.
[0077] Generally, the aforementioned outbound call script generation platform updates the emotion change data after each round of interaction and triggers strategy adjustments when the changes reach preset conditions. For example, when the emotional state shows a negative increasing trend for several consecutive rounds, the emotion change data can be used to prompt an adjustment of the outbound call script from a pushing type to a reassuring type.
[0078] By following the above steps, interactive data can be continuously collected and integrated during outbound calls, simultaneously obtaining the target user's intent and emotion data. Based on this, a script generation strategy can be dynamically determined to generate target outbound call script data that matches the user's needs and emotional state. This reduces the risk of misjudgment and communication interruption, improves outbound call interaction efficiency and conversion stability, and reduces the cost of maintaining fixed scripts.
[0079] Optionally, in the step of determining the script generation strategy based on intent data and emotion data, the following steps can be taken: determining the current dialogue emotion pattern based on current emotion state data and emotion change data; determining the first text adjustment parameters of the target outbound script data based on the current dialogue emotion pattern; determining the intent processing mode based on explicit demand information and potential demand information; determining the second text adjustment parameters of the target outbound script data based on the intent processing mode; and determining the corresponding script generation strategy based on the first and second text adjustment parameters.
[0080] In this embodiment of the invention, the aforementioned current dialogue emotion mode can refer to the comprehensive judgment result of the target user's overall emotional characteristics in the current outbound call stage based on the current emotional state data and emotional change data, and can be different modes such as stable, fluctuating, or confrontational.
[0081] In this embodiment, after acquiring emotional state data, the outbound call script generation platform also refers to the emotional change trend of adjacent rounds to classify and judge the emotions. For example, if the emotional state remains stable and the change is small, it can be judged as a stable mode, while if the emotional state changes continuously from stable to resistant, it can be judged as a fluctuating mode, thus providing a basis for subsequent script style adjustment.
[0082] The aforementioned first text adjustment parameters refer to a set of parameters generated around the current emotional pattern of the dialogue, used to control the expression and tone of the rhetoric. These parameters can be adjusted based on the emotional aspect of the rhetoric, constraining aspects such as tone intensity, the proportion of reassuring expressions, the pace of progression, and the level of information disclosure. Generally, when the current emotional pattern of the dialogue is stable, the first text adjustment parameters allow for normal progression and focused information expression; when the current emotional pattern is fluctuating or confrontational, the first text adjustment parameters tend to reduce the intensity of expression and increase empathetic and explanatory content to reduce communication pressure.
[0083] The aforementioned intent processing patterns refer to the judgment results on the target user's current intent structure and processing priority based on explicit and potential demand information. These patterns are generally manifested as explicit demands, uncertain demands, vague rejections, or demand shifts.
[0084] In one possible embodiment, when analyzing intent data, the aforementioned outbound call script generation platform combines explicit demand information and potential demand information to determine whether the user currently needs a direct response, further guidance, or a postponement. For example, when explicit demand information points to a specific request while potential demand information is weak, it can be determined as a explicit demand mode. When explicit demand information is missing and potential demand information shows an avoidance tendency, it can be determined as a vague rejection mode.
[0085] The aforementioned second text adjustment parameters refer to a set of parameters generated around the intent processing mode to control the content structure and information organization of the dialogue. These parameters primarily affect the content presentation of the dialogue, constraining the order of information presentation, the priority of responses to objections, the method of explaining benefits, and the closing confirmation path. Generally, in a clearly defined demand mode, the second text adjustment parameters emphasize directly responding to core needs and quickly concluding the dialogue; in a mode with uncertain demands or ambiguous rejections, the second text adjustment parameters emphasize clarification, explanation, and providing alternative solutions to reduce the user's decision-making burden.
[0086] By using the above methods and steps, the first text adjustment parameters and the second text adjustment parameters are integrated to form a complete script generation strategy. This strategy ensures that the generated script matches the user's emotional state in terms of tone and rhythm, and matches the user's intention characteristics in terms of content and structure. As a result, more natural and adaptive communication can be achieved in the same outbound call process.
[0087] Optionally, in the step of generating target outbound call script data based on the script generation strategy, the outbound call script text for the current round can also be generated based on the script generation strategy; during the multi-round outbound call interaction, the target user's emotional data is acquired in real time, and emotional change data is determined based on the emotional data of adjacent rounds; when the emotional change data meets the preset change conditions, the script generation strategy is adjusted based on the updated emotional data to obtain an updated script generation strategy; based on the updated script generation strategy, the outbound call script text for at least one subsequent round is updated and generated to obtain the target outbound call script data and output it.
[0088] In this embodiment of the invention, the aforementioned preset change conditions refer to a set of judgment rules used to determine whether an emotional change reaches a threshold requiring adjustment of the speech generation strategy. It is understood that a comprehensive judgment can be made by combining the direction, magnitude, and duration of the emotional change to distinguish between short-term fluctuations and trend changes. Generally, when emotional change data shows only slight fluctuations within a single round, strategy adjustment may not be triggered; when emotional changes continuously intensify across multiple adjacent rounds, or when the magnitude of the change exceeds a preset threshold, the preset change conditions are deemed met. For example, if the emotional state changes from stable and continuous to resistant, and the intensity level continues to rise, it can be considered that the preset change conditions are met, thereby triggering the subsequent adjustment process.
[0089] In one possible embodiment, when the emotion change data meets preset change conditions, the aforementioned outbound call script generation platform can adjust the original script generation strategy based on the updated emotion data to form an updated script generation strategy. This updated script generation strategy can be understood as the result of reconfiguring the text adjustment parameters based on the original strategy. Its core lies in rebalancing the tone, rhythm, and content organization of the script, making subsequent scripts more suitable for the target user's current communication state. Generally, the updated script generation strategy will reduce the original push intensity, increase reassuring and confirmatory expressions, or adjust the order of information disclosure to alleviate dialogue pressure.
[0090] After receiving the updated script generation strategy, the aforementioned outbound call script generation platform regenerates or modifies the outbound call script text for at least one subsequent round based on the updated strategy, and outputs new target outbound call script data. This allows the script expression within the same call to dynamically adapt to changes in the target user's emotional state. For example, if a standard approach is used at the beginning of the conversation, but subsequent deterioration in mood is detected, the outbound call script generation platform updates the script generation strategy, changing the script for subsequent rounds from direct advancement to a reassurance followed by confirmation, thereby reducing user resistance and maintaining communication continuity.
[0091] Optionally, in the step after generating target outbound call script data based on the script generation strategy, customer feedback data from target users regarding the target outbound call script data can also be obtained; based on the customer feedback data, the feedback evaluation result is determined; and based on the feedback evaluation result, the target user profile database, the preset scenario analysis model, and the preset semantic recognition model are updated.
[0092] In this embodiment of the invention, the aforementioned customer feedback data can be used to reflect the actual response of the target user to the effectiveness of this outbound communication. It generally includes result information such as whether the call was completed, whether it was interrupted, whether the business goal was achieved, and whether negative feedback was generated. It can also be obtained by combining confirmation operations, follow-up results, or subsequent behaviors after the call ends.
[0093] The aforementioned outbound call script generation platform analyzes and processes the customer feedback data to generate feedback evaluation results, such as different evaluation conclusions like successful communication, partial success, or communication failure. If the target user confirms key information during the outbound call without showing obvious resistance, the feedback evaluation result can be judged as positive; if the target user interrupts the call prematurely or expresses dissatisfaction multiple times, the feedback evaluation result can be judged as negative.
[0094] After generating feedback evaluation results, the aforementioned outbound call script generation platform updates relevant data and models based on the evaluation conclusions. Specifically, the target user profile database supplements or corrects user preferences, sensitivity levels, and response habits based on the feedback evaluation results, enabling subsequent outbound calls to more accurately match user characteristics. The preset scenario analysis model adjusts emotion judgment parameters based on the feedback evaluation results to improve the accuracy of recognizing specific expressions or voice features. The preset semantic recognition model optimizes intent recognition parameters based on the feedback evaluation results, making the semantic understanding results more consistent with the actual outbound call context.
[0095] By employing the methods and steps described above, user profiles and analysis models can be continuously refined through multiple outbound call interactions, gradually bringing the script generation strategy closer to real user behavior. Understandably, the feedback and evaluation results can also be incorporated as samples into subsequent model optimization processes, thereby improving the overall script generation effectiveness and adaptability without affecting the stability of online outbound calls.
[0096] In one embodiment, an outbound call script generation device is provided, which corresponds one-to-one with the outbound call script generation method described in the above embodiments. For example... Figure 3 As shown, the outbound call script generation device includes: The first data acquisition module 301 is used to collect the interaction data of the target user in real time during the outbound call process; The first processing module 302 is used to perform semantic processing on the interaction data through a preset semantic recognition model to obtain the intention data of the target user during the outbound call process; The second processing module 303 is used to perform emotion judgment processing on the interaction data through a preset scenario analysis model to obtain the emotional data of the target user during the outbound call process. The first determining module 304 is used to determine the speech generation strategy based on the intent data and the emotion data; The first generation module 305 is used to generate target outbound call script data based on the script generation strategy.
[0097] Optionally, the first acquisition module 301 is further configured to: During the outbound call process, the voice information of the target user is acquired, and the voice information is subjected to speech recognition to obtain text information; Obtain the interaction context information corresponding to the current outbound call round, wherein the interaction context information includes at least the outbound call text information and the corresponding outbound call script text information input by the target user in the historical rounds; Obtain historical interaction information associated with the target user, including historical login request information and historical outbound call record information; The voice information, text information, interaction context information, and historical interaction information are integrated to obtain the target user's interaction data.
[0098] Optionally, the first processing module 302 is further configured to: Based on the interaction data, semantic prompt information is constructed; Based on the semantic prompt information, a request is sent to the preset semantic recognition model to obtain the intent data of the target user during the outbound call process returned by the preset semantic recognition model based on the semantic prompt information. The intent data includes the target user's explicit demand information and potential demand information.
[0099] Optionally, the second processing module 303 is further configured to: Determine the interaction context information and historical interaction information of the target user; Based on the interaction context information and historical interaction information, the voice feature data and historical text content data of the target user are determined respectively. Based on the target user's voice feature data and historical text content data, construct emotion feature prompt information; Based on the emotional feature prompt information, a request is sent to the preset scenario analysis model to obtain the emotional data of the target user during the outbound call process returned by the preset scenario analysis model based on the emotional feature prompt information. The emotional data includes current emotional state data and emotional change data.
[0100] Optionally, the first determining module 304 is further configured to: Based on the current emotional state data and emotional change data, the current dialogue emotional pattern is determined; Based on the current dialogue emotion pattern, determine the first text adjustment parameter of the target outbound call script data; Based on the explicit and potential demand information, the intent processing mode is determined. Based on the intent processing mode, the second text adjustment parameters of the target outbound call script data are determined; Based on the first text adjustment parameters and the second text adjustment parameters, the corresponding speech generation strategy is determined.
[0101] Optionally, the first generation module 305 is further configured to: Based on the script generation strategy, generate the outbound call script text for the current round; During the multi-round outbound call interaction, the target user's emotional data is acquired in real time, and emotional change data is determined based on the emotional data of adjacent rounds; When the emotion change data meets the preset change conditions, the script generation strategy is adjusted based on the updated emotion data to obtain an updated script generation strategy. Based on the updated script generation strategy, the outbound script text for at least one subsequent round is updated and generated to obtain the target outbound script data and output it.
[0102] Optionally, the device is further used for: Obtain customer feedback data from the target user regarding the target outbound call script data; Based on the customer feedback data, the feedback evaluation results are determined; Based on the feedback evaluation results, the target user profile database, the preset scenario analysis model, and the preset semantic recognition model are updated.
[0103] Specific limitations regarding the outbound call script generation device can be found in the limitations of the outbound call script generation method described above, and will not be repeated here. Each module in the aforementioned outbound call script generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0104] In one embodiment, a computer device is provided, which may be a terminal device, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, and network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a readable storage medium storing computer-readable instructions. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer-readable instructions implement an outbound call script generation method. The readable storage medium provided in this embodiment includes both non-volatile and volatile readable storage media.
[0105] In this application embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, it implements the steps of the outbound call script generation method described above.
[0106] In one embodiment of the application, a readable storage medium is provided, which stores computer-readable instructions. When the computer-readable instructions are executed by a processor, they implement the steps of the outbound call script generation method described above.
[0107] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When executed, these computer-readable instructions can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0108] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0109] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for generating outbound call scripts, characterized in that, The method includes: During outbound calls, real-time collection of interaction data from target users is conducted. The interaction data is semantically processed by a preset semantic recognition model to obtain the target user's intent data during the outbound call process; The interaction data is processed by a preset scenario analysis model to determine the emotion of the target user during the outbound call process. Based on the intent data and the emotion data, a script generation strategy is determined; Based on the aforementioned script generation strategy, target outbound call script data is generated.
2. The outbound call script generation method as described in claim 1, characterized in that, The interaction data includes voice information, text information, interaction context information, and historical interaction information. During the outbound call process, the real-time collection of the target user's interaction data includes: During the outbound call process, the voice information of the target user is acquired, and the voice information is subjected to speech recognition to obtain text information; Obtain the interaction context information corresponding to the current outbound call round, wherein the interaction context information includes at least the outbound call text information and the corresponding outbound call script text information input by the target user in the historical rounds; Obtain historical interaction information associated with the target user, including historical login request information and historical outbound call record information; The voice information, text information, interaction context information, and historical interaction information are integrated to obtain the target user's interaction data.
3. The outbound call script generation method as described in claim 1, characterized in that, The step of performing semantic processing on the interaction data using a preset semantic recognition model to obtain the target user's intent data during the outbound call process includes: Based on the interaction data, semantic prompt information is constructed; Based on the semantic prompt information, a request is sent to the preset semantic recognition model to obtain the intent data of the target user during the outbound call process returned by the preset semantic recognition model based on the semantic prompt information. The intent data includes the target user's explicit demand information and potential demand information.
4. The outbound call script generation method as described in claim 1, characterized in that, The step of performing emotion judgment processing on the interaction data through a preset scenario analysis model to obtain the target user's emotion data during the outbound call process includes: Determine the interaction context information and historical interaction information of the target user; Based on the interaction context information and historical interaction information, the voice feature data and historical text content data of the target user are determined respectively. Based on the target user's voice feature data and historical text content data, construct emotion feature prompt information; Based on the emotional feature prompt information, a request is sent to the preset scenario analysis model to obtain the emotional data of the target user during the outbound call process returned by the preset scenario analysis model based on the emotional feature prompt information. The emotional data includes current emotional state data and emotional change data.
5. The outbound call script generation method as described in claim 3 or 4, characterized in that, The step of determining the script generation strategy based on the intent data and the emotion data includes: Based on the current emotional state data and emotional change data, the current dialogue emotional pattern is determined; Based on the current dialogue emotion pattern, determine the first text adjustment parameter of the target outbound call script data; Based on the explicit and potential demand information, the intent processing mode is determined. Based on the intent processing mode, the second text adjustment parameters of the target outbound call script data are determined; Based on the first text adjustment parameters and the second text adjustment parameters, the corresponding speech generation strategy is determined.
6. The outbound call script generation method as described in claim 1, characterized in that, The step of generating target outbound call script data based on the script generation strategy includes: Based on the script generation strategy, generate the outbound call script text for the current round; During the multi-round outbound call interaction, the target user's emotional data is acquired in real time, and emotional change data is determined based on the emotional data of adjacent rounds; When the emotion change data meets the preset change conditions, the script generation strategy is adjusted based on the updated emotion data to obtain an updated script generation strategy. Based on the updated script generation strategy, the outbound script text for at least one subsequent round is updated and generated to obtain the target outbound script data and output it.
7. The outbound call script generation method as described in claim 1, characterized in that, After generating the target outbound call script data based on the script generation strategy, the method further includes: Obtain customer feedback data from the target user regarding the target outbound call script data; Based on the customer feedback data, the feedback evaluation results are determined; Based on the feedback evaluation results, the target user profile database, the preset scenario analysis model, and the preset semantic recognition model are updated.
8. An outbound call script generation device, characterized in that, The device includes: The first data collection module is used to collect the target user's interaction data in real time during the outbound call process; The first processing module is used to perform semantic processing on the interaction data through a preset semantic recognition model to obtain the intent data of the target user during the outbound call process; The second processing module is used to perform emotion judgment processing on the interaction data through a preset scenario analysis model to obtain the emotional data of the target user during the outbound call process; The first determining module is used to determine the speech generation strategy based on the intent data and the emotion data; The first generation module is used to generate target outbound call script data based on the script generation strategy.
9. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and running on the processor, characterized in that, When the processor executes the computer-readable instructions, it implements the outbound call script generation method as described in any one of claims 1 to 7.
10. A readable storage medium having computer-readable instructions stored thereon, characterized in that, When the computer-readable instructions are executed by a processor, they implement the outbound call script generation method as described in any one of claims 1 to 7.