Method and system for dynamically adjusting insurance renewal call-out strategy based on emotion recognition

By processing customer voice streams in real time and performing multimodal analysis, the insurance renewal outbound call strategy is dynamically adjusted, solving the problem of rigid dialogue strategies in traditional systems. This enables accurate understanding and personalized responses to customer emotions and intentions, improving communication efficiency and customer experience.

CN121884864APending Publication Date: 2026-04-17BEIJING ZHIBAO HUIZHONG DIGITAL TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZHIBAO HUIZHONG DIGITAL TECHNOLOGY CO LTD
Filing Date
2026-02-02
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional insurance renewal outbound call systems lack the ability to effectively perceive customers' real-time emotions and intentions, resulting in rigid dialogue strategies that cannot be flexibly adjusted according to the call context, thus affecting communication efficiency and customer experience.

Method used

By receiving and processing customer voice streams in real time, combining acoustic features and semantic analysis, and using pre-trained emotion classification and intent classification models, outbound calling strategies are dynamically adjusted to generate accurate voice responses, and the interaction process is recorded to optimize the model.

Benefits of technology

It achieves a comprehensive and precise understanding of customers' psychological state and true intentions. The system can assess the situation in real time during a call, select the optimal communication strategy, and improve communication efficiency and personalized interaction capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884864A_ABST
    Figure CN121884864A_ABST
Patent Text Reader

Abstract

The invention relates to an insurance renewal call-out strategy dynamic adjustment method and system based on emotion recognition, and the method comprises the steps: receiving a customer voice stream in a call link in real time, carrying out the noise reduction processing, and separating a customer voice signal; extracting acoustic features, inputting the acoustic features into a pre-trained emotion classification model, and outputting a customer emotion state tag and confidence; converting the client voice signal into text data and carrying out semantic analysis to obtain a semantic analysis result; fusing the customer emotional state tag and the semantic analysis result, inputting a pre-trained intention classification model, and outputting a deep intention tag; querying a preset dynamic strategy map, and selecting a target dialogue strategy node and a corresponding script according to a strategy node jump condition; executing the script content of the selected target dialogue strategy node, generating response voice and outputting the response voice to the call link; and generating a structured communication report and optimizing the sentiment classification model or the dynamic strategy graph. The policy accuracy and the situation adaptability of the insurance renewal call-out system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent voice interaction technology, and in particular to a method and system for dynamically adjusting insurance renewal outbound call strategies based on emotion recognition. Background Technology

[0002] As the digital transformation of the insurance industry deepens, outbound calling systems have become crucial tools for customer relationship maintenance, renewal reminders, and service notifications. Traditional outbound calling methods are gradually evolving from purely manual dialing towards automation and intelligence, especially in renewal business scenarios with massive customer data and high timeliness requirements. Insurance companies are increasingly introducing intelligent outbound calling systems with automated dialing, script guidance, and basic interactive capabilities. These systems improve dialing efficiency to some extent and attempt to filter potential customers by analyzing historical policy data or simple keyword matching, aiming to optimize outreach strategies. Furthermore, some systems are beginning to integrate basic customer profiles or business tags, aiming to provide differentiated script content for customer groups with different characteristics.

[0003] However, most outbound calling systems rely on pre-set, fixed dialogue flows or scripts, lacking the ability to effectively perceive the real-time status of customers. During a call, the system struggles to dynamically capture and understand the complex feedback generated by customers due to personal preferences, immediate emotions, or differing opinions on product content. This leads to rigid dialogue strategies that cannot be flexibly adjusted according to the call context. For example, when customers show doubt or resistance during policy renewal communication, a fixed script may not be able to effectively reassure or provide targeted explanations, thus missing a crucial communication opportunity. Furthermore, even if some systems can obtain customer speech content through speech-to-text technology, their analysis is often limited to superficial keyword recognition. This makes it difficult for the system to accurately determine the customer's true intentions, and consequently, it cannot provide agents with the most suitable response strategies or script suggestions, impacting communication efficiency and customer experience. Summary of the Invention

[0004] To address the aforementioned technical issues, this application provides a method and system for dynamically adjusting insurance renewal outbound calling strategies based on emotion recognition.

[0005] Firstly, this application provides a method for dynamically adjusting an insurance renewal outbound call strategy based on emotion recognition, employing the following technical solution: The system receives customer voice streams from the call link in real time, performs noise reduction processing on the customer voice streams, and separates the customer voice signals. Acoustic features are extracted from the separated customer voice signals, input into a pre-trained emotion classification model, and the customer's emotional state label and confidence level are output. The customer's voice signal is converted into text data, and semantic analysis is performed on the text data to obtain the semantic analysis results; By integrating the customer emotional state tags with the semantic analysis results, the data is input into a pre-trained intent classification model, and the output is a deep intent tag. Based on the customer's emotional state tags and deep intent tags, a preset dynamic strategy graph is queried, and the target dialogue strategy node and corresponding script are selected according to the strategy node jump conditions in the dynamic strategy graph. Execute the script content of the selected target dialogue strategy node, generate the response voice and output it to the call link; Record customer emotional state tags, deep intent tags, strategy node jump paths, and response execution logs throughout the entire call process, generate a structured communication report, and optimize the aforementioned emotion classification model or dynamic strategy graph.

[0006] By employing the aforementioned technical solution, acoustic emotion recognition and natural language semantic analysis are deeply integrated, surpassing the limitations of traditional outbound calling systems that rely solely on keyword triggers. This enables a comprehensive and precise understanding of customers' psychological states and true intentions. Furthermore, relying on a dynamic strategy graph with conditional jump logic, the system can assess the situation in real time during a call, flexibly select the optimal communication strategy, and translate the strategy into a precise voice response. Crucially, the entire interaction process is fully recorded and structured for analysis, forming the data fuel that drives continuous iteration and optimization of the model and strategies. Ultimately, this technical solution endows the insurance renewal outbound calling system with high contextual adaptability, strategy precision, and autonomous evolution capabilities, providing a solid technical foundation for achieving intelligent and personalized interaction in complex interpersonal communication scenarios.

[0007] Secondly, this application provides a dynamic adjustment system for insurance renewal outbound call strategies based on emotion recognition, employing the following technical solution: The voice signal preprocessing module is used to receive the customer's voice stream in the call link in real time, perform noise reduction processing on the customer's voice stream and separate the customer's voice signal. The emotion recognition module is used to extract acoustic features from the separated customer voice signals, input them into a pre-trained emotion classification model, and output customer emotion state labels and confidence scores. The semantic analysis module is used to convert the customer's voice signal into text data, and perform semantic analysis on the text data to obtain semantic analysis results; The intent depth calculation module is used to fuse the customer emotional state labels with semantic analysis results, input them into a pre-trained intent classification model, and output deep intent labels. The dynamic strategy decision module is used to query a preset dynamic strategy graph based on the customer's emotional state tags and deep intent tags, and select the target dialogue strategy node and corresponding script according to the strategy node jump conditions in the dynamic strategy graph. The intelligent response execution module is used to execute the script content of the selected target dialogue strategy node, generate response voice, and output it to the call link; The closed-loop optimization module is used to record customer emotional state tags, deep intent tags, strategy node jump paths, and response execution logs throughout the entire call process, generate a structured communication report, and optimize the emotional classification model or dynamic strategy graph.

[0008] Thirdly, this application provides a computer device, which adopts the following technical solution: A computer device includes a memory, a processor, and a computer program stored in the memory, the processor executing the computer program to perform the steps of the method as described in the first aspect.

[0009] Fourthly, this application provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as in any of the methods in the first aspect. Attached Figure Description

[0010] Figure 1 This is a schematic diagram of the first process of a dynamic adjustment method for insurance renewal outbound call strategy based on emotion recognition, according to one embodiment of this application.

[0011] Figure 2 This is a schematic diagram of the second process of a dynamic adjustment method for insurance renewal outbound call strategy based on emotion recognition, which is one embodiment of this application.

[0012] Figure 3 This is a schematic diagram of the third process of a dynamic adjustment method for insurance renewal outbound call strategy based on emotion recognition, according to one embodiment of this application.

[0013] Figure 4 This is a schematic diagram of the fourth process of the method for dynamically adjusting the outbound call strategy for insurance renewal based on emotion recognition, which is one embodiment of this application.

[0014] Figure 5 This is a schematic diagram of the fifth process of a dynamic adjustment method for insurance renewal outbound call strategy based on emotion recognition, according to one embodiment of this application.

[0015] Figure 6 This is a schematic diagram of the sixth process of a dynamic adjustment method for insurance renewal outbound call strategy based on emotion recognition, according to one embodiment of this application.

[0016] Figure 7 This is a schematic diagram of the seventh process of a dynamic adjustment method for insurance renewal outbound call strategy based on emotion recognition, according to one embodiment of this application. Detailed Implementation

[0017] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figures 1-7 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.

[0018] This application discloses a method for dynamically adjusting insurance renewal outbound call strategies based on emotion recognition.

[0019] Reference Figure 1 A method for dynamically adjusting insurance renewal outbound calling strategies based on emotion recognition, specifically including: Step S101: Receive the customer's voice stream in the call link in real time, perform noise reduction processing on the customer's voice stream, and separate the customer's voice signal. In real-time telephone conversations, the raw customer voice stream is often mixed with ambient noise, line noise, and may even contain echoes from the agent or system. Directly analyzing this raw signal would severely interfere with the accuracy of acoustic feature extraction, leading to distortion in emotion recognition. Therefore, preprocessing using digital signal processing techniques is essential.

[0020] Specifically, noise reduction typically employs adaptive filtering algorithms (such as spectral subtraction or deep learning-based noise reduction models). The principle is to estimate the spectral characteristics of the noise and suppress or remove it from the mixed signal in the frequency domain, thereby obtaining a cleaner speech signal. The subsequent "separation of customer speech signal" operation is particularly crucial in full-duplex call scenarios. It uses technologies such as voice activity detection and voiceprint comparison to precisely segment the mixed audio stream in the call link into independent customer speech channels and agent speech channels, ensuring that subsequent analysis focuses solely on the customer's single voice source.

[0021] Step S102: Extract acoustic features from the separated customer speech signals, input them into the pre-trained emotion classification model, and output the customer emotion state label and confidence level. This involves calculating physical parameters in the speech signal that are closely related to human emotions to achieve objective quantitative recognition of the customer's subjective emotional state. Speech is not only a carrier of semantics, but its acoustic characteristics (such as pitch, loudness, speech rate, and tone quality) are also an important channel for conveying the speaker's emotional state.

[0022] Specifically, the system extracts multidimensional acoustic feature sequences from the preprocessed speech signal, such as fundamental frequency (reflecting pitch changes and related to emotional intensity), Mel-frequency cepstral coefficients (simulating human auditory characteristics, characterizing the short-time spectral shape of sound, and related to tone and timbre), and short-time energy (reflecting speech intensity changes and related to the intensity of emotion). These features constitute a mathematical vector describing the emotional content of the speech.

[0023] Subsequently, this feature vector is input into a "pre-trained sentiment classification model." This model is typically a machine learning model based on a deep neural network (such as a Convolutional Neural Network (CNN) or a Long Short-Term Memory (LSTM) network). It has been trained on massive speech datasets labeled with sentiment tags (such as "positive-pleasant," "neutral-calm," "negative-anxious," "negative-resistant," etc.), learning a complex mapping relationship from the acoustic feature space to the sentiment category space. The model calculates the input real-time speech features, outputting not only the most likely "customer sentiment state label" but also a "confidence" score to quantify the reliability of the classification, providing a reliability reference for subsequent decisions.

[0024] Step S103: Convert the customer's voice signal into text data, and perform semantic analysis on the text data to obtain the semantic analysis results; The logic behind this step lies in understanding the objective semantic information of the customer's dialogue content, serving as an important supplement to the analysis of subjective sentiment.

[0025] First, continuous customer speech signals are converted into discrete text data using Automatic Speech Recognition (ASR) technology. The ASR engine, based on acoustic and language models, maps acoustic feature sequences to the most probable word sequences. However, the original transcribed text is merely a collection of characters and requires further "semantic analysis" to extract its meaning.

[0026] Specifically, this includes: (1) Entity recognition: locating and classifying key information units from the text, such as identifying specific entities like "insurance product name," "premium price," and "coverage period" in an insurance scenario; (2) Keyword / topic extraction: identifying words that express the customer's core concerns or attitudes, such as objection or intention keywords like "too expensive," "think about it," and "slow claims processing"; (3) Dependency syntax or shallow semantic analysis: understanding the modification, negation, and other relationships between words in a sentence to more accurately grasp the semantics. The semantic analysis results are ultimately structured into feature representations containing the above information, revealing what the customer "said" and the key objective elements in the discourse.

[0027] Step S104: Integrate customer emotional state labels with semantic analysis results, input them into the pre-trained intent classification model, and output deep intent labels; Specifically, acoustic sentiment analysis alone may not be able to distinguish between "anger due to price dissatisfaction" and "anger due to poor service experience"; semantic analysis alone may only identify the literal meaning of "price objection" but cannot determine whether the customer is "strongly opposed" or "casually complaining." Therefore, it is necessary to integrate the two.

[0028] Specifically, customer emotional state labels (usually converted into one-hot encodings or embedded vectors) are used as an additional feature representing psychological state. These labels are then concatenated with text feature vectors extracted from semantic analysis results (such as concatenated word vectors or sentence vectors) or fused through an attention mechanism to form a joint feature vector. This fused vector simultaneously contains the customer's "psychological state (how to say it)" and "discourse content (what to say)".

[0029] This joint feature is then fed into another "pre-trained intent classification model" (such as a Transformer-based text classification model), which learns more complex patterns and outputs deeper intent labels. For example, it can distinguish deeper, more business-guided customer intents such as "strong price objection and potential churn," "general price inquiry," and "dissatisfaction with service but recoverable." This is more accurate than analysis based solely on text or voice.

[0030] Step S105: Based on the customer's emotional state tags and deep intent tags, query the preset dynamic strategy graph, and select the target dialogue strategy node and corresponding script according to the strategy node jump conditions in the dynamic strategy graph. The system does not employ a fixed dialogue flow but instead relies on a pre-built "dynamic strategy graph." This graph is a directed graph structure where nodes represent different dialogue strategy stages (such as "icebreakers," "product value enhancement," "price objection handling," "risk warning to close the deal," and "reassurance and recovery"). Each node stores the standard response script, key points of the dialogue, and possible interaction logic for that scenario. The directed edges between nodes represent possible turns in the dialogue flow, and each edge is bound to a logical judgment condition (jump condition) composed of a combination of "customer emotional state label" and "deep intent label."

[0031] For example, an edge originating from the "Facilitate" node might be conditional: IF Emotional Label = "Negative-Resistant" AND Deep Intent Label = "Strong Price Objection" THEN Jump to the "Reassurance and Explanation" node. During the call, the system monitors the "Deep Intent Label" and "Customer Emotional State Label" in real time, matching them against the conditions of all edges originating from the current strategy node. Once a match is found, the system immediately jumps from the current node to the "Target Dialogue Strategy Node" indicated by the condition and loads the "Corresponding Script" for that node. This enables real-time, data-driven dynamic navigation of dialogue strategies, ensuring that the agent's responses accurately adapt to the customer's current psychological state and true intent.

[0032] Step S106: Execute the script content of the selected target dialogue strategy node, generate response voice and output it to the call link; Once the target dialogue strategy node and script are selected, the system needs to convert the text-based script content into perceptible voice actions. The execution process includes fusing the text in the strategy script (which may contain variables such as customer name and product price) with the context information of the current call to generate the final response text.

[0033] Subsequently, the response text is converted into natural and fluent audio through speech synthesis technology. The system can use pre-recorded high-quality speech segments for splicing, or a more advanced end-to-end neural network speech synthesis model for generation. Finally, the generated response speech is injected into the call link in real time and delivered to the customer, thus completing a complete interactive loop from perception, understanding, decision-making to execution. This ensures that every speech by the agent (or intelligent voice robot) has a clear strategic purpose and contextual adaptability.

[0034] Step S107: Record the customer's emotional state tags, deep intent tags, strategy node jump paths, and response execution logs throughout the entire call, generate a structured communication report, and optimize the emotion classification model or dynamic strategy graph.

[0035] The system meticulously records the complete interaction trajectory of each call, generating a structured communication report. This report records the dynamic changes in customer emotions, the evolution sequence of deep intent, the jump paths of system strategy nodes, and the execution status of each response step along a timeline. This data, when correlated with the final call outcome (such as successful renewal, failure, or need for follow-up), becomes valuable optimization material.

[0036] Specifically, aggregate analysis based on a large number of such reports can quantitatively assess the actual conversion effect of different strategy nodes under different emotion-intent combinations. For example, it can be analyzed whether jumping from the "standard product introduction" node to the "risk empathy and reassurance" node can significantly improve the subsequent closing success rate when customers are in a "negative-anxious" and "deep consultation" state.

[0037] Furthermore, optimization can be carried out in two aspects: First, "optimize the dynamic strategy graph", such as adjusting the threshold of specific jump conditions, or adding, deleting, or modifying strategy nodes and edges to make the decision logic more in line with actual business rules; Second, "optimize the emotion classification model", using newly added call data as training samples to incrementally train or fine-tune the model, so that its emotion recognition ability continues to evolve with the accumulation of data, thereby forming a reinforcing cycle of "data-driven decision-making, decision-making generating data, and data feedback optimization".

[0038] In the above implementation, acoustic emotion recognition and natural language semantic analysis are deeply integrated, surpassing the limitations of traditional outbound calling systems that rely solely on keyword triggers. This achieves a comprehensive and precise understanding of the customer's psychological state and true intentions. Furthermore, relying on a dynamic strategy graph with conditional jump logic, the system can assess the situation in real time during a call, flexibly select the optimal communication strategy, and translate the strategy into a precise voice response. Crucially, the entire interaction process is fully recorded and structured for analysis, forming the data fuel that drives continuous iteration and optimization of the model and strategy. Ultimately, this technical solution endows the insurance renewal outbound calling system with high contextual adaptability, strategy precision, and autonomous evolution capabilities, providing a solid technical foundation for achieving intelligent and personalized interaction in complex interpersonal communication scenarios.

[0039] Reference Figure 2 As one implementation of step S102, the steps of extracting acoustic features from the separated customer speech signal, inputting them into a pre-trained emotion classification model, and outputting customer emotion state labels and confidence scores include: Step S201: Receive the noise-reduced and separated customer voice signal; Noise reduction refers to suppressing or eliminating steady-state and non-steady-state noise from the environment and equipment circuits through digital signal processing algorithms (such as spectral subtraction or deep learning-based noise reduction networks). The principle is to construct a noise model and remove it from the mixed signal spectrum, thereby improving the signal-to-noise ratio. Separation specifically refers to accurately separating the customer's and agent's voice signals in a full-duplex call scenario using voiceprint recognition, blind source separation, or channel-based voice activity detection technology. This ensures that the signal received in this step is a pure sound source independent of the customer and unaffected by others' voices.

[0040] This preprocessing is crucial because the calculation of all subsequent acoustic features is highly dependent on the accuracy of the signal. Any residual noise or crosstalk from other people's speech will directly contaminate the feature values, leading to fundamental deviations in emotion recognition. Therefore, this step does not receive the original recording, but rather a standardized artifact that has undergone preliminary purification, marking the formal start of the dedicated emotion analysis process.

[0041] Step S202: Perform frame segmentation processing on the customer's voice signal, dividing the continuous voice signal into voice frames of fixed duration; Speech signals exhibit typical non-stationary characteristics, with their statistical properties changing rapidly over time. However, within a very short time window (e.g., 10-30 milliseconds), the vibration patterns of the vocal cords and the resonance shape of the oral cavity can be approximated as relatively stable, exhibiting "short-term stationarity." Frame segmentation is based on this principle, using a sliding window (i.e., a "frame") of fixed length to extract speech signals along the time axis with a certain overlap rate (frame shift). For example, using a frame length of 25 milliseconds and a frame shift of 10 milliseconds means analyzing a 25-millisecond speech segment every 10 milliseconds, with a 15-millisecond overlap between adjacent frames.

[0042] Understandably, this overlapping design is intended to ensure a smooth transition between frames and avoid losing important edge information due to truncation. Before further analysis of each frame's signal, a window function (such as a Hamming window) is usually applied to reduce spectral energy leakage caused by signal truncation, allowing the signal at the frame edges to smoothly decay to zero, thereby obtaining clearer spectral characteristics. Frame segmentation transforms the continuous time-domain waveform into a time series composed of multiple short time segments, preparing data for subsequent frame-by-frame extraction of stable feature parameters.

[0043] Step S203: Extract acoustic feature parameters from each speech frame; This process transforms each frame of speech signal from its original waveform sampling points into a series of low-dimensional mathematical vectors that efficiently and compactly represent the emotional content of human speech. The extracted feature parameters are not randomly selected, but rather physical quantities that are strongly correlated with the speaker's emotional state, chosen based on phonetics and psychology research.

[0044] In some embodiments, the extracted acoustic feature parameters include fundamental frequency, Mel-frequency cepstral coefficients, and short-time energy. These three types of features, from three orthogonal and complementary dimensions of pitch (fundamental frequency), spectral structure (MFCCs), and intensity (energy), together constitute a comprehensive acoustic feature vector describing the emotional content of speech, providing the model with rich and comprehensive information input.

[0045] Specifically, the fundamental frequency reflects the basic frequency of vocal cord vibration and is the physical basis for perceiving pitch. When emotionally excited, the fundamental frequency value and its range of variation (i.e., pitch fluctuation) usually increase significantly; conversely, when emotionally depressed, the fundamental frequency tends to be flat and low. The fundamental frequency is often calculated using the autocorrelation method or the cepstral method, both of which aim to accurately estimate this most basic periodic parameter from the complex speech waveform.

[0046] Mel-frequency cepstral coefficients are parameters that characterize the shape of the short-time power spectrum of a speech signal, simulating the nonlinear auditory characteristics of the human ear (more sensitive to low frequencies). The calculation process typically involves performing a Fast Fourier Transform (FFT) on the speech frame to obtain the spectrum, filtering and integrating it using a set of triangular filters distributed according to the Mel-frequency scale, taking the logarithm to simulate the human ear's perception of loudness, and finally performing a Discrete Cosine Transform (DCT) to decorrelated and compress the data, retaining the top few coefficients that best represent the spectral envelope. Mel-frequency cepstral coefficients can effectively capture changes in oral and nasal resonance caused by emotional variations, thus reflecting subtle differences in "timbre."

[0047] Short-time energy, or the sum of the squares of the sampling points of a single frame of speech signal, directly reflects the intensity or loudness of the speech. The intensity of emotions (such as a roar of anger versus a whisper of calm) is directly reflected in the magnitude of the energy value.

[0048] Step S204: Input the extracted acoustic feature parameter sequence into the pre-trained sentiment classification model; This study utilizes machine learning models to learn a complex, non-linear mapping from a low-level acoustic feature space to a high-level human emotional semantic category space. Continuous speech signals are transformed into a temporally ordered sequence of acoustic feature parameters, where each time step corresponds to a feature vector for one frame. However, the expression and perception of human emotions are dynamic and time-dependent processes; for example, the emotional intensity of a complaining utterance may accumulate as the syllables progress.

[0049] Therefore, the "pre-trained emotion classification model" used is specifically defined as a deep learning model based on temporal neural networks, such as Long Short-Term Memory networks or gated recurrent units. This type of model is essentially a special recurrent neural network with an ingenious "gating" mechanism that allows it to selectively remember or forget historical information, thus excelling at processing and modeling long-distance dependencies in time-series data. This model has been trained on a massive dataset of manually labeled emotional speech data. The training process involves continuously adjusting millions of connection weights within the network, enabling it to automatically capture temporal patterns and feature combinations associated with specific emotion categories (such as "positive" or "negative-resistant") from the input acoustic feature sequences. During the inference phase, the model receives real-time feature sequences, and its internal hidden states evolve continuously with each time step, integrating information from the entire speech segment to ultimately form a comprehensive judgment of the speaker's current emotional state.

[0050] Step S205: Calculate and output the customer's emotional state label and corresponding confidence level using the emotion classification model.

[0051] The final layer of the emotion classification model is usually a Softmax classifier, which receives the comprehensive feature vector containing the contextual information of the entire sequence from the last time step of the LSTM or GRU network, and calculates the probability distribution of the speech segment belonging to each preset emotion category through a series of linear and nonlinear transformations.

[0052] Specifically, the category with the highest probability value is selected as the output "customer emotional state label". The labels exemplified in the claims, such as "positive", "negative-impatient", and "negative-resistant", are carefully defined based on common customer reactions in insurance outbound calling scenarios, making the output results highly business-oriented and operable.

[0053] Simultaneously, the model outputs the "confidence score" corresponding to the label, which is typically the highest probability value obtained for that category. Confidence score is a crucial piece of meta-information, quantifying the model's degree of certainty in its judgment. High confidence score indicates a high degree of match between the acoustic features and the pattern of a particular emotion category; low confidence score may suggest that the current speech segment's emotional expression is ambiguous, confusing, or in a transitional phase of emotion shift. This confidence score information is vital for subsequent decision fusion steps (such as fusion with the semantic intent in the original claim). Based on the confidence score, the system can decide whether to fully adopt the acoustic emotion judgment or rely more on information from other modalities (such as textual semantics) for a comprehensive decision, thereby improving the robustness of the entire system in handling complex situations.

[0054] In the above implementation, non-stationary speech signals are standardized based on rigorous frame-segmentation processing. Then, starting from physiological and auditory principles, three core acoustic features—fundamental frequency, Mel-frequency cepstral coefficients, and short-time energy—are systematically extracted, comprehensively capturing the mapping of emotions across pitch, timbre, and intensity dimensions. Subsequently, utilizing the powerful sequence modeling capabilities of a pre-trained temporal neural network model, this model can deeply explore the dynamic patterns of acoustic feature evolution over time, thereby achieving nuanced and accurate identification and classification of customer emotional states. The final output not only includes highly contextualized emotion labels but also incorporates the model's own confidence assessment, providing upper-level decision-making with dual information input that considers both the judgment result and the reliability of the result.

[0055] Reference Figure 3 As one implementation of step S103, the steps of converting the customer's voice signal into text data and performing semantic analysis on the text data to obtain the semantic analysis result include: Step S301: Receive the noise-reduced and separated customer voice signal; Among them, the voice signal is no longer the original call flow, but a finished product after deep cleaning and channel separation. The specific processing process can be referred to the detailed description in step S201 above.

[0056] Step S302, convert the customer voice signal into initial text data through automatic speech recognition technology; Among them, this step preferably uses an end-to-end deep learning model, which is usually based on architectures such as Transformer or RNN-T. Its working principle is to directly map the input sequence of speech frames (i.e., the time series of acoustic features) to the output sequence of characters or subwords. It omits multiple components such as independent acoustic models, pronunciation dictionaries, and language models in traditional ASR, and through a unified neural network, it is trained end-to-end on a large amount of "speech-text" paired data to learn the optimal mapping relationship from acoustic features to text symbols.

[0057] In the embodiment of the present application, for the insurance outbound call scenario, the model can be pre-trained for domain adaptation. The training data contains a large number of conversation recordings and transcribed texts in the insurance domain, enabling the model to not only learn general language patterns but also specifically learn the pronunciation habits, liaison methods, and contexts unique to financial insurance terms such as "premium", "cooling-off period", and "beneficiary", thereby significantly improving the recognition accuracy of professional vocabulary. Although the initially output text data may have homophone errors or inaccurate sentence breaks, it has transformed the physical expression of sound into a symbol sequence that can be further processed by algorithms.

[0058] Step S303, preprocess the initial text data, including word segmentation, stop word removal, and error correction; Specifically, first perform word segmentation. For languages such as Chinese without obvious word boundary delimiters, this operation is crucial. It uses a word segmentation model based on statistics or deep learning to split a continuous character sequence into the smallest units with independent semantics - words (or subwords). For example, correctly segment "I don't want to renew the insurance" into "I / don't want / to renew / the insurance", which is the cornerstone of all subsequent morphological and syntactic analyses. Then remove stop words. Stop words refer to those words that appear extremely frequently in the text but carry very little specific information content, such as "de", "le", "ma", "wo", etc. Removing these words can significantly reduce data noise, lower the subsequent computational complexity, and make keyword and entity information more prominent. Finally, perform error correction. Since ASR may produce recognition errors due to accents, speech rates, or residual noise (such as misrecognizing "premium" as "baofei"), this step checks and corrects the text through a language model that combines context or a pre-trained error correction model.

[0059] Step S304, based on the insurance domain scenario, extract keywords and entity information from the preprocessed text; This step delves into the vertical field of insurance renewal outbound calls, accurately extracting structured information that directly drives business decisions from the cleaned text. This is not a simple keyword matching, but a multi-layered semantic information extraction process.

[0060] Specifically, the system matches a pre-defined keyword database for the insurance industry. This database includes not only standard product names (such as "critical illness insurance" and "million-dollar medical insurance"), but also common expressions used in business scenarios. Furthermore, it employs a named entity recognition (NER) model, a sequence labeling technique that intelligently identifies words or phrases belonging to specific categories within text. In the insurance context, the NER model is specifically trained to recognize entities such as "premium amount" (e.g., "five thousand yuan"), "policy number," and "date and time" (e.g., "next month").

[0061] In addition, the system identifies keywords related to customer objections. By analyzing massive amounts of historical dialogues, the system has constructed an objection classification system (such as price objections, service inquiries, and information inquiries) and trained a model to mark key expressions reflecting these objections from customer discourse. For example, "too expensive" corresponds to a price objection, and "too slow claims processing" corresponds to a service inquiries.

[0062] Step S305: Integrate contextual semantics and output structured semantic analysis results.

[0063] Analyzing only a single sentence can lead to misunderstandings. For example, a customer saying "I'll take another look" might seem like an information inquiry in isolation, but if the agent has just quoted a price, the deeper intention is more likely a subtle price objection. Therefore, this step goes beyond superficial analysis of isolated statements. By introducing dialogue history and strategic context, it achieves a deep understanding and structured representation of the customer's true intentions, completing the final loop of semantic analysis.

[0064] In this embodiment, the fusion operation first performs dependency parsing. By analyzing the subject-verb, verb-object, and modifier-relative grammatical dependencies between words in the sentence, a logical structure diagram of the sentence is constructed, thereby understanding "who did what" and accurately grasping the semantic focus. A deeper level of fusion involves combining the strategy types of preceding dialogue nodes. The system interprets the current customer's utterance within the entire dialogue process, considering the strategies previously implemented by the agent (such as the customer saying "wait" after a strong push), thereby inferring the potential motivation behind the customer's current response.

[0065] Ultimately, all the analytical results were integrated into a structured semantic analysis result. This result is not a simple text summary, but a machine-readable data object containing a "keyword list," a "sequence of entity labels" (indicating the types of entities and their positions in the text), and a "contextual intent confidence score."

[0066] In the above implementation, the system deeply integrates with insurance business scenarios, accurately extracting key entity and intent information such as product, amount, and objection type. Ultimately, by fusing the context of the dialogue with historical strategy logic, it deciphers the true motivations behind customer statements and outputs structured semantic data that allows for direct machine decision-making. This technical solution provides crucial language understanding input for the entire outbound call strategy dynamic adjustment system, forming a core pillar alongside acoustic emotion recognition in multimodal intelligent decision-making. It significantly enhances the system's cognitive depth and accuracy in responding to complex business dialogues.

[0067] Reference Figure 4 As one implementation of step S104, the steps of fusing customer emotional state labels and semantic analysis results, inputting them into a pre-trained intent classification model, and outputting deep intent labels include: Step S401: The confidence vector of the customer's emotional state label is concatenated with the text feature vector of the semantic analysis result to generate a fused feature representation. Among them, the algorithm organically integrates information from different dimensions and sources, namely emotion and semantics, at the original feature level. The feature-level fusion algorithm used is a deep fusion strategy, which is different from decision-level fusion.

[0068] Specifically, the system first transforms the emotion state labels into a machine-processable form. For example, it converts label categories (such as "negative-resistant") into a probability distribution vector composed of their confidence scores. This vector encodes the emotion recognition result and its certainty. Simultaneously, it transforms the textual information in the semantic analysis results into a dense text feature vector through word embedding or sentence embedding techniques.

[0069] Then, the two feature vectors are concatenated dimensionally to form a longer, fused feature vector. The advantage of this concatenation method is that it preserves all the information from the original two modalities to the greatest extent possible, without performing any selection or filtering that might cause information loss before fusion. This allows the subsequent intent classification model to autonomously learn the complex interaction relationships and joint patterns between sentiment and semantic features from the data. For example, the model can learn that when the entity "price" appears in the text and is accompanied by a high-confidence "negative" sentiment, it should be mapped to a different deep intent than when "price" appears in the text but the sentiment is "neutral."

[0070] Step S402: Input the fused feature representation into the pre-trained intent classification model; Among them, by utilizing the ability of complex nonlinear function mapping, it can automatically mine and identify deep customer intent patterns from fused features that cannot be reliably inferred from single modal information alone.

[0071] In this embodiment, the pre-trained intent classification model is a deep learning-based neural network model. Its core mission is to complete a classification task: mapping the multimodal fusion feature representation generated in the previous step to a "deep intent label" that best summarizes the customer's current complex state. The input layer dimension of the model must precisely match the size of the fusion feature representation to ensure lossless information injection. The model itself (such as a convolutional neural network or Transformer architecture) can construct an extremely complex decision boundary through its multi-layered nonlinear transformation structure and a large number of internal parameters. This model has been pre-trained on massive amounts of labeled "fusion feature-deep intent" pairing data.

[0072] During training, the model continuously adjusts its weights through backpropagation, gradually learning to identify patterns such as: high negative sentiment coupled with specific price objection words (e.g., "too expensive") likely pointing to "strong price objection"; while moderate negative sentiment coupled with service-related words (e.g., "slow claims processing") likely pointing to "service doubts and concerns." The model functions similarly to a high-level pattern recognizer, with its input being a mixture of surface-level sentiment and semantics, and its output being abstracted and interpreted deep conclusions with direct business guidance.

[0073] Step S403: Infer deep intent labels through intent classification model reasoning.

[0074] Among them, the deep intent tags include at least one type of price objection, service questioning, and information inquiry.

[0075] During the inference phase, the fused feature representation is forward-propagated through a pre-trained intent classification model. The final layer of the model is typically a Softmax classification layer, which transforms the abstract features output from the final hidden layer of the neural network into probability distributions belonging to various pre-defined "deep intent labels." The category with the highest probability is determined as the deep intent label for this inference. The depth of this label is reflected in the fact that it is no longer a simple sentiment classification or keyword matching result, but a concise expression of the customer's true motivation and attitude under the interplay of emotion and semantics.

[0076] For example, in an insurance renewal scenario, deep intent tags might be precisely defined as "strong price objection," "hesitation requiring value reinforcement," "misunderstanding requiring clarification," and "service dissatisfaction requiring reassurance," etc. These tags directly correspond to different strategy branches that need to be adopted in outbound call dialogues. Outputting such a tag means that the system has completed the entire chain from multimodal perception of raw signals to feature extraction, then to feature fusion, and finally to advanced cognitive judgment, elevating the understanding of customer status to an operational cognitive level that can directly drive business decisions.

[0077] In the above implementation, the emotional state identified by the acoustic channel and the semantic information parsed by the text channel are deeply fused at the feature level to generate a joint feature that is more comprehensive and richer in representation. Subsequently, a pre-trained deep learning model is used to deeply mine and recognize patterns in this fused feature, ultimately outputting a deep intent label that goes beyond surface information and directly points to the customer's intrinsic motivation and attitude. This deep intent label serves as a direct and accurate input condition for subsequent strategy graph transitions, enabling the entire system to make the most context-appropriate interaction strategy selection based on a deep understanding of the customer's psychology and needs. This elevates the intelligence level of human-computer dialogue from simple keyword responses to a new level of strategic dialogue based on multimodal fusion perception, possessing deep understanding and empathy.

[0078] Reference Figure 5 As one implementation of step S105, the steps of querying a preset dynamic strategy graph based on customer emotional state tags and deep intent tags, and selecting the target dialogue strategy node and corresponding script according to the strategy node jump conditions in the dynamic strategy graph include: Step S501: Based on the preset dynamic strategy graph, the real-time received customer emotional state tags and deep intent tags are logically compared with the jump conditions of each side in the graph. The preset dynamic strategy graph is a directed graph data structure, where nodes represent different strategy stages or links in insurance outbound call dialogues (such as value enhancement and objection handling), and directed edges between nodes represent the "jump conditions" that must be met to transition from one strategy link to another. These jump conditions are not in a fixed order, but are precisely defined as a set of logical combination expressions (such as Boolean expressions) based on "customer emotional state labels" and "deep intent labels".

[0079] During the call, the system uses real-time received tag combinations (e.g., emotion = "negative-resistant", intent = "strong price objection") as a query vector, comparing it with the logical conditions attached to all outgoing edges from the currently active node in the dynamic strategy graph. This process is essentially a real-time pattern matching and rule triggering determination, aiming to find the entry point of the path that best matches the current customer's psychological state and needs from numerous potential dialogue paths preset in the graph. The query and matching results identify the specific jump rule that should be activated in the current context, thus providing a clear instruction signal for subsequent path transitions.

[0080] Step S502: When the tag combination satisfies the jump condition of any edge, select the target dialogue strategy node from the dynamic strategy graph according to the matched jump condition. Once a certain jump condition is met (i.e., the logical expression evaluates to "true"), the system executes the graph traversal operation bound to that condition, jumping from the "current node" in the graph to the "target dialogue strategy node" pointed to by the directed edge. For example, if the current node is "value facilitation" and the rule "IF Emotion = Negative-Resistant AND Intent = Strong Price Objection THEN Jump to the reassurance and explanation node" is matched, the system will switch the dialogue's "strategy context" from the original facilitation logic to a strategy node centered on reassurance and explanation.

[0081] Understandably, this selection mechanism abandons fixed, linear dialogue processes, allowing the dialogue path to branch, backtrack, or leap based on real-time customer feedback. This is analogous to an experienced agent dynamically selecting the most appropriate response script from their strategy library when faced with different customer reactions. The selected "target dialogue strategy node" not only represents the core objective of the next communication step (such as reassurance, clarification, or facilitation), but also embodies the standardized communication framework and logic pre-set to achieve that objective.

[0082] Step S503: Access the pre-stored script database, retrieve the associated script content according to the selected target dialogue strategy node identifier, and output it; wherein, the script content includes text-based dialogue content or voice templates.

[0083] Specifically, through predefined mapping relationships, node identifiers representing strategic intent are instantiated into standardized interactive content that can be directly delivered to the execution engine. Each target dialogue strategy node in the system is associated with one or more "dialogue scripts," which are pre-stored in a script library or database. Scripts are the embodiment of the strategy, and their content can be validated, efficient dialogue text, voice templates, or structured dialogue components (such as greetings, value point lists, rhetorical questions, etc.) under that strategy node.

[0084] In this embodiment, once a target node is selected, the system initiates a retrieval request to the script database based on the node identifier (such as node ID or name) to obtain the corresponding script bound to it. The output script content is the core expressive material that the agent (or speech synthesis system) should follow or transform in the next round of interaction. For example, the script corresponding to the "reassurance and explanation node" may include standard reassurance phrases such as "I fully understand your concern about the price," as well as a concise explanation framework for the product's value points.

[0085] In the above implementation, accurate dual tags of emotion and deep intent are received, and real-time navigation is performed within a structured strategy graph, achieving a fundamental shift in dialogue strategy from static preset to dynamic response. By transforming the complex art of sales communication into a computable and navigable graph model, the system can intelligently select the most suitable path from multiple preset optimization paths based on the customer's immediate psychological state and true needs, and instantly generate a matching strategy script. This technical solution overcomes the shortcomings of traditional automated outbound calls or fixed script processes, which are rigid and unable to flexibly respond to complex customer emotions. It endows the system with a high degree of situational awareness and strategic adaptability, ensuring that every customer interaction is conducted within the optimal strategic framework, thus providing a core technical decision-making mechanism for achieving personalized communication and improving overall service efficiency.

[0086] Reference Figure 6 As one implementation of step S107, the steps of recording customer emotional state tags, deep intent tags, strategy node jump paths, and response execution logs throughout the call, generating a structured communication report, and optimizing the emotion classification model or dynamic strategy graph include: Step S601: Collect customer emotional state tags, deep intent tags, strategy node jump paths, and response execution logs in real time throughout the entire call process; The time sequence of customer emotional state tags and deep intent tags records the dynamic evolution of the customer's psychological state and core demands throughout the conversation, providing both emotional and rational dimensions for understanding customer reactions. The strategy node jump path precisely records every strategic decision made by the system based on the above understanding; that is, from the beginning to the end of the conversation, the system sequentially activates which strategy nodes (e.g., from "value reinforcement" to "price objection handling"), and this path is a complete presentation of the system's thought process. The "response execution log" records operational details such as whether the specific script corresponding to each strategy node was fully executed and the execution duration.

[0087] Step S602: Store the collected data in the evaluation database in a time-series structure. The system creates a separate record for each call session, with key fields in each record arranged and stored strictly according to time sequence. For example, the "timestamp field" marks the precise moment each key event occurred; the "emotion tag sequence field" stores the evolution of emotional states in chronological order; and the "node jump path field" records the transition chain of strategy nodes in sequence. This structured storage according to time sequence can discretize a dynamic, continuous dialogue process into an event sequence with strict temporal relationships.

[0088] Step S603: Based on the preset evaluation dimension rules, perform field mapping and correlation analysis on the structured data; Specifically, by using predefined business logic and evaluation rules, the system performs deep correlation and cross-calculation on the original structured data, thereby revealing the hidden causal relationship between interaction effects and the decision-making process. The system executes a series of complex correlation analyses based on preset evaluation dimension rules.

[0089] For example, by establishing a mapping table between emotion tags and strategy nodes, the success rate difference between the system jumping to the "soothing and explaining" node and the system jumping to the "facilitating" node can be statistically analyzed when a customer exhibits "negative-resistant" emotions. This allows for the quantification of the effectiveness of different strategies in addressing specific emotions. By calculating the objection handling response delay time metric (the time difference between identifying a specific intent and the system executing the corresponding processing script), the agility of the system's decision-making and execution can be assessed.

[0090] Furthermore, by correlating the final "renewal status" with the complete "strategy path" experienced, a matching score can be calculated to evaluate which type of strategy decision tree (i.e., a specific sequence of node jumps) is more likely to lead to a successful business outcome. This correlation analysis goes beyond the statistics of single-point data, aiming to establish quantifiable connections between multidimensional data (sentiment, intent, strategy, and outcome), thereby elevating interactive data into strategic knowledge that can guide decision-making.

[0091] Step S604: Output a structured communication report containing the emotional interaction curve, strategy execution path, and result indicators; The structured communication report is a comprehensive data product. Its "emotional interaction curve," plotted on the horizontal axis with emotional intensity or type on the vertical axis, visually illustrates the fluctuations in customer emotions throughout the call, clearly marking key moments of emotional shifts. The "strategy execution path," typically presented as a decision tree or state diagram, graphically displays the entirety of the strategy nodes the system experienced during the call, aligned with the emotional curve on the timeline, making it immediately clear "where, why, and what strategy was adopted." The "outcome metrics" include various quantitative evaluation results of the call, such as the aforementioned response delay, intent recognition confidence level, and strategy matching score.

[0092] Step S605: Based on the correlation analysis results in the structured communication report, iteratively optimize the sentiment classification model or dynamic strategy map.

[0093] The optimization process is divided into two parallel channels. For the emotion classification model, the system can automatically extract ambiguous voice data and their true labels (which can be inferred through a small amount of manual verification or by using high-confidence samples) from call segments marked in the report as having "emotion label confidence below the threshold," forming an incremental training sample set. After data augmentation, the model is retrained using these new samples, and its weight parameters are fine-tuned to specifically improve the model's recognition accuracy in areas with ambiguous original classification boundaries.

[0094] For dynamic strategy graphs, optimization becomes even more strategic. The system can statistically analyze the path data of all calls in the report to identify strategy edges with "conversion rates below preset values" (for example, conversations jumping from node A to node B have significantly lower renewal success rates). Based on this correlation analysis, the system can automatically "correct the logical expression of the jump condition," such as tightening or loosening the emotion-intent combination condition that triggers the jump; or "add / delete optional script options in the strategy graph," replacing inefficient scripts with more effective ones.

[0095] In the above implementation, multimodal data from the entire interaction chain is systematically collected and structured into an analyzable time series. Then, in-depth correlation analysis is performed using preset business rules, and the insights derived from the analysis are automatically fed back into the optimization process of the sentiment recognition model and the strategy decision graph. This closed-loop mechanism not only ensures the system's adaptability to the current context but also ensures continuous improvement in its long-term performance and the ongoing refinement of its strategies through data-driven iteration.

[0096] Reference Figure 7 As a further implementation of the method for dynamically adjusting outbound call strategies for insurance renewal, before executing the script content of the selected target dialogue strategy node, the method further includes: Step S701: Retrieve risk warning entries associated with deep intent tags from the pre-built insurance knowledge graph; The pre-built insurance knowledge graph is not a simple keyword database, but a domain knowledge base organized with a graph data structure. Its nodes represent entities in the insurance domain (such as specific insurance clauses, regulatory entries, risk types, and product names), while the edges represent semantic relationships between entities (such as "trigger", "violation", "belong to", and "related cases").

[0097] In this embodiment, once the system obtains a deep intent tag (e.g., "strong price objection" or "service challenge"), it initiates a graph query and reasoning process. First, entity linking is performed, mapping the semantics of the intent tag to the most relevant entity nodes in the knowledge graph. For example, "strong price objection" is mapped to the regulatory node "Negative List of Life Insurance Products – Opaque Rates".

[0098] Subsequently, the system performs graph traversal and relational reasoning along the various relational edges connected to the node to retrieve all associated "risk warning entries." These entries may include specific prohibitions, mandatory notification obligations, and characteristics of similar customer objections leading to complaints or lawsuits in the past.

[0099] Step S702: Extract the compliance rule constraints from the risk warning entries; Among these, legal provisions in natural language (such as "sales personnel shall not promise policy benefits") cannot be directly used by computers for comparison and verification; they must be transformed into formal logical expressions. This step involves constraint deconstruction, that is, using natural language processing technology to break down the provisions into logical pairs of "condition-behavior" or "condition-state".

[0100] For example, "No promises of returns" can be broken down into: IF (the dialogue script contains predicates such as "promise," "guarantee," or "definitely") AND (the script object involves concepts such as "returns," "interest," or "rate of return") THEN Status = Violation. Simultaneously, the system prioritizes and weights the extracted rules based on their source authority and severity, enabling the system to distinguish between serious violations and general advice during subsequent verification. Ultimately, all relevant risk warning items are transformed into a prioritized, executable set of compliance rule constraints, which serve as a precise benchmark for measuring the compliance of the script content.

[0101] Step S703: Match and verify the script content of the target dialogue strategy node with the compliance rule constraints; Specifically, the first step is syntactic layer verification. The system uses regular expressions and a sensitive word database to quickly scan the script text for explicitly prohibited words or phrase combinations. This is an efficient but relatively superficial check. The second step is semantic layer verification. To address more subtle and disguised violations (e.g., rewriting "promised returns" as "previous clients have received good returns"), the system uses a pre-trained language model like BERT. This model encodes the overall sentence meaning and the deep semantics of each sentence into high-dimensional vectors, while also encoding the restrictive semantics expressed by the "compliance rule constraints" into vectors. By calculating the cosine similarity between the two or performing implication reasoning, the system determines whether the script's substantive meaning violates the spirit of the rules.

[0102] For example, even if the word "guarantee" does not appear in the script, if semantic analysis determines that its overall expression constitutes a deterministic commitment, the verification will still fail. The system combines the results of both syntactic and semantic verification, and gives a "satisfied" or "not satisfied" judgment for each high-priority constraint, ultimately forming the overall verification conclusion.

[0103] Step S704: When the matching verification fails, select alternative script content that meets the verification conditions from the pre-configured backup strategy node library; When the primary script is deemed to pose a compliance risk, the system can intelligently and quickly select the best alternative from the preset "strategy arsenal" that achieves similar business objectives while fully complying with regulatory requirements.

[0104] In this application embodiment, the backup strategy node library is a script resource pool that has undergone pre-compliance review and scenario labeling. Each backup script is associated with its applicable business objectives (such as "value reinforcement" and "objection appeasement"), the appropriate emotional state (such as "neutral" and "negative"), and a compliance label that has passed verification.

[0105] Specifically, the selection process is a multi-objective optimization decision: First, the system filters out a set of candidate scripts that match the business stage (original target node) and the customer's emotional state. Then, for each script in the candidate set, a rapid compliance rule matching verification is performed again (or its pre-stored compliance tags are invoked) to ensure its absolute safety. Finally, the system uses an optimal selection algorithm to calculate the comprehensive score of each candidate script across multiple dimensions, including "business objective matching degree," "emotional fit," and "rule coverage completeness," and selects the script with the highest weighted total score as the final replacement. This ensures that the replacement is not random, but rather the most strategic and smoothest transition within the compliance framework.

[0106] Step S705: Generate script replacement instructions and update the response execution log.

[0107] The system generates a structured script replacement instruction package, which acts like a precise surgical plan. It explicitly includes the identifier of the original script to be replaced, the identifier of the newly selected replacement script, the specific rule identifier that triggered the replacement (accurate to the legal clause number), and the replacement trigger timestamp accurate to the millisecond. This instruction is sent to the speech synthesis and broadcasting engine in real time, driving it to switch the broadcast content.

[0108] At the same time, the system will also record in the response execution log that the "compliance status" has changed to "replaced", and will also record in detail the rule ID that caused the replacement, the hash value of the script content before and after the replacement (used for content integrity verification), and the associated session context.

[0109] In the above implementation, an intelligent compliance verification and self-correction subsystem driven by an insurance knowledge graph and based on real-time rule matching is embedded in the dynamic outbound call strategy execution chain, realizing a fundamental shift from "post-event compliance review" to "real-time risk control during the event." By automatically identifying potential compliance risks in the wording and intelligently switching to safe and appropriate alternative content, regulatory penalties and legal risks caused by inappropriate language are proactively intercepted while ensuring smooth and continuous business dialogue. This technical solution not only significantly reduces reliance on manual legal review and improves operational efficiency, but also effectively resolves the long-standing contradiction between business agility and regulatory rigidity in the insurance telemarketing field.

[0110] This application also discloses a dynamic adjustment system for insurance renewal outbound call strategies based on emotion recognition.

[0111] A dynamic adjustment system for insurance renewal outbound calling strategies based on emotion recognition, specifically including: The voice signal preprocessing module is used to receive the customer's voice stream in the call link in real time, perform noise reduction processing on the customer's voice stream, and separate the customer's voice signal. The emotion recognition module is used to extract acoustic features from the separated customer speech signals, input them into a pre-trained emotion classification model, and output customer emotion state labels and confidence scores. The semantic analysis module is used to convert customer voice signals into text data and perform semantic analysis on the text data to obtain semantic analysis results. The intent depth calculation module is used to integrate customer emotional state labels and semantic analysis results, inputting them into a pre-trained intent classification model and outputting deep intent labels. The dynamic strategy decision-making module is used to query the preset dynamic strategy graph based on customer emotional state tags and deep intent tags, and select the target dialogue strategy node and corresponding script according to the jump conditions of the strategy node in the dynamic strategy graph. The intelligent response execution module is used to execute the script content of the selected target dialogue strategy node, generate response voice, and output it to the call link; The closed-loop optimization module records customer emotional state tags, deep intent tags, strategy node jump paths, and response execution logs throughout the entire call process, generates structured communication reports, and optimizes the emotion classification model or dynamic strategy graph.

[0112] The emotional recognition-based insurance renewal outbound call strategy dynamic adjustment system of this application embodiment can implement any of the above methods, and the specific working process of each module in the system can refer to the corresponding process in the above method embodiment.

[0113] In the several embodiments provided in this application, it should be understood that the provided methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for example, the division of a certain module is merely a logical functional division, and in actual implementation there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0114] This application also discloses a computer device.

[0115] A computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a dynamic adjustment method for an insurance renewal outbound call strategy based on emotion recognition, as described above.

[0116] This application also discloses a computer-readable storage medium.

[0117] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above in any of the methods for dynamically adjusting an insurance renewal outbound call strategy based on emotion recognition.

[0118] The computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device; the program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0119] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.

Claims

1. A method for dynamically adjusting insurance renewal outbound calling strategies based on emotion recognition, characterized in that, The method includes: The system receives customer voice streams from the call link in real time, performs noise reduction processing on the customer voice streams, and separates the customer voice signals. Acoustic features are extracted from the separated customer voice signals, input into a pre-trained emotion classification model, and the customer's emotional state label and confidence level are output. The customer's voice signal is converted into text data, and semantic analysis is performed on the text data to obtain the semantic analysis results; By integrating the customer emotional state tags with the semantic analysis results, the data is input into a pre-trained intent classification model, and the output is a deep intent tag. Based on the customer's emotional state tags and deep intent tags, a preset dynamic strategy graph is queried, and the target dialogue strategy node and corresponding script are selected according to the strategy node jump conditions in the dynamic strategy graph. Execute the script content of the selected target dialogue strategy node, generate the response voice and output it to the call link; Record customer emotional state tags, deep intent tags, strategy node jump paths, and response execution logs throughout the entire call process, generate a structured communication report, and optimize the aforementioned emotion classification model or dynamic strategy graph.

2. The method for dynamically adjusting insurance renewal outbound calling strategies based on emotion recognition according to claim 1, characterized in that, The steps of extracting acoustic features from the separated customer speech signals, inputting them into a pre-trained emotion classification model, and outputting customer emotion state labels and confidence scores include: Receive the noise-reduced and separated customer voice signal; The customer's voice signal is processed by frame segmentation, dividing the continuous voice signal into voice frames of fixed duration; Extract acoustic feature parameters from each speech frame; The extracted acoustic feature parameter sequence is input into a pre-trained sentiment classification model; The emotion classification model is used to calculate and output customer emotion state labels and corresponding confidence levels.

3. The method for dynamically adjusting insurance renewal outbound calling strategies based on emotion recognition according to claim 1, characterized in that, The steps of converting the customer's voice signal into text data and performing semantic analysis on the text data to obtain the semantic analysis results include: Receive the noise-reduced and separated customer voice signal; The customer's voice signal is converted into initial text data using automatic speech recognition technology; The initial text data is preprocessed, including word segmentation, stop word removal, and error correction. Based on the insurance industry scenario, keywords and entity information are extracted from preprocessed text; By integrating contextual semantics, the output is a structured semantic analysis result.

4. The method for dynamically adjusting insurance renewal outbound calling strategies based on emotion recognition according to claim 3, characterized in that, The steps of integrating the customer emotional state labels and semantic analysis results, inputting them into a pre-trained intent classification model, and outputting deep intent labels include: The confidence vector of the customer's emotional state label is concatenated with the text feature vector of the semantic analysis result to generate a fused feature representation; The fused feature representation is input into a pre-trained intent classification model; The intent classification model is used to infer and output a deep intent label.

5. The method for dynamically adjusting insurance renewal outbound calling strategies based on emotion recognition according to claim 4, characterized in that, Based on the customer's emotional state tags and deep intent tags, the steps of querying a preset dynamic strategy graph and selecting the target dialogue strategy node and corresponding script according to the strategy node jump conditions in the dynamic strategy graph include: Based on the preset dynamic strategy graph, the real-time received customer emotional state tags and deep intent tags are logically compared with the jump conditions of each side in the graph. When the tag combination satisfies the jump condition of any edge, the target dialogue strategy node is selected from the dynamic strategy graph according to the matched jump condition. Access the pre-stored script database, retrieve the associated script content based on the selected target dialogue strategy node identifier, and output it; wherein, the script content includes text-based dialogue content or voice templates.

6. The method for dynamically adjusting insurance renewal outbound calling strategy based on emotion recognition according to claim 1, characterized in that, The steps for recording customer emotional state tags, deep intent tags, strategy node jump paths, and response execution logs throughout the entire call, generating a structured communication report, and optimizing the emotional classification model or dynamic strategy graph include: Real-time collection of customer emotional state tags, deep intent tags, strategy node jump paths, and response execution logs throughout the entire call process; The collected data is stored in the evaluation database in a time-series structure. Based on preset evaluation dimension rules, field mapping and correlation analysis are performed on structured data; The output includes a structured communication report containing emotional interaction curves, strategy execution paths, and outcome metrics; Based on the correlation analysis results in the structured communication report, the sentiment classification model or dynamic strategy graph is iteratively optimized.

7. A method for dynamically adjusting an insurance renewal outbound call strategy based on emotion recognition, as described in any one of claims 1 to 6, characterized in that, Before executing the script content of the selected target dialogue strategy node, the following is also included: Retrieve risk warning entries associated with the deep intent tags from the pre-built insurance knowledge graph; Extract the compliance rule constraints from the risk warning entries; The script content of the target dialogue strategy node is matched and verified against the compliance rule constraints. When the matching verification fails, select an alternative script content that meets the verification conditions from the pre-configured backup strategy node library; Generate script replacement instructions and update the response execution log.

8. A dynamic adjustment system for insurance renewal outbound calling strategies based on emotion recognition, characterized in that, The system includes: The voice signal preprocessing module is used to receive the customer's voice stream in the call link in real time, perform noise reduction processing on the customer's voice stream and separate the customer's voice signal. The emotion recognition module is used to extract acoustic features from the separated customer voice signals, input them into a pre-trained emotion classification model, and output customer emotion state labels and confidence scores. The semantic analysis module is used to convert the customer's voice signal into text data, and perform semantic analysis on the text data to obtain semantic analysis results; The intent depth calculation module is used to fuse the customer emotional state labels with semantic analysis results, input them into a pre-trained intent classification model, and output deep intent labels. The dynamic strategy decision module is used to query a preset dynamic strategy graph based on the customer's emotional state tags and deep intent tags, and select the target dialogue strategy node and corresponding script according to the strategy node jump conditions in the dynamic strategy graph. The intelligent response execution module is used to execute the script content of the selected target dialogue strategy node, generate response voice, and output it to the call link; The closed-loop optimization module is used to record customer emotional state tags, deep intent tags, strategy node jump paths, and response execution logs throughout the entire call process, generate a structured communication report, and optimize the emotional classification model or dynamic strategy graph.

9. A computer device, characterized in that: The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Intelligent customer service system based on natural language processing

    CN120316234A

  • Intelligent agent-based spoken dialogue data processing method

    CN120726996A

  • Conversational AI intelligent outbound service automatic control method in combination with knowledge graph

    CN120849556A

  • Robot dialogue intelligent early warning system based on voice outbound

    CN121012896A

  • Customer service information generation method and system based on multi-modal intention recognition

    CN121303279A