Methods, systems, equipment, and media for intelligent generation of outbound call scripts using LLM.

By generating personalized outbound call scripts through LLM, the problems of consistency and personalization in traditional outbound marketing scripts are solved, enabling accurate analysis of user emotions and intentions, and improving user experience and sales results.

CN120123472BActive Publication Date: 2025-11-14GUANGZHOU JIUSI INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510175610.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-11-14
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

Traditional outbound marketing methods rely on human customer service, which is inefficient and makes it difficult to ensure consistency and personalization of the scripts, thus failing to effectively improve sales results.

Method used

The LLM-based intelligent generation method for outbound call scripts is adopted. By collecting user voice data and text dialogue data in real time, a time-aligned text sequence matrix is ​​generated. Combined with multimodal sentiment analysis and intent recognition models, personalized outbound call scripts are generated.

Benefits of technology

It enables precise analysis of user emotions and intentions, improving user experience and satisfaction, and increasing business conversion rates and market competitiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123472B_ABST
    Figure CN120123472B_ABST
Patent Text Reader

Abstract

This invention relates to the technical field of customer service, and in particular to a method, system, device, and medium for intelligent generation of outbound call scripts using LLM (Limited Language Management). The method includes real-time acquisition of user voice data and text dialogue data; conversion of the user voice data and text dialogue data into a time-aligned text sequence matrix; inputting the time-aligned text sequence matrix into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value; extracting intent keywords based on the time-aligned text sequence matrix; inputting the intent keywords into a preset intent recognition model to output the user's purchase intent probability; and generating personalized outbound call scripts using LLM based on the user's sentiment tendency value and purchase intent probability, and outputting them to a dialogue system. This application achieves highly customized and dynamically adaptable customer service, significantly improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of customer service, and in particular to a method, system, device, and medium for intelligent generation of outbound call scripts using LLM. Background Technology

[0002] In today's highly competitive market environment, businesses are increasingly eager to improve customer service quality and sales conversion rates. Traditional outbound marketing methods often rely on the experience and intuition of human customer service representatives. This approach is not only inefficient but also struggles to ensure consistency and personalization of the scripts, thus limiting the improvement of sales results.

[0003] With the rapid development of artificial intelligence technology, especially the continuous advancement of natural language processing (NLP) and machine learning (ML) technologies, new opportunities for change have been brought to outbound marketing. Through natural language processing technology, the emotions and key information input by users can be analyzed more accurately to understand their true needs and purchasing intentions. Machine learning technology can predict users' purchasing tendencies based on a large amount of historical data, providing strong support for developing more precise communication strategies.

[0004] However, relying solely on NLP and ML technologies is insufficient to achieve highly customized and dynamically adaptable customer service. In practical applications, factors such as the user's communication style, tone, and conversational context significantly influence the choice of response techniques. Therefore, determining the most appropriate response techniques based on the customer's communication style has become a pressing issue. Summary of the Invention

[0005] To achieve highly customized and dynamically adaptable customer service and significantly improve user experience, this application provides a method, system, device, and medium for intelligent generation of outbound call scripts using LLM.

[0006] Firstly, the above-mentioned inventive objective of this application is achieved through the following technical solution:

[0007] A method for intelligently generating outbound call scripts using LLM, the method comprising the following steps:

[0008] Real-time acquisition of user voice data and text dialogue data, and conversion of the user voice data and text dialogue data into a time-aligned text sequence matrix;

[0009] The time-aligned text sequence matrix is ​​input into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value;

[0010] Intent keywords are extracted based on the time-aligned text sequence matrix, and the intent keywords are input into a preset intent recognition model to output the probability of user purchase intent.

[0011] Based on the user's sentiment tendency value and the probability of the user's purchase intention, the LLM is invoked to generate a personalized outbound call script, which is then output to the dialogue system.

[0012] By adopting the above technical solution, user voice data (including intonation, speech rate, pauses, etc.) and text dialogue content are collected in real time during the outbound call script generation process, ensuring the immediacy and accuracy of the data. This provides a reliable foundation for subsequent sentiment analysis and intent recognition. The data is converted into a time-aligned text sequence matrix using speech recognition technology. This time-aligned text sequence matrix ensures that the voice and text data are consistent in time, facilitating subsequent model processing and ensuring accurate matching of semantics and speech rhythm. The obtained time-aligned text sequence matrix is ​​then input into a pre-set multimodal sentiment analysis model. This multimodal sentiment analysis model is used to determine and analyze the user's emotional tendency, comprehensively considering the emotional content in both voice and text data. This system utilizes a multimodal sentiment analysis model to calculate user sentiment inclination values, thereby enabling the analysis of user emotions. By extracting keywords from a time-aligned text sequence matrix, it can more effectively capture key information in user expressions. The extracted intent keywords are then input into a pre-defined intent recognition model, which accurately determines the user's purchase intent and outputs the probability of the user's intended purchase. Combining the user's sentiment inclination value and the purchase intent probability, it generates dialogue strategies that better match user needs and emotions, outputting personalized messages. This achieves precise understanding and efficient visualization of user needs, not only improving user experience and satisfaction but also contributing to increased business conversion rates and market competitiveness.

[0013] In a preferred embodiment, this application can be further configured as follows: the real-time acquisition of user voice data and text dialogue data, and the conversion of the user voice data and text dialogue data into a time-aligned text sequence matrix, specifically includes:

[0014] Mel spectrum feature extraction is performed on the user voice data to generate a spectrogram matrix;

[0015] The text dialogue data is segmented into words to generate a sequence of word vectors;

[0016] The speech spectral features are matched with the text word vector sequence using a timestamp alignment algorithm, and the fused time-aligned text sequence matrix is ​​output.

[0017] By employing the aforementioned technical solution, Mel-spectral feature extraction is performed on the collected user speech data, followed by speech signal processing. This extracts frequency features from the speech data, reflecting the spectral structure of the speech signal, including the intensity distribution of different frequency components and their changes over time. A spectrogram matrix is ​​generated, containing information about the speech signal in both time and frequency dimensions. This provides a rich feature foundation for subsequent text-speech feature alignment and fusion. The corresponding speech features are extracted using the spectrogram matrix and aligned with the text dialogue data, outputting the aligned speech features. This dynamically adjusts the correspondence between text and speech features, making them more semantically consistent and improving the efficiency and effectiveness of feature fusion. The aligned speech features and text dialogue data are then jointly fused to generate a time-aligned text sequence matrix. This time-aligned text sequence matrix not only contains information in the time dimension but also incorporates semantic features, providing strong support for sentiment analysis and intent recognition.

[0018] In a preferred embodiment, this application can be further configured such that: inputting the time-aligned text sequence matrix into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value specifically includes:

[0019] Based on the user's voice data, voice segment features are obtained, wherein the voice segment features include the mean fundamental frequency of the voice segment, the fundamental frequency of the background noise, and the amplitude fluctuation coefficient.

[0020] Based on the features of the speech segment, an emotion intensity calibration formula is constructed, and the emotion intensity in the speech data is output:

[0021] Where μ p The fundamental frequency mean of the speech segment, μ b As the fundamental frequency of the background noise, A(υ) j () represents the amplitude fluctuation coefficient;

[0022] Text keywords are obtained based on the time-aligned text sequence matrix, and word frequency weights and keyword sentiment intensity are obtained based on the text keywords;

[0023] The emotional intensity, word frequency weight, and keyword emotional intensity from the voice data are input into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value:

[0024] Wherein s(t) i ) represents the sentiment intensity of the keyword, ω i The term frequency weight of the keyword, I(υ) j ) represents the emotional intensity in the speech data, and α is the modality fusion coefficient.

[0025] By adopting the above technical solution and combining information from both speech and text modalities for sentiment analysis, the accuracy and robustness of sentiment recognition can be significantly improved. By capturing speech features such as fundamental frequency and amplitude fluctuations, as well as text keywords and their emotional intensity, the multimodal fusion method can more comprehensively understand users' emotional expressions, reduce the impact of poor single-modal data quality on the results, and improve the effect of user sentiment tendency analysis.

[0026] In a preferred embodiment, this application can be further configured as follows: after inputting the emotional intensity, word frequency weight, and keyword emotional intensity from the voice data into a preset multimodal sentiment analysis model to calculate the user's emotional tendency value, the outbound call script intelligent generation method using LLM further includes:

[0027] Based on preset emotional conflict detection rules, it is determined whether the emotional intensity of the keyword is opposite in sign to the emotional intensity in the speech data. If so, modal reweighting is triggered, and the second modal fusion coefficient is output.

[0028] α'=α-γ·|s(t i )-sign(I(υ i ))|, where γ is the weighting coefficient;

[0029] When the second modality fusion coefficient is less than the coefficient threshold, a voice emotion-driven strategy is adopted to generate a second user emotion value.

[0030] By adopting the above technical solution, the analysis of user sentiment tendencies considers both the sentiment in voice data and the sentiment in text data. Significant differences in sentiment expression can occur between the two. Specifically, when the sentiment intensity of a keyword is opposite to the sentiment intensity in the voice data, a modal reweighting mechanism is triggered. By adjusting the modal fusion coefficient, it can more flexibly handle situations of emotional conflict. When the reweighted second modal fusion coefficient is less than the coefficient threshold, the voice sentiment-driven strategy is prioritized. This setting is based on the fact that voice sentiment often more directly and realistically reflects the user's immediate emotional state. This not only improves the accuracy of sentiment analysis and avoids misjudgments of sentiment due to textual ambiguity or misunderstanding, but also helps improve the user experience, ensuring that the system can more accurately understand user emotions and provide users with more considerate and personalized services.

[0031] In a preferred embodiment, this application can be further configured as follows: extracting intent keywords based on the time-aligned text sequence matrix, inputting the intent keywords into a preset intent recognition model, and outputting the probability of user purchase intent, specifically includes:

[0032] Intent keywords are extracted from the time-aligned text sequence matrix to form a dynamic keyword lexicon;

[0033] Based on the dynamic keyword library, keyword group weights are calculated, and keywords with keyword group weights higher than a threshold are input into a preset intent recognition model to output the probability of user purchase intent.

[0034] By adopting the above technical solution and dynamically constructing a keyword library, the core intent expressed by users can be captured in real time, improving the accuracy and timeliness of intent recognition. At the same time, by combining the calculation of keyword group weights, the expression of user intent is further refined, enabling the intent recognition model to understand user purchasing needs more precisely. This not only helps to improve the intelligent recommendation effect of products and provide users with more personalized and demand-matched product recommendations, but also optimizes the user experience and improves user satisfaction and loyalty.

[0035] In a preferred example, this application can be further configured as follows: Based on the user's sentiment tendency value and the user's purchase intent probability, the step of calling LLM to generate an adapted dialogue strategy and outputting personalized scripts specifically includes:

[0036] The emotional level is divided according to the user's emotional tendency value. The user's purchase intention probability and emotional level are input into the preset strategy matching rule table to map out the speech template.

[0037] The LLM is invoked to generate natural language scripts based on the script template. After filtering out illegal content through a compliance filter, personalized scripts are output.

[0038] By adopting the above technical solution, and comprehensively considering users' emotional inclination and purchase intent probability, and inputting them into a preset strategy matching rule table, accurate mapping of dialogue templates is achieved. This not only greatly improves the matching degree between dialogue and users' actual needs, enhancing the personalization and targeting of interactions, but also, through the division of emotional levels, allows response dialogue to more delicately match users' emotional states, effectively improving the user communication experience. Furthermore, by using a large language model (LLM) to generate natural language dialogue based on the selected dialogue template, not only are the forms of dialogue expression enriched, but the natural fluency and logical coherence of the dialogue are also ensured. At the same time, the built-in compliance filter can screen and block illegal content in real time, ensuring that the output personalized dialogue not only complies with business specifications but also respects user feelings, laying a solid foundation for building a safe, healthy, and efficient interactive environment.

[0039] In a preferred embodiment, this application can be further configured as follows: after generating personalized outbound call scripts using LLM based on the user sentiment tendency value and the user purchase intent probability, and outputting them to the dialogue system, the intelligent generation method of outbound call scripts using LLM further includes:

[0040] Real-time monitoring of feedback signals in user conversations, including response time, frequency of affirmative words, and duration of silence;

[0041] Based on the feedback signal, the speech effectiveness score is calculated and the speech score value is output:

[0042]

[0043] The script score is compared with a preset score threshold. When the script score is less than the score threshold, a script switching strategy is generated.

[0044] The intent recognition model parameters are updated based on the aforementioned dialogue switching strategy.

[0045] By adopting the above technical solution, after the personalized script is output to the user's dialog box, the feedback signals in the user's dialogue are monitored in real time. The feedback signals include key indicators such as response time, affirmative word frequency, and silence duration. The effect of the script is evaluated and quantitatively scored in real time, which enhances the sensitivity and responsiveness of the dialogue system. It can also accurately capture subtle changes in the user's dialogue process. When the script score is lower than the preset score threshold, a script switching strategy is automatically generated. In response to the script switching strategy, the parameters of the intent recognition model are automatically updated, which enhances the intelligence of the dialogue and makes it more in line with the user's intent.

[0046] Secondly, the above-mentioned inventive objective of this application is achieved through the following technical solutions:

[0047] A system for intelligently generating outbound call scripts using LLM (Library Management Model), the system comprising:

[0048] The dialogue data acquisition module is used to acquire user voice data and text dialogue data in real time, and convert the user voice data and text dialogue data into a time-aligned text sequence matrix.

[0049] The sentiment analysis module is used to input the time-aligned text sequence matrix into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value.

[0050] The purchase intent analysis module is used to extract intent keywords based on the time-aligned text sequence matrix, input the intent keywords into a preset intent recognition model, and output the probability of the user's purchase intent.

[0051] The script generation module is used to generate personalized outbound call scripts by calling LLM based on the user's sentiment tendency value and the probability of the user's purchase intention, and output them to the dialogue system.

[0052] By adopting the above technical solution, user voice data (including intonation, speech rate, pauses, etc.) and text dialogue content are collected in real time during the outbound call script generation process, ensuring the immediacy and accuracy of the data. This provides a reliable foundation for subsequent sentiment analysis and intent recognition. The data is converted into a time-aligned text sequence matrix using speech recognition technology. This time-aligned text sequence matrix ensures that the voice and text data are consistent in time, facilitating subsequent model processing and ensuring accurate matching of semantics and speech rhythm. The obtained time-aligned text sequence matrix is ​​then input into a pre-set multimodal sentiment analysis model. This multimodal sentiment analysis model is used to determine and analyze the user's emotional tendency, comprehensively considering the emotional content in both voice and text data. This system utilizes a multimodal sentiment analysis model to calculate user sentiment inclination values, thereby enabling the analysis of user emotions. By extracting keywords from a time-aligned text sequence matrix, it can more effectively capture key information in user expressions. The extracted intent keywords are then input into a pre-defined intent recognition model, which accurately determines the user's purchase intent and outputs the probability of the user's intended purchase. Combining the user's sentiment inclination value and the purchase intent probability, it generates dialogue strategies that better match user needs and emotions, outputting personalized messages. This achieves precise understanding and efficient visualization of user needs, not only improving user experience and satisfaction but also contributing to increased business conversion rates and market competitiveness.

[0053] Thirdly, the above-mentioned objectives of this application are achieved through the following technical solutions:

[0054] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described intelligent generation method for outbound call scripts using LLM.

[0055] Fourthly, the above-mentioned objectives of this application are achieved through the following technical solutions:

[0056] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described intelligent generation method for outbound call scripts using LLM.

[0057] In summary, this application includes at least one of the following beneficial technical effects:

[0058] 1. During the outbound call script generation process, user voice data (including intonation, speech rate, pauses, etc.) and text dialogue content are collected in real time, ensuring the immediacy and accuracy of the data. This provides a reliable foundation for subsequent sentiment analysis and intent recognition. The data is converted into a time-aligned text sequence matrix using speech recognition technology. This time-aligned text sequence matrix ensures that the voice and text data are consistent in time, facilitating subsequent model processing and ensuring accurate matching of semantics and speech rhythm. The resulting time-aligned text sequence matrix is ​​then input into a pre-defined multimodal sentiment analysis model. This model is used to determine and analyze the user's emotional tendencies, comprehensively considering the emotional information in both voice and text data. By utilizing a multimodal sentiment analysis model to calculate user sentiment inclination values, the system can analyze user emotions. By extracting keywords from a time-aligned text sequence matrix, intent keywords can be extracted, which can more effectively capture key information in user expressions. The extracted intent keywords are input into a preset intent recognition model, which accurately determines the user's purchase intent and outputs the user's purchase probability. Combining the user's sentiment inclination value and purchase intent probability, a dialogue strategy that better matches the user's needs and emotions is generated, and personalized messages are output. This achieves accurate grasp and efficient display of user needs, which not only improves user experience and satisfaction but also helps to improve business conversion rate and market competitiveness.

[0059] 2. Combining information from both speech and text modalities for sentiment analysis can significantly improve the accuracy and robustness of sentiment recognition. By capturing speech features such as fundamental frequency and amplitude fluctuations, as well as text keywords and their emotional intensity, the multimodal fusion method can more comprehensively understand users' emotional expressions, reduce the impact of poor single-modal data quality on the results, and improve the effectiveness of user sentiment analysis.

[0060] 3. In the process of analyzing user sentiment, both the sentiment in voice data and the sentiment in text data are considered. There will be significant differences in the emotional expression between the two. That is, when the sentiment intensity of keywords is opposite to the sentiment intensity in voice data, a modality reweighting mechanism is triggered. By adjusting the modality fusion coefficient, it can more flexibly deal with emotional conflicts. When the reweighted second modality fusion coefficient is less than the coefficient threshold, the voice sentiment-driven strategy is given priority. This setting is based on the fact that voice sentiment can often more directly and realistically reflect the user's real-time emotional state. This not only improves the accuracy of sentiment analysis and avoids misjudgment of sentiment due to text ambiguity or misunderstanding, but also helps to improve the user experience, ensure that the system can more accurately understand user emotions, and provide users with more considerate and personalized services.

[0061] 4. By dynamically constructing a keyword lexicon, the core intent expressed by users can be captured in real time, improving the accuracy and timeliness of intent recognition. At the same time, by combining the calculation of keyword group weights, the expression of user intent is further refined, enabling the intent recognition model to understand user purchasing needs more precisely. This not only helps to improve the intelligent recommendation effect of products and provide users with more personalized and demand-matched product recommendations, but also optimizes the user experience and improves user satisfaction and loyalty. Attached Figure Description

[0062] Figure 1 This is a flowchart of an embodiment of the outbound call script intelligent generation method using LLM in this application;

[0063] Figure 2 This is a flowchart illustrating the implementation of step S10 in an embodiment of the outbound call script intelligent generation method using LLM in this application.

[0064] Figure 3 This is a flowchart illustrating the implementation of step S20 in an embodiment of the outbound call script intelligent generation method using LLM in this application.

[0065] Figure 4 This is another implementation flowchart of the outbound call script intelligent generation method using LLM in one embodiment of this application;

[0066] Figure 5 This is a flowchart illustrating the implementation of step S30 in an embodiment of the outbound call script intelligent generation method using LLM in this application.

[0067] Figure 6 This is a flowchart illustrating the implementation of step S40 in an embodiment of the outbound call script intelligent generation method using LLM in this application.

[0068] Figure 7 This is another implementation flowchart of the outbound call script intelligent generation method using LLM in one embodiment of this application;

[0069] Figure 8 This is a principle block diagram of an outbound call script intelligent generation system using LLM in one embodiment of this application;

[0070] Figure 9 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0071] The present application will be further described in detail below with reference to the accompanying drawings.

[0072] In one embodiment, such as Figure 1 As shown, this application discloses a method for intelligent generation of outbound call scripts using LLM, which specifically includes the following steps:

[0073] S10: Real-time acquisition of user voice data and text dialogue data, and conversion of the user voice data and text dialogue data into a time-aligned text sequence matrix.

[0074] Specifically, during the outbound call script generation process, user voice data (including intonation, speech rate, pauses, etc.) and text dialogue content are collected in real time, ensuring the immediacy and accuracy of the data. This provides a reliable foundation for subsequent sentiment analysis and intent recognition. The data is then converted into a time-aligned text sequence matrix using speech recognition technology. This time-aligned text sequence matrix ensures that the voice and text data are consistent in time, facilitating subsequent model processing and ensuring accurate matching of semantics and speech rhythm.

[0075] S20: Input the time-aligned text sequence matrix into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value.

[0076] Specifically, a multimodal sentiment analysis model is a model used to determine and analyze a user's sentiment tendency. The obtained time-aligned text sequence matrix is ​​input into the multimodal sentiment analysis model, and the model uses voice sentiment features and text keyword sentiment values ​​to calculate the user's sentiment tendency value, thereby realizing the function of analyzing user sentiment.

[0077] S30: Extract intent keywords based on the time-aligned text sequence matrix, input the intent keywords into a preset intent recognition model, and output the probability of user purchase intent.

[0078] Specifically, keyword extraction is performed on the time-aligned text sequence matrix to extract intent keywords, which can more effectively capture key information in user expression. The extracted intent keywords are then input into a preset intent recognition model, which accurately determines the user's purchase intent and outputs the probability of the user's intended purchase.

[0079] S40: Based on the user's sentiment tendency value and the probability of the user's purchase intention, call LLM to generate a personalized outbound call script and output it to the dialogue system.

[0080] Specifically, by combining user sentiment scores and purchase intent probabilities, LLM adjusts the tone and wording of the script to generate dialogue strategies that better suit user needs and emotions. For example, when a user exhibits a positive sentiment, the script will be more enthusiastic, emphasizing the product's advantages and benefits; while when a user exhibits a negative sentiment, the script will be more gentle and patient, attempting to alleviate the user's doubts and dissatisfaction. Simultaneously, LLM also adjusts the content and focus of the script based on the user's purchase intent probability. For users with strong purchase intent, the script will be more direct and explicit, highlighting product purchase information and promotional activities; while for users with weak purchase intent, the script will be more tactful and guiding, attempting to stimulate the user's interest and desire to buy.

[0081] In this embodiment, during the outbound call script generation process, user voice data (including intonation, speech rate, pauses, etc.) and text dialogue content are collected in real time, ensuring the immediacy and accuracy of the data. This provides a reliable foundation for subsequent sentiment analysis and intent recognition. The data is then converted into a time-aligned text sequence matrix using speech recognition technology. This time-aligned text sequence matrix ensures that the voice and text data are consistent in time, facilitating subsequent model processing and ensuring precise matching of semantics and speech rhythm. The resulting time-aligned text sequence matrix is ​​then input into a preset multimodal sentiment analysis model. This multimodal sentiment analysis model is used to determine and analyze the user's emotional tendencies, comprehensively considering the emotional information in both voice and text data. This system utilizes a multimodal sentiment analysis model to calculate user sentiment inclination values, thereby enabling the analysis of user emotions. By extracting keywords from a time-aligned text sequence matrix, it can more effectively capture key information in user expressions. The extracted intent keywords are then input into a pre-defined intent recognition model, which accurately determines the user's purchase intent and outputs the probability of the user's intended purchase. Combining the user's sentiment inclination value and the purchase intent probability, it generates dialogue strategies that better match user needs and emotions, outputting personalized messages. This achieves precise understanding and efficient visualization of user needs, not only improving user experience and satisfaction but also contributing to increased business conversion rates and market competitiveness.

[0082] In one embodiment, such as Figure 2 As shown, in step S10, user voice data and text dialogue data are collected in real time, and then converted into a time-aligned text sequence matrix based on the user voice data and text dialogue data. Specifically, this includes:

[0083] S11: Extract Mel spectrum features from the user's voice data to generate a spectrogram matrix.

[0084] Specifically, Mel spectrum feature extraction is performed on the collected user speech data, and speech signal processing is carried out to extract the frequency features in the speech data, which can reflect the spectral structure of the speech signal, including the intensity distribution of different frequency components and their changes over time, and generate a spectrogram matrix. The spectrogram matrix contains information about the speech signal in both time and frequency dimensions.

[0085] S12: Generate a word vector sequence after segmenting the text dialogue data.

[0086] Specifically, for text dialogue data, word segmentation technology is used to break it down into independent lexical units, and word embedding models are used to convert these lexical units into word vector sequences in a high-dimensional space, which facilitates subsequent numerical calculations and analysis.

[0087] S13: Match the speech spectrum features with the text word vector sequence using a timestamp alignment algorithm, and output the fused time-aligned text sequence matrix.

[0088] Specifically, the timestamp alignment algorithm is used to precisely match the speech spectral features with the text word vector sequence in time based on the timestamps in the speech and text data. This ensures that the speech features at each moment strictly correspond to the corresponding text content, thereby generating a fused time-aligned text sequence matrix that contains both speech information and text semantics, providing reliable basic data for subsequent multimodal analysis and intent analysis.

[0089] In one embodiment, such as Figure 3 As shown, in step S20, the time-aligned text sequence matrix is ​​input into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value, specifically including:

[0090] S21: Obtain speech segment features based on the user's speech data, wherein the speech segment features include the speech segment's fundamental frequency mean, background noise fundamental frequency, and amplitude fluctuation coefficient.

[0091] Specifically, advanced speech processing technology is used to analyze user speech data, accurately extract the fundamental frequency mean of speech segments to reflect the speaker's average pitch level, and separate and calculate the fundamental frequency characteristics of background noise to measure the degree of interference of environmental noise on speech signals. In addition, the amplitude stability of speech segments is quantitatively evaluated by calculating the fluctuation coefficient of amplitude over time.

[0092] S22: Construct an emotion intensity calibration formula based on the features of the speech segment, and output the emotion intensity in the speech data.

[0093] Specifically, the emotional intensity in the speech data is calculated using the following formula:

[0094]

[0095] Where μ p The fundamental frequency mean of the speech segment, μ b As the fundamental frequency of the background noise, A(υ) j ) represents the amplitude fluctuation coefficient.

[0096] S23: Obtain text keywords based on the time-aligned text sequence matrix, and obtain word frequency weights and keyword sentiment intensity based on the text keywords.

[0097] Specifically, text fragments are extracted from a time-aligned text sequence matrix. The TF-IDF algorithm combined with text mining techniques is used to identify text keywords with significant distinguishability. The frequency of each keyword in the text is counted, and word frequency weights are calculated based on the frequency distribution to quantify the importance of the keywords. At the same time, a sentiment lexicon is used to analyze the sentiment of the keywords and assign each keyword a corresponding sentiment intensity value, thereby comprehensively revealing the key information, importance, and emotional tone of the text content.

[0098] S24: Input the emotional intensity, word frequency weight and keyword emotional intensity in the voice data into a preset multimodal sentiment analysis model to calculate the user's emotional tendency value.

[0099] Specifically, the user's sentiment tendency value is calculated based on the following formula, combined with the emotional intensity, word frequency weight, and keyword sentiment intensity in the voice data:

[0100] Wherein s(t) i ) represents the sentiment intensity of the keyword, ω i For keyword frequency weights, I(υ) j ) represents the emotional intensity in the speech data, and α is the modality fusion coefficient (0.6≤α≤0.8).

[0101] In one embodiment, such as Figure 4 As shown, after step S24, that is, after inputting the emotional intensity, word frequency weight, and keyword emotional intensity from the voice data into a preset multimodal sentiment analysis model to calculate the user's emotional tendency value, the LLM-based outbound call script intelligent generation method further includes:

[0102] S25: Based on the preset emotional conflict detection rules, determine whether the emotional intensity of the keyword is opposite in sign to the emotional intensity in the voice data. If so, trigger modal reweighting and output the second modal fusion coefficient.

[0103] Specifically, in the process of analyzing user sentiment, the sentiment in the voice data and the sentiment in the text data are considered. There will be significant differences in the emotional expression between the two. That is, when the sentiment intensity of the keywords is opposite in sign to the sentiment intensity in the voice data, the modality reweighting mechanism is triggered, and the modality fusion coefficients are adjusted using the following formula:

[0104] α'=α-γ·|s(t i )-sign(I(υ i ))|, where γ is the weighting coefficient.

[0105] S26: When the second modality fusion coefficient is less than the coefficient threshold, the voice emotion-driven strategy is adopted to generate the second user emotion value.

[0106] Specifically, when the reweighted second modality fusion coefficient is less than the coefficient threshold, i.e., α′<0.5, based on the fact that voice emotion can often more directly and realistically reflect the user's real-time emotional state, the voice emotion-dominated strategy is prioritized, and the second modality fusion coefficient is updated into the calculation formula of step S24, outputting a new user emotional tendency value. This avoids emotional misjudgment caused by text ambiguity or misunderstanding, helps improve the user experience, and ensures that the system can more accurately understand the user's emotions.

[0107] In one embodiment, such as Figure 5 As shown, in step S30, which involves extracting intent keywords based on the time-aligned text sequence matrix, inputting the intent keywords into a preset intent recognition model, and outputting the probability of the user's purchase intent, specifically includes:

[0108] S31: Extract intent keywords from the time-aligned text sequence matrix to form a dynamic keyword lexicon.

[0109] Specifically, a dynamic keyword library is built, and domain keywords are dynamically loaded according to industry type to capture the core intent in user expressions in real time, thereby improving the accuracy and timeliness of intent recognition.

[0110] S32: Calculate keyword group weights based on the dynamic keyword library, input keywords with keyword group weights higher than the threshold into the preset intent recognition model, and output the probability of user purchase intent.

[0111] Specifically, based on the weighting formula, the weight of the keyword combination is calculated, further refining the user's intent expression and enabling the intent recognition model to understand the user's purchase needs more precisely.

[0112]

[0113] Where TF(k) is the term frequency and IDF(k) is the reverse document frequency;

[0114] Further, input the keywords with weights higher than the threshold into a preset intent recognition model to output the probability of the user's purchase intent:

[0115] P intent = σ(W h ·h t + b), where h t is the sequential hidden state, and W h is the keyword with a weight higher than the threshold; it can provide users with more personalized and demand-compliant product recommendations, optimize the user experience, and improve user satisfaction and loyalty.

[0116] In one embodiment, as Figure 6 shown, in step S40, according to the user's emotional tendency value and the probability of the user's purchase intent, call the LLM to generate an adapted dialogue strategy and output personalized words, specifically including:

[0117] S41: Divide the emotional level according to the user's emotional tendency value, input the probability of the user's purchase intent and the emotional level into a preset strategy matching rule table, and map out a word template.

[0118] Specifically, divide the emotional level according to the emotional tendency value: positive (E≥0.5E≥0.5): adopt a promotion strategy; neutral (-0.5 < E < 0.5-0.5 < E < 0.5): adopt a guiding strategy; negative (E≤-0.5E≤-0.5): adopt a soothing strategy; input it into a preset strategy matching rule table, map out a word template, and through the division of the emotional level, the response words can more delicately fit the user's emotional state, effectively improving the user communication experience.

[0119] S42: Call the LLM to generate natural language words based on the word template, and output personalized words after shielding illegal content through a compliance filter.

[0120] Specifically, use a large language model (LLM) to generate natural language words based on the selected word template, which not only enriches the expression form of the words, but also ensures the natural fluency and logical coherence of the words.

[0121] Further, the built-in compliance filter can screen and shield illegal content in real time. There is a sensitive word library and compliance regular expressions set during the compliance filtering period. According to the words generated by the LLM, calculate the word risk value, and judge whether the generated words meet the regulations according to the word risk value, ensuring that the output personalized words comply with business norms and respect the user's feelings, laying a solid foundation for building a safe, healthy and efficient interaction environment, such as:

[0122] When R is less than 1, it means that the generated natural language conforms to the regulations.

[0123] In one embodiment, such as Figure 7 As shown, after step S40, that is, after generating personalized outbound call scripts using LLM based on the user sentiment tendency value and the user purchase intention probability, and outputting them to the dialogue system, the intelligent generation method of outbound call scripts using LLM further includes:

[0124] S50: Monitors feedback signals in user conversations in real time, including response time, affirmative word frequency, and silence duration.

[0125] S60: Calculate the speech effect score based on the feedback signal, output the speech score value, compare the speech score value with a preset score threshold, and generate a speech switching strategy when the speech score value is less than the score threshold.

[0126] Specifically, after the personalized script is output to the user's dialog box, the system monitors the user's feedback signals in real time. These feedback signals include key indicators such as response time, frequency of affirmative words, and duration of silence. The effectiveness of the script is then evaluated and quantitatively scored in real time.

[0127] When the script score is lower than the preset score threshold, a script switching strategy is automatically generated. Generally, the score threshold is 0.3, that is, when S is less than 0.3, the script strategy is automatically switched.

[0128] S70: Update the intent recognition model parameters based on the aforementioned dialogue switching strategy.

[0129] Specifically, it responds to dialogue switching strategies, automatically updates the parameters of the intent recognition model, enhances the intelligence of the dialogue, and makes it more aligned with user intent.

[0130] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0131] In one embodiment, an intelligent outbound call script generation system using LLM is provided, which corresponds one-to-one with the intelligent outbound call script generation method using LLM described in the above embodiments. For example... Figure 8 As shown, this LLM-based outbound call script intelligent generation system includes a dialogue data acquisition module, a sentiment analysis module, a purchase intent analysis module, and a script generation module. Detailed descriptions of each functional module are as follows:

[0132] The dialogue data acquisition module is used to acquire user voice data and text dialogue data in real time, and convert the user voice data and text dialogue data into a time-aligned text sequence matrix.

[0133] The sentiment analysis module is used to input the time-aligned text sequence matrix into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value.

[0134] The purchase intent analysis module is used to extract intent keywords based on the time-aligned text sequence matrix, input the intent keywords into a preset intent recognition model, and output the probability of the user's purchase intent.

[0135] The script generation module is used to generate personalized outbound call scripts by calling LLM based on the user's sentiment tendency value and the probability of the user's purchase intention, and output them to the dialogue system.

[0136] Specific limitations regarding the intelligent outbound call script generation system using LLM can be found in the limitations of the intelligent outbound call script generation method using LLM described above, and will not be repeated here. Each module in the aforementioned intelligent outbound call script generation system using LLM can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the electronic device, or stored in the memory of the electronic device as software, so that the processor can call and execute the corresponding operations of each module.

[0137] In one embodiment, an electronic device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9 As shown, the electronic device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores the database. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements an intelligent outbound call script generation method using LLM (Limited Language Management).

[0138] In one embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0139] Real-time acquisition of user voice data and text dialogue data, and conversion of the user voice data and text dialogue data into a time-aligned text sequence matrix;

[0140] The time-aligned text sequence matrix is ​​input into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value;

[0141] Intent keywords are extracted based on the time-aligned text sequence matrix, and the intent keywords are input into a preset intent recognition model to output the probability of user purchase intent.

[0142] Based on the user's sentiment tendency value and the probability of the user's purchase intention, the LLM is invoked to generate a personalized outbound call script, which is then output to the dialogue system.

[0143] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0144] Real-time acquisition of user voice data and text dialogue data, and conversion of the user voice data and text dialogue data into a time-aligned text sequence matrix;

[0145] The time-aligned text sequence matrix is ​​input into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value;

[0146] Intent keywords are extracted based on the time-aligned text sequence matrix, and the intent keywords are input into a preset intent recognition model to output the probability of user purchase intent.

[0147] Based on the user's sentiment tendency value and the probability of the user's purchase intention, the LLM is invoked to generate a personalized outbound call script, which is then output to the dialogue system.

[0148] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0149] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0150] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for intelligent generation of outbound call scripts using LLM, characterized in that, The method for intelligently generating outbound call scripts using LLM includes the following steps: Real-time acquisition of user voice data and text dialogue data, and conversion of the user voice data and text dialogue data into a time-aligned text sequence matrix; The time-aligned text sequence matrix is ​​input into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value; Intent keywords are extracted based on the time-aligned text sequence matrix, and the intent keywords are input into a preset intent recognition model to output the probability of user purchase intent. Based on the user's sentiment tendency value and the probability of the user's purchase intention, the LLM is invoked to generate a personalized outbound call script, which is then output to the dialogue system. After generating personalized outbound call scripts using LLM based on the user sentiment tendency value and the probability of user purchase intent, and outputting them to the dialogue system, the intelligent generation method of outbound call scripts using LLM further includes: Real-time monitoring of feedback signals in user conversations, including response time, frequency of affirmative words, and duration of silence; The speech effectiveness score is calculated based on the feedback signal, and the speech score value is output. The script score is compared with a preset score threshold. When the script score is less than the score threshold, a script switching strategy is generated. Update the intent recognition model parameters based on the aforementioned dialogue switching strategy; The step of generating an appropriate dialogue strategy based on the user's sentiment index and purchase intent probability, and outputting personalized scripts, specifically includes: The emotional level is divided according to the user's emotional tendency value. The user's purchase intention probability and emotional level are input into the preset strategy matching rule table to map out the speech template. The LLM is invoked to generate natural language scripts based on the script template. After filtering out illegal content through a compliance filter, personalized scripts are output.

2. The method for intelligent generation of outbound call scripts using LLM according to claim 1, characterized in that, The real-time acquisition of user voice data and text dialogue data, and the conversion of the user voice data and text dialogue data into a time-aligned text sequence matrix, specifically includes: Mel spectrum feature extraction is performed on the user voice data to generate a spectrogram matrix; The text dialogue data is segmented into words to generate a sequence of word vectors; The speech spectral features are matched with the text word vector sequence using a timestamp alignment algorithm, and the fused time-aligned text sequence matrix is ​​output.

3. The method for intelligent generation of outbound call scripts using LLM according to claim 1, characterized in that, The step of inputting the time-aligned text sequence matrix into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value specifically includes: Based on the user's voice data, voice segment features are obtained, wherein the voice segment features include the mean fundamental frequency of the voice segment, the fundamental frequency of the background noise, and the amplitude fluctuation coefficient. Based on the features of the speech segment, an emotion intensity calibration formula is constructed, and the emotion intensity in the speech data is output: Where, μ p The fundamental frequency mean of the speech segment, μ b As the fundamental frequency of the background noise, A(v j () represents the amplitude fluctuation coefficient; Text keywords are obtained based on the time-aligned text sequence matrix, word frequency weights are obtained based on the text keywords, and sentiment tendency analysis is performed on the text keywords using a sentiment dictionary to obtain the sentiment intensity of the keywords; The emotional intensity, word frequency weight, and keyword emotional intensity from the voice data are input into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value: Wherein s(t) i ) represents the sentiment intensity of the keyword, ω i For keyword frequency weights, I(v) j ) represents the emotional intensity in the speech data, and α is the modality fusion coefficient.

4. The method for intelligent generation of outbound call scripts using LLM according to claim 3, characterized in that, After inputting the emotional intensity, word frequency weight, and keyword emotional intensity from the voice data into a preset multimodal sentiment analysis model to calculate the user's emotional tendency value, the LLM-based intelligent generation method for outbound call scripts further includes: Based on preset emotional conflict detection rules, it is determined whether the emotional intensity of the keywords matches the emotional intensity in the voice data. If the sign is negative, then modal reweighting is triggered, and the second modal fusion coefficients are output: a=α-γ·s(t i )-sign(I(v i )) Where γ is the weighting coefficient; When the second modality fusion coefficient is less than the coefficient threshold, the voice emotion-driven strategy is adopted, the second modality fusion coefficient is updated into the formula used to calculate the user emotion tendency value, and the new user emotion tendency value is output as the second user emotion value.

5. The method for intelligent generation of outbound call scripts using LLM according to claim 1, characterized in that, The step of extracting intent keywords based on the time-aligned text sequence matrix, inputting the intent keywords into a preset intent recognition model, and outputting the probability of user purchase intent specifically includes: Intent keywords are extracted from the time-aligned text sequence matrix to form a dynamic keyword lexicon; The keyword group weights are calculated based on the aforementioned dynamic keyword database, specifically by using the TF-IDF algorithm. The calculation formula is as follows: Where TF(k) is the term frequency and IDF(k) is the reverse document frequency; The keyword group weights are input into a preset intent recognition model, and the probability of the user's purchase intent is output.

6. A system for intelligent generation of outbound call scripts using LLM, characterized in that, The outbound call script intelligent generation system using LLM includes: The dialogue data acquisition module is used to acquire user voice data and text dialogue data in real time, and convert the user voice data and text dialogue data into a time-aligned text sequence matrix. The sentiment analysis module is used to input the time-aligned text sequence matrix into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value. The purchase intent analysis module is used to extract intent keywords based on the time-aligned text sequence matrix, input the intent keywords into a preset intent recognition model, and output the probability of the user's purchase intent. The script generation module is used to generate personalized outbound call scripts by calling LLM based on the user's sentiment tendency value and the probability of the user's purchase intention, and output them to the dialogue system; After generating personalized outbound call scripts using LLM based on the user sentiment tendency value and the probability of user purchase intent, and outputting them to the dialogue system, the intelligent generation method of outbound call scripts using LLM further includes: Real-time monitoring of feedback signals in user conversations, including response time, frequency of affirmative words, and duration of silence; The speech effect score is calculated based on the feedback signal, and the speech score value is output. The speech score value is compared with a preset score threshold. When the speech score value is less than the score threshold, a speech switching strategy is generated. Update the intent recognition model parameters based on the aforementioned dialogue switching strategy; The step of generating an appropriate dialogue strategy based on the user's sentiment index and purchase intent probability, and outputting personalized scripts, specifically includes: The emotional level is divided according to the user's emotional tendency value. The user's purchase intention probability and emotional level are input into the preset strategy matching rule table to map out the speech template. The LLM is invoked to generate natural language scripts based on the script template. After filtering out illegal content through a compliance filter, personalized scripts are output.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the outbound call script intelligent generation method using LLM as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the outbound call script intelligent generation method using LLM as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Voice conversion method and device, electronic equipment and storage medium

    CN116312617A

  • Management method, system and equipment integrating large language model and voice recognition

    CN118350859A