Method, system and device for intelligently generating outbound call skill by using LLM (Logical Link Model) and medium
By applying LLM to generate personalized out-of-call speech in out-of-call marketing, the problems of inefficiency and consistency of traditional out-of-call marketing are solved, efficient and personalized customer service is achieved, and user experience and business conversion rate are improved.
Patent Information
- Application Number
- CN202510175610.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The traditional outbound marketing method relies on manual customer service, which is inefficient and difficult to ensure the consistency and personalization of the speech, limiting the improvement of sales results.
LLM (large language model) is used to generate personalized out-of-call speech, and by collecting user voice data and text dialogue data in real time, it is converted into a time-aligned text sequence matrix, and combining multi-modal sentiment analysis and intention recognition model, personalized speech is generated.
It realizes highly personalized and dynamic adaptable customer service, improves user experience and satisfaction, and improves business conversion rate and market competitiveness.
Smart Images

Figure CN120123472A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of customer service, and in particular, to a method, system, device, and medium for intelligently generating outbound call scripts using an LLM. Background Art
[0002] In today's highly competitive market environment, enterprises are increasingly eager to improve the quality of customer service and the sales conversion rate. Traditional outbound marketing methods often rely on the experience and intuition of human customer service representatives. This approach is not only inefficient but also difficult to ensure the consistency and personalization of the scripts, thus limiting the improvement of sales effectiveness.
[0003] With the rapid development of artificial intelligence technology, especially the continuous progress of natural language processing (NLP) and machine learning technology (ML), new opportunities for transformation have been brought to outbound marketing. Through natural language processing technology, the emotions and key information input by users can be parsed more accurately, and the true needs and purchase intentions of users can be understood. Machine learning technology can predict users' purchase tendencies based on a large amount of historical data, providing strong support for formulating more accurate script strategies.
[0004] However, relying solely on NLP and ML technologies is still not sufficient to achieve highly customized and dynamically adaptable customer service. In practical applications, factors such as the user's communication style, tone, and conversation context will have an important impact on the selection of scripts. Therefore, how to determine the most appropriate response script skills according to the customer's communication style has become an urgent problem to be solved. Summary of the Invention
[0005] In order to achieve highly customized and dynamically adaptable customer service and significantly improve the user experience, the present application provides a method, system, device, and medium for intelligently generating outbound call scripts using an LLM.
[0006] In a first aspect, the above-mentioned inventive object of the present application is achieved through the following technical solutions:
[0007] A method for intelligently generating outbound call scripts using an LLM, the method for intelligently generating outbound call scripts using an LLM includes the steps of:
[0008] Real-time collecting user voice data and text conversation data, and converting the user voice data and text conversation data into a temporally aligned text sequence matrix;
[0009] Inputting the temporally aligned text sequence matrix into a preset multi-modal sentiment analysis model to calculate the user sentiment tendency value;
[0010] Extract intent keywords based on the temporally aligned text sequence matrix, input the intent keywords into a preset intent recognition model, and output the probability of the user's purchase intent;
[0011] According to the user's emotional tendency value and the probability of the user's purchase intent, call the LLM to generate personalized outbound call scripts and output them to the dialogue system.
[0012] By adopting the above technical solution, during the generation of outbound call scripts, real-time collection of user voice data (including intonation, speech rate, pauses, etc.) and text conversation content is carried out, ensuring the immediacy and accuracy of the data, providing a reliable basis for subsequent sentiment analysis and intent recognition. The voice data is converted into a temporally aligned text sequence matrix through speech recognition technology. The temporally aligned text sequence matrix enables the voice and text data to be consistent in time, facilitating subsequent model processing and ensuring an accurate match between semantics and speech rhythm. The obtained temporally aligned text sequence matrix is input into a preset multi-modal sentiment analysis model. The multi-modal sentiment analysis model is a model used to judge and analyze the user's emotional tendency, which can comprehensively consider the emotional information in voice and text data. Using the multi-modal sentiment analysis model to calculate the user's emotional tendency value, thereby realizing the function of analyzing the user's emotions. By extracting keywords from the temporally aligned text sequence matrix, intent keywords are extracted, which can more effectively capture the key information in the user's expression. The extracted intent keywords are input into a preset intent recognition model, and the intent recognition model is used to accurately judge the user's purchase intent and output the probability of the user's intended purchase. Combining the user's emotional tendency value and the probability of purchase intent, a dialogue strategy that better fits the user's needs and emotions is generated, and personalized scripts are output, realizing the accurate grasp and efficient manifestation of the user's needs. This not only improves the user experience and satisfaction but also helps to increase the business conversion rate and market competitiveness.
[0013] In a preferred example of the present application, it can be further configured as follows: The real-time collection of user voice data and text conversation data, and the conversion into a temporally aligned text sequence matrix based on the user voice data and text conversation data specifically includes:
[0014] Extract Mel-spectrum features from the user voice data to generate a spectrogram matrix;
[0015] After segmenting the text conversation data, generate a sequence of word vectors;
[0016] Match the voice spectrogram features and the text word vector sequence through a timestamp alignment algorithm, and output the fused temporally aligned text sequence matrix.
[0017] By adopting the above technical solution, Mel spectrum features are extracted from the collected user speech data, speech signal processing is performed, and the frequency features in the speech data are extracted, which can reflect the spectral structure of the speech signal, including the intensity distribution of different frequency components and the change over time, generating a spectrogram matrix. The spectrogram matrix contains information about the speech signal in both the time and frequency dimensions, providing a rich feature basis for subsequent text and speech feature alignment and fusion. The corresponding speech features are extracted using the spectrogram matrix, the speech features in the spectrogram matrix are aligned with the text dialogue data, and the aligned speech features are output, which can dynamically adjust the corresponding relationship between the text and speech features, making the two more consistent semantically and improving the efficiency and effect of feature fusion. The aligned speech features and text dialogue data are used for joint fusion to generate a temporally aligned text sequence matrix. The temporally aligned text sequence matrix not only contains information in the time dimension but also integrates semantic features, providing strong support for sentiment analysis and intent recognition.
[0018] In a preferred example of the present application, it can be further configured that: inputting the temporally aligned text sequence matrix into a preset multi-modal sentiment analysis model to calculate the user sentiment tendency value, specifically including:
[0019] Obtaining speech segment features based on the user speech data, where the speech segment features include the average fundamental frequency of the speech segment, the fundamental frequency of the background noise, and the amplitude fluctuation coefficient;
[0020] Constructing an emotion intensity calibration formula according to the speech segment features and outputting the emotion intensity in the speech data:
[0021] where μ p is the average fundamental frequency of the speech segment, μ b is the fundamental frequency of the background noise, A(υ j ) is the amplitude fluctuation coefficient;
[0022] Obtaining text keywords based on the temporally aligned text sequence matrix, and obtaining the word frequency weight and keyword emotion intensity according to the text keywords;
[0023] Inputting the emotion intensity, word frequency weight, and keyword emotion intensity in the speech data into a preset multi-modal sentiment analysis model to calculate the user sentiment tendency value:
[0024] where, s(t i ) is the keyword emotion intensity, ω i is the word frequency weight of the keyword, I(υ j ) is the emotion intensity in the speech data, and α is the modal fusion coefficient.
[0025] By adopting the above technical solution and conducting sentiment analysis by combining information in two modalities, speech and text, the accuracy of sentiment recognition can be significantly improved and the robustness can be enhanced. By capturing speech features such as fundamental frequency, amplitude fluctuation, etc., as well as text keywords and their sentiment intensities, the multimodal fusion method can more comprehensively understand the user's sentiment expression, reduce the impact of poor quality of single-modal data on the results, and improve the effect of user sentiment tendency analysis.
[0026] In a preferred example, this application can be further configured as follows: after inputting the sentiment intensity, word frequency weight, and keyword sentiment intensity in the speech data into a preset multimodal sentiment analysis model and calculating the user sentiment tendency value, the outbound call script intelligent generation method using the LLM further includes:
[0027] Based on a preset sentiment conflict detection rule, determine whether the keyword sentiment intensity and the sentiment intensity in the speech data have opposite signs. If so, trigger modal reweighting and output a second modal fusion coefficient:
[0028] α' = α - γ·|s(t i ) - sign(I(υ i ))|, where γ is a weight coefficient;
[0029] When the second modal fusion coefficient is less than the coefficient threshold, adopt a speech sentiment dominant strategy to generate a second user sentiment value.
[0030] By adopting the above technical solution, during the process of analyzing the user's sentiment tendency, the sentiment in the speech data and the sentiment in the text data are considered. There may be obvious differences in sentiment expression between the two. That is, when the keyword sentiment intensity and the sentiment intensity in the speech data have opposite signs, at this time, trigger the modal reweighting mechanism. By adjusting the modal fusion coefficient, it can more flexibly handle the sentiment conflict situation. When the reweighted second modal fusion coefficient is less than the coefficient threshold, preferentially adopt the speech sentiment dominant strategy. This setting is based on the fact that speech sentiment can often more directly and truly reflect the user's immediate emotional state. It can not only improve the accuracy of sentiment analysis, avoid sentiment misjudgment caused by text ambiguity or misunderstanding, help improve the user experience, ensure that the system can more accurately understand the user's sentiment, and provide more considerate and personalized services for users.
[0031] In a preferred example, this application can be further configured as follows: extracting intent keywords based on the temporally aligned text sequence matrix, and inputting the intent keywords into a preset intent recognition model to output the probability of the user's purchase intent, specifically including:
[0032] Extract intent keywords according to the temporally aligned text sequence matrix to form a dynamic keyword library;
[0033] Calculate the keyword group weights based on the dynamic keyword library, input the keywords with keyword group weights higher than the threshold into a preset intention recognition model, and output the probability of the user's purchase intention.
[0034] By adopting the above technical solution, by dynamically constructing a keyword library, the core intention in the user's expression can be captured in real time, improving the accuracy and timeliness of intention recognition. At the same time, combined with the calculation of keyword group weights, the expression of the user's intention is further refined, enabling the intention recognition model to more precisely understand the user's purchase needs. This not only helps to improve the intelligent recommendation effect of goods, providing more personalized and demand-compliant product recommendations for users, but also can optimize the user experience, enhancing user satisfaction and loyalty.
[0035] In a preferred example, this application can be further configured as: according to the user's emotional tendency value and the probability of the user's purchase intention, call the LLM to generate an adapted dialogue strategy and output personalized conversation words, specifically including:
[0036] Divide the emotional levels according to the user's emotional tendency value, input the user's purchase intention probability and emotional levels into a preset policy matching rule table, and map out the conversation word template;
[0037] Call the LLM to generate natural language conversation words based on the conversation word template, and output personalized conversation words after shielding illegal content through a compliance filter.
[0038] By adopting the above technical solution, by comprehensively considering the user's emotional tendency value and the probability of purchase intention and inputting them into a preset policy matching rule table, the accurate mapping of the conversation word template is realized. This not only greatly improves the matching degree of the conversation words with the user's actual needs, enhancing the personalization and pertinence of the interaction, but also through the division of emotional levels, enables the response conversation words to more delicately fit the user's emotional state, effectively improving the user communication experience. Further, using the large language model (LLM) to generate natural language conversation words based on the selected conversation word template not only enriches the expression form of the conversation words, but also ensures the natural fluency and logical coherence of the conversation words. At the same time, the built-in compliance filter can screen and shield illegal content in real time, ensuring that the output personalized conversation words comply with business specifications and respect the user's feelings, laying a solid foundation for building a safe, healthy and efficient interaction environment.
[0039] In a preferred example, this application can be further configured as: after calling the LLM to generate personalized outbound call conversation words according to the user's emotional tendency value and the probability of the user's purchase intention and outputting them to the dialogue system, the method for intelligently generating outbound call conversation words using the LLM further includes:
[0040] Monitor the feedback signals in the user conversation in real time, including response time, frequency of affirmative words, and duration of silence;
[0041] Calculate the speech effect score based on the feedback signals and output the speech score value:
[0042]
[0043] Compare the speech score value with a preset score threshold. When the speech score value is less than the score threshold, generate a speech switching strategy;
[0044] Update the parameters of the intent recognition model based on the speech switching strategy.
[0045] By adopting the above technical solutions, after the personalized speech is output to the user's dialog box, the feedback signals in the user conversation are monitored in real time. The feedback signals include key indicators such as response time, frequency of affirmative words, and duration of silence. The speech effect is evaluated and quantitatively scored immediately, enhancing the sensitivity and response speed of the dialogue system. It can also accurately capture the subtle changes during the conversation with the user. When the speech score value is lower than the preset score threshold, a speech switching strategy is automatically generated, and the parameters of the intent recognition model are automatically updated in response to the speech switching strategy, enhancing the intelligence of the dialogue and making it more in line with the user's intent.
[0046] In a second aspect, the above-mentioned invention object of the present application is achieved through the following technical solutions:
[0047] An outbound speech intelligent generation system using LLM, the outbound speech intelligent generation system using LLM includes:
[0048] A dialogue data collection module for collecting user voice data and text dialogue data in real time and converting them into a text sequence matrix with time series alignment based on the user voice data and text dialogue data;
[0049] An emotion analysis module for inputting the text sequence matrix with time series alignment into a preset multi-modal emotion analysis model to calculate the user emotion tendency value;
[0050] A purchase intention analysis module for extracting intention keywords based on the text sequence matrix with time series alignment, inputting the intention keywords into a preset intent recognition model, and outputting the user purchase intention probability;
[0051] A speech generation module for calling LLM to generate personalized outbound speech according to the user emotion tendency value and the user purchase intention probability and outputting it to the dialogue system.
[0052] By adopting the above technical solutions, during the generation process of outbound call scripts, user voice data (including intonation, speech rate, pauses, etc.) and text conversation content are collected in real time, ensuring the immediacy and accuracy of the data, providing a reliable basis for subsequent sentiment analysis and intent recognition. Through speech recognition technology, it is converted into a time-series aligned text sequence matrix. The time-series aligned text sequence matrix enables the voice and text data to be consistent in time, facilitating subsequent model processing and ensuring an accurate match between semantics and speech rhythm. The obtained time-series aligned text sequence matrix is input into a preset multi-modal sentiment analysis model. The multi-modal sentiment analysis model is a model used to judge and analyze the user's sentiment tendency, which can comprehensively consider the sentiment information in the voice and text data. Using the multi-modal sentiment analysis model, the user sentiment tendency value is calculated, and then the function of analyzing the user's sentiment is realized. By extracting keyword for the time-series aligned text sequence matrix, the intent keywords are extracted, which can more effectively capture the key information in the user's expression. The extracted intent keywords are input into a preset intent recognition model, and the intent recognition model is used to accurately judge the user's purchase intent and output the user's intent to purchase probability. Combining the user sentiment tendency value and the purchase intent probability, a dialogue strategy that better fits the user's needs and emotions is generated, and personalized scripts are output, achieving an accurate grasp and efficient manifestation of the user's needs. This not only improves the user experience and satisfaction, but also helps to improve the business conversion rate and market competitiveness.
[0053] In a third aspect, the above object of the present application is achieved by the following technical solutions:
[0054] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned outbound call script intelligent generation method using LLM are implemented.
[0055] In a fourth aspect, the above object of the present application is achieved by the following technical solutions:
[0056] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned outbound call script intelligent generation method using LLM are implemented.
[0057] In summary, the present application includes at least one of the following beneficial technical effects:
[0058] 1. During the generation process of outbound call scripts, real-time collection of user voice data (including intonation, speech rate, pauses, etc.) and text conversation content ensures the immediacy and accuracy of the data, providing a reliable basis for subsequent sentiment analysis and intent recognition. It is converted into a time-aligned text sequence matrix through speech recognition technology. The time-aligned text sequence matrix keeps the speech and text data consistent in time, facilitating subsequent model processing and ensuring an accurate match between semantics and speech rhythm. The obtained time-aligned text sequence matrix is input into a preset multi-modal sentiment analysis model. The multi-modal sentiment analysis model is a model used to judge and analyze the user's sentiment tendency, which can comprehensively consider the sentiment information in speech and text data. Using the multi-modal sentiment analysis model, the user sentiment tendency value is calculated, and then the function of analyzing the user's sentiment is realized. By extracting keyword of intent from the time-aligned text sequence matrix, the key information in the user's expression can be captured more effectively. The extracted keyword of intent is input into a preset intent recognition model, and the intent recognition model is used to accurately judge the user's purchase intent and output the user's intent to purchase probability. Combining the user sentiment tendency value and the purchase intent probability, a dialogue strategy that better suits the user's needs and emotions is generated, and personalized scripts are output, achieving an accurate grasp and efficient manifestation of the user's needs. This not only improves the user experience and satisfaction but also helps to increase the business conversion rate and market competitiveness;
[0059] 2. Conducting sentiment analysis by combining information from both speech and text modalities can significantly improve the accuracy of sentiment recognition and enhance robustness. By capturing speech features such as fundamental frequency, amplitude fluctuations, etc., as well as text keywords and their sentiment intensity, the multi-modal fusion method can more comprehensively understand the user's sentiment expression, reduce the impact of poor single-modal data quality on the results, and improve the effect of user sentiment tendency analysis;
[0060] 3. During the analysis of the user's sentiment tendency, the sentiment in the voice data and the sentiment in the text data are considered. There may be an obvious difference in sentiment expression between the two, that is, when the sentiment intensity of the keyword is opposite to the sentiment intensity in the voice data, the modal reweighting mechanism is triggered at this time. By adjusting the modal fusion coefficient, it is possible to more flexibly handle the situation of sentiment conflict. When the reweighted second-modal fusion coefficient is less than the coefficient threshold, the speech sentiment dominant strategy is preferentially adopted. This setting is based on the fact that speech sentiment can often more directly and truly reflect the user's immediate emotional state. It can not only improve the accuracy of sentiment analysis, avoid sentiment misjudgment caused by text ambiguity or misunderstanding, help to improve the user experience, ensure that the system can more accurately understand the user's sentiment, and provide more considerate and personalized services for users;
[0061] 4. By dynamically constructing a keyword library, the core intention in the user's expression can be captured in real time, improving the accuracy and timeliness of intention recognition. At the same time, combined with the calculation of keyword group weights, the expression of the user's intention is further refined, enabling the intention recognition model to more precisely understand the user's purchase needs. This not only helps to improve the intelligent recommendation effect of goods, providing more personalized and demand-compliant product recommendations for users, but also can optimize the user experience and improve user satisfaction and loyalty. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is a flowchart of a method for intelligently generating outbound call scripts using an LLM in an embodiment of the present application;
[0063] Figure 2 is a flowchart for implementing step S10 in the method for intelligently generating outbound call scripts using an LLM in an embodiment of the present application;
[0064] Figure 3 is a flowchart for implementing step S20 in the method for intelligently generating outbound call scripts using an LLM in an embodiment of the present application;
[0065] Figure 4 is another flowchart for implementing the method for intelligently generating outbound call scripts using an LLM in an embodiment of the present application;
[0066] Figure 5 is a flowchart for implementing step S30 in the method for intelligently generating outbound call scripts using an LLM in an embodiment of the present application;
[0067] Figure 6 is a flowchart for implementing step S40 in the method for intelligently generating outbound call scripts using an LLM in an embodiment of the present application;
[0068] Figure 7 is another flowchart for implementing the method for intelligently generating outbound call scripts using an LLM in an embodiment of the present application;
[0069] Figure 8 is a schematic block diagram of the principle of a system for intelligently generating outbound call scripts using an LLM in an embodiment of the present application;
[0070] Figure 9 is a schematic diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] The present application will be further described in detail below with reference to the accompanying drawings.
[0072] In one embodiment, as Figure 1 shown, the present application discloses a method for intelligently generating outbound call scripts using an LLM, which specifically includes the following steps:
[0073] S10: Real-time collect user voice data and text conversation data, and convert the user voice data and text conversation data into a temporally aligned text sequence matrix.
[0074] Specifically, during the outbound call script generation process, real-time collect user voice data (including intonation, speech rate, pauses, etc.) and text conversation content, ensuring the immediacy and accuracy of the data, providing a reliable basis for subsequent sentiment analysis and intent recognition. Convert it into a temporally aligned text sequence matrix through speech recognition technology. The temporally aligned text sequence matrix keeps the voice and text data consistent in time, facilitating subsequent model processing and ensuring an accurate match between semantics and speech rhythm.
[0075] S20: Input the temporally aligned text sequence matrix into a preset multi-modal sentiment analysis model to calculate the user sentiment tendency value.
[0076] Specifically, the multi-modal sentiment analysis model is a model used to judge and analyze the user's sentiment tendency. Input the obtained temporally aligned text sequence matrix into the multi-modal sentiment analysis model, and use the multi-modal sentiment analysis model to calculate the user sentiment tendency value by combining voice sentiment features and text keyword sentiment values, thereby realizing the function of analyzing the user's sentiment.
[0077] S30: Extract intent keywords based on the temporally aligned text sequence matrix, input the intent keywords into a preset intent recognition model, and output the user purchase intent probability.
[0078] Specifically, extract keywords from the temporally aligned text sequence matrix to extract intent keywords, which can more effectively capture the key information in the user's expression. Input the extracted intent keywords into a preset intent recognition model, and use the intent recognition model to accurately judge the user's purchase intent and output the user's intent to purchase probability.
[0079] S40: According to the user sentiment tendency value and the user purchase intent probability, call the LLM to generate a personalized outbound call script and output it to the dialogue system.
[0080] Specifically, by combining the user's emotional tendency value and the probability of purchase intention, the LLM adjusts the tone and diction of the conversation strategy according to the user's emotional tendency value to generate a conversation strategy that better suits the user's needs and emotions. For example, when the user shows a positive emotional tendency, the conversation strategy will be more enthusiastic, emphasizing the advantages and strengths of the product; while when the user shows a negative emotional tendency, the conversation strategy will be more gentle and patient, trying to alleviate the user's doubts and dissatisfaction. At the same time, the LLM also adjusts the content and focus of the conversation strategy according to the probability of the user's purchase intention. For users with a strong purchase intention, the conversation strategy will be more direct and clear, highlighting the purchase information and promotional activities of the product; while for users with a weak purchase intention, the conversation strategy will be more euphemistic and guiding, trying to stimulate the user's purchase interest and desire.
[0081] In this embodiment, during the generation of the outbound call conversation strategy, the user's voice data (including tone, speech rate, pauses, etc.) and text conversation content are collected in real time, ensuring the immediacy and accuracy of the data, providing a reliable basis for subsequent emotional analysis and intention recognition. The voice data is converted into a text sequence matrix aligned in time series through speech recognition technology. The text sequence matrix aligned in time series enables the voice and text data to be consistent in time, facilitating subsequent model processing and ensuring an accurate match between semantics and speech rhythm. The obtained text sequence matrix aligned in time series is input into a preset multi-modal emotional analysis model. The multi-modal emotional analysis model is a model used to judge and analyze the user's emotional tendency, which can comprehensively consider the emotional information in the voice and text data. The user's emotional tendency value is calculated using the multi-modal emotional analysis model, thereby realizing the function of analyzing the user's emotions. By extracting the intention keywords from the text sequence matrix aligned in time series, the key information in the user's expression can be captured more effectively. The extracted intention keywords are input into a preset intention recognition model, and the intention recognition model is used to accurately judge the user's purchase intention and output the probability of the user's intended purchase. By combining the user's emotional tendency value and the probability of purchase intention, a conversation strategy that better suits the user's needs and emotions is generated, and personalized conversation strategies are output, realizing the accurate grasp and efficient visualization of the user's needs, not only improving the user experience and satisfaction, but also helping to improve the business conversion rate and market competitiveness.
[0082] In one embodiment, as Figure 2 shown, in step S10, that is, the user's voice data and text conversation data are collected in real time, and are converted into a text sequence matrix aligned in time series based on the user's voice data and text conversation data, which specifically includes:
[0083] S11: Extract the Mel-spectrum features of the user's voice data to generate a spectrogram matrix.
[0084] Specifically, perform Mel-spectrum feature extraction on the collected user voice data, conduct speech signal processing, extract the frequency features in the voice data, which can reflect the spectral structure of the speech signal, including the intensity distribution of different frequency components and their changes over time, and generate a spectrogram matrix. The spectrogram matrix contains information about the speech signal in both the time and frequency dimensions.
[0085] S12: After tokenizing the text dialogue data, generate a sequence of word vectors.
[0086] Specifically, for the text dialogue data, use the tokenization technique to break it down into independent lexical units, and utilize the word embedding model to convert these lexical units into a sequence of word vectors in a high-dimensional space, facilitating subsequent numerical calculations and analyses.
[0087] S13: Match the speech spectral features with the text word vector sequence through the timestamp alignment algorithm, and output the fused temporally aligned text sequence matrix.
[0088] Specifically, use the timestamp alignment algorithm to precisely match the speech spectral features with the text word vector sequence according to the time stamps in the speech and text data, ensuring that the speech features at each moment strictly correspond to the corresponding text content, thereby generating the fused temporally aligned text sequence matrix that contains both speech information and text semantics, providing reliable basic data for subsequent multimodal analysis and intention analysis.
[0089] In one embodiment, as Figure 3 shown, in step S20, input the temporally aligned text sequence matrix into a preset multimodal sentiment analysis model to calculate the user sentiment tendency value, which specifically includes:
[0090] S21: Obtain the speech segment features based on the user voice data, where the speech segment features include the mean fundamental frequency of the speech segment, the fundamental frequency of the background noise, and the amplitude fluctuation coefficient.
[0091] Specifically, parse the user voice data through advanced speech processing techniques, precisely extract the mean fundamental frequency of the speech segment to reflect the average pitch level of the speaker, simultaneously separate and calculate the fundamental frequency features of the background noise to measure the interference degree of the environmental noise on the speech signal; in addition, also calculate the fluctuation coefficient of the amplitude change over time to quantitatively evaluate the amplitude stability of the speech segment.
[0092] S22: Construct an emotion intensity calibration formula based on the speech segment features, and output the emotion intensity in the voice data.
[0093] Specifically, calculate the emotion intensity in the voice data according to the following formula:
[0094]
[0095] where μ p is the average fundamental frequency of the speech segment, and μ b is the fundamental frequency of the background noise, and A(υ j ) is the amplitude fluctuation coefficient.
[0096] S23: Obtain text keywords based on the temporally aligned text sequence matrix, and obtain the word frequency weight and keyword sentiment intensity according to the text keywords.
[0097] Specifically, extract each text segment from the temporally aligned text sequence matrix, use the TF-IDF algorithm combined with text mining technology to identify text keywords with significant discrimination, count the occurrence frequency of each keyword in the text, calculate the word frequency weight based on the frequency distribution to quantify the importance of the keyword. At the same time, perform sentiment tendency analysis on the keywords with the help of a sentiment dictionary, and assign each keyword a corresponding sentiment intensity value, so as to comprehensively reveal the key information, importance and sentiment color of the text content.
[0098] S24: Input the sentiment intensity, word frequency weight and keyword sentiment intensity in the speech data into a preset multi-modal sentiment analysis model to calculate the user sentiment tendency value.
[0099] Specifically, calculate the user sentiment tendency value according to the following formula and combine the sentiment intensity, word frequency weight and keyword sentiment intensity in the speech data:
[0100] where s(t i ) is the keyword sentiment intensity, ω i is the word frequency weight of the keyword, I(υ j ) is the sentiment intensity in the speech data, and α is the modal fusion coefficient (0.6 ≤ α ≤ 0.8).
[0101] In one embodiment, as Figure 4 shown, after step S24, that is, after inputting the sentiment intensity, word frequency weight and keyword sentiment intensity in the speech data into a preset multi-modal sentiment analysis model to calculate the user sentiment tendency value, the method for intelligently generating the outbound call script of the LLM further includes:
[0102] S25: Based on the preset sentiment conflict detection rule, determine whether the keyword sentiment intensity is opposite in sign to the sentiment intensity in the speech data. If so, trigger modal reweighting and output the second modal fusion coefficient.
[0103] Specifically, during the analysis of the user's emotional tendency, the emotions in the voice data and the emotions in the text data are considered. There will be obvious differences in emotional expressions between the two. That is, when the emotional intensity of the keyword is opposite to the emotional intensity in the voice data, the modal reweighting mechanism is triggered at this time, and the following formula is used to adjust the modal fusion coefficient:
[0104] α' = α - γ·|s(t i ) - sign(I(υ i ))|, where γ is the weight coefficient.
[0105] S26: When the second modal fusion coefficient is less than the coefficient threshold, the voice emotion dominant strategy is adopted to generate the second user emotion value.
[0106] Specifically, when the reweighted second modal fusion coefficient is less than the coefficient threshold, that is, α′ < 0.5, based on the fact that voice emotion can often more directly and truly reflect the user's immediate emotional state, the voice emotion dominant strategy is preferentially adopted, and the second modal fusion coefficient is updated into the calculation formula of step S24, and a new user emotional tendency value is output, avoiding emotional misjudgment caused by text ambiguity or misunderstanding, which helps to improve the user experience and ensure that the system can more accurately understand the user's emotion.
[0107] In one embodiment, as Figure 5 shown, in step S30, that is, based on the temporally aligned text sequence matrix, intention keywords are extracted, and the intention keywords are input into a preset intention recognition model to output the user's purchase intention probability, which specifically includes:
[0108] S31: Extract intention keywords according to the temporally aligned text sequence matrix to form a dynamic keyword library.
[0109] Specifically, a dynamic keyword library is constructed, domain keywords are dynamically loaded according to the industry type, and the core intention in the user's expression is captured in real time, improving the accuracy and timeliness of intention recognition.
[0110] S32: Calculate the keyword group weight based on the dynamic keyword library, and input the keywords with the keyword group weight higher than the threshold into a preset intention recognition model to output the user's purchase intention probability.
[0111] Specifically, according to the weight calculation formula, the weight of the keyword combination is calculated, further refining the user's intention expression, so that the intention recognition model can more precisely understand the user's purchase demand:
[0112]
[0113] where TF(k) is the term frequency and IDF(k) is the inverse document frequency;
[0114] Further, input the keywords with weights higher than the threshold into a preset intention recognition model to output the probability of the user's purchase intention:
[0115] P intent = σ(W h ·h t + b), where h t is the sequential hidden state, and W h are the keywords with weights higher than the threshold; it can provide users with more personalized and demand-compliant product recommendations, and can also optimize the user experience, improving user satisfaction and loyalty.
[0116] In one embodiment, as Figure 6 shown, in step S40, that is, according to the user's emotional tendency value and the probability of the user's purchase intention, call the LLM to generate an adapted dialogue strategy and output personalized words, specifically including:
[0117] S41: Divide the emotional level according to the user's emotional tendency value, input the user's purchase intention probability and emotional level into a preset strategy matching rule table, and map out a word template.
[0118] Specifically, divide the emotional level according to the emotional tendency value: positive (E≥0.5): adopt a promotion strategy; neutral (-0.5 < E < 0.5): adopt a guiding strategy; negative (E≤-0.5): adopt a soothing strategy; input it into a preset strategy matching rule table, and map out a word template. Through the division of the emotional level, the response words can more delicately fit the user's emotional state, effectively improving the user communication experience.
[0119] S42: Call the LLM to generate natural language words based on the word template, and output personalized words after shielding illegal content through a compliance filter.
[0120] Specifically, use a large language model (LLM) to generate natural language words based on the selected word template, which not only enriches the expression form of the words, but also ensures the natural fluency and logical coherence of the words.
[0121] Further, the built-in compliance filter can screen and shield illegal content in real time. There is a sensitive word library and compliance regular expressions set during the compliance filtering period. According to the words generated by the LLM, calculate the word risk value, and judge whether the generated words meet the regulations according to the word risk value, ensuring that the output personalized words comply with business specifications and respect the user's feelings, laying a solid foundation for building a safe, healthy and efficient interaction environment, such as:
[0122] When R is less than 1, it means that the generated natural speech meets the requirements.
[0123] In one embodiment, if Figure 7 As shown, after step S40, that is, after calling LLM to generate personalized outbound call speech according to the user's emotional tendency value and the user's purchase intention probability, and outputting it to the dialogue system, the outbound call speech intelligent generation method using LLM also includes:
[0124] S50: monitor the feedback signals in the user conversation in real time, including response time, frequency of affirmative words and silence duration.
[0125] S60: Calculate the speech effect score based on the feedback signal, output the speech score value, compare the speech score value with a preset score threshold, and generate a speech switching strategy when the speech score value is less than the score threshold.
[0126] Specifically, after the personalized speech is output to the user's dialog box, the feedback signals in the user's dialogue are monitored in real time. The feedback signals include key indicators such as response time, frequency of affirmative words and silence duration, and the effect of the dialogue is instantly evaluated and quantitatively scored:
[0127] When the speech score value is lower than the preset score threshold, the speech switching strategy is automatically generated. Generally, the score threshold is 0.3, that is, when S is less than 0.3, the speech strategy is automatically switched.
[0128] S70: Update the intent recognition model parameters based on the speech switching strategy.
[0129] Specifically, it responds to the speech switching strategy, automatically updates the parameters of the intent recognition model, enhances the intelligence of the conversation, and is more in line with user intent.
[0130] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0131] In one embodiment, a system for intelligently generating outbound call speech using LLM is provided, and the system for intelligently generating outbound call speech using LLM corresponds one-to-one to the method for intelligently generating outbound call speech using LLM in the above embodiment. Figure 8 As shown in FIG. 1 , the outbound call script intelligent generation system using LLM includes a dialogue data collection module, a sentiment analysis module, a purchase intention analysis module, and a script generation module. The detailed description of each functional module is as follows:
[0132] A conversation data collection module, used for collecting user voice data and text conversation data in real time, and converting the user voice data and text conversation data into a time-aligned text sequence matrix;
[0133] A sentiment analysis module, used to input the time-aligned text sequence matrix into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value;
[0134] A purchase intention analysis module, used to extract intention keywords based on the time-aligned text sequence matrix, input the intention keywords into a preset intention recognition model, and output the user's purchase intention probability;
[0135] The speech generation module is used to call LLM to generate personalized outbound call speech according to the user's emotional tendency value and the user's purchase intention probability, and output it to the dialogue system.
[0136] For the specific limitations of the outbound call speech intelligent generation system using LLM, please refer to the limitations of the outbound call speech intelligent generation method using LLM above, which will not be repeated here. Each module in the above-mentioned outbound call speech intelligent generation system using LLM can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the electronic device in the form of hardware, or can be stored in the memory of the electronic device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0137] In one embodiment, an electronic device is provided. The electronic device may be a server, and its internal structure diagram may be as follows: Figure 9 As shown. The electronic device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is used to store a database. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for intelligently generating outbound call speech using LLM is implemented.
[0138] In one embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:
[0139] Collecting user voice data and text conversation data in real time, and converting the user voice data and text conversation data into a time-aligned text sequence matrix;
[0140] Input the time-aligned text sequence matrix into a preset multi-modal sentiment analysis model to calculate the user sentiment tendency value;
[0141] Extract intent keywords based on the time-aligned text sequence matrix, input the intent keywords into a preset intent recognition model, and output the user purchase intent probability;
[0142] According to the user sentiment tendency value and the user purchase intent probability, call the LLM to generate personalized outbound call scripts and output them to the dialogue system.
[0143] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0144] Real-time collect user voice data and text dialogue data, and convert them into a time-aligned text sequence matrix based on the user voice data and text dialogue data;
[0145] Input the time-aligned text sequence matrix into a preset multi-modal sentiment analysis model to calculate the user sentiment tendency value;
[0146] Extract intent keywords based on the time-aligned text sequence matrix, input the intent keywords into a preset intent recognition model, and output the user purchase intent probability;
[0147] According to the user sentiment tendency value and the user purchase intent probability, call the LLM to generate personalized outbound call scripts and output them to the dialogue system.
[0148] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0149] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0150] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for intelligently generating outbound call scripts using LLM, characterized in that: The method for intelligently generating outbound call scripts using LLM comprises the following steps: Collecting user voice data and text conversation data in real time, and converting the user voice data and text conversation data into a time-aligned text sequence matrix; Inputting the time-aligned text sequence matrix into a preset multimodal sentiment analysis model to calculate the user sentiment tendency value; Extracting intent keywords based on the time-aligned text sequence matrix, inputting the intent keywords into a preset intent recognition model, and outputting a user purchase intention probability; According to the user's emotional tendency value and the user's purchase intention probability, LLM is called to generate personalized outbound call scripts and output them to the dialogue system.
2. The method for intelligently generating outbound call scripts using LLM according to claim 1, characterized in that: The real-time collection of user voice data and text conversation data, and conversion of the user voice data and text conversation data into a time-aligned text sequence matrix specifically includes: Performing Mel-spectrogram feature extraction on the user voice data to generate a spectrogram matrix; Generate word vector sequences after segmenting text conversation data; The speech spectrum features are matched with the text word vector sequence through the timestamp alignment algorithm, and the fused time-aligned text sequence matrix is output.
3. The method for intelligently generating outbound call scripts using LLM according to claim 1, characterized in that: The step of inputting the time-aligned text sequence matrix into a preset multimodal sentiment analysis model to calculate the user sentiment tendency value specifically includes: Acquire speech segment features based on the user speech data, wherein the speech segment features include a speech segment fundamental frequency mean, a background noise fundamental frequency, and an amplitude fluctuation coefficient; Construct an emotion intensity calibration formula based on the speech segment features, and output the emotion intensity used in the speech data: where μ p is the mean fundamental frequency of the speech segment, μ b is the fundamental frequency of background noise, A(υ j ) is the amplitude fluctuation coefficient; Acquire text keywords based on the time-aligned text sequence matrix, and acquire word frequency weights and keyword sentiment strengths according to the text keywords; The emotion intensity, word frequency weight and keyword emotion intensity in the voice data are input into a preset multimodal emotion analysis model to calculate the user emotion tendency value: Among them, s(t i ) is the keyword sentiment intensity, ω i is the keyword frequency weight, I(υ j ) is the emotion intensity in the speech data, and α is the modal fusion coefficient.
4. The method for intelligently generating outbound call scripts using LLM according to claim 3 is characterized in that: After inputting the emotion intensity, word frequency weight and keyword emotion intensity in the voice data into a preset multimodal emotion analysis model and calculating the user emotion tendency value, the method for intelligently generating outbound call speech using LLM further includes: Based on the preset emotion conflict detection rules, it is determined whether the emotion intensity of the keyword is opposite to the emotion intensity in the voice data. If so, the modal reweighting is triggered and the second modal fusion coefficient is output: α'=α-γ·s(t i )-sign(I(υ i ))|, where γ is the weight coefficient; When the second modal fusion coefficient is less than the coefficient threshold, a speech emotion-dominant strategy is adopted to generate a second user emotion value.
5. The method for intelligently generating outbound call scripts using LLM according to claim 1, characterized in that: The extracting intent keywords based on the time-series aligned text sequence matrix, inputting the intent keywords into a preset intent recognition model, and outputting the user purchase intention probability specifically includes: Extracting intent keywords according to the time-aligned text sequence matrix to form a dynamic keyword vocabulary; The keyword group weights are calculated based on the dynamic keyword word library, the keyword group weights are input into a preset intention recognition model, and the user purchase intention probability is output.
6. The method for intelligently generating outbound call scripts using LLM according to claim 1, characterized in that: According to the user's emotional tendency value and the user's purchase intention probability, calling LLM to generate an adaptive dialogue strategy and outputting personalized speech specifically includes: The emotional level is divided according to the emotional tendency value of the user, and the user's purchase intention probability and emotional level are input into a preset strategy matching rule table to map out a speech template; Call LLM to generate natural language speech based on the speech template, and output personalized speech after blocking illegal content through compliance filters.
7. The method for intelligently generating outbound call scripts using LLM according to claim 1, characterized in that: After calling LLM to generate personalized outbound call scripts according to the user's emotional tendency value and the user's purchase intention probability, and outputting them to the dialogue system, the outbound call script intelligent generation method using LLM also includes: Real-time monitoring of feedback signals in user conversations, including response time, frequency of affirmative words, and duration of silence; Calculate the speech effect score based on the feedback signal and output the speech score value: Compare the speech scoring value with a preset scoring threshold, and when the speech scoring value is less than the scoring threshold, generate a speech switching strategy; Update the intent recognition model parameters based on the speech switching strategy.
8. An intelligent generation system for outbound call scripts using LLM, characterized in that: The outbound call script intelligent generation system using LLM includes: A conversation data collection module, used for collecting user voice data and text conversation data in real time, and converting the user voice data and text conversation data into a time-aligned text sequence matrix; A sentiment analysis module, used to input the time-aligned text sequence matrix into a preset multimodal sentiment analysis model to calculate the user's sentiment tendency value; A purchase intention analysis module, used to extract intention keywords based on the time-aligned text sequence matrix, input the intention keywords into a preset intention recognition model, and output the user's purchase intention probability; The speech generation module is used to call LLM to generate personalized outbound call speech according to the user's emotional tendency value and the user's purchase intention probability, and output it to the dialogue system.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method for intelligently generating outbound call scripts using LLM as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of a method for intelligently generating outbound call scripts using LLM as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Speech skill recommendation method and device based on semantic recognition, equipment and storage medium
CN112732911A
Intention recognition method and device, electronic equipment and computer readable storage medium
CN113094481A
Intelligent verbal skill recommendation method, system and device and storage medium
CN115810345A
Voice conversion method and device, electronic equipment and storage medium
CN116312617A
Intelligent outbound client intention prediction and analysis system
CN117834780A
Cited By
Voice call real-time transcription system and method
CN120526774A
Emotion analysis model establishment method and device and response generation method and device
CN120783801A
Customer service dialogue sentiment analysis method and system based on multi-modal time sequence perception
CN120853620A
Voice outbound method and device based on intention recognition and storage medium
CN121098987A
Psychological counseling real-time speech recognition method based on multi-modal data
CN121337358A