Call center verbal skill recommendation method and device
By constructing user behavior diagrams and multi-modal sentiment analysis, and optimizing speech generation, the complexity of speech recommendations in the financial dispute scenarios in the existing technology is solved, accurate, compliant and efficient communication effects are achieved, and user experience and mediation success rate is improved.
Patent Information
- Application Number
- CN202510356688.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-08-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology is difficult to adapt to complex communication needs in financial dispute scenarios, cannot accurately understand the deep emotions of users, poses compliance risks, insufficient data integration and low human-computer coordination efficiency, resulting in poor speech recommendation results.
By constructing user behavior diagrams, generating portrait vectors, performing multi-modal sentiment analysis, combining preset thesaurus and mediator adjustments, optimizing speech generation, ensuring compliance and emotional resonance, and realizing multi-channel data fusion and human-machine collaborative optimization.
It improves the accuracy and compliance of speech recommendations, enhances the emotional resonance of communication and the matching of regulations, improves the success rate and user experience of mediation, and reduces the work burden of mediators.
Smart Images

Figure CN120470078A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech generation, and in particular to a speech recommendation method and device for a call center. Background Art
[0002] Traditional methods for recommending conversational dialogue in financial and economic mediation rely on pre-set templates, making them ill-suited for the complex financial dispute scenarios encountered by call centers. For example, in the context of multiple debt disputes or regulatory adjustments, templated conversational dialogue often fails to meet the individual needs of clients, resulting in poor communication effectiveness. Furthermore, most existing methods can only identify simple surface emotions but fail to accurately understand the client's deeper psychological state, making it difficult to provide truly effective soothing dialogue.
[0003] Secondly, compliance issues are a major challenge in recommending mediation scripts. Some scripts may contain implicit incentives or legal loopholes, such as misleading expressions, which may infringe on user rights and trigger legal risks. At the same time, insufficient data integration also limits the effectiveness of recommended scripts. Information from multiple channels such as phone calls and emails is often processed in a decentralized manner, lacking a global perspective. As a result, the recommended scripts fail to fully consider the user's complete historical interactions. Finally, the low efficiency of human-computer collaboration is also a major shortcoming of current methods. Although artificial intelligence can generate preliminary scripts, the lack of an effective feedback loop mechanism makes it difficult to continuously improve the AI model.
[0004] Therefore, although the application of artificial intelligence and natural language processing technologies has made certain progress, there are still problems such as insufficient dynamic adaptability, limitations in sentiment analysis, high compliance risks, difficulties in data integration, and low efficiency in human-computer collaboration, which affect the accuracy and practicality of recommended speech techniques. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and device for recommending speech techniques in a call center to improve the above-mentioned problem. To achieve the above-mentioned purpose, the technical solution adopted by the present invention is as follows:
[0006] In a first aspect, the present application provides a method for recommending speech techniques in a call center, comprising:
[0007] Obtain historical recommendation data and user information to be recommended through the call center, wherein the user information to be recommended includes basic information, current behavior data, text data and voice data;
[0008] Build a user behavior graph through historical recommendation data and information of users to be recommended, and generate a portrait vector of the user to be recommended based on the user behavior graph;
[0009] Perform multimodal sentiment analysis based on the portrait vector to obtain sentiment features;
[0010] Build a speech recommendation model based on historical recommendation data, and input the portrait vector and sentiment features into the speech recommendation model to generate preliminary recommended speech;
[0011] The preliminary recommendation words are reviewed and optimized based on the preset vocabulary to obtain the optimized recommendation words, and the mediator adjusts the optimized recommendation words to obtain the final recommendation words.
[0012] In a second aspect, the present application further provides a call center speech recommendation device, comprising:
[0013] An acquisition module is used to acquire historical recommendation data and user information to be recommended through a call center, wherein the user information to be recommended includes basic information, current behavior data, text data, and voice data;
[0014] A construction module is used to construct a user behavior graph based on historical recommendation data and information about the user to be recommended, and generate a portrait vector of the user to be recommended based on the user behavior graph;
[0015] Sentiment analysis module, used to perform multimodal sentiment analysis based on portrait vectors to obtain sentiment features;
[0016] The recommendation module is used to build a recommendation model for a specific topic based on historical recommendation data, and input the portrait vector and sentiment features into the recommendation model to generate preliminary recommended topics.
[0017] The optimization module is used to review and optimize the preliminary recommendation words based on the preset vocabulary to obtain the optimized recommendation words, and adjust the optimized recommendation words through the mediator to obtain the final recommendation words.
[0018] The beneficial effects of the present invention are as follows: by constructing a user behavior graph, the present invention enhances the accuracy of user portraits, and based on portrait vector analysis, explores users' emotions and underlying needs, generating emotionally and professionally engaging dialogue. Furthermore, through compliance testing, the generated dialogue is replaced and optimized to avoid implicit interest tendencies or legal loopholes in the generated dialogue, which infringes on user rights. At the same time, multi-channel data fusion and human-machine collaborative optimization closed loops are achieved, facilitating professional and efficient communication for mediators, reducing the burden of dialogue adjustment on mediators, and enhancing the emotional resonance and regulatory matching of communication, thereby improving mediation success rates and user experience.
[0019] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the embodiments of the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 This is a flowchart of a method for recommending speech techniques in a call center according to an embodiment of the present invention;
[0022] Figure 2 This is a user behavior diagram of historical user 1 in an embodiment of the present invention. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0024] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are used only to distinguish the description and should not be understood as indicating or implying relative importance.
[0025] Example 1:
[0026] This embodiment provides a method for recommending speech techniques in a call center.
[0027] See also Figure 1 , the figure shows that the method includes step S100, step S200, step S300, step S400 and step S500.
[0028] Step S100: Obtain historical recommendation data and user information to be recommended through a call center, wherein the user information to be recommended includes basic information, current behavior data, text data, and voice data;
[0029] In this embodiment, the basic information of the user to be recommended is filled in by the user himself, including age, gender, occupation, income level, credit score, credit history, etc., while the current behavior data includes the number of recent interactions with the mediator, emotional fluctuation information, number of complaints and consultation type, etc.
[0030] Text data includes text information from multiple channels, such as the text of the call content, the mediation needs, the specific circumstances of the required mediation and historical text data, while voice data is the voice call information between the user and the mediator.
[0031] Step S200: constructing a user behavior graph based on historical recommendation data and information of the user to be recommended, and generating a portrait vector of the user to be recommended based on the user behavior graph;
[0032] In this embodiment, the user behavior graph is based on a graph structure, representing the user's past behaviors and interactive relationships, and is used to predict the user's demand patterns and mediation preferences.
[0033] In step S200, constructing a user behavior graph includes:
[0034] Step A100: Obtaining basic information, historical interaction records, historical recommendation words, and historical relevant regulations of each historical user through historical recommendation data, wherein the historical interaction records include historical behavior data, historical sentiment trends, and high-frequency keywords;
[0035] Step A200: Construct a user behavior graph for each historical user based on the historical user's basic information, historical interaction records, historical recommendation words, and historical relevant regulations;
[0036] In this embodiment, each historical user, each historical interaction record, each historical recommended phrase, and each historical relevant regulation is treated as a node. The connection between each historical user and the corresponding historical interaction record is used as an edge, the connection between the historical recommended phrase and the corresponding historical relevant regulation is used as an edge, and the connection between historical interaction records at different time points in each history is used as an edge to form a behavior sequence. The historical behavior data in the historical interaction record includes the number of phone calls, emails, and online consultations.
[0037] like Figure 2 As shown, this is the user behavior diagram of historical user 1, where C i Represents historical user i, I ij represents the jth historical interaction record of historical user i, j = 1, 2, ..., n, n represents the total number of historical interaction records, Q ij represents the historical recommendation words corresponding to the jth historical interaction record of historical user i, P kIndicates the kth regulation document used in historical relevant regulations, k = 1, 2,…, m, and m represents the total number of regulation documents.
[0038] At the same time, each node and edge in the user behavior graph is vectorized to facilitate subsequent use and save computing time.
[0039] Step A300: Calculate the similarity between the user to be recommended and each historical user through the user behavior graph of the historical users;
[0040] In this embodiment, since the user may be mediating for the first time and has no historical data, we find similar users through the user's basic information and extract these users' historical interaction records as supplementary information to avoid the cold start problem. Even first-time mediators can get personalized recommendation words to improve communication effectiveness.
[0041] In step A300, the similarity calculation steps are:
[0042] Step A301: Obtain historical interaction records of the user to be recommended;
[0043] Step A302: If the historical interaction record of the user to be recommended is empty, then the cosine similarity is calculated based on the basic information of the user to be recommended and the historical user, and then a second similar user is selected;
[0044] In this embodiment, the calculated cosine similarities are sorted, and according to actual conditions, a certain number of historical users with the highest cosine similarities are selected as second similar users.
[0045] Step A303: Generate pseudo historical interaction records of the user to be recommended by using nearest neighbor interpolation and the historical interaction records of the second similar user;
[0046] In this embodiment, since the new user has no historical interaction records, the behavior patterns of similar users can be used to generate pseudo historical interaction records, and the historical interaction records of the second similar user are subjected to nearest neighbor interpolation calculation to obtain pseudo historical interaction records.
[0047] Step A304: Calculate the similarity between the user to be recommended and each historical user through dynamic time warping, pseudo historical interaction records, historical interaction records, and basic information;
[0048] Step A305: If the historical interaction record of the user to be recommended is not empty, the similarity between the user to be recommended and each historical user is calculated through dynamic time warping, historical interaction records and basic information.
[0049] In this embodiment, since the user's historical interaction records may not be aligned in time, dynamic time warping is used to calculate similarity. Dynamic time warping processes data with time series, which is more suitable for data with variable length, uneven sampling, and time offset than direct Euclidean distance or cosine similarity. Therefore, through dynamic time warping, although the time points of events are different, users with similar behavior patterns are considered similar. At the same time, by calculating similarity through dynamic time warping and selecting similar users to supplement the data of new users, the accuracy of speech recommendation can be enhanced.
[0050] In step A300, the similarity calculation formula is:
[0051] S(u,v)=α·Sim DTW (u,v)+β·Sim basic (u,v)
[0052]
[0053] In the formula, S(u,v) represents the similarity between the user u to be recommended and the historical user v, α and β are weight parameters, Sim DTW (u,v) represents the similarity between the interaction records of the recommended user u and the historical user v, Sim basic (u,v) represents the similarity between the basic information of the recommended user u and the historical user v, DTW(·) represents the dynamic time warping distance, H u Represents the historical interaction records of the user u to be recommended, H v Represents the historical interaction records of historical user v, H' u represents the pseudo historical interaction record of the user u to be recommended, B u Indicates the basic information of the user u to be recommended, B v Represents the basic information of historical user v.
[0054] Step A400: Select a preset number of historical users as first similar users based on similarity, and generate a user behavior graph of the users to be recommended based on the first similar users.
[0055] In step S200, generating a portrait vector of the user to be recommended based on the user behavior graph includes:
[0056] Step B100: Use a graph neural network to extract features from the user behavior graph of the recommended user to obtain historical features;
[0057] In this example, a graph neural network aggregates the influence of historical behaviors to extract a global historical profile, rather than simply viewing each behavior separately. This modeling of relationships between various data allows information such as voice emotion, text keywords, and historical behaviors to mutually enhance and produce a more accurate user profile. Furthermore, by smoothing features using adjacency relationships, the interference of individual abnormal behaviors is reduced, resulting in a more stable user behavior profile.
[0058] Step B200: extracting text keywords from the text data to obtain text semantic features, extracting speech features from the speech data using a pre-trained speech model to obtain speech features, and labeling the text keywords with sentiment scores using a sentiment dictionary;
[0059] In this example, word embeddings are extracted using BERT to obtain semantic features of the text, which are then clustered to extract the most important keywords. Simultaneously, the Wav2Vec 2.0 model is used to extract features from the speech data to obtain high-dimensional speech features. These features include time-series information such as speech rate, pauses, pitch, and intensity.
[0060] Step B300: Perform feature engineering processing on the current behavior data to obtain current behavior features;
[0061] In this embodiment, the current behavior data represents the user's status regarding a mediation problem, and feature engineering is performed on it, that is, converting it into vectorized features that can be processed by the model for further use in speech recommendation.
[0062] Step B400: Multimodally combine historical features, current behavior features, text semantic features, and voice features to obtain a portrait vector of the user to be recommended.
[0063] Step S300: performing multimodal sentiment analysis based on the portrait vector to obtain sentiment features;
[0064] The step S300 includes:
[0065] Step S301: learning context information of historical features through a sequence model to obtain a context vector;
[0066] In this embodiment, LSTM is used to learn the contextual information of historical features, which can capture the long-term dependencies in the time series, that is, how the user's historical behavior and emotional change trends affect the current state.
[0067] Step S302: concatenate the context vector and the portrait vector to obtain a temporal concatenation feature;
[0068] Step S303: adjusting the feature weight of the temporal splicing feature through the attention mechanism to obtain enhanced features;
[0069] Step S304: inputting the enhanced features into a sentiment classification model for sentiment analysis to obtain sentiment features, where the sentiment features include a sentiment type and a sentiment score.
[0070] In this embodiment, since existing sentiment analysis only recognizes surface emotions and cannot understand the deep psychological needs of customers, deep needs mining is carried out through multimodal portrait vectors to obtain emotional characteristics, and then the portrait vectors and emotional characteristics are used to generate more empathetic recommendation words to improve the communication experience and mediation success rate.
[0071] Step S400: Build a speech recommendation model based on historical recommendation data, input the portrait vector and sentiment features into the speech recommendation model to generate preliminary recommended speech;
[0072] In step S400, the portrait vector and the emotional features are input into the speech recommendation model to generate preliminary recommended speech, including:
[0073] Step C100: performing a nearest neighbor search using the portrait vector and a case knowledge base to obtain relevant cases, wherein the case knowledge base is constructed using historical recommendation data and current regulations;
[0074] In this embodiment, the case knowledge base stores each historical recommendation data and the corresponding portrait vector, and the nearest neighbor search is used to match cases. By setting the number of cases, the corresponding number of similar cases is returned.
[0075] The case knowledge base includes historical cases and various regulations. In the case knowledge base, regulations will be updated according to actual conditions. For example, if previous regulations are adjusted, they will be replaced with new regulations to achieve dynamic updates of regulations, so that the latest regulations can be given when matching applicable regulations in the future.
[0076] Step C200: performing reasoning based on relevant cases and the pre-trained TransE model and RotatE model to obtain matching requirements for the user to be recommended;
[0077] The step C200 includes:
[0078] Step C201: Create triples through related cases to obtain a historical knowledge graph;
[0079] Step C202: establishing a triplet to be recommended based on the information of the user to be recommended, wherein the triplet to be recommended is specified as an entity to be predicted;
[0080] Step C203: Perform knowledge reasoning using the triples to be recommended, the historical knowledge graph, and the pre-trained TransE model to obtain the predicted specifications of the user to be recommended;
[0081] Step C204: Correct the predicted specification vector using the pre-trained RotatE model, and obtain the matching specification of the user to be recommended based on multi-hop reasoning.
[0082] In this example, key entities are extracted from relevant cases to construct triples, such as (case claim, applicable, regulation), to generate a historical knowledge graph. For example, a case claim might be a client applying for an extension due to financial hardship. The case type and claim are then derived from the user information to be recommended, and a recommendation triple is constructed: (recommended case claim, applicable, ?). The last entity in the recommendation triple is actually the entity to be predicted, and the ? in (recommended case claim, applicable, ?) represents the entity to be predicted.
[0083] The TransE model is used to calculate the potential position of the entity to be predicted, and the provision closest to the potential position is selected from the provision set of the historical knowledge graph as the predicted provision. Then, multi-hop reasoning is performed through the RotatE model to ensure that the predicted provision is consistent with the case context, and multiple applicable provisions are obtained as matching provisions.
[0084] Step C300: Generate a prompt project based on the portrait vector, sentiment features, related cases, and matching regulations, and input the prompt project into the large language model to generate preliminary recommendation words.
[0085] In this embodiment, a prompting project is generated to convert various information into understandable input, as shown in Table 1.
[0086] Table 1 Schematic diagram of the prompt project for user A to be recommended
[0087]
[0088] In addition, through a large language model trained with historical cases, such as the GPT-4 model, the prompt project will be input into the trained large language model, combined with emotional characteristics to adjust the tone of the speech, while controlling the length of the speech to avoid it being too lengthy, and obtaining a preliminary recommended speech that complies with regulations and is empathetic.
[0089] Step S500: Review and optimize the preliminary recommended words based on the preset vocabulary to obtain optimized recommended words, and adjust the optimized recommended words through the mediator to obtain the final recommended words.
[0090] In this embodiment, the preset vocabulary is very important in the call center's speech recommendation to ensure that the recommended speech meets compliance requirements and avoids the use of inappropriate, misleading or sensitive expressions.
[0091] In step S500, the preliminary recommended speech is reviewed and optimized based on the preset vocabulary to obtain the optimized recommended speech, including:
[0092] Step D100: Establish a preset vocabulary based on banned words, sensitive terms, misleading expressions, and compliance requirements. The preset vocabulary includes illegal expressions and corresponding replacement expressions.
[0093] In this embodiment, the preset vocabulary includes non-compliant words, which may involve offensive language and sensitive language, as well as language that causes misunderstanding or misleading. Table 2 is a partial illustration of the preset vocabulary.
[0094] Table 2 Schematic diagram of the preset vocabulary
[0095]
[0096] As shown in Table 2, each of the pre-set lexicon's illegal expressions has a corresponding replacement. This helps mediators avoid using inappropriate language and ensures that call center communications meet compliance requirements. By replacing illegal expressions, potential customer dissatisfaction can be effectively avoided.
[0097] Step D200: Detect violations of the preliminary recommended speech based on natural language processing, a preset vocabulary, and keyword matching to obtain violation sentences;
[0098] Step D300: Replace the offending statement with a replacement expression to obtain an optimized recommendation statement.
[0099] In this embodiment, through speech generation and compliance review, recommended speech that takes into account factors such as compliance and emotion can be generated, as shown in Table 3, which shows some of the generated optimized recommended speech.
[0100] Table 3 Schematic diagram of optimized recommendation script generation
[0101]
[0102] As shown in Table 3, corresponding soothing words can be generated according to customer demands, and they are professional, making it easier for users to accept.
[0103] This embodiment also displays information such as user requests, emotions, emotion scores, and generated optimized recommendation phrases on a visual interface. This allows moderators to adjust and optimize the recommended phrases to obtain the final recommended phrases. Adjustments made by the moderator are annotated, with each final recommended phrase generated after adjustment being labeled as a positive example, while the initial phrases before adjustment are labeled as negative examples. This supervised learning of the large language model is then optimized through a closed-loop feedback mechanism.
[0104] In summary, this invention collects multimodal user data through a call center, constructs a user behavior graph, and utilizes graph neural networks to extract features and enhance the accuracy of user profiles. Secondly, dynamic time warping is introduced to calculate the behavioral similarity of historical users to address the cold start problem for new users. Nearest neighbor interpolation is then used to generate pseudo-historical interaction records to improve data integrity. Subsequently, LSTM is used to learn contextual information, combined with multimodal sentiment analysis to uncover deeper user needs and enhance understanding of user demands.
[0105] During the script generation phase, the case knowledge base is combined with TransE and RotatE model reasoning to ensure compliance with regulations. Furthermore, a large language model, combined with prompt engineering, automatically generates context-appropriate and empathetic scripts. Violations are reviewed and optimized based on a pre-set vocabulary, enhancing the professionalism and security of communication. Finally, the recommended scripts are further adjusted by moderators on a visual interface, and a closed-loop feedback loop optimizes the large model to continuously improve recommendation quality.
[0106] Therefore, the present invention not only improves the accuracy of speech recommendation and reduces the burden of speech adjustment on mediators, but also enhances the emotional resonance and regulation matching of communication, thereby improving the mediation success rate and user experience.
[0107] Example 2:
[0108] like Figure 2 As shown, this embodiment provides a speech recommendation device for a call center, the device comprising:
[0109] An acquisition module is used to acquire historical recommendation data and user information to be recommended through a call center, wherein the user information to be recommended includes basic information, current behavior data, text data, and voice data;
[0110] A construction module is used to construct a user behavior graph based on historical recommendation data and information about the user to be recommended, and generate a portrait vector of the user to be recommended based on the user behavior graph;
[0111] Sentiment analysis module, used to perform multimodal sentiment analysis based on portrait vectors to obtain sentiment features;
[0112] The recommendation module is used to build a recommendation model for a specific topic based on historical recommendation data, and input the portrait vector and sentiment features into the recommendation model to generate preliminary recommended topics.
[0113] The optimization module is used to review and optimize the preliminary recommendation words based on the preset vocabulary to obtain the optimized recommendation words, and adjust the optimized recommendation words through the mediator to obtain the final recommendation words.
[0114] It should be noted that, regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.
[0115] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
[0116] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A method for recommending speech techniques in a call center, characterized in that: include: Obtain historical recommendation data and user information to be recommended through the call center, wherein the user information to be recommended includes basic information, current behavior data, text data and voice data; Build a user behavior graph through historical recommendation data and information of users to be recommended, and generate a portrait vector of the user to be recommended based on the user behavior graph; Perform multimodal sentiment analysis based on the portrait vector to obtain sentiment features; Build a speech recommendation model based on historical recommendation data, and input the portrait vector and sentiment features into the speech recommendation model to generate preliminary recommended speech; The preliminary recommendation words are reviewed and optimized based on the preset vocabulary to obtain the optimized recommendation words, and the mediator adjusts the optimized recommendation words to obtain the final recommendation words.
2. The call center speech recommendation method according to claim 1 is characterized in that ,The construction of the user behavior graph includes: Obtain each historical user's basic information, historical interaction records, historical recommendation scripts, and historical relevant regulations through historical recommendation data. The historical interaction records include historical behavior data, historical sentiment trends, and high-frequency keywords. Build a user behavior graph for each historical user based on their basic information, historical interaction records, historical recommendation phrases, and historical relevant regulations; Calculate the similarity between the recommended user and each historical user through the user behavior graph of historical users; A preset number of historical users are selected as first similar users based on similarity, and a user behavior graph of the users to be recommended is generated based on the first similar users.
3. The call center speech recommendation method according to claim 2, characterized in that , the similarity calculation steps are: Obtain historical interaction records of the user to be recommended; If the historical interaction record of the user to be recommended is empty, the cosine similarity is calculated based on the basic information of the user to be recommended and the historical user, and then the second similar user is selected; Generate pseudo historical interaction records of the user to be recommended through nearest neighbor interpolation and the historical interaction records of the second most similar user; Calculate the similarity between the recommended user and each historical user through dynamic time warping, pseudo-historical interaction records, historical interaction records and basic information; If the historical interaction record of the user to be recommended is not empty, the similarity between the user to be recommended and each historical user is calculated through dynamic time warping, historical interaction records and basic information.
4. The call center speech recommendation method according to claim 3 is characterized in that , the calculation formula of the similarity is: S(u,v)=α·Sim DTW (u,v)+β·Sim basic (u,v) In the formula, S(u,v) represents the similarity between the user u to be recommended and the historical user v, α and β represent weight parameters, Sim DTW (u,v) represents the similarity between the interaction records of the recommended user u and the historical user v, Sim basic (u,v) represents the similarity between the basic information of the user u to be recommended and the historical user v, DTW(·) represents the dynamic time warping distance, Hu represents the historical interaction record of the user u to be recommended, and H v Represents the historical interaction records of historical user v, H' u represents the pseudo historical interaction record of the user u to be recommended, B u Indicates the basic information of the user u to be recommended, B v Represents the basic information of historical user v.
5. The call center speech recommendation method according to claim 1, characterized in that The process of generating a user portrait vector based on the user behavior graph includes: Use graph neural networks to extract features from the user behavior graph of recommended users and obtain historical features; Extracting text keywords from text data to obtain text semantic features, extracting speech features from speech data using a pre-trained speech model to obtain speech features, and labeling the text keywords with sentiment scores using a sentiment dictionary; Perform feature engineering on the current behavior data to obtain the current behavior features; Historical features, current behavior features, text semantic features and voice features are multimodally spliced to obtain the portrait vector of the user to be recommended.
6. The call center speech recommendation method according to claim 1, characterized in that ,The multimodal sentiment analysis is performed based on the portrait vector to obtain the ,sentiment features, including: Learn the context information of historical features through the sequence model to obtain the context vector; Concatenate the context vector and the portrait vector to obtain the temporal concatenation feature; The feature weights of the temporal splicing features are adjusted through the attention mechanism to obtain enhanced features; The enhanced features are input into a sentiment classification model for sentiment analysis to obtain sentiment features, which include sentiment types and sentiment scores.
7. The call center speech recommendation method according to claim 1, characterized in that ,The portrait vector and sentiment features are input into the speech recommendation model to generate preliminary recommended speech, including: Performing nearest neighbor search using the portrait vector and a case knowledge base to obtain relevant cases, wherein the case knowledge base is constructed using historical recommendation data and current regulations; Reasoning is performed using relevant cases and pre-trained TransE and RotatE models to obtain matching requirements for users to be recommended. The portrait vector, sentiment features, relevant cases and matching regulations are used to generate a prompt project, which is then input into the large language model to generate preliminary recommendation words.
8. The call center speech recommendation method according to claim 7, characterized in that The matching requirements for obtaining the user to be recommended include: Establish triples through related cases to obtain a historical knowledge graph; Establishing a triplet to be recommended based on the information of the user to be recommended, wherein the provisions in the triplet to be recommended are entities to be predicted; By performing knowledge reasoning on the triples to be recommended, the historical knowledge graph, and the pre-trained TransE model, the predicted specifications of the users to be recommended are obtained. The predicted regulation vector is corrected by the pre-trained RotatE model, and the matching regulation of the user to be recommended is obtained based on multi-hop reasoning.
9. The call center speech recommendation method according to claim 1, characterized in that The preliminary recommendation words are reviewed and optimized based on the preset word library to obtain the optimized recommendation words, including: Establish a preset vocabulary based on banned words, sensitive terms, misleading expressions, and compliance requirements. The preset vocabulary includes illegal expressions and corresponding replacement expressions. Based on natural language processing, preset vocabulary and keyword matching, the preliminary recommended words are detected for violations and illegal sentences are obtained; Replace the offending statements with replacement expressions to obtain optimized recommendation phrases.
10. A speech recommendation device for a call center, characterized in that: include: An acquisition module is used to acquire historical recommendation data and user information to be recommended through a call center, wherein the user information to be recommended includes basic information, current behavior data, text data, and voice data; A construction module is used to construct a user behavior graph based on historical recommendation data and information about the user to be recommended, and generate a portrait vector of the user to be recommended based on the user behavior graph; Sentiment analysis module, used to perform multimodal sentiment analysis based on portrait vectors to obtain sentiment features; The recommendation module is used to build a recommendation model for a specific topic based on historical recommendation data, and input the portrait vector and sentiment features into the recommendation model to generate preliminary recommended topics. The optimization module is used to review and optimize the preliminary recommendation words based on the preset vocabulary to obtain the optimized recommendation words, and adjust the optimized recommendation words through the mediator to obtain the final recommendation words.
Citation Information
Cited By
AI Agent-based low-code outbound call skill process configuration method and system
CN121279349A
A Low-Code Outbound Call Script Configuration Method and System Based on AI Agent
CN121279349B