Model-constructed virtual customer service response system and data processing method
By building a virtual customer service response system and utilizing natural language processing and generative AI technologies, we have solved the problem that existing systems are unable to handle complex customer interactions, achieved personalized services and natural interactions, and improved customer satisfaction and sales conversion rates.
Patent Information
- Application Number
- CN202510811723.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-26
AI Technical Summary
Existing virtual customer service systems lack deep semantic understanding, are unable to handle complex customer interactions, and are unable to provide personalized services and natural interactions.
The virtual customer service response system built using the model includes a user interaction module, an NLP understanding module, a generative AI module, an intelligent recommendation module, a speech recognition and synthesis module, a scenario-adaptive speech generation library, and a feedback analysis module. It achieves deep semantic understanding and personalized responses through natural language processing and generative AI technologies.
Provide natural and coherent multi-round conversations, generate personalized product or service recommendations, improve customer satisfaction and sales conversion rate, and simulate the interactive effects of real customer service.
Smart Images

Figure CN120707154A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of customer service systems, and in particular to a model-based virtual customer service response system and a data processing method. Background Art
[0002] Virtual customer service systems have become a crucial tool for customer service and sales in modern enterprises. Leveraging automation and artificial intelligence (AI), many companies are using virtual customer service to handle a large volume of customer requests, reducing labor costs and improving service efficiency. However, most existing virtual customer service systems rely on rule-based or search-based interaction methods, failing to fully meet customer demands for personalized service, natural interactions, and complex problem resolution.
[0003] Traditional customer service systems rely on keyword matching, preset rules, and templated responses. While these systems are effective for handling standardized questions (such as order inquiries and FAQs), they often struggle with dynamic user needs, open-ended questions, or complex sales scenarios. This is primarily due to their lack of flexibility and inability to understand context or generate personalized responses. Summary of the Invention
[0004] The purpose of the present invention is to provide a model-based virtual customer service response system and data processing method, aiming to solve the problem that existing systems lack deep semantic understanding and are unable to handle complex customer interactions.
[0005] To achieve the above objectives, in a first aspect, the present invention provides a model-built virtual customer service response system, comprising a user interaction module, an NLP understanding module, a generative AI module, an intelligent recommendation module, a speech recognition and synthesis module, a scenario-adaptive speech generation library, and a feedback analysis module, wherein the user interaction module is connected to the feedback analysis module and the NLP understanding module, the generative AI module and the intelligent recommendation module are respectively connected to the NLP understanding module, and the speech recognition and synthesis module and the scenario-adaptive speech generation library are respectively connected to the generative AI module;
[0006] The user interaction module is used to process user input and convert the user's text or voice request into instructions that can be recognized by the system;
[0007] The NLP understanding module performs semantic analysis, intent recognition, and context understanding on the text or voice input by the user based on natural language processing technology;
[0008] The generative AI module uses a generative model to generate natural and fluent language responses based on user input and contextual information, and intelligently responds to customer needs.
[0009] The intelligent recommendation module generates personalized product or service recommendations by analyzing users' historical behavior, real-time interactions, and current market trends;
[0010] The speech recognition and synthesis module converts the user's speech into text for system processing, and also supports converting the generated text replies into speech output;
[0011] The scenario-adaptive speech generation library generates personalized sales speech based on different sales scenarios.
[0012] The feedback analysis module monitors user feedback and adjusts generation strategies and recommendation logic.
[0013] The user interaction module includes a text input unit, a voice output unit, a conversion unit and a reply unit, wherein the text input unit and the voice output unit are respectively connected to the conversion unit;
[0014] The text input unit is used to collect text content input by the user;
[0015] The voice output unit is used to collect user input voice data;
[0016] The conversion unit is used to convert text content or voice data into instructions that can be recognized by the system;
[0017] The reply unit is used to reply to the user according to the generated reply content.
[0018] The NLP understanding module includes a semantic parsing unit, an analysis unit, and a multi-round dialogue unit;
[0019] The semantic parsing unit performs word segmentation, part-of-speech tagging, and dependency analysis on the text input by the user to extract key semantics and context information;
[0020] The analysis unit is used to analyze the emotional state of the user during the conversation;
[0021] The multi-round dialogue unit is used to keep a memory of the user's historical input during the dialogue process, ensuring a coherent response in complex, multi-round interactions.
[0022] The generative AI module includes an emotion perception generation unit and a personalized generation unit;
[0023] The emotion perception generation unit is used to generate an emotion-adaptive response based on the emotion analysis result;
[0024] The personalized generation unit is used to integrate user personalized data to generate customized conversation content.
[0025] The intelligent recommendation module includes a collaborative filtering unit, a content recommendation unit and a scene adaptive recommendation unit;
[0026] The collaborative filtering unit is used to make recommendations based on user behavior data;
[0027] The content recommendation unit is used to generate recommendations based on real-time user interaction data and preferences;
[0028] The scenario-adaptive recommendation unit is used to dynamically generate recommended content according to the conversation scenario.
[0029] Wherein, the speech recognition and synthesis module includes a speech recognition unit and a speech synthesis unit;
[0030] The speech recognition unit is used to convert speech into text;
[0031] The speech synthesis unit converts text into speech.
[0032] The scenario-adaptive speech generation library includes a speech template management unit and a generative speech expansion unit;
[0033] The speech template management unit is used to store and manage speech templates for various sales scenarios;
[0034] The generative speech expansion unit is used to generate new candidate speech when there is no matching template.
[0035] The feedback analysis module includes a behavior data collection unit, a model self-optimization unit and a continuous learning unit;
[0036] The behavior data collection unit is used to record user behavior data;
[0037] The model self-optimization unit is used to adjust the recommendation logic and generation strategy based on user feedback;
[0038] The continuous learning unit is used to continuously optimize the generative large model through user feedback.
[0039] In a second aspect, a data processing method for a model-based virtual customer service response system is provided, which is used in the model-based virtual customer service response system described in the first aspect, and includes the following steps:
[0040] Receive user input through the user interaction module and convert the user's text or voice request into instructions that the system can recognize;
[0041] Use the NLP understanding module to perform semantic analysis, intent recognition, and context understanding on user input text or voice;
[0042] Based on the generative AI module, the generative model is called to generate natural and fluent language responses based on user input and contextual information, and intelligently respond to customer needs;
[0043] The intelligent recommendation module analyzes users' historical behaviors, real-time interactions, and current market trends to intelligently generate personalized product or service recommendations.
[0044] If the user input is voice, the speech recognition and synthesis module converts the speech into text for system processing, and converts the generated text response into voice output;
[0045] Based on different sales scenarios, the scenario-adaptive script generation library is called to generate personalized sales scripts;
[0046] Monitor user feedback through the feedback analysis module and adjust the generation strategy and recommendation logic.
[0047] A virtual customer service response system constructed by a model of the present invention includes a user interaction module, an NLP understanding module, a generative AI module, an intelligent recommendation module, a speech recognition and synthesis module, a scene-adaptive speech generation library and a feedback analysis module. The user interaction module is connected to the feedback analysis module and the NLP understanding module, the generative AI module and the intelligent recommendation module are respectively connected to the NLP understanding module, the speech recognition and synthesis module and the scene-adaptive speech generation library are respectively connected to the generative AI module; the user interaction module is used to process the user's input and convert the user's text or voice request into instructions that the system can recognize; the NLP understanding module, based on natural language processing technology, processes the user's input The system performs semantic analysis, intent recognition, and contextual understanding on input text or speech. The generative AI module, based on user input and contextual information, invokes a generative model to generate natural, fluent verbal responses and intelligently responds to customer needs. The intelligent recommendation module generates personalized product or service recommendations by analyzing user historical behavior, real-time interactions, and current market trends. The speech recognition and synthesis module converts user speech into text for system processing and supports converting generated text responses into speech output. The scenario-adaptive speech generation library uses the speech library to generate personalized sales scripts based on different sales scenarios. The feedback analysis module monitors user feedback and adjusts the generation strategy and recommendation logic. This invention deeply understands user needs and, through contextual semantic analysis, generates natural, coherent multi-round conversations, solving complex problems and providing high-quality customer service. Based on user behavior data and real-time interaction information, the system generates personalized product or service recommendations, flexibly adjusts sales strategies, and improves customer satisfaction and sales conversion rates. By leveraging generative AI's natural language generation technology and advanced speech recognition and synthesis capabilities, the system provides a more natural and smooth human-computer interaction experience, simulating the interaction effects of real-person customer service. This solves the problem that existing systems lack deep semantic understanding and are unable to handle complex customer interactions. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0049] Figure 1 This is a schematic diagram of a virtual customer service response system constructed by a model provided by the present invention.
[0050] Figure 2 It is a schematic diagram of the user interaction module.
[0051] Figure 3 It is a schematic diagram of the NLP understanding module.
[0052] Figure 4 This is a schematic diagram of the generative AI module.
[0053] Figure 5 This is a schematic diagram of the intelligent recommendation module.
[0054] Figure 6 This is a schematic diagram of the speech recognition and synthesis module.
[0055] Figure 7 This is a schematic diagram of the scenario-adaptive speech generation library.
[0056] Figure 8 It is a schematic diagram of the feedback analysis module.
[0057] Figure 9 The present invention provides a data processing method for a virtual customer service response system based on a model construction.
[0058] In the figure: 1-user interaction module, 2-NLP understanding module, 3-generative AI module, 4-intelligent recommendation module, 5-speech recognition and synthesis module, 6-scenario adaptive speech generation library, 7-feedback analysis module, 11-text input unit, 12-speech output unit, 13-conversion unit, 14-reply unit, 21-semantic parsing unit, 22-analysis unit, 23-multi-round dialogue unit, 31-emotional perception generation unit, 32-personalized generation unit, 41-collaborative filtering unit, 42-content recommendation unit, 43-scenario adaptive recommendation unit, 51-speech recognition unit, 52-speech synthesis unit, 61-speech template management unit, 62-generative speech expansion unit, 71-behavioral data collection unit, 72-model self-optimization unit, 73-continuous learning unit. DETAILED DESCRIPTION
[0059] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.
[0060] See also Figures 1 to 8In the first aspect, the present invention provides a model-built virtual customer service response system, comprising a user interaction module 1, an NLP understanding module 2, a generative AI module 3, an intelligent recommendation module 4, a speech recognition and synthesis module 5, a scene-adaptive speech generation library 6, and a feedback analysis module 7. The user interaction module 1 is connected to the feedback analysis module 7 and the NLP understanding module 2, the generative AI module 3 and the intelligent recommendation module 4 are respectively connected to the NLP understanding module 2, and the speech recognition and synthesis module 5 and the scene-adaptive speech generation library 6 are respectively connected to the generative AI module 3;
[0061] The user interaction module 1 is used to process user input and convert the user's text or voice request into instructions that can be recognized by the system;
[0062] The NLP understanding module 2 performs semantic analysis, intent recognition, and context understanding on the text or voice input by the user based on natural language processing technology;
[0063] The generative AI module 3 uses a generative model to generate natural and fluent language responses based on user input and contextual information, and intelligently responds to customer needs;
[0064] The intelligent recommendation module 4 generates personalized product or service recommendations by analyzing users' historical behaviors, real-time interactions, and current market trends;
[0065] The speech recognition and synthesis module 5 converts the user's speech into text for system processing, and also supports converting the generated text reply into speech output;
[0066] The scenario-adaptive speech generation library 6, according to different sales scenarios, the system calls the speech library to generate personalized sales speech;
[0067] The feedback analysis module 7 monitors user feedback and adjusts generation strategies and recommendation logic.
[0068] In this embodiment, the User Interaction Module 1 processes user input, supporting both text and voice input. It converts user text or voice requests into system-recognizable commands and outputs generated responses to the user. The NLP Understanding Module 2 leverages natural language processing technology to perform semantic analysis, intent recognition, and contextual understanding on user text or voice input, laying the foundation for generating appropriate responses or recommendations. The Generative AI Module 3 utilizes a generative model based on user input and contextual information to generate natural, fluent responses and intelligently address customer needs. The Intelligent Recommendation Module 4 analyzes user historical behavior, real-time interactions, and current market trends to intelligently generate personalized product or service recommendations. The Speech Recognition and Synthesis Module 5 recognizes user voice input, converts speech into text for system processing, and converts generated text responses into speech output, enhancing the user experience. The Scenario-Adaptive Script Generation Library 6 uses a script library to generate personalized sales scripts based on different sales scenarios, improving recommendation effectiveness. The Feedback Analysis Module 7 monitors user feedback, adjusts generation strategies and recommendation logic, and optimizes the user experience. This invention can deeply understand user needs and, through contextual semantic analysis, generate natural, coherent multi-round conversations to solve complex problems and provide high-quality customer service. Based on user behavior data and real-time interaction information, the system generates personalized product or service recommendations, flexibly adjusts sales strategies, and improves customer satisfaction and sales conversion rates. Through generative AI's natural language generation technology and advanced speech recognition and synthesis capabilities, the system provides a more natural and fluid human-computer interaction experience, simulating the interactive effects of real-person customer service. This solves the problem that existing systems lack deep semantic understanding and are unable to handle complex customer interactions.
[0069] Furthermore, the user interaction module 1 includes a text input unit 11, a voice output unit 12, a conversion unit 13 and a reply unit 14, wherein the text input unit 11 and the voice output unit 12 are respectively connected to the conversion unit 13;
[0070] The text input unit 11 is used to collect text content input by the user;
[0071] The voice output unit 12 is used to collect user input voice data;
[0072] The conversion unit 13 is used to convert text content or voice data into system-recognizable instructions;
[0073] The reply unit 14 is configured to reply to the user according to the generated reply content.
[0074] In this embodiment, the user interaction module 1 is the core entry point for the system to interact with the user, and supports two-way input and output of text and voice. This module integrates multiple input methods: Text mode: Users can interact with the system by typing or clicking on predefined options. Text input is captured in real time and sent to the NLP understanding module 2 for analysis. Voice mode: The user's voice input is processed by the voice recognition module. Voice recognition uses an acoustic model to convert the voice signal into high-precision text in real time to ensure that the user's intention is accurately captured. The system converts the generated text back into voice output through speech synthesis technology (Tacotron or WaveNet), providing a smooth voice interaction experience. Real-time feedback mechanism: After receiving user input, the system will immediately call the generative large model for processing, quickly generate personalized replies, and ensure that users receive real-time responses. This module focuses on multi-channel support, allowing users to interact seamlessly on different platforms and devices, and dynamically adjusts the output form according to the user's input method (voice or text) to ensure flexibility and convenience of interaction. User Interaction Module 1 serves as the core interface for system-user interaction, supporting both text and voice interaction. The specific technical implementation is as follows: Data Acquisition: In text mode, the front-end page's input box monitors user input events and transmits the input content to the back-end service in real time. In voice mode, a microphone is used to capture audio data, and the audio stream is transmitted to the server using WebRTC technology. Data Processing: Text data is directly passed to NLP Understanding Module 2. After entering the speech recognition module, the audio data undergoes preprocessing, including noise reduction and normalization. Taking the DeepSpeech model as an example, the processing flow is as follows: Audio data is converted into a mel-spectrogram and input into a neural network consisting of convolutional, recurrent, and fully connected layers. The CTC (Connectionist Temporal Classification) algorithm calculates the most likely text sequence. Data Output: After the system generates a response, the text response is displayed directly on the front-end page. Voice responses are converted into an audio stream using the speech synthesis module, which is then played to the user via the Web Audio API.
[0075] Furthermore, the NLP understanding module 2 includes a semantic parsing unit 21, an analysis unit 22 and a multi-round dialogue unit 23;
[0076] The semantic parsing unit 21 performs word segmentation, part-of-speech tagging, and dependency analysis on the text input by the user to extract key semantics and context information;
[0077] The analysis unit 22 is used to analyze the user's emotional state during the conversation;
[0078] The multi-round dialogue unit 23 is used to keep a memory of the user's historical input during the dialogue process, ensuring a coherent response in complex, multi-round interactions.
[0079] In this embodiment, the NLP understanding module 2 is the key component in the system responsible for semantic analysis and intent recognition. It adopts deep learning models such as BERT and Transformer architecture to ensure deep understanding of user input: Semantic parsing: Through deep learning models, the text input by the user is segmented, POS tagged, and dependency analyzed to extract key semantics and contextual information. Pre-trained models such as BERT have strong context capture capabilities when processing natural language understanding tasks, and can accurately identify the user's intention based on the input conversation content. Sentiment analysis: Based on the sentiment classification model, the system uses LSTM or Transformer to classify the input sentiment and analyze the user's emotional state in the conversation, such as anger, satisfaction, confusion, etc. This emotional information will be passed to the generative AI module 3 to generate a response that is more in line with the user's emotional state. Multi-round dialogue processing: The NLP understanding module 2 supports context tracking of multi-round dialogues. By keeping a memory of the user's historical input in the dialogue process, it ensures a coherent response in complex, multi-round interactions. The NLP understanding module 2 implements semantic parsing, sentiment analysis and multi-round dialogue processing based on a deep learning model. The specific implementation method is as follows:
[0080] Word segmentation and syntactic analysis:
[0081] Use the Spacy library for word segmentation, part-of-speech tagging (POStagging), and dependency syntax analysis to extract subject-verb-object structures and named entities (NER).
[0082] For example: "I want to book a flight to Beijing on Friday" → Entity recognition is time: Friday, location: Beijing, and dependency analysis extracts the subject-object relationship booking-flight ticket.
[0083] Context Modeling:
[0084] The BERT model (pre-trained version is bert-base-chinese) is used. The input is the current sentence + the last three rounds of dialogue history (separated by [SEP]), and the output is a 768-dimensional representation.
[0085] Intent classification: Add a fully connected layer + Softmax on top of BERT to output intent labels (such as "buy a ticket" and "complaint").
[0086] Sentiment Analysis:
[0087] Model architecture: Transformer-based bidirectional LSTM, input is BERT sentence vector, and output is sentiment polarity (positive / negative) and intensity (0-1 score).
[0088] Training data: Fine-tuning is performed using the public dataset ChnSentiCorp, and the loss function is cross entropy.
[0089] Multi-turn dialogue management:
[0090] Use Redis to cache the conversation state, using session_id as the key to store the following information:
[0091] json
[0092] {
[0093] "history":["User: Booking Ticket","System: What is your destination?",...],
[0094] "slots":{"time":"Friday","location":"Beijing"}
[0095] }
[0096] The DQN (deep reinforcement learning) model is used to decide whether to ask further questions or jump to other intentions.
[0097] Furthermore, the generative AI module 3 includes an emotion perception generation unit 31 and a personalized generation unit 32;
[0098] The emotion perception generating unit 31 is used to generate an emotion adaptive response according to the emotion analysis result;
[0099] The personalized generation unit 32 is used to integrate user personalized data to generate customized conversation content.
[0100] In this implementation, the LLaMA framework, a lightweight, large-scale language model architecture, generates high-quality text content with minimal computing resources. Through supervised fine-tuning on a variety of conversational datasets from customer service scenarios, the model can quickly adapt to various user needs and generate accurate and contextually appropriate responses.
[0101] Emotion-aware generation: The model not only understands the user's specific needs but also generates emotionally adaptive conversational content based on sentiment analysis results (provided by the NLP module). For example, when a user expresses confusion or dissatisfaction, the system generates more soothing and solution-oriented responses.
[0102] Personalized Generation: The LLaMA model’s fine-tuning process incorporates users’ personalized data, enabling it to generate customized conversational content based on user preferences, thereby providing a more engaging customer experience.
[0103] Generative AI module 3 implements semantic, emotional, and multi-round dialogue context understanding based on an autoregressive generative model. The specific implementation is as follows:
[0104] Model architecture, based on the LLaMA-7B architecture:
[0105] Added a cross-modal attention layer and supported sentiment vectors (from the NLP module) as additional input.
[0106] Use LoRA (low-rank adaptation) technology to fine-tune customer service data to reduce training overhead.
[0107] Emotional perception generation
[0108] During generation, the emotional label (such as "confused") is converted into a prefix vector (Prompt-tuning) and concatenated to the input text to guide the model to generate soothing responses.
[0109] Example: User input: [Confused]: Why hasn't my order arrived? → Output: We apologize for the inconvenience. We have contacted the logistics department and expect to update the information within 2 hours.
[0110] Personalized generation:
[0111] User portraits are stored in MongoDB, including historical orders, click behaviors, etc. When generated, relevant data is dynamically retrieved through the Key-Value Memory Network (Key-Value Memory Networks) as a generation condition.
[0112] Furthermore, the intelligent recommendation module 4 includes a collaborative filtering unit 41, a content recommendation unit 42 and a scene adaptive recommendation unit 43;
[0113] The collaborative filtering unit 41 is used to make recommendations based on user behavior data;
[0114] The content recommendation unit 42 is used to generate recommendations based on real-time user interaction data and preferences;
[0115] The scenario-adaptive recommendation unit 43 is configured to dynamically generate recommended content according to the conversation scenario.
[0116] In this embodiment, the intelligent recommendation module 4 provides users with real-time, personalized product or service recommendations by combining user behavior data analysis and recommendation algorithms:
[0117] Based on collaborative filtering: The system uses a collaborative filtering algorithm to analyze the behavioral data of similar users and recommend items that other users like. This algorithm quickly discovers user preferences by building a correlation matrix between users and items.
[0118] Content recommendation: For new users or special scenarios, the system adopts a content-based recommendation strategy, combining users' real-time interaction data and personalized preferences to generate product recommendations.
[0119] Scenario-adaptive recommendation: Based on the current conversation scenario (such as promotions, holiday events, user needs), the system dynamically generates recommendations suitable for the current context through the generative AI module 3. The recommendation module works in conjunction with the scenario-adaptive speech generation library 6 to ensure that the recommended content is consistent with the speech. Specific implementation methods:
[0120] Collaborative filtering:
[0121] Construct a user-item matrix, decompose it into a latent factor matrix (dimension k=64) using the ALS algorithm (alternating least squares), and calculate similar user groups in real time.
[0122] Cold start processing: For new users, popular items (Top-N, sorted by weekly sales) are recommended for the first time.
[0123] Recommended content:
[0124] Item feature extraction: Use TF-IDF to extract text features (product descriptions), and ResNet-50 to extract image features, which are then fused into a 512-dimensional vector.
[0125] Calculate cosine similarity to match keywords in users' real-time conversations (e.g., "summer dress" → recommend similar styles).
[0126] Scene Adaptation:
[0127] Define scenario conditions through the rule engine (Drools), for example:
[0128] rule"Double Eleven Promotion"
[0129] when
[0130] $s:Scene(type=="promotion",timein"2023-11-11")
[0131] then
[0132] recommendTemplate.set("discount_template");
[0133] end
[0134] Combining templates with generative models: First, retrieve the placeholders (such as ${product name}) in the template discount_template, and then call LLaMA to generate the complete script.
[0135] Furthermore, the speech recognition and synthesis module 5 includes a speech recognition unit 51 and a speech synthesis unit 52;
[0136] The speech recognition unit 51 is used to convert speech into text;
[0137] The speech synthesis unit 52 converts text into speech.
[0138] In this embodiment, the speech recognition and synthesis module 5 integrates DeepSpeech and WaveNet / Tacotron technologies to ensure efficient speech processing capabilities:
[0139] Speech Recognition: The speech recognition architecture combines deep neural networks (DNN) and convolutional neural networks (CNN), which can effectively handle background noise and provide accurate speech-to-text services in complex environments.
[0140] Speech synthesis: Using WaveNet or Tacotron technology, the generated text responses are converted into natural and fluent speech output. The synthesis module can adjust the timbre and intonation of the speech based on the user's context or emotional state, making the generated voice responses more realistic and humane. Specific implementations include:
[0141] Speech recognition optimization
[0142] Noise robustness: A Squeeze-Excitation Network (SE) module is added to the DeepSpeech front-end to dynamically enhance the voice band.
[0143] Streaming processing: Block inference is used (recognition is triggered every 500ms of audio stream) and context coherence is maintained through CTC prefix beam search (PrefixBeamSearch).
[0144] Speech synthesis emotion adaptation
[0145] The emotional label (such as "happy") is spliced at the output layer of the Tacotron2 encoder, and the pitch and speaking rate are controlled (adjusted by DurationPredictor).
[0146] The GAN-trained adversarial network HiFi-GAN improves the naturalness of synthesis, with the MOS score increasing from 3.8 to 4.2.
[0147] Furthermore, the scenario-adaptive speech generation library 6 includes a speech template management unit 61 and a generative speech expansion unit 62;
[0148] The speech template management unit 61 is used to store and manage speech templates for various sales scenarios;
[0149] The generative speech expansion unit 62 is used to generate new candidate speech when there is no matching template.
[0150] In this embodiment, the scenario-adaptive speech generation library 6 is an important module in this system, which can generate flexible recommendation speech based on different sales scenarios:
[0151] Design of speech script library: The speech script generation library pre-stores speech script templates for various common sales scenarios (such as promotional discounts, holiday-specific activities, etc.), and can be dynamically called according to actual dialogue scenarios.
[0152] Generative AI Assistance: With the support of generative AI models (such as LLaMA), the system can not only call preset dialogues, but also generate new dialogues based on the current conversation. For example, when a user expresses interest in a specific product, the system can generate personalized dialogue templates in real time. The specific implementation is:
[0153] Script template management
[0154] json
[0155] {
[0156] "scene":"promotion",
[0157] "template":"Limited time special offer! ${item} direct discount of ${discount}, click to buy → ${link}",
[0158] "rules":{"discount":"0.3","link":"retrieve URL from product table"}
[0159] }
[0160] Dynamic filling: Use SQL to query the product database to obtain ${link}, and LLaMA generates the description text of ${product}.
[0161] Generative speech extension:
[0162] If there is no matching scenario in the template library, LLaMA is called to generate candidate speech (top-3 sampling), which are then put into the database after manual review, forming a closed-loop iteration.
[0163] Furthermore, the feedback analysis module 7 includes a behavior data collection unit 71, a model self-optimization unit 72 and a continuous learning unit 73;
[0164] The behavior data collection unit 71 is used to record user behavior data;
[0165] The model self-optimization unit 72 is used to adjust the recommendation logic and generation strategy according to user feedback;
[0166] The continuous learning unit 73 is used to continuously optimize the generative large model through user feedback.
[0167] In this embodiment, the feedback analysis module 7 continuously optimizes the system's recommendation and generation strategies by real-time monitoring of user behavior and interaction data:
[0168] Behavioral data collection: The system records all behavioral data of users during the interaction process, including clicks, browsing time, conversation content, emotional changes, etc. These data will be transmitted to the recommendation algorithm and generative model in real time for updating.
[0169] Model self-optimization: By analyzing user feedback, the system adjusts its recommendation logic and generation strategy. For example, if a certain category of recommended products has a low click-through rate, the system will reduce the frequency of recommendations for that category, and vice versa.
[0170] Continuous Learning: Generative large models are continuously fine-tuned and optimized through feedback data, enabling them to handle different types of user needs, specifically:
[0171] Data collection
[0172] Tracking design: Monitor click, mouseover, and other events on the front end, transmit them to the Kafka queue through Logstash, and finally store them in Elasticsearch.
[0173] Conversation log: records the complete conversation flow (including timestamp, user ID, and sentiment label) for offline analysis.
[0174] Real-time optimization
[0175] A / B testing: Divide users into groups with different recommendation strategies (e.g., Group A uses collaborative filtering, Group B uses content recommendation), and calculate click-through rate (CTR) and conversion rate (CVR).
[0176] Model update: Every morning, Spark is used to calculate user behavior data, generate a new collaborative filtering matrix, and then hot-update it to the recommendation module.
[0177] Continuous Learning
[0178] Active Learning: Generate responses with low confidence (the variance is calculated through Monte Carlo Dropout) and push them to the manual annotation platform, where they are added to the training set after annotation.
[0179] Incremental training: LLaMA is incrementally fine-tuned based on new data every week (learning rate is set to 1e-5, batch size is 32). It shows higher flexibility and accuracy.
[0180] See also Figure 9 In a second aspect, a data processing method for a model-based virtual customer service response system is provided, which is used in the model-based virtual customer service response system described in the first aspect, and includes the following steps:
[0181] S1 receives user input through the user interaction module 1 and converts the user's text or voice request into a command that the system can recognize;
[0182] Specifically, the user interaction module 1 serves as the core entry point for the system to interact with the user, supporting two-way input and output of text and voice, and includes a text input unit 11, a voice output unit 12, a conversion unit 13, and a reply unit 14. The text input unit 11 is used to collect text content input by the user, and the voice output unit 12 is used to collect voice data input by the user. The conversion unit 13 is responsible for converting text content or voice data into instructions that the system can recognize. In text mode, the user input event is monitored through the input box of the front-end page, and the input content is transmitted to the back-end service in real time; in voice mode, the microphone is used to collect audio data, and the audio stream is transmitted to the server using WebRTC technology. For text data, it is directly passed to the NLP understanding module 2; after the audio data enters the speech recognition module, it is first preprocessed, including noise reduction, normalization and other operations. Taking the DeepSpeech model as an example, the audio data is converted into a Mel-spectrogram and input into a neural network containing a convolutional layer, a recurrent layer, and a fully connected layer. The most likely text sequence is calculated using the CTC algorithm. After the system generates a reply, the text reply is directly displayed on the front-end page; the voice reply uses the speech synthesis module to convert the text into an audio stream, which is then played to the user through the Web audio API.
[0183] S2 uses NLP understanding module 2 to perform semantic analysis, intent recognition, and context understanding on the text or voice input by the user;
[0184] Specifically, the NLP understanding module 2 is a key component in the system responsible for semantic analysis and intent recognition. It uses deep learning models such as BERT and Transformer architecture to ensure a deep understanding of user input. It includes a semantic parsing unit 21, an analysis unit 22, and a multi-round dialogue unit 23. The semantic parsing unit 21 performs word segmentation, part-of-speech tagging, and dependency analysis on the text input by the user, extracts key semantics and contextual information, and uses the Spacy library for word segmentation, part-of-speech tagging, and dependency syntax analysis to extract subject-verb-object structures and named entities. The analysis unit 22 is used to analyze the user's emotional state in the conversation. Based on the sentiment classification model, it uses LSTM or Transformer to perform sentiment classification on the input. The model architecture is a bidirectional LSTM based on Transformer, with BERT sentence vectors as input and sentiment polarity and intensity as output. It is fine-tuned using the public dataset ChnSentiCorp, and the loss function is cross entropy. The multi-round dialogue unit 23 is used to maintain memory of user historical input during the dialogue process to ensure coherent responses in complex, multi-round interactions. It uses the BERT model, with the input being the current sentence + the history of the last three rounds of dialogue, separated by [SEP], and outputting a 768-dimensional representation. A fully connected layer + Softmax is added to the top layer of BERT to output the intent label. Redis is used to cache the dialogue state, and the dialogue history and slot information are stored with session_id as the key. The DQN model is used to decide whether to ask a follow-up question or jump to other intents.
[0185] S3 is based on the generative AI module 3. Based on the user's input and contextual information, it calls the generative model to generate natural and fluent language responses and intelligently respond to customer needs.
[0186] Specifically, the generative AI module 3 is based on the LLaMA framework. By fine-tuning a variety of conversation data sets in customer service scenarios in a supervised manner, it can quickly adapt to various user needs and generate accurate and contextual responses. It includes an emotion perception generation unit 31 and a personalized generation unit 32. The emotion perception generation unit 31 generates an emotion-adaptive response based on the results of the emotion analysis, converts the emotion label into a prefix vector, and splices it before the input text to guide the model to generate a soothing response. The personalized generation unit 32 integrates user personalized data to generate customized conversation content. The user portrait is stored in MongoDB, including historical orders, click behaviors, etc. During generation, relevant data is dynamically retrieved through a key-value memory network as a generation condition. The model architecture is based on the LLaMA-7B architecture, adds a cross-modal attention layer, supports emotion vectors as additional input, and uses LoRA technology to fine-tune customer service field data to reduce training overhead.
[0187] S4 analyzes users’ historical behaviors, real-time interactions, and current market trends through the intelligent recommendation module 4 to intelligently generate personalized product or service recommendations;
[0188] Specifically, the intelligent recommendation module 4 combines user behavior data analysis and recommendation algorithms to provide users with product or service recommendations that better meet their personalized needs. It includes a collaborative filtering unit 41, a content recommendation unit 42, and a scene-adaptive recommendation unit 43. The collaborative filtering unit 41 makes recommendations based on user behavior data, constructs a user-item matrix, decomposes it into a latent factor matrix using the ALS algorithm, calculates similar user groups in real time, and recommends popular items to new users for the first time during cold start processing. The content recommendation unit 42 generates recommendations based on real-time user interaction data and preferences, uses TF-IDF to extract text features, ResNet-50 to extract image features, fuses them into a 512-dimensional vector, calculates cosine similarity, and matches keywords in the user's real-time conversation. The scene-adaptive recommendation unit 43 dynamically generates recommended content based on the conversation scene, defines scene conditions through a rule engine, combines templates with generative models, first retrieves placeholders in the template, and then calls LLaMA to generate complete conversations.
[0189] S5 If the user input is voice, the voice recognition and synthesis module 5 converts the voice into text for system processing, and converts the generated text reply into voice output;
[0190] Specifically, the speech recognition and synthesis module 5 ensures efficient speech processing capabilities by integrating DeepSpeech and WaveNet / Tacotron technologies. The speech recognition unit 51 adopts a speech recognition architecture that combines deep neural networks and convolutional neural networks. It can effectively handle background noise and provide accurate speech-to-text services in complex environments. The front-end adds an SE module to dynamically enhance the voice band, adopts block reasoning, and maintains context coherence through CTC prefix beam search. The speech synthesis unit 52 converts the generated text responses into natural and fluent speech output through WaveNet or Tacotron technology, splices emotional labels at the encoder output layer of Tacotron2, controls pitch and speaking speed, and uses GAN training to train the adversarial network HiFi-GAN to improve the naturalness of synthesis.
[0191] S6 calls the scenario-adaptive speech generation library 6 to generate personalized sales speech according to different sales scenarios;
[0192] Specifically, the scenario-adaptive speech generation library 6 is a key module in the system, capable of generating flexible recommended speech based on different sales scenarios. It includes a speech template management unit 61 and a generative speech expansion unit 62. The speech template management unit 61 pre-stores speech templates for various common sales scenarios and can dynamically call them based on the actual conversation scenario. It queries the product database through SQL to obtain links, and LLaMA generates product descriptions. The generative speech expansion unit 62 generates new candidate speech when no matching template exists, calls LLaMA to generate candidate speech, and stores them after manual review, forming a closed-loop iteration.
[0193] S7 monitors user feedback through the feedback analysis module 7 and adjusts the generation strategy and recommendation logic.
[0194] Specifically, the feedback analysis module 7 continuously optimizes the system's recommendation and generation strategies through real-time monitoring of user behavior and interaction data. It includes a behavior data collection unit 71, a model self-optimization unit 72, and a continuous learning unit 73. The behavior data collection unit 71 records all behavioral data of users during the interaction process, including clicks, browsing time, conversation content, emotional changes, etc. These data will be transmitted to the recommendation algorithm and generative model in real time for updating. The model self-optimization unit 72 adjusts the recommendation logic and generation strategy based on user feedback. When the click-through rate of a certain type of recommended product is low, the recommendation frequency of this type of product is reduced, and vice versa. The continuous learning unit 73 continuously optimizes the generative large model through user feedback, uses active learning, pushes low-confidence generated responses to the manual annotation platform, adds them to the training set after annotation, and performs incremental fine-tuning on LLaMA based on new data every week.
[0195] The above disclosure is merely a preferred embodiment of a virtual customer service response system and data processing method constructed by a model of the present invention. It is certainly not intended to limit the scope of rights of the present invention. A person skilled in the art can understand that all or part of the processes of the above embodiments and equivalent changes made in accordance with the claims of the present invention still fall within the scope of the invention.
Claims
1. A virtual customer service response system built by a model, characterized in that: It includes a user interaction module, an NLP understanding module, a generative AI module, an intelligent recommendation module, a speech recognition and synthesis module, a scene-adaptive speech generation library and a feedback analysis module. The user interaction module is connected to the feedback analysis module and the NLP understanding module, the generative AI module and the intelligent recommendation module are respectively connected to the NLP understanding module, and the speech recognition and synthesis module and the scene-adaptive speech generation library are respectively connected to the generative AI module; The user interaction module is used to process user input and convert the user's text or voice request into instructions that can be recognized by the system; The NLP understanding module performs semantic analysis, intent recognition, and context understanding on the text or voice input by the user based on natural language processing technology; The generative AI module uses a generative model to generate natural and fluent language responses based on user input and contextual information, and intelligently responds to customer needs. The intelligent recommendation module generates personalized product or service recommendations by analyzing users' historical behavior, real-time interactions, and current market trends; The speech recognition and synthesis module converts the user's speech into text for system processing, and also supports converting the generated text replies into speech output; The scenario-adaptive speech generation library generates personalized sales speech based on different sales scenarios. The feedback analysis module monitors user feedback and adjusts generation strategies and recommendation logic.
2. The virtual customer service response system constructed by the model according to claim 1, characterized in that: The user interaction module includes a text input unit, a voice output unit, a conversion unit and a reply unit, wherein the text input unit and the voice output unit are respectively connected to the conversion unit; The text input unit is used to collect text content input by the user; The voice output unit is used to collect user input voice data; The conversion unit is used to convert text content or voice data into instructions that can be recognized by the system; The reply unit is used to reply to the user according to the generated reply content.
3. The virtual customer service response system constructed by the model according to claim 1, characterized in that: The NLP understanding module includes a semantic parsing unit, an analysis unit, and a multi-round dialogue unit; The semantic parsing unit performs word segmentation, part-of-speech tagging, and dependency analysis on the text input by the user to extract key semantics and context information; The analysis unit is used to analyze the emotional state of the user during the conversation; The multi-round dialogue unit is used to keep a memory of the user's historical input during the dialogue process, ensuring a coherent response in complex, multi-round interactions.
4. The virtual customer service response system constructed by the model according to claim 1, characterized in that: The generative AI module includes an emotion perception generation unit and a personalized generation unit; The emotion perception generation unit is used to generate an emotion-adaptive response based on the emotion analysis result; The personalized generation unit is used to integrate user personalized data to generate customized conversation content.
5. The virtual customer service response system constructed by the model according to claim 1, characterized in that: The intelligent recommendation module includes a collaborative filtering unit, a content recommendation unit and a scene adaptive recommendation unit; The collaborative filtering unit is used to make recommendations based on user behavior data; The content recommendation unit is used to generate recommendations based on real-time user interaction data and preferences; The scenario-adaptive recommendation unit is used to dynamically generate recommended content according to the conversation scenario.
6. The virtual customer service response system constructed by the model according to claim 1, characterized in that: The speech recognition and synthesis module includes a speech recognition unit and a speech synthesis unit; The speech recognition unit is used to convert speech into text; The speech synthesis unit converts text into speech.
7. The virtual customer service response system constructed by the model according to claim 1, characterized in that: The scenario-adaptive speech generation library includes a speech template management unit and a generative speech expansion unit; The speech template management unit is used to store and manage speech templates for various sales scenarios; The generative speech expansion unit is used to generate new candidate speech when there is no matching template.
8. The virtual customer service response system constructed by the model according to claim 1, characterized in that: The feedback analysis module includes a behavior data collection unit, a model self-optimization unit and a continuous learning unit; The behavior data collection unit is used to record user behavior data; The model self-optimization unit is used to adjust the recommendation logic and generation strategy based on user feedback; The continuous learning unit is used to continuously optimize the generative large model through user feedback.
9. A data processing method for a model-based virtual customer service response system, used in the model-based virtual customer service response system according to any one of claims 1 to 8, characterized in that: The following steps are involved: Receive user input through the user interaction module and convert the user's text or voice request into instructions that the system can recognize; Use the NLP understanding module to perform semantic analysis, intent recognition, and context understanding on user input text or voice; Based on the generative AI module, the generative model is called to generate natural and fluent language responses based on user input and contextual information, and intelligently respond to customer needs; The intelligent recommendation module analyzes users' historical behaviors, real-time interactions, and current market trends to intelligently generate personalized product or service recommendations. If the user input is voice, the speech recognition and synthesis module converts the speech into text for system processing, and converts the generated text response into voice output; Based on different sales scenarios, the scenario-adaptive script generation library is called to generate personalized sales scripts; Monitor user feedback through the feedback analysis module and adjust the generation strategy and recommendation logic.