Online customer service response method and device, equipment, storage medium and product

By combining classification and summary models with large models to generate response text, the problem of online customer service being unable to cover all issues is solved, and accurate understanding and response to user intentions are achieved, improving response accuracy and user experience.

CN120687562APending Publication Date: 2025-09-23CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510727693.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing online customer service relies on a pre-set question and answer library, which cannot cover all user questions and makes it impossible to provide accurate answers.

Method used

By obtaining user input information, using classification models to identify information types, and combining text extraction models, summary extraction models and large models to generate reply text, we can achieve accurate understanding and response to user intentions and avoid relying on pre-configured question and answer libraries.

Benefits of technology

The accuracy of online customer service responses has been improved, enabling the large model to respond based on any question information, generate more accurate response text, and enhance user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687562A_ABST
    Figure CN120687562A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an online customer service response method and device, equipment, a storage medium and a product. According to the embodiment of the invention, user input information is obtained, an information type corresponding to the user input information is identified through a classification model, and the user input information and the information type corresponding to the user input information are sent to a character extraction model; performing character extraction on the user input information and the information type corresponding to the user input information by using the character extraction model to obtain text total information, sending the user input information to the abstract extraction model, performing abstract extraction on the text total information by using the abstract extraction model to obtain abstract information, and sending the abstract information to the user input information. The summary information and the text total information are sent to the large model, so that the large model obtains the reply text and the response template according to the summary information and the text total information, the reply text is determined as the target reply according to the response template, and the reply accuracy of the online customer service can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, and in particular relates to an online customer service response method, device, equipment, storage medium and product. Background Art

[0002] Currently, online customer service on mobile applications can receive user questions and output corresponding answers. For example, after receiving a user's question, the customer service will query the answer corresponding to the question from a pre-set question and answer library and output it for display.

[0003] In the prior art, online customer service relies on a pre-set question-and-answer library to answer user questions. When a user asks a question to the online customer service, the online customer service provides a reply by matching the user's question with the questions in the library.

[0004] However, the content of the Q&A database is limited and cannot cover all questions asked by users. Therefore, when users ask new questions, online customer service cannot give accurate answers. Summary of the Invention

[0005] Embodiments of the present invention provide an online customer service response method, apparatus, device, storage medium, and product, which can improve the accuracy of online customer service responses.

[0006] In a first aspect, an embodiment of the present invention provides an online customer service response method, the method comprising:

[0007] Obtaining user input information, inputting the user input information into a classification model, and identifying the information type of the user input information through the classification model;

[0008] Inputting the user input information and the information type into a text extraction model, performing text extraction using the text extraction model to obtain full text information;

[0009] Inputting the user input information into a summary extraction model, performing summary extraction by the summary extraction model to obtain summary information;

[0010] Input the summary information and the full text information into a large model, and generate a reply text and a response template through the large model;

[0011] A target reply is generated based on the reply text and the response template.

[0012] In one possible implementation, the method further includes:

[0013] Inputting the summary information and the full text information into the large model, and determining the user intention through the large model;

[0014] Based on an extended database and a question-and-answer database, converting the user intention into the reply text, wherein the extended database includes extended information related to the communication field, and the question-and-answer database includes pre-set question information and a plurality of answer information corresponding to the question information;

[0015] Based on a response template library, a response template corresponding to the user intention is determined, wherein the response template library is pre-set with a correspondence between the user intention and the response template.

[0016] In one possible implementation, the method further includes:

[0017] Determining the extended information corresponding to the user intention according to the extended database;

[0018] Determining the question information and the answer information corresponding to the user intention according to the question and answer database;

[0019] Generate reply prompt information based on the extended information, the question information and the response information;

[0020] Inputting the user intention and the reply prompt information into the macro model, and extracting user intention features and reply prompt features through the macro model;

[0021] The user intention feature and the reply prompt feature are fused through a large model, and a response is made based on the fused feature to obtain the reply text.

[0022] In one possible implementation, the method further includes:

[0023] When the response template includes a reserved picture position, searching for a picture corresponding to the response text from a picture library according to the response text;

[0024] According to the response template, the picture and the response text are determined as the target response.

[0025] In one possible implementation, the method further includes:

[0026] Obtain training set data and validation set data, wherein the training set data includes a labeled data set and an unlabeled data set, and both the labeled data set and the unlabeled data set include: text input, voice input, and image input;

[0027] Inputting the labeled data into an initial multimodal multi-task model for training to obtain a first multimodal multi-task model;

[0028] Inputting the unlabeled data into the first multimodal multi-task model to perform pseudo label prediction to obtain a pseudo label corresponding to each unlabeled data in the unlabeled dataset;

[0029] Calculating the similarity between the pseudo label of each unlabeled data and the label of each labeled data, and if the similarity is higher than a first preset threshold, adding the unlabeled data and the corresponding pseudo label to the labeled data set to obtain a first labeled data set;

[0030] performing feature enhancement on the labeled data in the first labeled data set based on a spectral clustering algorithm to obtain a feature-enhanced first labeled data set;

[0031] Training the first multimodal multi-task model based on the feature-enhanced first labeled dataset to obtain a second multimodal multi-task model;

[0032] Input the validation set data into the second multimodal multi-task model to obtain a loss function value. If the loss function value is less than or equal to a second preset threshold, the second multimodal multi-task model is a trained second multimodal multi-task model. If the loss function value is greater than the second preset threshold, the network parameters of the second multimodal multi-task model are adjusted according to the loss function value, and the step of inputting the unlabeled data set into the first multimodal multi-task model for pseudo-label prediction is returned to obtain a pseudo-label corresponding to each unlabeled data in the unlabeled data set.

[0033] In one possible implementation, the method further includes:

[0034] Based on the upper bound of the confidence interval, determine the weights corresponding to text input, voice input, and image input during each round of training;

[0035] Based on the weights, the data ratios of text input, voice input, and image input are determined during each round of training.

[0036] In a second aspect, an embodiment of the present invention provides an online customer service response device, the device comprising:

[0037] an acquisition module, configured to acquire user input information, input the user input information into a classification model, and identify the information type of the user input information through the classification model;

[0038] A first processing module is configured to input the user input information and the information type into a text extraction model, and perform text extraction using the text extraction model to obtain full text information;

[0039] A second processing module is configured to input the user input information into a summary extraction model, perform summary extraction using the summary extraction model, and obtain summary information;

[0040] A third processing module is configured to input the summary information and the full text information into a large model, and generate a reply text and a response template through the large model;

[0041] The fourth processing module is used to generate a target reply based on the reply text and the response template.

[0042] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: a processor, and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the online customer service response method as described in any one of the first aspects.

[0043] In a fourth aspect, an embodiment of the present invention provides a computer storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, the online customer service response method as described in any one of the first aspects is implemented.

[0044] In a fifth aspect, an embodiment of the present invention provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the online customer service response method as described in any one of the first aspects.

[0045] An online customer service response method, apparatus, device, storage medium, and product according to an embodiment of the present invention can obtain user input information, input the user input information into a classification model, identify the information type corresponding to the user input information through the classification model, input the user input information and the information type into a text extraction model, and use the text extraction model to perform text extraction based on the user input information and the information type corresponding to the user input information to obtain full text information to provide responses to image input and voice input. The user input information is then input into a summary extraction model, which performs summary extraction on the user input information through the summary extraction model to obtain summary information. The summary information and the full text information are input into a large model, so that the large model obtains a response text and a response template based on the summary information and the full text information. The large model responds to the user input information rather than based on a preconfigured question and answer library. The large model can respond based on any question information, so that regardless of whether the question raised by the user is in the question and answer library, the large model can provide an accurate response. The large model of the present invention responds based on the summary information and the full information. The summary information guides the large model, allowing it to focus on the core content and key points of the full text information, resulting in a more accurate response text, thereby improving the accuracy of the online customer service response. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0047] Figure 1 This is a flow chart of Example 1 of the online customer service response method provided by the present invention;

[0048] Figure 2 This is a flow chart of Example 2 of the online customer service response method provided by the present invention;

[0049] Figure 3 This is a flow chart of an online customer service response method provided by the present invention;

[0050] Figure 4 This is a structural diagram of an online customer service response device provided by an embodiment of the present invention;

[0051] Figure 5 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The features and exemplary embodiments of various aspects of the present invention will be described in detail below. In order to make the objects, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can be implemented without the need for some of these specific details. The following description of the embodiments is merely intended to provide a better understanding of the present invention by illustrating examples of the present invention.

[0053] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.

[0054] In the prior art, online customer service relies on a pre-configured question-and-answer library to answer user questions. When a user's question does not exist in the answer library, the online customer service will determine an answer to a similar question and output it, but the answer output at this time is inaccurate.

[0055] Based on this, the inventive concept of the present invention is to provide an online customer service response method, device, equipment, storage medium and product that are not limited to a pre-configured question and answer database, so that regardless of whether the question raised by the user is in the pre-configured question and answer database, the online customer service can still provide an accurate answer.

[0056] In order to solve the problems of the prior art, embodiments of the present invention provide an online customer service response method, apparatus, device, storage medium and product.

[0057] The following first introduces the online customer service response method provided by an embodiment of the present invention.

[0058] Figure 1 This is a flow chart of the first embodiment of the online customer service response method provided by the present invention. Figure 1 As shown, the method may include the following steps:

[0059] S101: Obtain user input information, input the user input information into a classification model, and identify the information type of the user input information through the classification model.

[0060] S102: Input the user input information and information type into the text extraction model, perform text extraction through the text extraction model, and obtain the full text information.

[0061] S103: Input the user input information into the summary extraction model, perform summary extraction through the summary extraction model, and obtain summary information.

[0062] S104: Input the summary information and the full text information into the big model, and generate the reply text and answer template through the big model.

[0063] S105: Generate a target reply based on the reply text and the response template.

[0064] The specific implementation methods of the above steps are introduced below.

[0065] Step S101: obtaining user input information, inputting the user input information into a classification model, and identifying the information type of the user input information through the classification model.

[0066] In this embodiment, user input information is obtained. In one example, the user input information includes text input, and / or voice input, and / or image input. The information type corresponding to the user input information is identified by a classification model. In one example, the classification model pre-stores a correspondence between information type and file format. When the classification model receives the user input information, it identifies the corresponding information type by the file format of the input information. For example, when the file format of the user input information is a jpg or png image format, the classification model determines the user input information as an image input. It is understandable that the user input information type is not single, that is, the user input information can include text input, voice input, or image input at the same time.

[0067] Step S102: input the user input information and information type into a text extraction model, and perform text extraction through the text extraction model to obtain the full text information.

[0068] In this embodiment, after the classification model identifies the information type corresponding to the user input information, the classification model sends the user input information and the information type corresponding to the user input information to a text extraction model, which then uses the text extraction model to extract text from the user input information to obtain the full text information. For example, the classification model includes an OCR (Optical Character Recognition, OCR) image recognition model and an ASR (Automatic Speech Recognition, ASR) speech recognition model. The OCR image recognition model can extract text information from images to obtain the full text information, while the ASR speech recognition model can extract text information from audio to obtain the full text information.

[0069] It can be understood that when the user input information type is not single, that is, when the user input information includes text input, voice input and picture input, the classification model will send the picture input part of the user input information to the OCR image recognition model in the text extraction model, and extract the text information in the picture input through the OCR image recognition model. The classification model will send the voice input part of the user input information to the ASR speech recognition model in the text extraction model, and extract the text information in the voice input through the ASR speech recognition model. Finally, the text information in the text input, the text information in the picture input and the text information in the voice input are combined to obtain the full text information.

[0070] In this embodiment, the classification model classifies the user input information and can identify whether the user input is text input, image input or voice input, and then performs text extraction according to the text extraction model to obtain the full text information. Therefore, the online customer service response method of the present invention can not only answer text questions, but also receive user's picture questions and voice questions, thereby improving the user experience.

[0071] Step S103: input the user input information into a summary extraction model, and extract the summary through the summary extraction model to obtain summary information.

[0072] In this embodiment, the user input information may be input into a summary extraction model, and the user input information may be extracted by the summary extraction model to obtain summary information corresponding to the user input information.

[0073] In one example, the summary extraction model can be a separate multimodal multitask model (Multimodal Multitask Learning with a Unified Transformer, UniT). After receiving user input information, the user input information is sent to the multimodal multitask model for summary extraction to obtain summary information corresponding to the user input information. The multimodal multitask model consists of a shared Transformer encoder, a Transformer decoder corresponding to each modality, and classifiers corresponding to multiple anomaly detection tasks. Compared with previous Transformer-based multitask models, the multimodal multitask model is free from fine-tuning methods. Model parameters can be shared between different fields, and common knowledge can be shared between different tasks. For a set of tasks with strong correlation, the multimodal multitask model achieves stronger performance with fewer parameters.

[0074] In another example, the summary extraction model integrates a classification model, a text extraction model and a multimodal multi-task model. That is, when the summary extraction model receives user input information, the classification model in the summary extraction model classifies the user input information, and inputs the user input information that is voice input and picture input into the text extraction model. The text extraction model extracts the full text information based on the user input information, and then the multimodal multi-task model receives the full text information and extracts summary information based on the full text information.

[0075] Step S104: Input the summary information and the full text information into the big model, and generate the reply text and the answer template through the big model.

[0076] In this embodiment, the full text and the summary information extracted by the multimodal, multi-task model are sent to the large model, allowing the large model to generate a response text and a response template based on the summary information and the full text. The summary information serves as a guide, allowing the large model to more accurately respond to the full text, resulting in a more accurate response text based on the summary information and the full text.

[0077] In one example, the large model is an artificial intelligence large model, and the large model can be fine-tuned based on question text and answer text in a historical database to obtain a large model capable of answering questions in the communications field. Specifically, question text and answer text pairs are obtained from the historical database, and data cleaning is performed on the question text and answer text pairs to remove null values ​​and error values ​​to obtain cleaned question text and answer text pair data. The cleaned question text and answer text pair data is then used to fine-tune the pre-trained artificial intelligence large model. In one example, the fine-tuning method includes full fine-tuning, partial fine-tuning, or parameter-efficient fine-tuning.

[0078] In order to make the reply text obtained by the large model based on the summary information and the full text information more accurate, as another implementation of the present invention, a specific implementation of S104 may further include the following steps:

[0079] The summary information and full text information are input into the big model, and the user intention is determined through the big model.

[0080] In this embodiment, the big model jointly determines the user intention based on the summary information and the full text information. The summary information is the information that is refined and summarized from the full text information, retaining the core content and key points of the full text information. The big model is guided by the summary information, so that the big model focuses on the core content and key points of the full text information rather than secondary information. Therefore, the user intention obtained by the big model is more accurate.

[0081] For example, for a full text input from a user: "I'm going on a week-long business trip to Europe next month. How do I activate international roaming on my mobile data?", the corresponding summary might be: "Activate a European international roaming data plan." The large model encodes the full text to obtain a corresponding feature vector, then encodes the summary to obtain a corresponding feature vector. Using a cross-attention mechanism, the model weights the feature vectors for the full text and the summary to generate a fused feature vector. Finally, the model outputs the user intent based on the fused feature vector. Because redundant content in the full text is filtered out with low weights, the model can more accurately capture the user intent.

[0082] Based on the extended database and the question-answer database, the user intention is converted into a reply text, wherein the extended database includes extended information related to the communication field, and the question-answer database includes pre-set question information and multiple answer information corresponding to the question information.

[0083] In this embodiment, the big model combines the extended database and the question-and-answer database to convert the user intention into the reply text. On the basis of obtaining the user intention based on the summary information and the full text information, the extended database and the question-and-answer database are introduced to guide the big model to obtain the reply text according to the user intention, which can further improve the accuracy of the reply text generated by the big model.

[0084] The extended database includes extended information related to the communications field, such as data package information, voice package information, card activation information, number selection information, etc., or a knowledge graph corresponding to the communications field. The process of constructing the knowledge graph corresponding to the communications field is as follows: obtaining unstructured raw communication data from communications-related websites or APIs, extracting structured raw communication data from the operator's internal database, performing data cleansing on the raw communication data to obtain target communication data, and then, based on data modeling methodology, such as defining entity-relationship-entity according to Schema design, and then extracting data from the target communication data according to entity-relationship-entity using triple extraction technology, obtaining multiple entity-relationship-entity triples, and finally fusing the multiple triples to obtain the knowledge graph corresponding to the communications field.

[0085] In addition, the question and answer database includes pre-set question information and multiple answer information corresponding to each question information, and the question information and answer information in the question and answer database are constantly updated. For example, after the big model obtains the reply text based on the user input information, the big model will store the user input information and reply text in the question and answer database to update the question and answer database, so that when the big model encounters similar questions in the future, it can refer to the existing replies and generate more accurate and comprehensive replies. In one example, when storing multiple answer versions for the same question, the generation time and user satisfaction of each answer version are recorded, and the answer information with a later answer version time and higher user satisfaction is given priority.

[0086] Based on the response template library, a response template corresponding to the user intention is determined, and the response template library is pre-set with a correspondence between the user intention and the response template.

[0087] In this embodiment, the response template library pre-stores the correspondence between user intent and response templates. Once the macro model determines the user's intent, it searches the response template library for the corresponding response template based on the user's intent. In one example, the correspondence between user intent and response templates stored in the response template library is continuously updated. For example, response templates are updated based on user satisfaction. If users repeatedly give the same response template low ratings, manual review and intervention are triggered, and the corresponding response template is modified.

[0088] In this embodiment, user intent is determined based on summary information and the full text. The summary information is used to guide the large model, allowing it to focus on the core content and key points of the full text, thereby obtaining a more accurate user intent. Based on this accurate user intent, the generated reply text is then adapted to the communications field by leveraging the support of the communication domain knowledge in the extended database and reusing the experience of the question-and-answer database. This, combined with highly rated historical question information and response information, improves the accuracy of the reply text and user satisfaction. Furthermore, by combining a response template library to determine the response template that corresponds to the user's intent, the response style can be unified, eliminating the need for users to adapt to changing language and allowing them to more accurately understand the online customer service's target response.

[0089] In order to further improve the accuracy of the reply text generated by the large model, a specific implementation method of converting the user intention into the reply text is as follows, combining the extended database and the question-answer database:

[0090] Determine the extended information corresponding to the user's intention based on the extended database.

[0091] In this embodiment, extended information corresponding to the user intention is searched in the extended database. For example, if the user intention obtained by the large model is European international roaming processing, the extended information related to European international roaming processing is searched in the extended database based on the user intention. The extended information is the tariff standards and traffic packages for European international roaming processing. In one example, the extended information can be obtained based on "Europe" and "International roaming processing" as query conditions.

[0092] Based on the question and answer database, determine the question information and answer information corresponding to the user's intention.

[0093] In this embodiment, the question information closest to the user's intention is determined in the question and answer database, and then the answer information is determined based on the question information. In one example, the user's intention can be encoded as a semantic vector, and then the question information whose similarity with the user's intention is greater than a preset threshold is searched in the question and answer database, and the answer information corresponding to the question information is queried based on the question information.

[0094] Generate reply prompt information based on the extended information, question information and answer information.

[0095] In this embodiment, reply prompt information is generated based on the extended information, question information, and response information.

[0096] The user intention and reply prompt information are input into the big model, and the user intention features and reply prompt features are extracted through the big model.

[0097] In this embodiment, the user intention and the reply prompt information are sent to the big model, and the big model extracts the user intention feature according to the user intention and extracts the reply prompt feature according to the prompt information.

[0098] The user intention features and the reply prompt features are fused through the large model, and the response is made based on the fused features to obtain the reply text.

[0099] In this embodiment, the large model fuses the obtained user intention features and reply prompt features to obtain fused features, and then the large model responds based on the fused features to obtain a reply text.

[0100] In this embodiment, based on the extended database, the extended information corresponding to the user's intention is determined, and based on the question and answer database, the question information and answer information corresponding to the user's intention are determined, and the extended information, question information and answer information are determined as prompt information. The prompt information combines the information of the extended database and the question and answer database. The extended database includes extended information related to the communication field, so that the large model is subject to knowledge constraints in the process of generating the reply text, that is, the reply text must be related to the communication field, and the question and answer database includes pre-set question information and multiple answer information corresponding to the question information, so that the large model can output a reply text similar to the historical question information and answer information based on the historical question information and answer information. Then the large model performs feature fusion according to the user intention feature and the reply prompt feature to obtain the fused feature, and obtains the reply text based on the fused feature. In this way, on the basis of ensuring the accuracy of the user's intention, the prompt information is introduced to guide the process of the large model generating the reply text, which can transform the process of the large model generating the reply text from "free association" to "directed reasoning", thereby further improving the accuracy of the large model generating the reply text.

[0101] Step S105: Generate a target reply based on the reply text and the response template.

[0102] In this embodiment, after obtaining the reply text and the response template, the reply text is organized into a target response in the form of the response template.

[0103] To improve user experience, when a user asks how to handle a service, detailed steps for handling the service are output and pictures related to the steps are displayed. As another implementation of the present invention, a specific implementation of S105 may further include the following steps:

[0104] When the answer template includes a reserved image location, the image library is searched for an image corresponding to the answer text based on the answer text.

[0105] In this embodiment, when the obtained response template includes a reserved image position, the image corresponding to the response text is searched from the gallery based on the response text. In an example, the response template is an operation step type response template, and the screenshot corresponding to the business processing is searched from the gallery, for example, a screenshot of the login application interface and a screenshot of the package selection interface.

[0106] According to the answer template, the picture and answer text are determined as the target answer.

[0107] In this embodiment, according to the response template, the picture and the reply text are jointly determined as the target reply. If the reply text determined by the large model is: "Hello, you can change the package by following the steps: 1. Log in to the application interface; 2. Enter the 'My Package' page; 3. Select 'Change Package'; 4. Select the package you like and confirm the change." Then attach a screenshot of the login application interface next to "1. Log in to the application interface", and attach a screenshot of the interface for selecting the package next to "2. Enter the 'My Package' page;".

[0108] In this embodiment, by outputting detailed steps for handling business and displaying relevant pictures, users can understand how to operate more intuitively without having to search or consult manual customer service on their own, thereby improving the user experience.

[0109] In this embodiment, user input information is obtained, and the information type corresponding to the user input information is identified using a classification model. The user input information and the information type corresponding to the user input information are sent to a text extraction model. The text extraction model is used to perform text extraction based on the user input information and the information type corresponding to the user input information to obtain full text information to provide responses to image input and voice input. The user input information is then sent to a multimodal multi-task model, and the multimodal multi-task model is used to perform summary extraction on the full text information to obtain summary information. The summary information and the full text information are sent to a large model, so that the large model obtains a reply text and a response template based on the summary information and the full text information. The large model responds to the user input information rather than based on a preconfigured question and answer library. The large model can respond based on any question information, so that regardless of whether the question raised by the user is in the question and answer library, the large model can provide an accurate response. The large model of the present invention responds based on the summary information and the full information. The summary information guides the large model, allowing it to focus on the core content and key points of the full text information, resulting in a more accurate response text, thereby improving the accuracy of the online customer service response.

[0110] Figure 2 This is a flow chart of Example 2 of the online customer service response method provided by the present invention, as shown in FIG. Figure 2 As shown, when the summary extraction model is a multimodal multi-task model, before step S103, the following steps are further included:

[0111] S201: Acquire training set data and validation set data. The training set data includes a labeled data set and an unlabeled data set. Both the labeled data set and the unlabeled data set include text input, voice input, and image input.

[0112] In an embodiment of the present application, since the input types include text input, voice input, and image input, there is a high cost in labeling data, resulting in high training costs for the multimodal multi-task model. Therefore, in order to reduce the model training cost, the present invention uses a consistent semi-supervised learning method to train the multimodal multi-task model.

[0113] Specifically, the training set data obtained includes: labeled data sets and unlabeled data sets, and the proportion of labeled data sets is smaller than that of unlabeled data. For example, 20% of labeled data is used to perform consistent semi-supervised learning training on a multimodal multi-task model. It should be noted that the label of the labeled data is the annotated summary information. For example, "My data flow is insufficient. What package do you recommend for this month?" The corresponding summary information is "Recommended monthly data package." Among them, both labeled data and unlabeled data include: text input, voice input, and image input.

[0114] S202: Input the labeled data into the initial multimodal multi-task model for training to obtain a first multimodal multi-task model.

[0115] In this embodiment, the initial multimodal multi-task model is trained using a labeled dataset to obtain a first multimodal multi-task model.

[0116] S203: Input the unlabeled data into the first multimodal multi-task model to perform pseudo-label prediction to obtain a pseudo-label corresponding to each unlabeled data in the unlabeled dataset.

[0117] In this embodiment, the unlabeled data is input into the first multimodal multi-task model, and the first multimodal multi-task model generates corresponding pseudo labels based on the unlabeled data, and finally obtains the pseudo labels corresponding to each unlabeled data in the unlabeled data set.

[0118] S204: Calculate the similarity between the pseudo label of each unlabeled data and the label of each labeled data. If the similarity is higher than a first preset threshold, add the unlabeled data and the corresponding pseudo label to the labeled data set to obtain a first labeled data set.

[0119] In this embodiment, the similarity between the pseudo-label of each unlabeled data and the label of each labeled data is calculated. In one example, the similarity between the pseudo-label and the label of the labeled data is calculated based on the JS divergence algorithm. If the similarity between the two is higher than a first preset threshold, the unlabeled data and the corresponding pseudo-label are added to the labeled data set to expand the labeled data set and obtain a first labeled data set.

[0120] S205: Perform feature enhancement on the labeled data in the first labeled data set based on a spectral clustering algorithm to obtain a feature-enhanced first labeled data set.

[0121] In this embodiment, feature enhancement is performed on the labeled data in the first labeled dataset based on a spectral clustering algorithm to obtain a feature-enhanced first labeled dataset. In one example, for example, if there are m labeled data in the first labeled dataset, where m is a positive integer, feature extraction is performed on the m labeled data to obtain m first features. Spectral clustering is performed on the m first features to obtain a plurality of class prototype features and first feature groups corresponding to the plurality of class prototype features. The first feature groups include the plurality of first features. Among the plurality of class prototype features, a target class prototype feature closest to each first feature is determined. For each first feature, feature enhancement is performed on the first feature using the target class prototype feature to obtain the feature-enhanced first labeled dataset.

[0122] S206: Training the first multimodal multi-task model according to the feature-enhanced first labeled data set to obtain a second multimodal multi-task model.

[0123] In this embodiment, the first multimodal multi-task model is trained based on the feature-enhanced first labeled data set to obtain a second multimodal multi-task model. In one example, the first multimodal multi-task model is used to summarize the feature-enhanced first labeled data to obtain a predicted label, and the consistency loss function value and the cross-entropy loss function value corresponding to the predicted label are determined. The consistency loss function value is used to measure the consistency between the predicted label and the pseudo label to promote unsupervised learning, and the cross-entropy loss function value is used to measure the difference between the predicted label and the true label to ensure the accuracy of supervised learning. The consistency loss and the cross-entropy loss are weighted and combined to determine the target loss function value. According to the target loss function value, the training parameters of the first multimodal multi-task model are adjusted. The above steps are repeated until the target loss function value converges to obtain the second multimodal multi-task model.

[0124] To prevent overfitting or underfitting of the multimodal multitask model during training, the following steps are performed when training the first modality multitask model:

[0125] Based on the upper limit of the confidence interval, the weights corresponding to text input, voice input, and image input are determined during each round of training.

[0126] In this embodiment, because the multimodal, multitask model fits different tasks (text input, image input, and voice input) at different speeds, a task selector, based on a confidence interval upper bound algorithm, is designed to automatically adjust the proportion of different tasks in the next training round to prevent overfitting or underfitting of some tasks. Driven by model validation results on the validation set, the task selector dynamically adjusts the ratio of task iterations to ensure that the model has a relatively balanced recognition capability for different tasks.

[0127] In one example, after each round of training, the multimodal multi-task model calculates the recognition accuracy of each task on the validation set and then calculates the upper bound of the confidence interval:

[0128]

[0129] Among them, UCB i represents the upper bound of the confidence interval of task i, μ i represents the historical average accuracy of task i, N i represents the number of rounds that task i has been trained, T represents the total number of training rounds, lnT represents the logarithm of the total number of training rounds, and c represents the exploration coefficient.

[0130] Based on the weights, the data ratios of text input, voice input, and image input are determined during each round of training.

[0131] After normalizing the UCB values ​​corresponding to text input, image input, and voice input, determine the training weight of each task in the next round. Adjust the data ratio of text input, voice input, and image input in the training batch based on the weight, and give priority to increasing the training opportunities of low UCB tasks.

[0132] In this embodiment, based on the confidence interval upper bound algorithm, the proportion of different input type tasks in each training process is dynamically adjusted to balance the recognition ability of the first multimodal multi-task model for different input type tasks, so that even in the case of unbalanced input type samples, such as more text input and less picture input, the multimodal multi-task model can still have balanced recognition capabilities.

[0133] S207: Input the validation set data into the second multimodal multi-task model to obtain a loss function value. If the loss function value is less than or equal to the second preset threshold, the second multimodal multi-task model is a trained second multimodal multi-task model. If the loss function value is greater than the second preset threshold, adjust the network parameters of the second multimodal multi-task model according to the loss function value, and return to S203.

[0134] In an embodiment of the present application, the second multimodal multi-task model is verified using the validation set data to obtain a loss function value. If the loss function value is greater than a second preset threshold, in an example, the second preset threshold is a revised value of the first preset threshold based on data characteristics and scenario requirements, the hyperparameters of the model are adjusted according to the loss function value, and step S203 is returned to continue the pseudo-label prediction of the unlabeled data, that is, the pseudo-label of the unlabeled data is predicted using the second multimodal multi-task model, and the similarity between the pseudo-label of each unlabeled data and the labeled data is calculated, and the unlabeled data with a similarity higher than the first threshold is added to the labeled data set again, so as to continuously perform data division to obtain an expanded labeled data set, until the loss function value of the multimodal multi-task model converges and a trained multimodal multi-task model is obtained.

[0135] In this embodiment, a labeled dataset and an unlabeled dataset are obtained. An initial multimodal multi-task model is trained using the labeled dataset to obtain a first multimodal multi-task model. Pseudo-label predictions are performed on the unlabeled dataset based on the first multimodal multi-task model. Unlabeled data with a similarity to the labeled data exceeding a preset threshold is screened out and added to the labeled dataset to form a first labeled dataset. Label predictions are performed on the unlabeled dataset to expand the labeled dataset, reducing the cost of manual data annotation and, consequently, the cost of model training. Furthermore, a spectral clustering algorithm is used to feature enhance the labeled data in the first labeled dataset to obtain a feature-enhanced first labeled dataset. The first modal multi-task model is trained using the feature-enhanced first labeled dataset until the loss function converges, thereby obtaining a trained multimodal multi-task model. Feature enhancement of the training set data using the spectral clustering algorithm during model training can improve the training accuracy of the multimodal multi-task model.

[0136] Figure 3 A flowchart of an online customer service response method provided by the present invention is shown as follows: Figure 3 As shown, the terminal is an application 301, and the cloud includes: a pre-processing module 302, a response module 303, and a post-processing module 304, wherein the pre-processing module 302 includes: a classification model 3021, a text extraction submodule 3022, and a multimodal multi-task model 3023.

[0137] Application 301 receives user input and sends it to classification model 3021. Classification model 3021 then sends it to multimodal multitask model 3023. Classification model 3021 identifies the user input type and sends the image and voice input to text extraction submodule 3022 to extract the full text. The full text corresponding to the text input is then sent to response module 303. After receiving the voice and image input, text extraction submodule 3022 extracts the full text corresponding to the voice and image input and sends it to response module 303. After receiving the user input, multimodal multitask model 3023 extracts summary information based on the user input and sends it to response module 303. Response module 303 generates a response template and reply text based on the summary and full text information. These are then sent to post-processing module 304. Post-processing module 304 determines a target response based on the response template and reply text and sends the target response to application 301.

[0138] Figure 4 This is a schematic diagram of the structure of an online customer service response device provided by an embodiment of the present invention. Figure 4As shown, an online customer service response device 400 includes an acquisition module 401 , a first processing module 402 , a second processing module 403 , a third processing module 404 , and a fourth processing module 405 .

[0139] An acquisition module 401 acquires user input information, inputs the user input information into a classification model, and identifies the information type of the user input information through the classification model;

[0140] A first processing module 402 is configured to input the user input information and the information type into a text extraction model, and perform text extraction using the text extraction model to obtain full text information;

[0141] The second processing module 403 is configured to input the user input information into a summary extraction model, perform summary extraction using the summary extraction model, and obtain summary information;

[0142] The third processing module 404 is configured to input the summary information and the full text information into a large model, and generate a reply text and a response template using the large model;

[0143] The fourth processing module 405 is used to generate a target reply based on the reply text and the response template.

[0144] The online customer service answering device provided by the present invention also includes:

[0145] The third processing module 404 is further configured to input the summary information and the full text information into the large model and determine the user intent through the large model;

[0146] The third processing module 404 is further configured to convert the user's intention into a reply text based on an extended database and a question-and-answer database, wherein the extended database includes extended information related to the communication field, and the question-and-answer database includes pre-set question information and a plurality of corresponding answer information.

[0147] The third processing module 404 is further configured to determine a response template corresponding to the user intention based on the response template library, wherein the response template library is pre-set with a correspondence between the user intention and the response template.

[0148] The online customer service answering device provided by the present invention also includes:

[0149] The third processing module 404 is further configured to determine the extended information corresponding to the user's intention based on the extended database;

[0150] The third processing module 404 is further configured to determine question information and answer information corresponding to the user's intention based on the question and answer database;

[0151] The third processing module 404 is further configured to generate reply prompt information based on the extended information, question information, and answer information;

[0152] The third processing module 404 is further configured to input the user intention and the reply prompt information into the large model, and extract the user intention features and the reply prompt features through the large model;

[0153] The third processing module 404 is also used to fuse the user intention features and the reply prompt features through the large model, and respond based on the fused features to obtain the reply text.

[0154] The online customer service answering device provided by the present invention further includes:

[0155] The fourth processing module 405 is further configured to search for an image corresponding to the reply text from the image library according to the reply text when the reply template includes a reserved image location;

[0156] The fourth processing module 405 is further used to determine the picture and the reply text as the target reply according to the reply template.

[0157] The online customer service answering device provided by the present invention further includes:

[0158] An acquisition module is used to acquire training set data and validation set data, wherein the training set data includes a labeled data set and an unlabeled data set, and both the labeled data set and the unlabeled data set include: text input, voice input, and image input;

[0159] A first training module, configured to input the labeled data into an initial multimodal multi-task model for training to obtain a first multimodal multi-task model;

[0160] A prediction module, configured to input the unlabeled data into the first multimodal multi-task model to perform pseudo-label prediction, and obtain a pseudo-label corresponding to each unlabeled data in the unlabeled dataset;

[0161] a calculation module, configured to calculate the similarity between the pseudo-label of each unlabeled data and the label of each labeled data, and if the similarity is higher than a first preset threshold, add the unlabeled data and the corresponding pseudo-label to the labeled data set to obtain a first labeled data set;

[0162] a feature enhancement module, configured to perform feature enhancement on the labeled data in the first labeled data set based on a spectral clustering algorithm to obtain a feature-enhanced first labeled data set;

[0163] A second training module is configured to train the first multimodal multi-task model based on the feature-enhanced first labeled data set to obtain a second multimodal multi-task model;

[0164] The third training module is used to input the verification set data into the second multimodal multi-task model to obtain a loss function value. If the loss function value is less than or equal to a second preset threshold, the second multimodal multi-task model is a trained second multimodal multi-task model. If the loss function value is greater than the second preset threshold, the network parameters of the second multimodal multi-task model are adjusted according to the loss function value, and the step of inputting the unlabeled data set into the first multimodal multi-task model for pseudo-label prediction is returned to obtain a pseudo-label corresponding to each unlabeled data in the unlabeled data set.

[0165] The online customer service answering device provided by the present invention further includes:

[0166] The first determination module is used to determine the weights corresponding to text input, voice input, and image input in each round of training based on the upper limit of the confidence interval;

[0167] The second determination module is used to determine the data ratio of text input, voice input, and image input in each round of training based on the weight.

[0168] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present invention is shown.

[0169] The electronic device may include a processor 501 and a memory 502 storing computer program instructions.

[0170] Specifically, the processor 501 may include a central processing unit (CPU) or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiment of the present invention.

[0171] The memory 502 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 502 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In one example, the memory 502 may include a removable or non-removable (or fixed) medium, or the memory 502 may be a non-volatile solid-state memory. The memory 502 may be inside or outside the integrated gateway disaster recovery device.

[0172] In one example, the memory 502 may be a read-only memory (ROM). In one example, the ROM may be a mask-programmable ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.

[0173] The memory 502 may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical or other physical / tangible memory storage devices. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.

[0174] The processor 501 reads and executes the computer program instructions stored in the memory 502 to implement Figure 1 The online customer service response method in the illustrated embodiment.

[0175] In addition, in conjunction with the online customer service response method in the above embodiments, embodiments of the present invention may provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any of the online customer service response methods in the above embodiments is implemented.

[0176] An embodiment of the present invention further provides a computer program product, including a computer program, which, when executed, implements any one of the online customer service response methods in the above embodiments.

[0177] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.

[0178] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The programs or code segments can be stored in a machine-readable medium or transmitted on a transmission medium or communication link via a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, read-only memories (ROMs), flash memories, erasable read-only memories (EROMs), floppy disks, compact disc read-only memories (CD-ROMs), optical discs, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segments can be downloaded via computer networks such as the Internet and intranets.

[0179] It should also be noted that the exemplary embodiments described herein describe methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the steps described above. In other words, the steps may be performed in the order described in the embodiments, or in a different order, or several steps may be performed simultaneously.

[0180] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It is also understood that each box in the block diagram and / or flowchart and the combination of the boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0181] The above description is only a specific embodiment of the present invention. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present invention is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention.

Claims

1. An online customer service response method, characterized in that: include: Obtaining user input information, inputting the user input information into a classification model, and identifying the information type of the user input information through the classification model; Inputting the user input information and the information type into a text extraction model, performing text extraction using the text extraction model to obtain full text information; Inputting the user input information into a summary extraction model, performing summary extraction by the summary extraction model to obtain summary information; Input the summary information and the full text information into a large model, and generate a reply text and a response template through the large model; A target reply is generated based on the reply text and the response template.

2. The method according to claim 1, characterized in that Inputting the summary information and the full text information into a large model, and generating a reply text and a response template through the large model, including: Inputting the summary information and the full text information into the large model, and determining the user intention through the large model; Based on an extended database and a question-and-answer database, converting the user intention into the reply text, wherein the extended database includes extended information related to the communication field, and the question-and-answer database includes pre-set question information and a plurality of answer information corresponding to the question information; Based on a response template library, a response template corresponding to the user intention is determined, wherein the response template library is pre-set with a correspondence between the user intention and the response template.

3. The method according to claim 2, characterized in that The converting the user intention into the reply text based on the extended database and the question-answer database includes: Determining the extended information corresponding to the user intention according to the extended database; Determining the question information and the answer information corresponding to the user intention according to the question and answer database; Generate reply prompt information based on the extended information, the question information and the response information; Inputting the user intention and the reply prompt information into the macro model, and extracting user intention features and reply prompt features through the macro model; The user intention feature and the reply prompt feature are fused through a large model, and a response is made based on the fused feature to obtain the reply text.

4. The method according to claim 1, wherein The generating a target reply based on the reply text and the answer template includes: When the response template includes a reserved picture position, searching for a picture corresponding to the response text from a picture library according to the response text; According to the response template, the picture and the response text are determined as the target response.

5. The method according to claim 1, wherein The summary extraction model is a multimodal multitask model. Before inputting the user input information into the summary extraction model and extracting a summary through the summary extraction model, the method further includes: Obtain training set data and validation set data, wherein the training set data includes a labeled data set and an unlabeled data set, and both the labeled data set and the unlabeled data set include: text input, voice input, and image input; Inputting the labeled data into an initial multimodal multi-task model for training to obtain a first multimodal multi-task model; Inputting the unlabeled data into the first multimodal multi-task model to perform pseudo label prediction to obtain a pseudo label corresponding to each unlabeled data in the unlabeled dataset; Calculating the similarity between the pseudo label of each unlabeled data and the label of each labeled data, and if the similarity is higher than a first preset threshold, adding the unlabeled data and the corresponding pseudo label to the labeled data set to obtain a first labeled data set; performing feature enhancement on the labeled data in the first labeled data set based on a spectral clustering algorithm to obtain a feature-enhanced first labeled data set; Training the first multimodal multi-task model based on the feature-enhanced first labeled dataset to obtain a second multimodal multi-task model; Input the validation set data into the second multimodal multi-task model to obtain a loss function value. If the loss function value is less than or equal to a second preset threshold, the second multimodal multi-task model is a trained second multimodal multi-task model. If the loss function value is greater than the second preset threshold, the network parameters of the second multimodal multi-task model are adjusted according to the loss function value, and the step of inputting the unlabeled data set into the first multimodal multi-task model for pseudo-label prediction is returned to obtain a pseudo-label corresponding to each unlabeled data in the unlabeled data set.

6. The method according to claim 5, characterized in that The training of the first multimodal multi-task model based on the feature-enhanced first labeled dataset further includes: Based on the upper bound of the confidence interval, determine the weights corresponding to text input, voice input, and image input during each round of training; Based on the weights, the data ratios of text input, voice input, and image input are determined during each round of training.

7. An online customer service answering device, characterized in that: include: an acquisition module, configured to acquire user input information, input the user input information into a classification model, and identify the information type of the user input information through the classification model; A first processing module is configured to input the user input information and the information type into a text extraction model, and perform text extraction using the text extraction model to obtain full text information; A second processing module is configured to input the user input information into a summary extraction model, perform summary extraction using the summary extraction model, and obtain summary information; A third processing module is configured to input the summary information and the full text information into a large model, and generate a reply text and a response template through the large model; The fourth processing module is used to generate a target reply based on the reply text and the response template.

8. An electronic device, characterized in that: The device includes: a processor and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the online customer service response method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer storage medium stores computer program instructions, and when the computer program instructions are executed by the processor, the online customer service response method according to any one of claims 1 to 6 is implemented.

10. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the online customer service response method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Customer service dialogue abstract generation method based on large model and readable storage medium

    CN117874219A

  • Retrieval enhancement method and device based on large model, intelligent customer service system and medium

    CN118708705A

  • Question answering method and system based on question answering system

    CN119106119A

  • Customer service interaction method, interaction device, equipment, storage medium and program product

    CN119646137A

  • Intelligent question and answer method and device

    CN119917622A