Multi-round dialogue optimization method based on intelligent customer service
Through dynamic entity maps and automated audit models, multiple rounds of dialogue processing of intelligent customer service systems are optimized, which solves the problems of insufficient contextual correlation and low manual audit efficiency, and achieves efficient and real-time multi-round dialogue management and resource utilization.
Patent Information
- Application Number
- CN202510330207.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing intelligent customer service system has insufficient context dynamic correlation capabilities in multiple rounds of conversations, resulting in information loss or repeated inquiries, and low manual review efficiency, high operating costs, and low resource utilization, which cannot meet the real-time performance requirements of large concurrency scenarios.
The dynamic entity map is used to update entity relationships in real time, introduce automated audit model and task complexity calculation, entity extraction and intent identification are performed through the SpanBERT model, and computing resources are dynamically allocated in combination with audit confidence and task complexity, and use edge devices or cloud servers to handle tasks of different complexity.
It improves the context information correlation capability in multiple rounds of dialogue, reduces the frequency of manual review, ensures the accuracy and real-time response, optimizes resource utilization, reduces system burden, and meets real-time performance in large concurrency scenarios.
Smart Images

Figure CN120256570A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence and natural language processing, and specifically relates to a multi-turn dialogue optimization method based on intelligent customer service. Background Art
[0002] With the rapid development of artificial intelligence and natural language processing technologies, intelligent customer service systems have been widely applied in various industries to provide users with efficient and convenient services. However, there are still many deficiencies in the existing technologies when dealing with complex multi-turn dialogues. On the one hand, the system has a weak ability to manage context in multi-turn dialogues and is difficult to correctly associate entity references across turns, resulting in information loss or repeated inquiries. On the other hand, the current system highly relies on manual review and decision-making. Especially in complex dialogue tasks, the frequency of manual intervention is extremely high, causing response delays and high operating costs.
[0003] In addition, the existing single intent recognition model has insufficient ability to dynamically associate context, and there are bottlenecks in manual review efficiency, which are prone to missed detections or misidentifications, seriously affecting the user experience. At the same time, taking the multi-turn dialogue intelligent question-answering method in the medical field disclosed in patent CN112256825A as an example, by extracting the entity information and intent information of the previous turn of dialogue, it can better simulate human communication. Even if the user uses pronouns or hides some entities or intents, the dialogue can still be made more natural and fluent by inheriting the context. However, this patent still cannot effectively solve the problem of insufficient dynamic context association, and its hybrid model structure is complex, using a fixed inference path, with low resource utilization rate in large concurrency scenarios and inference delays that cannot meet real-time performance requirements.
[0004] Based on this, it is urgent to improve the existing multi-turn dialogue optimization method based on intelligent customer service to solve the above technical defects. Summary of the Invention
[0005] The purpose of this application is to provide a multi-turn dialogue optimization method for intelligent customer service that solves the problems of insufficient dynamic context association ability and bottlenecks in manual review efficiency in view of the deficiencies of the existing technologies.
[0006] To achieve the above-mentioned invention purpose, the following technical solutions are implemented in this application:
[0007] A multi-turn dialogue optimization method based on intelligent customer service includes the following steps:
[0008] S101. Receive the input data provided by the user. If the input data is of the voice input type, perform speech recognition on the input data of the voice input type to generate text input data, perform word segmentation on the text input data, perform noise filtering on the input data after word segmentation to generate a word segmentation sequence data, perform feature vector encoding on the word segmentation sequence data, and map the word segmentation sequence to a vector sequence;
[0009] S201. Extract entity data from the vector sequence based on the trained SpanBERT model, initialize the dynamic entity graph, and add the extracted entity data to the dynamic entity graph;
[0010] The user generates multiple rounds of vector sequences through the input data provided in multiple rounds and generates multiple rounds of entity data through the SpanBERT model. The different entities in the multiple rounds of entity data are output and input into the dynamic entity graph to update the dynamic entity graph;
[0011] S301. Perform intent recognition on the dynamic entity graph and generate one or more slots according to the intent. Fill the entities in the entity data into the corresponding slots according to the information in the dynamic entity graph;
[0012] S401. After completing the slot filling, generate the response code H of the quasi-response message according to the intent and the filled slot information audit , and then calculate the review confidence P of the response code according to the review confidence model audit ;
[0013] When the review confidence P audit exceeds the confidence threshold, execute step S501. When the review confidence P audit is lower than the confidence threshold, return to step S101;
[0014] S501. Calculate the task complexity based on the entities in the dynamic entity graph and output the task complexity C;
[0015] When the task complexity C is less than the complexity threshold, generate a natural language response according to the quasi-response message and update the context state in real time and feedback it to the user;
[0016] When the task complexity C is greater than the complexity threshold, send the dynamic entity graph to the cloud server to generate natural language using the full-function model and update the context state in real time and feedback it to the user.
[0017] The above technical solutions have the following technical effects:
[0018] On the one hand, by using the dynamic entity graph, the present application can update and maintain the relationship between entities in real time during the multi-round conversation process. This can not only effectively avoid information loss or repeated inquiries, but also enable the system to better understand and track the conversation history, improving the ability to associate context information in cross-round conversations. The dynamically updated graph can enhance the semantic understanding of the system in complex multi-round conversation scenarios, thereby providing a more accurate and smoother conversation experience.
[0019] On the other hand, by introducing an automated review model, this application reviews the responses generated by the system and determines whether human intervention is required based on the review confidence. This significantly reduces the frequency of manual review and response latency. Only when the review confidence is lower than the threshold will the system return to the process of regenerating the input data, thus ensuring the accuracy and real-time nature of the response. In addition, the technical solution of this application dynamically allocates computing resources according to the task complexity, directly using edge devices for processing in low-complexity tasks and using cloud servers for processing in high-complexity tasks. This dynamic routing method based on task complexity can effectively utilize resources, reduce computing overhead, and ensure real-time performance even in large-concurrency scenarios, reducing system burden and resource waste.
[0020] As a further improvement of a multi-turn conversation optimization method based on intelligent customer service in this application, in step S201, the initialized dynamic entity graph is:
[0021] G=(V, E)
[0022] where G is the dynamic entity graph, V is the set of nodes, representing all entities in the dynamic entity graph; E is the set of edges, representing the relationships between the entity nodes in the dynamic entity graph.
[0023] As a further improvement of a multi-turn conversation optimization method based on intelligent customer service in this application, for different entities in the multi-turn entity data, a new node v is added i ;
[0024] If the association degree between the new node v i and the existing node v j in the dynamic entity graph is higher than the threshold τ, then a corresponding edge is added between the node v i and the node v j , and the edge is added to the edge set E.
[0025] As a further improvement of a multi-turn conversation optimization method based on intelligent customer service in this application, for the nodes v i and the node v j between them, the edge association probability is calculated, and the edge weight of the edge between the nodes v i and the node v j is updated according to the edge probability calculation result. The edge weight represents the association strength between the nodes v i and the node v j ;
[0026] The calculation formula for the edge association probability P assoc (v i , v j ) is:
[0027] P assoc (vi , v j ) = σ(α·sim(v i , v j ) + β·co - occur(v i , v j ) + γ·contextual(v i , v j ))
[0028] Among them, σ(x) is the Sigmoid activation function; sim(v i , v j ) represents the similarity between node v i and node v j ; co - occur(v i , v j ) represents the frequency of co - occurrence of node v i and node v j in multiple rounds of conversations; contextual(v i , v j ) represents the correlation degree between node v i and node v j in multiple rounds of conversations calculated based on the attention calculation mechanism; α, β, γ are all weight coefficients, adjusting the importance of different calculation factors.
[0029] As a further improvement of a multi - round conversation optimization method based on intelligent customer service in this application, the audit confidence model is:
[0030] P audit = σ(W audit ·H audit + b audit )
[0031] Among them, σ is the Sigmoid activation function, W audit and b audit are the weight and bias term of the audit model respectively, and P audit is the audit confidence.
[0032] As a further improvement of a multi - round conversation optimization method based on intelligent customer service in this application, the task complexity calculation formula is:
[0033] C = β1·L input + β2·|E|
[0034] Among them, L input is the total length of entities in the dynamic entity graph, |E| is the number of entities in the dynamic entity graph, and β1, β2 are adjustment weights.
[0035] As a further improvement of a multi-turn dialogue optimization method based on intelligent customer service in this application, the steps of the intention recognition method in step S301 are as follows:
[0036] S302. Construct a context feature vector based on the existing entities in the dynamic entity graph, and use a classification model to map the context feature vector into different intention categories;
[0037] S303. When the recognized intention is a long-tail intention, use a meta-learning model to perform additional training and recognition on the intention.
[0038] As a further improvement of a multi-turn dialogue optimization method based on intelligent customer service in this application, in step S201, when the input data provided by the user in multiple turns cannot generate entity data through entity recognition but there is a pronoun reference, it is matched with the entities in the dynamic entity graph through a multi-head attention mechanism. After coreference resolution, slot filling is performed through step S301.
[0039] As a further improvement of a multi-turn dialogue optimization method based on intelligent customer service in this application, in step S101, if the input data is of the voice input type, the ASR module is called to perform speech recognition on the input data of the voice input type to generate text-based input data.
[0040] As a further improvement of a multi-turn dialogue optimization method based on intelligent customer service in this application, in step S501, the natural language output to the user is a reply in text or voice form, and it is displayed on the interface or played to the user through the TTS module to implement multi-turn voice interaction of the intelligent customer service. Description of the Drawings
[0041] The drawings described herein are used to provide a further understanding of this application and form a part of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0042] Figure 1 is one of the workflow diagrams of this application;
[0043] Figure 2 is the second workflow diagram of this application. Detailed Embodiments
[0044] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used in the specification of the present application herein are only for the purpose of describing specific embodiments and are not intended to limit the present application.
[0045] Although the present application is disclosed below in preferred embodiments, it is not used to limit the claims. Any person skilled in the art can make several possible changes and modifications without departing from the concept of the present application. Therefore, the protection scope of the present application should be determined by the scope defined by the claims of the present application.
[0046] Next, in combination with specific embodiments, the present application will be further described in detail, but the embodiments of the present application are not limited thereto.
[0047] Embodiment 1
[0048] As Figure 1 shown, it is known that existing intelligent customer service systems often cannot effectively manage and associate context information in multi-round conversations, especially cross-round entity references, and it is easy to have information loss or repeated inquiries, resulting in a poor user experience. Through the introduction of a dynamic entity graph, the technical solution of the present application enables the system to update entities and the relationships between entities in real time after each round of conversation, enabling the system to better capture context information and dynamically track and maintain entity relationships in multiple conversation rounds. The data of each round of conversation can update the entity graph to maintain the coherence of the conversation and the integrity of the context.
[0049] Furthermore, when facing complex conversations, the current intelligent customer service system requires a large amount of manual intervention to review the responses generated by the system, which leads to a high response delay and high operating costs. The present invention reduces the frequency of manual intervention by introducing an automated review model. When generating a response, the system calculates the review confidence to judge the quality of the response. If the review confidence exceeds the preset threshold, the system automatically generates a response and feeds it back to the user; if the review confidence is lower than the threshold, it will be regenerated and reviewed, greatly reducing the frequency of manual review and improving the response efficiency.
[0050] In addition, in the existing system under high-concurrency scenarios, the inference path is fixed, resulting in low resource utilization and inference latency that cannot meet the real-time performance requirements. In contrast, the present invention introduces a task complexity calculation mechanism to dynamically allocate different computing resources for tasks with different complexities. When the task complexity is low, a lightweight model on the edge device is used for processing; when the task complexity is high, it switches to the cloud server for processing using a full-featured model. Through this dynamic routing strategy, the system can efficiently utilize computing resources, optimize performance in different scenarios, and ensure the response speed and real-time performance under high concurrency.
[0051] Furthermore, the technical solution adopted by this application to solve the above technical defects is as follows: A multi-turn dialogue optimization method based on intelligent customer service, comprising the following steps:
[0052] S101. Receive the input data provided by the user. If the input data is of the voice input type, perform speech recognition on the voice input type of input data to generate text-based input data, perform word segmentation on the text-based input data, and perform noise filtering on the input data after word segmentation to generate a segmented sequence data. Perform feature vector encoding on the segmented sequence data and map the segmented sequence to a vector sequence;
[0053] S201. Based on the trained SpanBERT model, perform entity extraction on the vector sequence to obtain entity data, initialize a dynamic entity graph, and add the extracted entity data to the dynamic entity graph;
[0054] The user generates multi-turn vector sequences through multi-turn provided input data and generates multi-turn entity data through the SpanBERT model. The different entity outputs in the multi-turn entity data are input into the dynamic entity graph to update the dynamic entity graph;
[0055] S301. Perform intention recognition on the dynamic entity graph and generate one or more slots corresponding to the intention, and fill the entities in the entity data into the corresponding slots according to the information in the dynamic entity graph;
[0056] S401. After completing the slot filling, generate a reply code H for the reply information to be generated according to the intention and the filled slot information audit , and then calculate the review confidence P of the reply code according to the review confidence model audit ;
[0057] When the review confidence P audit exceeds the confidence threshold, execute step S501. When the review confidence P audit is lower than the confidence threshold, return to step S101;
[0058] S501. Calculate the task complexity C according to the entities in the dynamic entity graph and output the task complexity C;
[0059] When the task complexity C is less than the complexity threshold, a natural language response is generated based on the proposed response information and the context state is updated in real time and fed back to the user;
[0060] When the task complexity C is greater than the complexity threshold, the dynamic entity graph is sent to the cloud server to generate natural language using the full-functional model and the context state is updated in real time and fed back to the user.
[0061] Among them, for the preprocessing and feature extraction of the input data, after receiving the user input data, speech recognition is first performed (if the input is speech), then word segmentation is performed on the text input, then noise filtering is performed on the word segmentation result, and finally a feature vector is generated and mapped into a vector sequence. This step ensures that the system can correctly process and understand different types of user inputs and lays a foundation for subsequent entity recognition and intent recognition.
[0062] Furthermore, entity extraction and the dynamic entity graph initialization are to use the trained SpanBERT model to extract entities from the input vector sequence and initialize the extracted entity data into the dynamic entity graph. As the user provides more data through multiple rounds of conversations, the graph will be continuously updated. This process helps the system maintain and manage the entities and their relationships in multiple rounds, thus avoiding information loss or repeated questions. It should be noted that SpanBERT is an improvement of the BERT (Bidirectional Encoder Representations from Transformers) model, which is specifically optimized for cross-sentence entity relationship modeling. Its main goal is to improve the representation ability of entities, especially in tasks involving entity relationships (such as entity extraction and coreference resolution, etc.). Its key features include:
[0063] 1) Masking continuous word segments (Span): Different from the word masking in BERT, SpanBERT masks a continuous text segment during training (i.e., a span instead of a single word). This way can let the model learn to understand a longer range of context, which is particularly important for dealing with the relationships between entities.
[0064] 2) Optimizing cross-sentence modeling: SpanBERT pays more attention to the relationships and connections between words during the learning process and can better capture the deep dependencies between entities, especially the information dissemination in multi-round conversations or cross-passage texts.
[0065] Further, after the user inputs the information: "I want to book a meeting room at 2:00 pm tomorrow", the word segmentation result is: {I, want, book, tomorrow, afternoon, 2:00, of, meeting room}; further using pre-trained Word2Vec or BERT-Embedding, the word segmentation sequence is mapped to a vector sequence. For the vector sequence, SpanBERT is used for entity extraction to obtain entities:
[0066] E1 = {e1 = (time, 2:00 pm tomorrow), e2 = (intention, book meeting room)}
[0067] Further, initialize the dynamic entity graph:
[0068] G = (V, E), V = {v1, v2}, v1 = e1, v2 = e2
[0069] Further, when the user's next input is "The number of participants is about 10 people", then detect the new entity:
[0070] e3 = (number of people, 10 people)
[0071] Add node v3 and calculate the association probability P assov (v1, v3), if it is greater than the threshold τ, then establish an edge (v1, v3) and update the edge weight.
[0072] Further, if the user uses the pronoun "it" to refer to the "meeting room", the system calculates the reference score through the multi-head attention mechanism, selects the node that best matches the "meeting room" entity as the reference object, and completes the cross-turn reference. Through the lightweight EBTA model, intent recognition and slot filling are performed on the input; in case of long-tail intents, training and recognition are carried out through the meta-learning framework.
[0073] Further, combine the conversation history and the system's proposed response to obtain the response encoding H audit , and then calculate the review confidence P of the response encoding according to the review confidence model audit and the task complexity C; if the review confidence P audit is lower than the confidence threshold, then return to step S101 to re-perform entity extraction. If C is low and the review confidence P audit is higher than the confidence threshold, then directly generate a response to the user on the edge device; otherwise, send the request to the cloud for secondary judgment using a more powerful model.
[0074] Finally, the system generates a natural language response, such as "I have booked a meeting room for you at 2:00 pm tomorrow, which can accommodate 10 people to attend the meeting", and updates the conversation context status.
[0075] Embodiment 2
[0076] Different from Embodiment 1: In order to further improve the multi-turn dialogue optimization ability of the intelligent customer service of the present application, further, in step S201, the initialized dynamic entity graph is:
[0077] G=(V, E)
[0078] Where G is the dynamic entity graph, V is the set of nodes, representing all entities in the dynamic entity graph; E is the set of edges, representing the relationships between the entity nodes in the dynamic entity graph. For different entities in the multi-turn entity data, a new node v i is added; if the new node v i has a higher degree of association with the existing node v j in the dynamic entity graph than the threshold τ, then a corresponding edge is added between the node v i and the node v j , and the edge is added to the edge set E. For the nodes v i and the node v j , the edge association probability is calculated, and the edge weight of the edge between the node v i and the node v j is updated according to the calculation result of the edge probability. The edge weight represents the association strength between the node v i and the node v j ;
[0079] The formula for calculating the edge association probability P assoc (v i , v j ) is:
[0080] P assoc (v i , v j ) = σ(α·sim(v i , v j ) + β·co-occur(v i , v j ) + γ·contextual(v i , v j )
[0081] Where σ(x) is the Sigmoid activation function; sim(v i , v j ) represents the similarity between the node v i and the node v j ; co-occur(v i , v j ) represents the frequency of co-occurrence of the node v i and the node v j in the multi-turn dialogue; contextual(v i , v j)To represent node v in a multi-turn conversation calculated based on an attention calculation mechanism i and node v j correlation; α, β, and γ are all weight coefficients, adjusting the importance of different calculation factors.
[0082] Furthermore, σ(x) is the Sigmoid activation function, which normalizes the result to between [0, 1]. The specific calculation formula is:
[0083]
[0084] However, the similarity in sim(v i , v j ) can be calculated in different ways, such as based on the vector space model (e.g., using cosine similarity):
[0085]
[0086] In the above formula, v i and v j are the embedding vectors of node v i and node v j .
[0087] Suppose there are two nodes v1 and v3, where:
[0088] sim(v1, v3) = 0.8 (a relatively high cosine similarity, indicating that the semantics of these two entities are relatively close);
[0089] co-occur(v1, v3) = 5 (these two entities co-occur 5 times in the conversation);
[0090] contextual(v1, v3) = 0.6 (the correlation degree calculated through the context relationship).
[0091] Furthermore, setting the weight coefficients as α = 0.5, β = 0.3, and γ = 0.2, the edge correlation probability can be calculated:
[0092] P assoc (v1, v3) = σ(0.5 × 0.8 + 0.3 × 5 + 0.2 × 0.6)
[0093] First, calculate the weighted sum:
[0094] 0.5 × 0.8 = 0.4, 0.3 × 5 = 1.5, 0.2 × 0.6 = 0.12
[0095] The weighted sum is:
[0096] 0.4 + 1.5 + 0.12 = 2.02
[0097] Then, apply the Sigmoid function for normalization:
[0098]
[0099] Furthermore, the audit confidence model is:
[0100] P audit = σ(W audit ·H audit + b audit )
[0101] where σ is the Sigmoid activation function, W audit and b audit are the weight and bias term of the audit model respectively, and P audit is the audit confidence.
[0102] Furthermore, the task complexity calculation formula is:
[0103] C = β1·L input + β2·|E|
[0104] where L input is the total length of entities in the dynamic entity graph, |E| is the number of entities in the dynamic entity graph, and β1, β2 are adjustment weights.
[0105] For other parts that are the same as those in Embodiment 1, they will not be elaborated in this embodiment.
[0106] Embodiment 3
[0107] As Figure 1-2 shown, in order to further improve the ability of an intelligent customer service-based multi-turn dialogue optimization method of the present application for intent recognition and slot filling, further, construct a context feature vector according to the existing entities in the dynamic entity graph, and use a classification model to map the context feature vector to different intent categories; when the recognized intent is a long-tail intent, use a meta-learning model to perform additional training and recognition on the intent.
[0108] Specifically, it should be noted that: Intent recognition is to recognize the goal or requirement that the user input statement wants to express through a natural language understanding model (such as a classification model based on SpanBERT, EBTA model, Transformer model, etc.). Slot Filling refers to extracting key information (i.e., "slots") from the user input, such as parameters like "time", "number of people", "location", etc. when booking a service, so that the system can further perform corresponding actions.
[0109] Further, in step S201, when the input data provided by the user in multiple rounds cannot generate entity data through entity recognition but there is a pronoun reference, it is matched with the entities in the dynamic entity graph through the multi-head attention mechanism. After coreference resolution, slot filling is performed through step S301.
[0110] In addition, each intent corresponds to one or more slots. For example, the slots for "reserving a meeting room" may include:
[0111] Time, number of people, meeting room type, equipment requirements
[0112] According to the node information in the dynamic entity graph, the entities recognized from the user input are filled into the corresponding slots.
[0113] In the above example:
[0114] The system automatically fills the slots according to the entity graph node information: the time slot is filled with: "2:00 pm tomorrow"; the number of people slot is filled with: "10 people"; the equipment requirements and meeting room type are not mentioned, and the slots are left empty.
[0115] If the user input does not directly give a clear entity but there is a pronoun reference (such as "it", "this"), it is matched with the nodes in the dynamic entity graph through the multi-head attention mechanism. After coreference resolution, slot filling is performed.
[0116] Further, in step S101, if the input data is of the voice input type, the ASR module is called to perform speech recognition on the input data of the voice input type to generate text-based input data. In addition, in step S501, the natural language output to the user is a reply in text or voice form, and is displayed on the interface or played to the user through the TTS module to implement multi-round voice interaction of the intelligent customer service.
[0117] For those that are the same as those in Embodiment 1, they will not be elaborated in this embodiment.
[0118] Those skilled in the art should understand that the embodiments of the present invention can provide a method, a system, or a computer program product. Therefore, the present invention can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0119] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more flows and / or one or more blocks in the flowcharts and / or block diagrams. Figure 1 in one or more flows and / or one or more blocks Figure 1 in one or more blocks.
[0120] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in one or more flows and / or one or more blocks Figure 1 in one or more flows and / or one or more blocks Figure 1 in one or more blocks.
[0121] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or one or more blocks Figure 1 in one or more flows and / or one or more blocks Figure 1 in one or more blocks.
[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A multi-round dialogue optimization method based on intelligent customer service, characterized in that, It includes the following steps: S101. Receive the input data provided by the user. If the input data is of the voice input type, perform speech recognition on the input data of the voice input type to generate text-like input data, perform word segmentation processing on the text-like input data, perform noise filtering on the input data after word segmentation processing to generate a word segmentation sequence data, perform feature vector encoding on the word segmentation sequence data, and map the word segmentation sequence to a vector sequence; S201. Based on the trained SpanBERT model, perform entity extraction on the vector sequence to obtain entity data, initialize a dynamic entity graph, and add the extracted entity data to the dynamic entity graph; The user generates multiple rounds of vector sequences through the input data provided in multiple rounds and generates multiple rounds of the entity data through the SpanBERT model. For different entities in the multiple rounds of the entity data, the outputs are input into the dynamic entity graph to update the dynamic entity graph; S301. Perform intent recognition on the dynamic entity graph and generate one or more slots corresponding to the intent, and fill the entities in the entity data into the corresponding slots according to the information in the dynamic entity graph; S401. After slot filling is completed, generate a reply code H of the quasi-reply message according to the intention and the filled slot information audit , and then calculate the review confidence P of the reply code according to the review confidence model audit ; When the audit reliability P audit exceeds the reliability threshold, step S501 is executed. When the audit reliability P audit is lower than the reliability threshold, return to step S101; S501. Calculate the task complexity according to the entities in the dynamic entity graph and output the task complexity C; When the task complexity C is less than the complexity threshold, generate a natural language response according to the intended response information and update the context state in real time and feedback it to the user; When the task complexity C is greater than the complexity threshold, send the dynamic entity graph to the cloud server to generate natural language using a full-function model and update the context state in real time and feedback it to the user.
2. The multi-round dialogue optimization method based on intelligent customer service according to claim 1, wherein In the step S201, the initialized dynamic entity graph is: G=(V, E) where G is the dynamic entity graph, V is the set of nodes, representing all the entities in the dynamic entity graph; E is the set of edges, representing the relationships between the entity nodes in the dynamic entity graph.
3. The multi-round dialogue optimization method based on intelligent customer service according to claim 2, characterized in that, Add a new node v to different entities in multiple rounds of the entity data i ; If the new node v i has a higher degree of association with the existing node v j in the dynamic entity graph than the threshold τ, then add a corresponding edge between node v i and node v j , and add the edge to the edge set E.
4. A multi-round dialogue optimization method based on intelligent customer service according to claim 3, characterized in that, For node v i and node v j calculate the edge association probability between them, and update node v i and node v j according to the calculation result of the edge probability. The edge weight of the edge between them represents the association strength between node v i and node v j ; The edge association probability P assoc (v i ,v j ) is calculated by the formula: P assoc (v i ,v j ) = σ(α·sim(v i ,v j ) + β·co - occur(v i ,v j ) + γ·contextual(v i ,v j )) Among them, σ(x) is the Sigmoid activation function; sim(υ i , υ j ) represents the similarity between node v i and node v j ; co-occur(v i , v j ) represents the frequency of co-occurrence of node v i and node v j in multiple rounds of conversations; contextual(v i , v j ) represents the correlation degree between node v i and node v j in multiple rounds of conversations calculated based on the attention calculation mechanism; α, β, and γ are all weight coefficients to adjust the importance of different calculation factors.
5. The multi-round dialogue optimization method based on intelligent customer service according to claim 1, wherein The audit confidence model is: P audit = σ(W audit ·H audit + b audit ) Among them, σ is the Sigmoid activation function, W audit and b audit are the weight and bias terms of the review model respectively, and P audit is the review confidence level.
6. The multi-round dialogue optimization method based on intelligent customer service according to claim 1, wherein, The task complexity calculation formula is: C = β1·L input + β2·|E| where L input is the total length of the entities in the dynamic entity graph, |E| is the number of entities in the dynamic entity graph, and β1, β2 are adjustment weights.
7. A multi-round dialogue optimization method based on intelligent customer service according to claim 1, characterized in that, In the step S301, the steps of the intent recognition method are: S302. Construct a context feature vector according to the existing entities in the dynamic entity graph, and map the context feature vector to different intent categories by using a classification model; S303. When the recognized intent is a long-tail intent, use a meta-learning model to perform additional training and recognition on the intent.
8. A multi-round dialogue optimization method based on intelligent customer service according to claim 1, characterized in that, In the step S201, when the input data provided by the user in multiple rounds cannot generate entity data through entity recognition but there is a pronoun reference, match it with the entities in the dynamic entity graph through a multi-head attention mechanism, perform pronoun resolution, and then perform slot filling through the step S301.
9. The multi-round dialogue optimization method based on intelligent customer service according to claim 1, wherein In the step S101, if the input data is of the voice input type, call the ASR module to perform speech recognition on the input data of the voice input type to generate text-like input data.
10. A multi-round dialogue optimization method based on intelligent customer service according to claim 1, characterized in that, Output the response in the form of text or voice to the user in step S501, and display it on the interface or play it to the user through the TTS module to achieve multi-round voice interaction of the intelligent customer service.
Citation Information
Cited By
Front-end display method and system combining fixed text and streaming output
CN120804453A