Replay template recognition model training method, recognition method and device

By introducing training sets and topic distribution vectors, and utilizing technologies such as gating units and long short-term memory layers, response templates in customer service within the insurance industry are identified. This solves the problem of knowledge loss in unstructured dialogue records, achieving more accurate response template recognition and improving customer service quality.

CN120952137APending Publication Date: 2025-11-14CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511063126.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In the insurance industry, customer service agents' experience, scripts, and response strategies are scattered across massive amounts of unstructured dialogue records, making it difficult to systematically accumulate and resulting in serious knowledge loss. Traditional response template mining methods are not very effective.

Method used

By acquiring a training set, including dialogue texts from multiple topic sample sets, and training a pre-defined model, the topic distribution vectors of the dialogue texts are introduced, and gating units, fully connected layers, long short-term memory layers, and conditional random field layers are used to output a recognition response template.

Benefits of technology

It improves the accuracy of response template recognition and the generalization ability of the model, helping agents answer questions more accurately, reducing knowledge loss, and enhancing customer experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952137A_ABST
    Figure CN120952137A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a training method, a recognition method and a device for a recognition model of a reply template, and belongs to the technical field of artificial intelligence, the method can be applied to the fields of finance or health care and the like, and the method comprises the following steps: obtaining a training set; wherein the training set comprises dialogue text samples, each dialogue text sample comprises a dialogue text, a reply template tag corresponding to the dialogue text and a topic distribution vector of the dialogue text, and the topic distribution vector is used for representing the probability that the dialogue text belongs to each preset topic in different preset topics; training a preset model by adopting the training set to obtain a target recognition model; and inputting the target question content into the target identification model to obtain a target reply template. According to the embodiment of the invention, the model can be trained by introducing the topic distribution vector, so that the performance and generalization ability of the model can be improved, the recognition accuracy of the model on the reply template is improved, and a more effective target reply template is output during subsequent application of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a training method, recognition method, and apparatus for a response template recognition model. Background Technology

[0002] In the insurance industry, historical dialogue records between customer service agents and customers constitute a valuable tacit knowledge base for enterprises. However, this data typically exists in unstructured text form, making it difficult to analyze and utilize directly. The experience, techniques, and response strategies of excellent agents are often scattered across massive amounts of dialogue records, hindering systematic accumulation and leading to significant knowledge loss for the enterprise. As the scale of insurance business continues to expand and service scenarios become increasingly complex, efficiently extracting response templates from historical dialogues has become a core challenge for improving service efficiency, optimizing operational models, and enhancing customer experience.

[0003] Traditional methods for mining response templates mainly rely on the experience of operations personnel and manual analysis. This not only makes it difficult to cover massive amounts of historical data, but also makes it prone to biases in information extraction due to subjective factors. As a result, traditional methods for mining response templates have low effectiveness. Summary of the Invention

[0004] The main objective of this application is to propose a training method, recognition method, device, electronic device, and storage medium for a response template recognition model, aiming to improve the effectiveness of response template recognition.

[0005] To achieve the above objectives, a first aspect of this application proposes a method for training a response template recognition model, the method comprising:

[0006] Obtain a training set; wherein the training set includes multiple topic sample sets, each topic sample set belongs to a different preset topic, each topic sample set includes multiple dialogue text samples marked by corresponding preset topic identifiers, each dialogue text sample includes dialogue text, a reply template tag corresponding to the dialogue text, and a topic distribution vector of the dialogue text, the dialogue text includes the question content raised by the questioner and the reply content made by the respondent to the question content, and the topic distribution vector is used to characterize the probability that the dialogue text belongs to each preset topic in the different preset topics;

[0007] The preset model is trained using the training set until the preset model meets the training stopping condition to obtain the target recognition model; wherein, the target recognition model is used to identify a response template for assisting the respondent in responding to the question content.

[0008] In some embodiments, training a preset model using the training set until the preset model meets the training stopping condition to obtain a target recognition model includes: for each dialogue text sample in each topic sample set of the training set, performing word segmentation on the dialogue text in the dialogue text sample to obtain multiple target words of the dialogue text, and encoding the multiple target words to obtain word embedding vectors corresponding to the dialogue text; concatenating the word embedding vectors with the topic distribution vectors corresponding to the dialogue text to obtain a composite vector; and outputting a predicted response template based on the composite vector; determining a loss function value according to the predicted response templates corresponding to each dialogue text sample in each topic sample set of the training set and the response template labels in each dialogue text sample; obtaining a target recognition model when the loss function value reaches the training stopping condition; and adjusting the network parameters of the preset model when the loss function value does not reach the training stopping condition, and returning to the step of performing word segmentation on each dialogue text sample in each topic sample set of the training set to obtain multiple target words of the dialogue text, until the preset model meets the training stopping condition.

[0009] In some embodiments, the aforementioned preset model comprises a gating unit, a fully connected layer, a long short-term memory layer, and a conditional random field layer. The predicted response template based on the composite vector output recognition includes: inputting the composite vector into the gating unit and the long short-term memory layer; using the gating unit to extract temporal features from the composite vector to obtain a temporal feature vector; inputting the temporal feature vector into the fully connected layer; using the fully connected layer to map the temporal feature vector to a weight space with the same preset input embedding dimension to generate a weight vector; wherein the dimension of the weight vector is the input embedding dimension; inputting the weight vector into the long short-term memory layer; using the long short-term memory layer to perform element-wise multiplication of the weight vector and the composite vector to obtain a topic enhancement vector; inputting the topic enhancement vector into the conditional random field layer; and using the conditional random field layer to output the predicted response template based on the topic enhancement vector.

[0010] In some embodiments, the above-mentioned acquisition of training set includes: acquiring multiple dialogue texts; determining the target preset topic to which the multiple dialogue texts belong; wherein the target preset topic is determined from a plurality of preset topics; and dividing the multiple dialogue texts into multiple topic sample sets according to the target preset topic to which the multiple dialogue texts belong.

[0011] In some embodiments, the above-mentioned determination of the target preset topics to which the plurality of dialogue texts belong includes: for each dialogue text in the plurality of dialogue texts, inputting the dialogue text into a preset topic model to obtain the topic distribution corresponding to the dialogue text output by the preset topic model; wherein, the preset topic model is obtained by training a word matrix, the word matrix is ​​obtained by text vectorization processing of the plurality of dialogue text sets, each row of the word matrix represents the corresponding dialogue text set, each column represents the corresponding preset word, and the element value of the word matrix represents the number of times the corresponding preset word appears in the dialogue text set, the dialogue text set includes multiple rounds of dialogue between the questioner and the respondent, and the topic distribution is used to characterize the probability that the dialogue text belongs to each preset topic in the plurality of preset topics; for each dialogue text in the plurality of dialogue texts, the preset topic with the highest probability in the topic distribution corresponding to the dialogue text is determined as the target preset topic to which the dialogue text belongs.

[0012] To achieve the above objectives, a second aspect of this application proposes a method for identifying response templates, the method comprising:

[0013] Obtain the content of the target question input by the target questioner;

[0014] The target question content is input into the target recognition model to obtain the target response template output by the target recognition model; wherein, the target response template is used to assist the respondent in responding to the target question content, and the target recognition model is obtained by the training method of the recognition model of the response template as described in any one of the first aspects.

[0015] To achieve the above objectives, a third aspect of this application provides a training method apparatus for a response template recognition model, the apparatus comprising:

[0016] A training set acquisition module is used to acquire a training set; wherein, the training set includes multiple topic sample sets, each topic sample set belongs to a different preset topic, each topic sample set includes multiple dialogue text samples labeled with corresponding preset topic tags, each dialogue text sample includes dialogue text, a reply template tag corresponding to the dialogue text, and a topic distribution vector of the dialogue text, the dialogue text includes the question content raised by the questioner and the reply content of the respondent to the question content, and the topic distribution vector is used to characterize the probability that the dialogue text belongs to each preset topic in the different preset topics;

[0017] The model training module is used to train a preset model using the training set until the preset model meets the training stopping condition to obtain a target recognition model; wherein, the target recognition model is used to recognize a response template for assisting the respondent in responding to the question content.

[0018] To achieve the above objectives, a fourth aspect of this application provides a device for identifying response templates, the device comprising:

[0019] The question acquisition module is used to acquire the content of the target question input by the target questioner;

[0020] The template recognition module is used to input the target question content into the target recognition model to obtain the target response template output by the target recognition model; wherein, the target response template is used to assist the respondent in responding to the target question content, and the target recognition model is obtained by the training method of the above-mentioned response template recognition model.

[0021] To achieve the above objectives, a fifth aspect of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the training method for the recognition model of the response template described in the first aspect.

[0022] Alternatively, the method for recognizing the response template described in the second aspect above can be implemented.

[0023] To achieve the above objectives, a sixth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the training method for the recognition model of the response template described in the first aspect.

[0024] Alternatively, the method for recognizing the response template described in the second aspect above can be implemented.

[0025] The present application proposes a training method, recognition method, apparatus, device, and storage medium for a response template recognition model. This involves acquiring a training set, which includes multiple topic sample sets, each belonging to a different preset topic. Each topic sample set includes multiple dialogue text samples labeled with corresponding preset topic identifiers. Each dialogue text sample includes the dialogue text, a corresponding response template label, and a topic distribution vector. The dialogue text includes the question content raised by the questioner and the response content of the respondent. The topic distribution vector represents the probability that the dialogue text belongs to each of the different preset topics. The training set is used to train a preset model until the model meets the training stopping condition, resulting in a target recognition model. This target recognition model is used to recognize response templates used to assist the respondent in responding to questions. The target question content input by the target questioner is obtained; the target question content is input into the target recognition model to obtain the target response template output by the target recognition model. Thus, by introducing the topic distribution vector of the dialogue text to train the model, the performance and generalization ability of the preset model can be improved, thereby increasing the accuracy of the model in recognizing response templates and ultimately improving the output of more effective target response templates when the model is subsequently used. Attached Figure Description

[0026] Figure 1 This is a flowchart of the training method for the recognition model of the reply template provided in the embodiments of this application;

[0027] Figure 2 yes Figure 1 The flowchart of step S102 in the document;

[0028] Figure 3 yes Figure 1 The flowchart of step S201 in the text;

[0029] Figure 4 yes Figure 1 The flowchart of step S101 in the text;

[0030] Figure 5 yes Figure 4 The flowchart of step S401 in the text;

[0031] Figure 6 yes Figure 4 The flowchart of step S402 in the document;

[0032] Figure 7 yes Figure 6 A flowchart of the steps preceding step S601;

[0033] Figure 8 This is a flowchart of the method for identifying the response template provided in the embodiments of this application;

[0034] Figure 9 This is a schematic diagram of the structure of the training device for the recognition model of the response template provided in the embodiments of this application;

[0035] Figure 10 This is a schematic diagram of the structure of the response template recognition device provided in the embodiments of this application;

[0036] Figure 11 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0038] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0040] First, let's analyze some of the terms used in this application:

[0041] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0042] Natural Language Processing (NLP): NLP uses computers to process, understand, and utilize human language (such as Chinese and English). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. NLP includes syntactic analysis, semantic analysis, and discourse understanding. It is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, intent recognition, information extraction and filtering, text classification and clustering, sentiment analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computation.

[0043] Information extraction is a text processing technique that extracts factual information such as entities, relationships, and events from natural language text and outputs it as structured data. Information extraction is a technique for extracting specific information from text data. Text data is composed of specific units, such as sentences, paragraphs, and chapters. Text information is composed of smaller, specific units, such as characters, words, phrases, sentences, paragraphs, or combinations of these units. Extracting noun phrases, names of people, and place names from text data is an example of text information extraction. Of course, text information extraction techniques can extract information of various types.

[0044] Based on this, embodiments of this application provide a training method and apparatus for a response template recognition model, an electronic device, and a storage medium, aiming to improve the accuracy of response template recognition.

[0045] The methods, apparatus, electronic devices, and storage media provided in the embodiments of this application are specifically described through the following embodiments. First, the methods in the embodiments of this application are described.

[0046] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0047] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0048] The training method for the recognition model of the reply template provided in this application relates to the field of artificial intelligence technology. The training method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the training method for the recognition model of the reply template, etc., but is not limited to the above forms.

[0049] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0050] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0051] Figure 1 This is an optional flowchart of a training method for a response template recognition model provided in an embodiment of this application. Figure 1The method may include, but is not limited to, steps S101 to S102.

[0052] Step S101: Obtain the training set.

[0053] Step S102: Train the preset model using the training set until the preset model meets the training stopping condition to obtain the target recognition model.

[0054] The training set includes multiple topic sample sets, each belonging to a different preset topic, and each topic sample set includes multiple dialogue text samples labeled with corresponding preset topic identifiers.

[0055] Each topic sample set can correspond to the same domain or different domains. For example, each topic sample set can be data from the vehicle insurance domain, data from the health insurance domain, or data from the banking domain, and so on.

[0056] The preset topics mentioned above can be set according to needs and can be related to the corresponding fields. For example, the preset topics are divided into five themes: product consultation, complaint handling, claims process, renewal strategy, and service rights.

[0057] The aforementioned preset theme identifier can be a unique identifier for the preset theme. Each preset theme can have a corresponding unique preset theme identifier.

[0058] Each dialogue text sample includes the dialogue text, the corresponding reply template tags, and the topic distribution vector of the dialogue text. The dialogue text includes the question content raised by the questioner and the reply content of the respondent.

[0059] The questioner can be different or the same in different dialogue texts, and the respondent can be different or the same in different dialogue texts.

[0060] The training methods described above can be applied to question-and-answer / dialogue systems. The questioner can be a customer, and the respondent can be a customer service agent. A customer service agent is a staff member who serves customers and answers their questions. A customer service representative (CSR) is a frontline service personnel in an enterprise or organization who are dedicated to interacting directly with customers, resolving customer problems, and providing support through channels such as telephone, online chat, email, and social media. Their core responsibility is to represent the company's image, resolve customer needs through efficient and professional communication, and improve customer satisfaction and loyalty.

[0061] The above response template tags are response templates corresponding to the dialogue text, which can be used to calculate the loss between the response templates identified by the preset model during preset model training.

[0062] In one implementation, the aforementioned response template tags can be annotated using the BIOES tag system. In the extended BIOES tag system, each tag is followed by the specific entity type to provide more detailed information about the entity's category. For example:

[0063] B-CLAIM_TYPE: Indicates the beginning of the entity "Vehicle Damage Insurance" and that the type of this entity is "CLAIM_TYPE" (claim type);

[0064] I-CLAIM_TYPE: Represents the middle part of the entity "Vehicle Damage Insurance", and the type of this entity is "CLAIM_TYPE";

[0065] E-CLAIM_TYPE: Indicates the end part of the entity "Vehicle Damage Insurance", and the type of this entity is "CLAIM_TYPE";

[0066] S-CLAIM_TYPE: Represents a single-word entity, and the type of this entity is "CLAIM_TYPE";

[0067] O: Indicates an irrelevant character or word.

[0068] The aforementioned topic distribution vector is used to represent the probability that the dialogue text belongs to each of the aforementioned preset topics.

[0069] The aforementioned preset models can be Convolutional Neural Network (CNN), Recurrent Neural Network (RNN) and its variant Long Short-Term Memory Network (LSTM), etc. In addition, they can also be decision trees, random forests, K-Nearest Neighbor Algorithm (KNN), etc.

[0070] The target recognition model described above is used to identify response templates that assist respondents in answering questions. Response templates can be understood as tools designed to help respondents answer questions more accurately and effectively; they belong to key communication patterns and provide respondents with a design approach or framework for their responses.

[0071] The above response template can be understood as a set of scripts that, in specific scenarios, can directly trigger positive emotions in customers, drive the conversation toward the expected outcome, and reflect service standards. It can be a response strategy. For example, if a customer is worried about rising premiums next year and hesitates about making a claim, a customer service agent could reply: "I completely understand your concern about rising premiums. Let me calculate it for you using the back-end system: the claim amount this time is 2800 yuan. If you pay for the repairs yourself, the no-claims discount next year will be 850 yuan; going through insurance is more cost-effective. I'll give you a clear answer within 2 minutes, and you can decide later."

[0072] The aforementioned response template can also be understood as identifying the core vocabulary in the questioner's content to assist the respondent in strategic expression and goal-oriented dialogue, and to propose communication strategies for effectively resolving core issues. For example, in online customer service scenarios, customers may ask complex questions covering multiple aspects. For instance, a customer might ask, "Why hasn't my order been shipped yet? I saw that others have already received theirs, and I previously inquired about price discounts. Could you give me some more discounts?" In this case, the customer service representative's response template could be the extracted core issues (such as shipping delays and price discounts), and a strategic response based on core vocabulary such as shipping delays and price discounts. For example, "I'm very sorry to have kept you waiting. Regarding the shipping issue, I will check the order status for you immediately. As for price discounts, we currently have an exclusive member discount program, which I can apply for for you. Would it be convenient for you to provide your membership information?" This model not only quickly resolves the customer's core issues but also improves customer satisfaction through guided questioning and proactive responses.

[0073] Response templates are language expressions extracted from conversations that efficiently and accurately convey core information, resolve problems, or achieve communication goals. In customer service, the definition and application of response templates can help agents communicate more effectively with customers, improve service quality, reduce knowledge loss, shorten the training cycle for new agents, and ultimately enhance the customer experience.

[0074] Introducing topic distribution vectors (DDVs) into the dialogue text during training of the model for recognizing response templates can significantly improve model performance. DMVs provide the model with the overall semantic context of the dialogue, helping it to more accurately identify keywords and expressions relevant to the dialogue topic. For example, when handling customer inquiries about after-sales issues, DMVs allow the model to quickly focus on core words such as "returns" and "repairs," rather than being distracted by irrelevant high-frequency words. This enhanced semantic context not only improves the model's accuracy in identifying key patterns but also enhances its generalization ability across different topical dialogues. Furthermore, DMVs provide a clearer basis for the model's decisions, making its decision-making process easier to understand and interpret.

[0075] Steps S101 to S102 shown in the embodiments of the present application can improve the performance and generalization ability of the preset model by introducing the topic distribution vector of the dialogue text, thereby improving the recognition accuracy of the model for the reply template.

[0076] Please refer to Figure 2 , in some embodiments, step S202 may include, but is not limited to, steps S201 to S203:

[0077] Step S201, for each dialogue text sample in each topic sample set of the training set, perform word segmentation on the dialogue text in the dialogue text sample to obtain multiple target word segments of the dialogue text, and encode the multiple target word segments to obtain the word embedding vector corresponding to the dialogue text. Concatenate the word embedding vector with the topic distribution vector corresponding to the dialogue text to obtain a composite vector, and output the predicted reply template recognized based on the composite vector.

[0078] Step S202, determine the loss function value according to the predicted reply template corresponding to each dialogue text sample in each topic sample set of the training set, and the reply template label in each dialogue text sample.

[0079] Step S203, when the loss function value reaches the training stop condition, obtain the target recognition model. When the loss function value does not reach the training stop condition, adjust the network parameters of the preset model, and return to execute the step of performing word segmentation on the dialogue text in each dialogue text sample in each topic sample set of the training set to obtain multiple target word segments of the dialogue text, until the preset model meets the training stop condition.

[0080] In step S201 of some embodiments, the above-mentioned performing word segmentation on the dialogue text in the dialogue text sample to obtain multiple target word segments of the dialogue text can be to perform word segmentation on the dialogue text in the dialogue text sample by using any one of the preset word segmentation methods or word segmentation models to obtain multiple target word segments of the dialogue text. Among them, the above-mentioned word segmentation method can be the forward maximum matching method, the backward maximum matching method, etc., and the above-mentioned word segmentation model can be the hidden Markov model, the conditional random field, etc. Using the preset word segmentation method or word segmentation model, according to the preset dictionary or model parameters, the continuous dialogue text is segmented into independent word units, that is, target word segments.

[0081] For example, for the Chinese dialogue text "I went to the library to study today", the word segmentation tool may segment it into words such as "I / today / go / to the library / study / ". These target word segments can more clearly express the semantic structure of the text and provide basic data support for subsequent text analysis, semantic understanding or model training.

[0082] In some embodiments, the aforementioned preset model may include a word embedding layer. The above-mentioned word segmentation processing of the dialogue text sample to obtain multiple target words of the dialogue text, and encoding of the multiple target words to obtain word embedding vectors corresponding to the dialogue text, and concatenating the word embedding vectors with the topic distribution vectors corresponding to the dialogue text to obtain a composite vector, may be performed by using a word embedding layer to segment the dialogue text sample to obtain multiple target words of the dialogue text, using the Transformer encoder of the BERT model (Bidirectional Encoder Representations from Transformers) in the word embedding layer to encode the multiple target words to obtain word embedding vectors, and concatenating the word embedding vectors with the topic distribution vectors corresponding to the dialogue text to obtain a composite vector.

[0083] In step S202 of some embodiments, the determination of the loss function value based on the predicted response template corresponding to each dialogue text sample in each topic sample set in the training set, and the response template label in each dialogue text sample, can be achieved by using a preset loss function. The preset loss function can be Cross-Entropy Loss, Mean Squared Error (MSE), or Binary Cross-Entropy Loss, etc.

[0084] In step S203 of some embodiments, if the loss function value reaches the training stopping condition, the preset model at this time meets the preset requirements. Therefore, the current preset model can be a target recognition model. If the loss function value does not reach the training stopping condition, the preset model at this time does not meet the preset requirements and needs to continue training until the training stopping condition is reached.

[0085] In this embodiment, word embedding vectors are concatenated with topic distribution vectors corresponding to the dialogue text to form composite vectors. This provides the model with richer semantic information, thereby improving its performance in the response template recognition task. Word embedding vectors capture the semantic features and contextual relationships of words, while topic distribution vectors provide the overall semantic background and topic information of the dialogue. By combining the two, the model can not only understand the meaning of individual words but also accurately identify response templates within a broader dialogue topic context.

[0086] The aforementioned pre-defined model can consist of gated units (GRU units), fully connected layers, long short-term memory layers (LSTM layers), and conditional random field layers (CRF layers). Please refer to [link / reference]. Figure 3 In some embodiments, step S201 may include, but is not limited to, steps S301 to S304:

[0087] Step S301: Input the composite vector into the gating unit and the long short-term memory layer, and use the gating unit to extract the temporal features of the composite vector to obtain the temporal feature vector.

[0088] Step S302: Input the temporal feature vector into the fully connected layer, and use the fully connected layer to map the temporal feature vector to a weight space with the same preset input embedding dimension to generate a weight vector.

[0089] The dimension of the weight vector is the same as the input embedding dimension.

[0090] Step S303: Input the weight vector into the long short-term memory layer, and use the long short-term memory layer to multiply the weight vector and the composite vector element by element to obtain the topic enhancement vector.

[0091] Step S304: Input the topic enhancement vector into the conditional random field layer, and use the conditional random field layer to output the predicted response template based on the topic enhancement vector.

[0092] In step S301 of some embodiments, the aforementioned temporal feature vector can represent topic information and context information.

[0093] In step S303 of some embodiments, the above-mentioned input of the weight vector to the long short-term memory layer may be achieved by normalizing the weight vector to obtain a normalized weight vector, and then inputting the normalized weight vector into the long short-term memory layer. Normalization ensures that the weights are within the range [0,1], avoiding gradient explosion or vanishing.

[0094] The above method of using a long short-term memory layer to multiply the weight vector and the composite vector element-wise to obtain the topic enhancement vector can also be achieved by using a long short-term memory layer to multiply the normalized weight vector and the composite vector element-wise to obtain the topic enhancement vector.

[0095] Assume the weight vector and the composite vector are w and v, respectively:

[0096] w=[0.5,1.0,0.2],v=[3,4,5]

[0097] Then, the result of element-wise multiplication of w and v is u:

[0098] u=w⊙v=[0.5×3,1.0×4,0.2×5]=[1.5,4.0,1.0].

[0099] In step S304 of some embodiments, the predicted response template is labeled by the BIOES labeling system. Therefore, the predicted response template can be a label sequence output by the conditional random field layer.

[0100] In this embodiment, by introducing a GRU gating unit before the LSTM layer and dynamically adjusting the input weights of the LSTM in conjunction with topic augmentation, the model can more accurately capture key contextual information related to the topic. Furthermore, the constraint-driven CRF layer outputs the globally optimal label sequence, ensuring that the label sequence is more consistent and reasonable overall, conforming to business logic. The multi-layered model network structure complements each other and has a clear division of labor, jointly improving model performance and recognition accuracy.

[0101] Please see Figure 4 In some embodiments, step S101 may include, but is not limited to, steps S401 to S403:

[0102] Step S401: Obtain multiple dialogue texts.

[0103] Step S402: Determine the target preset topic to which multiple dialogue texts belong.

[0104] The target preset topic is determined from a plurality of preset topics.

[0105] Step S403: Based on the target preset topic to which the multiple dialogue texts belong, divide the multiple dialogue texts into multiple topic sample sets.

[0106] For example, there are multiple preset topics including five themes: product consultation, complaint handling, claims process, renewal strategy, and service benefits. Multiple dialogue texts include Dialogue Text A, Dialogue Text B, Dialogue Text C, Dialogue Text D, Dialogue Text E, Dialogue Text F, Dialogue Text G, Dialogue Text H, Dialogue Text J, Dialogue Text K, Dialogue Text L, and Dialogue Text M. The target preset topics corresponding to Dialogue Texts A, B, C, D, E, F, G, H, J, K, L, and M are respectively: product consultation, complaint handling, claims process, renewal strategy, and service benefits. For consultation, there are multiple topic sample sets, including the first topic sample set, the second topic sample set, the third topic sample set, the fourth topic sample set, and the fifth topic sample set. The preset topics corresponding to the first topic sample set, the second topic sample set, the third topic sample set, the fourth topic sample set, and the fifth topic sample set are product consultation, complaint handling, claims process, renewal strategy, and service rights, respectively. The first topic sample set includes dialogue text A, dialogue text F, dialogue text G, and dialogue text M; the second topic sample set includes dialogue text B and dialogue text H; the third topic sample set includes dialogue text C and dialogue text J; the fourth topic sample set includes dialogue text D and dialogue text K; and the fifth topic sample set includes dialogue text E and dialogue text L.

[0107] In this embodiment, dividing the dialogue text into multiple topic sample sets according to its subject allows the model to focus more on feature learning for specific topics and reduces noise interference. At the same time, this classification method helps enhance the model's generalization ability, enabling it to more accurately identify and process new dialogues on different topics.

[0108] Please see Figure 5 In some embodiments, step S401 may also include, but is not limited to, steps S501 to S502:

[0109] Step S501: Obtain the dialogue text set.

[0110] The aforementioned dialogue text set includes multiple rounds of dialogue between the questioner and the respondent.

[0111] Step S502: Perform single-turn dialogue segmentation on the dialogue text set to obtain multiple dialogue texts.

[0112] In step S501 of some embodiments, the aforementioned dialogue text set may include multiple dialogue text sets, each of which may include multiple rounds of dialogue between the questioner and the respondent. The questioners corresponding to different dialogue text sets may be different, and the respondents corresponding to different dialogue text sets may also be different.

[0113] The aforementioned dialogue text sets can be obtained by exporting customer consultation dialogues from a customer service system and by collecting dialogue data in specific scenarios using specialized data collection tools or services. These dialogue text sets can exist in document form; that is, one dialogue text set constitutes one document.

[0114] In some embodiments, step S502 can be implemented by recognizing line breaks in the dialogue, specific dialogue separators (such as "—" or ":"), or based on the alternating speech of the dialogue subjects (such as the user and the system). For example, in a dialogue record between a user and customer service, it can be divided into multiple single-turn dialogues based on the alternating speech of the user and customer service, each single-turn dialogue containing a question from the user and a response from the customer service.

[0115] In this embodiment, the dialogue text is segmented into single-turn dialogues, that is, continuous multi-turn dialogues are decomposed into multiple independent single-turn dialogue texts. This simplifies the complex dialogue structure, making each turn of dialogue clearer and easier to analyze, and facilitating the extraction of key information and semantic units. Secondly, single-turn dialogue segmentation helps improve the efficiency and accuracy of subsequent model training because the model can focus on the context of a single-turn dialogue, thereby better learning dialogue patterns and language features.

[0116] Please see Figure 6 In some embodiments, step S402 includes, but is not limited to, steps S601 to S602:

[0117] Step S601: For each dialogue text in multiple dialogue texts, input the dialogue text into a preset topic model to obtain the topic distribution corresponding to the dialogue text output by the preset topic model.

[0118] Step S602: For each dialogue text in multiple dialogue texts, determine the preset topic with the highest probability in the topic distribution corresponding to the dialogue text as the target preset topic to which the dialogue text belongs.

[0119] In step S601 of some embodiments, the aforementioned preset topic model is obtained by training a word matrix. The word matrix is ​​obtained by text vectorization processing of multiple dialogue text sets. Each row of the word matrix represents the corresponding dialogue text set, each column represents the corresponding preset word, and the element value of the word matrix represents the number of times the corresponding preset word appears in the dialogue text set. The dialogue text set includes multiple rounds of dialogue between the questioner and the respondent. The topic distribution is used to characterize the probability that the dialogue text belongs to each preset topic among multiple preset topics.

[0120] In step S602 of some embodiments, for example, multiple preset topics include product consultation, complaint handling, claims process, renewal strategy, and service benefits. The topic distribution corresponding to the dialogue text is such that the probability of the dialogue text belonging to product consultation, complaint handling, claims process, renewal strategy, and service benefits is 0.2, 0.3, 0.4, 0.5, and 0.6, respectively. Among them, the probability of the dialogue text belonging to service benefits is the highest. Therefore, the target preset topic to which the dialogue text belongs is service benefits. In another embodiment, if the topic distribution corresponding to the dialogue text is such that the probability of the dialogue text belonging to product consultation, complaint handling, claims process, renewal strategy, and service benefits is 0.2, 0.3, 0.6, 0.6, and 0.6, respectively, and the probability of the dialogue text belonging to claims process, renewal strategy, and service benefits is the highest, then the target preset topic to which the dialogue text belongs is three preset topics, and the target preset topics include service benefits, claims process, and renewal strategy.

[0121] In one implementation, the probability of a topic belonging to a preset topic in the above topic distribution is the sum of the occurrence probabilities of multiple preset keywords corresponding to the preset topic. The occurrence probability of the preset keywords can be the ratio of the number of times the preset keywords appear in the dialogue text to the total number of times all preset keywords appear in the dialogue text. For example, multiple preset topics include product consultation and claims process, and multiple preset keywords contain a total of 5 preset keywords. Among them, the 5 preset keywords for product consultation are car insurance, premium, NCD, product highlights, and discount, and the 5 preset keywords for claims process are claims, time limit, materials, damage assessment, and compensation. In the dialogue text, the probability of occurrence of the multiple preset keywords corresponding to product consultation is 0.1, 0, 0.6, 0.41, and 0, respectively, and the probability of occurrence of the multiple preset keywords corresponding to claims process is 0.2, 0.14, 0.3, 0, and 0, respectively. Then, the probability of belonging to product consultation in the topic distribution is 0.1+0+0.2+0.41+0=0.71, and the probability of belonging to claims process in the topic distribution is 0.2+0.14+0.3+0+0=0.64.

[0122] In another implementation, if there are multiple preset topics corresponding to the highest probability in the topic distribution, the preset topic with a probability greater than a preset probability threshold is determined as the target preset topic, that is, the probability corresponding to the target preset topic in the topic distribution is greater than the preset probability threshold. The preset probability threshold can be set according to needs, for example, it can be set to 0.3.

[0123] In this embodiment, the dialogue text is processed using a topic model to obtain the topic distribution corresponding to the dialogue text. The topic with the highest probability is determined as the topic to which the dialogue text belongs, providing a clear and specific topic label for the dialogue text, making subsequent dialogue text classification clearer and more accurate. Secondly, by utilizing the probability distribution characteristics of the topic model, the core topics in the text can be effectively and automatically identified.

[0124] Please see Figure 7 In some embodiments, prior to step S601, steps S701 to S702 may also be included, but are not limited to:

[0125] Step S701: Perform text vectorization on multiple dialogue text sets to obtain a word matrix.

[0126] Step S702: Train the target topic model using a word matrix to obtain the preset topic model.

[0127] In step S701 of some embodiments, the above-mentioned text vectorization processing of multiple dialogue text sets to obtain a word matrix can be performed by text cleaning of multiple dialogue text sets to obtain multiple text-cleaned dialogue text sets, and then text vectorization processing of the text-cleaned dialogue text sets to obtain a word matrix.

[0128] In step S702 of some embodiments, the target topic model can be a topic model to be trained with a determined number of topics K and each preset topic. The topic model is used to predict the topic distribution of the text. The topic model can be a Latent Dirichlet Allocation (LDA) model, a Correlated Topic Model (CTM) model, or a Hierarchical Dirichlet Process (HDP) model, etc.

[0129] The number of topics K and each preset topic can be determined by combining perplexity with business understanding; for example, K can be 5.

[0130] In this embodiment, training a topic model using a word matrix provides a structured representation of text data, enabling the model to more efficiently capture semantic information and topic distribution within the text. By presenting the relationships between documents and words in matrix form, the topic model can automatically discover potential topic structures based on statistical methods, thereby achieving automatic classification and topic extraction of text data and improving the effectiveness of topic distribution acquisition.

[0131] To better understand the above training method, this application provides a complete embodiment as follows:

[0132] In the insurance industry, historical dialogue records between customer service agents and customers constitute a valuable tacit knowledge base for enterprises. However, this data typically exists in unstructured text form, making it difficult to analyze and utilize directly. The experience, techniques, and strategies of excellent agents are often scattered across massive amounts of dialogue records, hindering systematic accumulation and leading to significant knowledge loss within the company. This forces extended training periods for new agents and results in inconsistent service quality. As the scale of insurance business continues to expand and service scenarios become increasingly complex, efficiently extracting key information from historical dialogues has become a core challenge for improving service efficiency, optimizing operational models, and enhancing customer experience.

[0133] Traditional data mining methods rely heavily on the experience of operations personnel and manual analysis, which is time-consuming, labor-intensive, and inefficient. This approach not only struggles to cover massive amounts of historical data but is also prone to biases in information extraction due to subjective factors, failing to comprehensively and accurately capture response templates and business logic within conversations. Furthermore, as customers' demands for service quality continue to rise, businesses need to respond quickly to customer needs and provide personalized services; traditional passive knowledge extraction methods are no longer sufficient to meet the demands of business development. Therefore, how to efficiently mine tacit knowledge from historical conversations through intelligent means to achieve systematic inheritance and rapid reuse of experience has become a key issue for the insurance industry to enhance its core competitiveness.

[0134] The topic-enhanced verbal knowledge mining method proposed in this invention is mainly reflected in the following three aspects: 1. By concatenating GRU to capture short-term dependencies and LSTM to process long-term context, the ability to model the temporal sequence of complex event descriptions in insurance verbal communication is enhanced; 2. By concatenating LDA topic distribution vectors and word embeddings, the model can simultaneously integrate global topic semantics and local word vector semantics, thereby improving the recognition accuracy of key patterns in insurance verbal communication; 3. By defining label transfer rules in the insurance domain, sequence labeling that conforms to business logic is achieved.

[0135] The main implementation process is as follows:

[0136] 1. Data preprocessing:

[0137] (1) Domain-adaptive text cleaning: Remove stop words and common insurance industry terms (such as "policy" and "claims"), and retain professional terms (such as "NCD coefficient" and "compulsory traffic accident liability insurance") to avoid losing semantics due to word segmentation errors; that is, perform text cleaning on the multiple text documents (the dialogue text set mentioned above).

[0138] (2) Text vectorization: Use TF-IDF to construct a document-word matrix (the word matrix above), where each row represents a document, each column represents a word, and the value represents the number of times the word appears in the document;

[0139] (3) The dialogue text between the customer and the agent is segmented into conversation units, i.e., single-turn dialogues, to ensure the semantic integrity of the topic division;

[0140] (4) Sequence labeling: For key patterns involved in the text, the extended BIOES labeling system (B-start, I-middle, E-end, S-word, O-irrelevant) is used, such as B-CLAIM_TYPE, I-CLAIM_TYPE, E-CLAIM_TYPE (labeled as "vehicle damage insurance").

[0141] 2. The preprocessed text is modeled using the LDA model to extract latent topic distributions:

[0142] (1) Determine the number of topics K by combining perplexity with business understanding, such as dividing them into 5 topics: product consultation, complaint handling, claims process, renewal strategy, and service rights;

[0143] (2) Train the LDA model using the preprocessed document-word matrix, setting the number of iterations and optimization algorithm for training;

[0144] (3) For each segmented conversation unit, the trained LDA model is used to calculate its probability distribution on each topic, and the topic with the highest probability is selected as the belonging label (topic explanation: extract the top-N keywords for each topic, and verify the rationality of the topic in combination with business understanding. For example, the top-5 keywords extracted for the product consultation topic are [car insurance, premium, NCD, product highlights, discount], and the top-5 keywords extracted for the claims process topic are [claims, timeliness, materials, damage assessment, compensation]. Subsequently, the topic classification is determined based on the extracted keywords with the highest probability). For complex dialogues with multiple topics, a threshold strategy is adopted (such as only keeping topics with a probability > 0.3) and splitting them into multiple topic datasets.

[0145] (4) Divide the original dialogue data (multiple dialogue texts) into K subsets (the above-mentioned topic sample sets) according to topic tags (K is the number of topics, i.e., 5). Each subset contains text data (dialogue texts) belonging to that topic. For data with uneven topic distribution, such as a small number of data samples for a certain topic, use oversampling to increase the sample data of that topic to ensure the robustness of the subsequent model.

[0146] 3. For each topic's text dataset, design and train an LSTM-CRF model (the aforementioned preset model) to extract key dialogue patterns (response templates):

[0147] (1) The text topic distribution vector (i.e., the 5-dimensional topic probability vector) obtained based on the LDA model (the topic distribution vector mentioned above) is concatenated to the word embedding layer (the word embedding layer here refers to the jieba word segmentation of each dialogue text, and the input of the segmented text into the BERT model, and the encoding of the segmented text by the Transformer encoder to obtain the word embedding vector of the text) to form a composite input vector (the composite vector mentioned above);

[0148] (2) Construct a lightweight GRU unit to generate a gating signal. The input is a word embedding vector (which contains a topic distribution vector), and the output contains topic information and context information.

[0149] (3) Use a fully connected layer to map the output of the GRU layer to a weight space with the same embedding dimension as the input, generate a weight vector, and normalize the generated weight vector to ensure that the weights are in the range of [0,1] to avoid gradient explosion or vanishing.

[0150] (4) Multiply the normalized weight vector with the input embedded word vector element by element to obtain the topic-enhanced input vector (the topic-enhanced vector mentioned above);

[0151] (5) Input the above-mentioned topic-enhanced input vector into the LSTM layer to capture and model context information;

[0152] (6) The output of the LSTM layer is used as the input of the CRF. The CRF maintains a learnable transition matrix, which represents the score of transitioning from one label to another. According to the insurance domain rules, constraints are imposed on the transition matrix, such as prohibiting the label "B-POLICY" from directly transitioning to "B-CLAIM". The globally optimal label sequence is output through the CRF layer. The output label sequence can be mapped to specific key conversational patterns according to the BIOES label system.

[0153] This invention proposes a topic-enhanced discourse knowledge mining method that concatenates LDA topic distribution vectors with original word embeddings, enabling the model to simultaneously integrate global topic semantics and local word vector semantics, thus achieving topic enhancement. Furthermore, by introducing a GRU gating unit before the LSTM layer, the input weights of the LSTM are dynamically adjusted in conjunction with topic enhancement, allowing the model to more accurately capture key contextual information related to the topic. Finally, a constraint-driven CRF layer outputs the globally optimal label sequence, ensuring that the overall label sequence is more consistent and reasonable, conforming to business logic. The multi-layered model network structure complements each other and has a clear division of labor, collectively improving model performance and recognition accuracy.

[0154] This application also provides a method for identifying reply templates. Figure 8 This is an optional flowchart of the response template identification method provided in the embodiments of this application. Figure 8 The method may include, but is not limited to, steps S801 to S802:

[0155] Step S801: Obtain the target question content input by the target questioner.

[0156] Step S802: Input the target question content into the target recognition model to obtain the target response template output by the target recognition model.

[0157] The aforementioned target response template is used to assist respondents in responding to the target question, and the aforementioned target recognition model is obtained by the aforementioned training method.

[0158] The target questioner mentioned above is the questioner corresponding to the target question content currently obtained.

[0159] In this embodiment, by inputting the target question content into the target recognition model and obtaining the target response template output by the model, the efficiency and quality of customer service can be improved. In this way, the model can quickly and accurately identify the core points of the user's question and provide an optimized response template. This effectively shortens response time and improves customer satisfaction. Furthermore, this model-based automated processing method can reduce manual intervention, lower operating costs, and provide standardized reference scripts for newly hired customer service personnel, helping them to quickly master effective communication skills, thereby improving the overall professionalism and reliability of the service.

[0160] Please see Figure 9 This application embodiment also provides a schematic diagram of the structure of a training device for a recognition model that identifies response templates, which can implement the above-mentioned training method for a recognition model that identifies response templates. The device 900 includes:

[0161] The training set acquisition module 901 is used to acquire the training set; wherein, the training set includes multiple topic sample sets, each topic sample set belongs to a different preset topic, each topic sample set includes multiple dialogue text samples labeled with corresponding preset topic tags, each dialogue text sample includes dialogue text, corresponding reply template tags for the dialogue text, and topic distribution vector of the dialogue text, the dialogue text includes the question content raised by the questioner and the reply content of the respondent to the question content, and the topic distribution vector is used to represent the probability of the dialogue text belonging to each preset topic in different preset topics;

[0162] The model training module 902 is used to train a preset model using a training set until the preset model meets the training stopping condition to obtain a target recognition model; wherein, the target recognition model is used to recognize a response template for assisting the respondent in responding to the question content.

[0163] In one embodiment, the training set acquisition module 901 is specifically used to: acquire multiple dialogue texts; determine the target preset topics to which the multiple dialogue texts belong; wherein the target preset topics are determined from a plurality of preset topics; and divide the multiple dialogue texts into multiple topic sample sets according to the target preset topics to which the multiple dialogue texts belong.

[0164] In one embodiment, the training set acquisition module 901 is specifically used to: input each dialogue text in multiple dialogue texts into a preset topic model to obtain the topic distribution corresponding to the dialogue text output by the preset topic model; wherein, the preset topic model is trained through a word matrix, the word matrix is ​​obtained by text vectorization processing of multiple dialogue text sets, each row of the word matrix represents the corresponding dialogue text set, each column represents the corresponding preset word, and the element value of the word matrix represents the number of times the corresponding preset word appears in the dialogue text set, the dialogue text set includes multiple rounds of dialogue between the questioner and the respondent, and the topic distribution is used to characterize the probability that the dialogue text belongs to each preset topic in multiple preset topics; for each dialogue text in multiple dialogue texts, the preset topic with the highest probability in the topic distribution corresponding to the dialogue text is determined as the target preset topic to which the dialogue text belongs.

[0165] In one embodiment, the model training module 902 is specifically used for: performing word segmentation on each dialogue text sample in each topic sample set of the training set to obtain multiple target words of the dialogue text, encoding the multiple target words to obtain word embedding vectors corresponding to the dialogue text, concatenating the word embedding vectors with the topic distribution vectors corresponding to the dialogue text to obtain composite vectors, and outputting the predicted response template based on the composite vectors; determining the loss function value according to the predicted response templates corresponding to each dialogue text sample in each topic sample set of the training set and the response template labels in each dialogue text sample; obtaining the target recognition model when the loss function value reaches the training stopping condition, and adjusting the network parameters of the preset model when the loss function value does not reach the training stopping condition, and returning to the step of performing word segmentation on each dialogue text sample in each topic sample set of the training set to obtain multiple target words of the dialogue text, until the preset model meets the training stopping condition.

[0166] In one embodiment, the aforementioned preset model comprises a gating unit, a fully connected layer, a long short-term memory layer, and a conditional random field layer. The model training module 902 is specifically used for: inputting a composite vector into the gating unit and the long short-term memory layer; using the gating unit to extract temporal features from the composite vector to obtain a temporal feature vector; inputting the temporal feature vector into the fully connected layer; using the fully connected layer to map the temporal feature vector to a weight space with the same preset input embedding dimension to generate a weight vector; wherein the dimension of the weight vector is the input embedding dimension; inputting the weight vector into the long short-term memory layer; using the long short-term memory layer to perform element-wise multiplication of the weight vector and the composite vector to obtain a topic enhancement vector; and inputting the topic enhancement vector into the conditional random field layer; using the conditional random field layer to output a predicted response template based on the topic enhancement vector.

[0167] The specific implementation of the training device for the recognition model of the recognition response template is basically the same as the specific implementation of the training method for the recognition model of the recognition response template described above, and will not be repeated here.

[0168] Please see Figure 10 This application embodiment also provides a structural schematic diagram of a response template identification device, which can implement the above-mentioned response template identification method. The device 100 includes:

[0169] The question acquisition module 101 is used to acquire the content of the target question input by the target questioner.

[0170] The template recognition module 102 is used to input the target question content into the target recognition model to obtain the target response template output by the target recognition model; wherein, the target response template is used to assist the respondent in responding to the target question content, and the target recognition model is obtained by the training method of the above-mentioned response template recognition model.

[0171] The specific implementation of the reply template recognition device is basically the same as the specific implementation of the reply template recognition method described above, and will not be repeated here.

[0172] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the training method for the recognition model of the above-mentioned response template; or, to implement the recognition method of the above-mentioned response template. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0173] Please see Figure 11 , Figure 11 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0174] The processor 110 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0175] The memory 111 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 111 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 111, and the processor 110 calls and executes the training method of the recognition model of the response template in the embodiments of this application, or executes the recognition method of the response template in the embodiments of this application.

[0176] Input / output interface 112 is used to implement information input and output;

[0177] The communication interface 113 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0178] Bus 114 transmits information between various components of the device (e.g., processor 110, memory 111, input / output interface 112, and communication interface 113);

[0179] The processor 110, memory 111, input / output interface 112 and communication interface 113 are connected to each other within the device via bus 114.

[0180] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements either the training method for the recognition model of the above-mentioned response template or the recognition method of the above-mentioned response template.

[0181] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0182] The present application provides a training method, recognition method, device, electronic device, and storage medium for a response template recognition model. This method involves acquiring a training set, which includes multiple topic sample sets, each belonging to a different preset topic. Each topic sample set includes multiple dialogue text samples labeled with corresponding preset topic identifiers. Each dialogue text sample includes dialogue text, a corresponding response template label, and a topic distribution vector. The dialogue text includes the question content raised by the questioner and the response content given by the respondent. The topic distribution vector represents the probability that the dialogue text belongs to each of the different preset topics. The training set is used to train a preset model until the preset model meets the training stopping condition, resulting in a target recognition model. The target recognition model is used to recognize response templates used to assist the respondent in responding to questions. The method involves acquiring the target question content input by the target questioner, inputting the target question content into the target recognition model, and obtaining the target response template output by the target recognition model. Thus, by introducing the topic distribution vector of the dialogue text to train the model, the performance and generalization ability of the preset model can be improved, thereby increasing the model's accuracy in recognizing response templates and, consequently, improving the output of more effective target response templates when the model is used in subsequent applications.

[0183] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0184] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

Claims

1. A method for training a response template recognition model, characterized in that, The method includes: Obtain a training set; wherein the training set includes multiple topic sample sets, each topic sample set belongs to a different preset topic, each topic sample set includes multiple dialogue text samples marked by corresponding preset topic identifiers, each dialogue text sample includes dialogue text, a reply template tag corresponding to the dialogue text, and a topic distribution vector of the dialogue text, the dialogue text includes the question content raised by the questioner and the reply content made by the respondent to the question content, and the topic distribution vector is used to characterize the probability that the dialogue text belongs to each preset topic in the different preset topics; The preset model is trained using the training set until the preset model meets the training stopping condition to obtain the target recognition model; wherein, the target recognition model is used to identify a response template for assisting the respondent in responding to the question content.

2. The method according to claim 1, characterized in that, The step of training a preset model using the training set until the preset model meets the training stopping condition to obtain a target recognition model includes: For each dialogue text sample in each topic sample set of the training set, the dialogue text in the dialogue text sample is segmented to obtain multiple target words of the dialogue text, and the multiple target words are encoded to obtain the word embedding vector corresponding to the dialogue text. The word embedding vector is concatenated with the topic distribution vector corresponding to the dialogue text to obtain a composite vector. Based on the composite vector, the predicted response template for recognition is output. The loss function value is determined based on the predicted response template corresponding to each dialogue text sample in each topic sample set in the training set, and the response template label in each dialogue text sample; If the loss function value reaches the training stopping condition, a target recognition model is obtained. If the loss function value does not reach the training stopping condition, the network parameters of the preset model are adjusted, and the process of segmenting the dialogue text in each dialogue text sample in each topic sample set of the training set to obtain multiple target words of the dialogue text is repeated until the preset model meets the training stopping condition.

3. The method according to claim 2, characterized in that, The preset model consists of a gating unit, a fully connected layer, a long short-term memory layer, and a conditional random field layer. The predicted response template based on the composite vector output recognition includes: The composite vector is input to the gating unit and the long short-term memory layer. The gating unit is used to extract the temporal features from the composite vector to obtain a temporal feature vector. The temporal feature vector is input into the fully connected layer, and the fully connected layer maps the temporal feature vector to a weight space with the same preset input embedding dimension to generate a weight vector; wherein, the dimension of the weight vector is the input embedding dimension; The weight vector is input into the long short-term memory layer, and the long short-term memory layer is used to multiply the weight vector and the composite vector element by element to obtain the topic enhancement vector. The topic enhancement vector is input into the conditional random field layer, and the conditional random field layer outputs the predicted response template based on the topic enhancement vector.

4. The method according to claim 1, characterized in that, The acquisition of the training set includes: Retrieve multiple dialogue texts; Each of the plurality of dialogue texts is determined to belong to a target preset topic; wherein, the target preset topic is determined from a plurality of preset topics; Based on the target preset topic to which the multiple dialogue texts belong, the multiple dialogue texts are divided into multiple topic sample sets.

5. The method according to claim 4, characterized in that, The step of determining the target preset topic to which the plurality of dialogue texts belong includes: For each dialogue text in the plurality of dialogue texts, the dialogue text is input into a preset topic model to obtain the topic distribution corresponding to the dialogue text output by the preset topic model; wherein, the preset topic model is trained through a word matrix, the word matrix is ​​obtained by text vectorization processing of the plurality of dialogue text sets, each row of the word matrix represents the corresponding dialogue text set, each column represents the corresponding preset word, and the element value of the word matrix represents the number of times the corresponding preset word appears in the dialogue text set, the dialogue text set includes multiple rounds of dialogue between the questioner and the respondent, and the topic distribution is used to characterize the probability that the dialogue text belongs to each preset topic in the plurality of preset topics; For each dialogue text in the plurality of dialogue texts, the preset topic with the highest probability in the topic distribution corresponding to the dialogue text is determined as the target preset topic to which the dialogue text belongs.

6. A method for identifying a response template, characterized in that, The method includes: Obtain the content of the target question input by the target questioner; The target question content is input into the target recognition model to obtain the target response template output by the target recognition model; wherein, the target response template is used to assist the respondent in responding to the target question content, and the target recognition model is obtained by the training method of the recognition model of the response template according to any one of claims 1-5.

7. A training device for a response template recognition model, characterized in that, The device includes: A training set acquisition module is used to acquire a training set; wherein, the training set includes multiple topic sample sets, each topic sample set belongs to a different preset topic, each topic sample set includes multiple dialogue text samples labeled with corresponding preset topic tags, each dialogue text sample includes dialogue text, a reply template tag corresponding to the dialogue text, and a topic distribution vector of the dialogue text, the dialogue text includes the question content raised by the questioner and the reply content of the respondent to the question content, and the topic distribution vector is used to characterize the probability that the dialogue text belongs to each preset topic in the different preset topics; The model training module is used to train a preset model using the training set until the preset model meets the training stopping condition to obtain a target recognition model; wherein, the target recognition model is used to recognize a response template for assisting the respondent in responding to the question content.

8. A device for recognizing reply templates, characterized in that, The device includes: The question acquisition module is used to acquire the content of the target question input by the target questioner; The template recognition module is used to input the target question content into the target recognition model to obtain the target response template output by the target recognition model; wherein, the target response template is used to assist the respondent in responding to the target question content, and the target recognition model is obtained by the training method of the above-mentioned response template recognition model.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the training method for the recognition model of the response template as described in any one of claims 1 to 5. Alternatively, the method for identifying the response template as described in claim 6 can be implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the training method for the recognition model of the response template as described in any one of claims 1 to 5; Alternatively, the method for identifying the response template as described in claim 6 can be implemented.