Model training method, response method and related products

By identifying and integrating the intent and entity information in the conversation text, combined with pre-trained models and user portraits, the problem of insufficient accuracy in conversation intent recognition is solved, and accurate and personalized intent recognition and response are achieved.

CN120725017APending Publication Date: 2025-09-30MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510703547.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing technologies find it difficult to achieve accurate and comprehensive semantic analysis in conversational intent recognition facing the diversity and ambiguity of user expressions, resulting in insufficient accuracy in intent recognition.

Method used

By identifying the intent and entity information in the conversation text, integrating them to generate more refined and comprehensive intents, combining them with pre-trained models for multi-dimensional learning, selecting models with higher recognition accuracy for intent recognition, and combining them with user portrait information for personalized responses.

Benefits of technology

It achieves the precise capture of intent information in complex conversation texts, improves the accuracy of intent recognition and the pertinence of responses, and enhances the personalization and relevance of responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725017A_ABST
    Figure CN120725017A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method, a response method and a related product, which are used for accurately and comprehensively analyzing session semantics so as to accurately identify a session intention. The model training method comprises the following steps: identifying a first intention and first entity information of a first session text; fusing the first intention and the first entity information to obtain a second intention of the first session text, and training a model based on the first session text and the second intention to obtain a first model; fusing the first intention, the first entity information and the second intention to obtain a third intention of the first session text, and training the model based on the first session text and the third intention to obtain a second model; and comparing the recognition accuracy of the first model and the recognition accuracy of the second model on the second session text, and determining a third model from the first model and the second model based on a comparison result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing technology, and in particular to a model training method, a response method, and related products. Background Art

[0002] As a key technology in the field of natural language processing (NLP), conversational intent recognition plays a vital role in promoting the development of application areas such as intelligent customer service and intelligent assistants.

[0003] Conversational intent recognition aims to accurately capture and understand the true intentions of users during conversations. In real-world scenarios, users express themselves in a variety of ways, often with ambiguous semantics and flexible expressions. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a model training method, a response method and related products for accurately and comprehensively analyzing conversation semantics, and then accurately identifying conversation intent.

[0005] In order to achieve the above objectives, the embodiments of the present application adopt the following technical solutions: In a first aspect, an embodiment of the present application provides a model training method, comprising: identifying a first intent and first entity information of a first conversation text; fusing the first intent and the first entity information to obtain a second intent of the first conversation text, and training a model based on the first conversation text and the second intent to obtain a first model; fusing the first intent, the first entity information, and the second intent to obtain a third intent of the first conversation text, and training the model based on the first conversation text and the third intent to obtain a second model; The recognition accuracy of the first model and the second model on the second conversation text are compared, and based on the comparison result, a third model is determined from the first model and the second model.

[0006] The model training method provided in the embodiment of the present application identifies a first intention and a first entity information from a first conversation text. The first entity information reflects a specific or abstract object with a clear referential meaning in the first conversation text, and semantics is constructed and expressed through these entity information. On the one hand, the first intention and the first entity information are integrated to obtain the second intention of the first conversation text. Compared with the first intention, the granularity of the second intention is finer, and it can more accurately portray the intention tendency contained in the first conversation text. On this basis, the model is trained based on the first conversation text and its second intention, so that the model can fully absorb the supervision signal provided by the second intention, and deeply learn the precise and comprehensive analysis of the semantics of the first conversation text from multiple dimensions, thereby having the ability to accurately identify conversation intentions. The model is capable of accurately capturing the hidden intent information in various complex conversation texts when faced with them. On the other hand, the first intent, the first entity information, and the second intent are integrated to obtain the third intent of the first conversation text. Compared with the first intention and the second intention, the third intent contains more comprehensive and complete information. On this basis, the model is trained based on the first conversation text and its third intent, which helps the model break through its own recognition preferences and limitations and improve the model's intent recognition ability. Finally, the recognition accuracy of the first model and the recognition accuracy of the second model are evaluated based on the second conversation text, and the model with better recognition accuracy is selected from the two as the third model for conversation intent recognition, thereby accurately and comprehensively analyzing the conversation semantics during the conversation and accurately identifying the conversation intent.

[0007] In a second aspect, an embodiment of the present application provides a response method, including: In response to the third conversation text of the first user, obtaining portrait information of the first user; Performing entity recognition on the third conversation text to obtain third entity information; performing intent recognition on the third conversation text using a third model to obtain the intent of the third conversation text, wherein the third model is trained based on the model training method provided in the first aspect; Based on the third entity information, the intention of the third conversation text and the portrait information, a first response text corresponding to the third conversation text is determined.

[0008] The answering method provided in the embodiment of the present application, during the conversation process, performs intent recognition on the user's conversation text through a third model, and accurately and completely extracts the intent of the conversation text; in addition, entity recognition is also performed on the user's conversation text to obtain corresponding entity information, and the user's portrait information is obtained. The entity information reflects the specific or abstract objects with clear referential meaning in the conversation text, and is the basis for constructing and expressing semantics, while the portrait information reflects the user's characteristics and preferences. These two types of information help to understand the user's real personalized needs more comprehensively and in-depth; further, the user's portrait information, the entity information of the conversation text, and the intent of the conversation text are integrated to determine the answer text corresponding to the conversation text, so that the response to the conversation text is more personalized, targeted, and more in line with user expectations, thereby improving the accuracy of the response.

[0009] In a third aspect, an embodiment of the present application provides a model training device, comprising: an identification module, configured to identify a first intent and first entity information of a first conversation text; A first training module is configured to fuse the first intent and the first entity information to obtain a second intent of the first conversation text, and to train a model based on the first conversation text and the second intent to obtain a first model; a second training module, configured to fuse the first intent, the first entity information, and the second intent to obtain a third intent of the first conversation text, and train the model based on the first conversation text and the third intent to obtain a second model; The first determination module is configured to compare the recognition accuracy of the first model and the second model on the second conversation text, and determine a third model from the first model and the second model based on the comparison result.

[0010] In a fourth aspect, an embodiment of the present application provides a response device, including: an acquisition module, configured to acquire portrait information of the first user in response to the third conversation text of the first user; an identification module, configured to perform entity recognition on the third conversation text to obtain third entity information; The recognition module is further configured to perform intent recognition on the third conversation text using a third model to obtain the intent of the third conversation text, wherein the third model is trained based on the model training method provided in the first aspect; The second determination module is used to determine the first response text corresponding to the third conversation text based on the third entity information, the intention of the third conversation text and the portrait information.

[0011] In a fifth aspect, an embodiment of the present application provides an electronic device, including: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the model training method provided in the first aspect or the response method provided in the second aspect.

[0012] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the model training method provided in the first aspect or the response method provided in the second aspect.

[0013] In the seventh aspect, an embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute part or all of the steps in the model training method provided in the first aspect or the response method provided in the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A flowchart of a model training method provided for one embodiment of the present application; Figure 2 A flowchart of a model training method provided in another embodiment of the present application; Figure 3 A flowchart of a response method provided in one embodiment of the present application; Figure 4 A schematic diagram of the structure of a model training device provided in one embodiment of the present application; Figure 5 A schematic structural diagram of a response device provided in one embodiment of the present application; Figure 6 A schematic structural diagram of an electronic device provided in accordance with an embodiment of the present application. DETAILED DESCRIPTION

[0015] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0016] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0017] Key terms explained: Bidirectional Encoder Representation from Transformers (BERT) is a pre-trained language representation model. It emphasizes the use of a new Masked Language Model (MLM) (Masked Language Model), which generates deep bidirectional language representations, rather than the traditional unidirectional language model or the shallow concatenation of two unidirectional language models for pre-training.

[0018] Natural Language Understanding (NLU): is a general term for all methods, models, or tasks that support machines to understand text content.

[0019] Natural Language Generation (NLG): is an automated process of generating language text through computers under specific interactive goals. Its main purpose is to automatically construct high-quality language text that can be understood by humans.

[0020] Prompt learning is a paradigm that uses handcrafted prompts to guide machine learning. It has achieved remarkable results, particularly in the field of natural language processing (NLP). It significantly improves performance by adding prompts to the input of pre-trained models. Prompt learning is particularly well-suited for low-resource scenarios, achieving good results even without a large number of samples.

[0021] As mentioned above, given the diversity, ambiguity, and flexibility of user expressions during conversations, there is an urgent need for a conversation intent recognition solution that can accurately and comprehensively analyze conversation semantics and thus accurately identify conversation intent.

[0022] To this end, an embodiment of the present application proposes a model training method to identify a first intention and a first entity information from a first conversation text. The first entity information reflects a specific or abstract object with a clear referential meaning in the first conversation text, and semantics is constructed and expressed through these entity information; on the one hand, the first intention and the first entity information are integrated to obtain the second intention of the first conversation text. Compared with the first intention, the granularity of the second intention is finer, and it can more accurately portray the intention tendency contained in the first conversation text. On this basis, the model is trained based on the first conversation text and its second intention, so that the model can fully absorb the supervision signal provided by the second intention, and deeply learn the precise and comprehensive analysis of the semantics of the first conversation text from multiple dimensions, thereby having the ability to accurately identify the conversation text. The ability of speech intent recognition can accurately capture the hidden intention information in various complex conversation texts; on the other hand, the first intention, the first entity information and the second intention are integrated to obtain the third intention of the first conversation text. Compared with the first intention and the second intention, the information contained in the third intention is more comprehensive and complete. On this basis, the model is trained based on the first conversation text and its third intention, which helps the model break through its own recognition preferences and limitations and improve the model's intention recognition ability; finally, the recognition effects of the first model and the second model are evaluated based on the second conversation text, and the model with better recognition effect is determined from the two as the third model for conversation intent recognition, so as to accurately and comprehensively analyze the conversation semantics during the conversation, and then accurately identify the conversation intent.

[0023] Based on the third model trained by the above-mentioned model training method, the embodiment of the present application also proposes a response method. During the conversation, the user's conversation text is subjected to intent recognition through the third model, and the intent of the conversation text is accurately and completely extracted; in addition, the user's conversation text is subjected to entity recognition to obtain corresponding entity information, and the user's portrait information is obtained. The entity information reflects the specific or abstract objects with clear referential meaning in the conversation text, and is the basis for constructing and expressing semantics, while the portrait information reflects the user's characteristics and preferences. These two types of information help to understand the user's real personalized needs more comprehensively and deeply; further, the user's portrait information, the entity information of the conversation text, and the intent of the conversation text are integrated to determine the response text corresponding to the conversation text, so that the response to the conversation text is more personalized, targeted, and more in line with user expectations, thereby improving the accuracy of the response.

[0024] It should be understood that the model training method and response method proposed in the embodiments of the present application can be executed by an electronic device. As an example, it can be executed by software in the electronic device. The so-called electronic devices here can include terminal devices, such as smart phones, tablet computers, laptops, desktop computers, intelligent voice interaction devices, smart home appliances, smart watches, vehicle terminals, aircraft, etc.; or, the electronic device can also include a server, such as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.

[0025] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.

[0026] Please refer to Figure 1 , is a flow chart of a model training method provided in one embodiment of the present application, the method comprising the following steps: S102: Identify the first intention and first entity information of the first conversation text.

[0027] The first conversation text can be any sample conversation text. There can be multiple first conversation texts. In the application, the first conversation text can be annotated with the speaker's role, such as user or customer service, to assist the model in considering the speaker's role in the first conversation text during the model training process for intent recognition.

[0028] First entity information reflects a specific or abstract object with a clear referential meaning in the first conversation text. First entity information may include, but is not limited to, entities in the first conversation text (herein, referred to as first entities), entity states, and relationships between different entities. The state of a first entity refers to the characteristics, attributes, or situation presented by the entity in a specific conversation context. For example, for the entity "apple," its state may be "fresh," "rotten," or "ripe"; for the entity "order," its state may be "paid," "unpaid," "shipped," "unshipped," or "completed"; and for the entity "income level," its state may be "high" or "low." Relationships between different entities may include, but are not limited to, action relationships, attribute relationships, temporal relationships, and spatial location relationships.

[0029] The above S102 can be implemented in various appropriate ways.

[0030] In one implementation, intent recognition is performed on the first conversation text by a model with intent recognition capability (i.e., intent recognition model) to obtain a first intent of the first conversation text, and entity recognition is performed on the first conversation text by a model with entity recognition capability (i.e., entity recognition model) to obtain first entity information of the first conversation text.

[0031] Exemplarily, the intent recognition model encodes the first conversation text to obtain a high-level semantic representation of the first conversation text, which is also called the representation vector of the first conversation text; then, the intent recognition model performs a linear transformation on the representation vector of the first conversation text to obtain a K-dimensional vector, where the value of K is equal to the number of preset intents, and each dimension corresponds to the unnormalized score of the first conversation text in the intent, which is also called logit; further, the intent recognition model normalizes the K-dimensional vector, such as converting the K-dimensional vector into a probability distribution form through a softmax function, to obtain the probability that the first conversation text corresponds to the preset K intents; finally, the intent with the highest probability among the K intents is determined as the first intention of the first conversation text.

[0032] The entity recognition model encodes each word in the first conversation text to obtain a high-level semantic representation of each word (i.e., a representation vector for each word). Then, the entity recognition model uses conditional random fields (CRFs) to classify the representation vector of each word to obtain a classification result for each word, which indicates whether the word belongs to an entity. Furthermore, for each entity in the first conversation text, the entity's features and the features of the entity's contextual information are extracted, and the entity's state is identified based on these features. In addition, the features of different entities themselves, the distances between entities, the grammatical relationships between entities, etc. in the first conversation text are extracted, and the relationships between different entities are predicted based on this information.

[0033] For example, the first conversation text is "I want to find a part-time job close to home with a high salary." The intent recognition model is used to perform intent recognition on the first conversation text, and the first intention "find a part-time job" is obtained. The entity recognition model is used to perform entity recognition on the first conversation text, and the first entity information obtained includes the first entity sequence [home, salary], the status of the first entity, and the relationship between the first entities [([close] home), ([high] salary)].

[0034] In the application, the type of intent recognition model and the type of entity recognition model can be set according to actual needs. For example, both the intent recognition model and the entity recognition model can be pre-trained language models such as Roberta, etc., and this embodiment of the application does not limit this.

[0035] In another implementation, the idea of ​​task fusion is respectively integrated into the intent recognition task and the entity recognition task, that is, the initial intent and initial entity information of the first conversation text are first identified using the traditional method, and then based on this, the initial entity information is integrated into the intent recognition task to utilize the initial entity information to promote the intent recognition task, thereby improving the accuracy of intent recognition, and the initial intent is integrated into the entity recognition task to utilize the initial intent to promote the entity recognition task, thereby improving the accuracy of entity recognition.

[0036] Specifically, the above S102 includes the following steps: S1021: Perform intent recognition on the first conversation text to obtain a fourth intent of the first conversation text.

[0037] The intent recognition model may be used to perform intent recognition on the first conversation text to obtain a fourth intent of the first conversation text, which is also called the initial intent.

[0038] S1022: Perform entity recognition on the first conversation text to obtain second entity information of the first conversation text.

[0039] Entity recognition can be performed on the first conversation text using an entity recognition model to obtain second entity information of the first conversation text. The second entity information is also called initial entity information.

[0040] S1023: Perform intent recognition on the first conversation text based on the fourth intent and the second entity information to obtain a first intent of the first conversation text.

[0041] As an example, the prompt word, the fourth intent, and the second entity information corresponding to the intent recognition task can be input into a large language model (LLM), and the natural language understanding capability of the LLM can be used to perform intent recognition on the first conversation text to obtain the first intent of the first conversation text.

[0042] As another example, the above-mentioned S1023 includes the following steps: Step A1, encoding and linearly transforming the first conversation text to obtain a first vector; Step A2, normalizing the first vector to obtain first probabilities that the first conversation text corresponds to multiple preset intentions; Step A3, fusing the first vector and the representation vector of the description text of the first intention to obtain a second vector, and normalizing the second vector to obtain second probabilities that the first conversation text corresponds to the multiple intentions; Step A4, fusing the first vector, the representation vector of the description text of the first intention, and the representation vector of the second entity information to obtain a third vector, and normalizing the third vector to obtain multiple third probabilities that the first conversation text corresponds to the multiple intentions; Step A5, for each intention, determining the fourth probability that the first conversation text corresponds to the intention based on the first probability, second probability, and third probability that the first conversation text corresponds to the intention; Step A6, determining the intention with the largest fourth probability among the multiple intentions as the first intention of the first conversation text.

[0043] In the above step A1, the first conversation text can be encoded using a convolutional neural network (CNN), a recurrent neural network (RNN), or BERT, and the encoded representation vector can be linearly transformed using a fully connected layer to obtain a first vector.

[0044] In the above step A2, the first vector may be normalized using a softmax function, thereby converting the first vector into a first probability that the first conversation text belongs to each intent.

[0045] In step A3, the descriptive text of the first intent can be text that describes the interpretation of the first text. By encoding the descriptive text, a representation vector for the descriptive text is obtained. For example, if the first conversation text is "No need, wrong call," "Not needed at this time," "No thanks," or "No need, hang up," and the first intent is "Mistakenly thought of as influence," the descriptive text of the first intent is "Customer mistakenly thought it was marketing, usually occurring at the beginning of a conversation." If the first conversation text is "Why is there my call?", "How do you have my phone number?", or "Where did my information come from?" and the first intent is "Inquiring / questioning the source of information," the descriptive text of the first intent is "Inquiring / questioning the agent's channel for obtaining user information."

[0046] The fusion of the representation vector of the first intent and the representation vector of the description text can be achieved through various fusion techniques known in the art, such as multiplying the two representation vectors, which is not limited in this embodiment of the present application. Furthermore, the resulting fused second vector can be normalized using a softmax function to convert the second vector into a second probability that the first conversation text belongs to each intent.

[0047] In step A4 above, the representation vector of the second entity information can be obtained by encoding the second entity information. The first vector, the representation vector of the first intent's description text, and the representation vector of the second entity information are multiplied together to fuse the three, resulting in a third vector. Furthermore, the third vector can be normalized using a softmax function to convert it into a third probability that the first conversation text belongs to each intent.

[0048] In the above step A5, for each intent, the first probability, the second probability and the third probability that the first conversation text corresponds to the intent may be added together to obtain a fourth probability that the first conversation text corresponds to the intent.

[0049] Integrating the third entity information into the intent recognition task in the above manner helps to accurately and comprehensively understand the semantics of the first conversation text from the second entity information and the description text of the first intent during the intent recognition process, thereby improving the accuracy of intent recognition.

[0050] S1024: Perform entity recognition on the first conversation text based on the fourth intent and the second entity information to obtain first entity information of the first conversation text.

[0051] As an example, the prompt word, the fourth intent and the second entity information corresponding to the entity recognition task can be input into the LLM, and the natural language understanding ability of the LLM can be used to perform entity recognition on the first conversation text to obtain the first entity information of the first conversation text.

[0052] As another example, the above S1024 includes the following steps: step B1, performing entity recognition on the first conversation text to obtain second entity information of the first conversation text; step B2, fusing the representation vector of the second entity information and the representation vector of the fourth intention to obtain a fourth vector; step B3, mapping the fourth vector to obtain the first entity information of the first conversation text.

[0053] In the above step B1, entity recognition can be performed on the first conversation text using CRF or the like to obtain second entity information.

[0054] In step B2 above, the representation vector of the second entity information can be obtained by encoding the second entity information using a CNN or RNN. The representation vector of the second entity information may include, but is not limited to, the representation vector of the entity in the first conversation text (the identified entity is referred to as the second entity here), the representation vector of the entity's state, and the representation vector of the relationship between different entities. In this case, the fourth vector = the representation vector of the second entity * the representation vector of the second entity's state + the representation vectors of different second entities * the representation vector of the relationship between different second entities + the representation vector of the fourth intent.

[0055] In the above step B3, the fourth vector may be mapped using a mapping function to obtain the first entity information of the first conversation text.

[0056] Alternatively, the representation vector of the first conversation text, the fourth vector, and the representation vector of the portrait information of the speaker of the first conversation text may be fused to obtain a new vector, which is then mapped to obtain the first entity information of the first conversation text. It is understood that because different speakers have different linguistic expressions or forms, the new vector incorporates the portrait information of the speaker of the first conversation text, which facilitates a deeper understanding of the semantics of the first conversation text from multiple dimensions, thereby improving the accuracy of entity recognition.

[0057] The above describes some implementations of the above S102. Of course, it should be understood that the above S102 can also be implemented in other ways, which are not limited in the present embodiment.

[0058] S104: Fusing the first intent and the first entity information to obtain a second intent of the first conversation text.

[0059] Since the second intent not only contains the first intent but also incorporates the first entity information, the second intent has finer granularity and can more accurately depict the intention tendency contained in the first conversation text, providing more refined supervision signals for the model training process.

[0060] In one implementation, the first intent and the first entity information are directly concatenated to obtain the second intent of the first conversation text.

[0061] In another implementation, a template for describing an intent is obtained, the template including at least one of the following slots: a slot for describing the cause of the intent, a slot for describing the result of the intent, a slot for describing the subject of the intent, and a slot for describing the object of the intent; the name of each slot in the template is embedded to obtain a representation vector for each slot; based on the similarity between the representation vector of the first intent and the representation vector of each slot, a first slot in the template that matches the first intent is determined; based on the similarity between the representation vector of the first entity information and the representation vector of each slot, a second slot in the template that matches the first entity information is determined; the first intent is filled into the first slot, and the first entity information is filled into the second slot to obtain the second intent of the first conversation text.

[0062] For example, a template is [cause intention, intention subject, intention object, result intention], where cause intention is the name of the slot used to describe the cause of the intention, intention subject is the name of the slot used to describe the subject of the intention, intention object is the name of the slot used to describe the object of the intention, and result intention is the name of the slot used to describe the result of the intention.

[0063] After embedding the name of each slot to obtain the representation vector of each slot, for the first entity information, multiply the representation vector of the first entity in the first entity information and the representation vector of the state of the first entity to obtain a new vector corresponding to the first entity, and calculate the similarity between the vector and the representation vector of each slot respectively. Assuming that the similarity between the new vector corresponding to the first entity and the representation vector of the slot "due to intention" is the largest, the first entity is filled into the slot "due to intention".

[0064] For the first intention, calculate the similarity between the representation vector of the first intention and the representation vector of each slot respectively. Assuming that the similarity between the representation vector of the first intention and the representation vector of the slot "fruit intention" is the largest, fill the first intention into the slot "fruit intention".

[0065] After completing the matching between the first intent and the first entity information and each slot, the second intent is obtained. Of course, it should be understood that in the application, the content of some slots is empty. For example, if the subject of the intent does not exist in the first conversation text, the content of the slot "intent subject" is empty.

[0066] It can be seen that by fusing the first intention and the first entity information through the above template, the second intention obtained can clearly reflect the cause and effect of the intention, the subject and object of the intention, which helps to improve the model's ability to perform a comprehensive semantic analysis of the first conversation text, thereby improving the model's intention recognition effect.

[0067] The above describes some implementations of the above S104. Of course, it should be understood that the above S104 can also be implemented in other ways, which are not limited in the present embodiment.

[0068] S106: Train the model based on the first conversation text and the second intention to obtain a first model.

[0069] The model here can be a model with preliminary intent recognition capabilities.

[0070] In one implementation, a supervised training approach is used, using the first conversation text as a sample and the second intent as the corresponding label. This allows the model to fully absorb the supervisory signal provided by the second intent and learn to accurately and comprehensively analyze the semantics of the first conversation text from multiple dimensions. This enables the model to accurately identify conversation intent and precisely capture the hidden intent information in complex conversation texts. The trained model is referred to as the first model.

[0071] For example, the model performs intent recognition on the first conversation text to obtain the intent of the first conversation text; then, based on the difference between the first intent and the second intent, the model loss is determined, and the backpropagation algorithm is used to adjust the model parameters based on the loss. This process is repeated multiple times until a preset training stop condition is met, thereby obtaining the first model. The training stop condition can be set according to actual needs, such as when the model loss converges or the number of times the model is adjusted reaches a threshold, and this embodiment of the application is not limited to this.

[0072] In another implementation, some repeated and less frequently occurring second intents and their corresponding first conversation texts may have a negative impact on the training effect of the model. To this end, the distribution of the frequency of occurrence of each second intent in all first conversation texts is first counted, and the repeated and less frequently occurring second intents and their corresponding first conversation texts are filtered out. The model is then trained using the remaining second intents and their corresponding first conversation texts, which can significantly improve the training effect of the model.

[0073] Specifically, based on the second intent of each first conversation text, the second intent that appears in multiple first conversation texts with a frequency greater than a frequency threshold is determined as the fifth intent; the model is trained based on the fifth intent and the first conversation texts with the fifth intent to obtain the first model.

[0074] The above describes some implementations of the above S106. Of course, it should be understood that the above S106 can also be implemented in other ways, which are not limited in the present embodiment.

[0075] S108: Fusing the first intent, the first entity information, and the second intent to obtain a third intent of the first conversation text.

[0076] Since the third intention integrates the first intention, the first entity information and the second intention, the information contained in the third intention is more comprehensive and complete than the first and second intentions. It can fully characterize the intention tendency contained in the first conversation text and provide a more comprehensive and complete supervision signal for the model training process.

[0077] In one implementation, the first intent, the first entity information, and the second intent are directly concatenated to obtain the third intent of the first conversation text.

[0078] In another implementation, Figure 2 As shown, the above S108 includes the following steps: S1081 , performing semantic decomposition on the first intention, the first entity information, and the second intention to obtain a plurality of semantic components.

[0079] Sememes are units of meaning (or content) in a language. Each text can be decomposed into at least one sememe, each of which is also called the smallest semantic unit of the text. For example, the sememes of the text "brother" include [immediate relative][sibling relationship][elder][male], and the sememes of the text "sister" include [immediate relative][sibling relationship][elder][female]. Sememe decomposition of text can be achieved using various sememe analysis methods commonly used in the art, and this embodiment of the present application does not limit this.

[0080] According to the source, the plurality of semantic elements include a second semantic element of the first intention, a third semantic element of the first entity information, and a fourth semantic element of the second intention.

[0081] S1082: Determine a first relevance between the plurality of sememes and a second relevance between each sememe and the first conversation text.

[0082] For each sememe, embedding processing is performed on the sememe to obtain the representation vector of the sememe; then, for each pair of semes, correlation analysis is performed on the representation vectors of the two semes, such as calculating the Pearson correlation coefficient, to obtain the first correlation between the two semes.

[0083] In addition, for each sememe, a correlation analysis is performed between the representation vector of the sememe and the representation vector of the first conversation text to obtain a second correlation between the sememe and the first conversation text.

[0084] S1083 : Determine a first sememe related to the first conversation text from the plurality of sememes based on the first relevance and the second relevance.

[0085] For example, a subset of semes is formed from multiple semes whose first correlation with each other is greater than a first threshold. Semes whose second correlation with the first conversation text is greater than a second threshold are then selected from the subset as first semes related to the first conversation text. In this way, the first semes can reflect the main content of the first conversation text, are consistent with each other, and can connect the first intent, the first entity information, and the second intent. Therefore, the first semes can also be called the backbone semes.

[0086] S1084: Determine a logical relationship between the first intention, the first entity information, and the second intention based on the first sememe, the second sememe, the third sememe, and the fourth sememe.

[0087] As an example, the first semantic element, the second semantic element, the third semantic element, the fourth semantic element and the prompt word for logical relationship identification are input into the LLM. By using the language understanding ability of the LLM, the semantic connection between the first intention, the first entity information and the second intention is captured from the semantics expressed by the first semantic element, the second semantic element, the third semantic element and the fourth semantic element, thereby obtaining the logical relationship between the three.

[0088] As another example, based on multiple operation methods, the representation vector of the second semantic element, the representation vector of the third semantic element, and the representation vector of the fourth semantic element are operated to obtain fifth vectors corresponding to the multiple operation methods; each fifth vector is mapped to obtain the fifth semantic element corresponding to each fifth vector; based on the difference between each fifth semantic element and the first semantic element and the operation method corresponding to each fifth semantic element, the logical relationship between the first intention, the first entity information, and the second intention is determined.

[0089] The difference between the fifth semantic element and the first semantic element corresponding to each operation mode manifests the semantic relevance and difference between the first intention, the first entity information and the second intention from the perspective of digital dimension representation. By analyzing this relevance and difference, the logical relationship between the first intention, the first entity information and the second intention can be obtained.

[0090] For example, the second, third, and fourth semes are individually subjected to the four operations of addition, subtraction, multiplication, and division to obtain four fifth vectors. Then, using the mapping relationship between semes and representation vectors, each fifth vector is mapped to obtain the corresponding fifth sememe. These fifth semes are then combined in a predetermined order (e.g., addition, subtraction, multiplication, and division) and compared with the first sememe to identify any discrepancies between the two. Furthermore, these discrepancies are input as prompts into the LLM, which uses its language understanding capabilities to identify the logical relationship between the first intent, the first entity information, and the second intent.

[0091] S1085: Based on the logical relationship, the second sememe, the third sememe, and the fourth sememe are combined to obtain a third intention of the first conversation text.

[0092] For example, a logical relationship might be: the second intent is created by adding the first entity representing the subject of the intent and the first entity representing the object of the intent in the first entity information to the first intent. Then, the second sememe of the first intent, the sememes belonging to this first entity in the third sememe, and the fourth sememe of the second intent can be combined to create the third intent of the first conversation text.

[0093] In the application, in order to avoid contradictory expressions in the third intent, which would provide erroneous supervision signals for the model training process and affect the model training effect, when combining the second, third and fourth semes, the contradictory semes can be eliminated first, and the remaining compatible semes can be combined to obtain the third intent.

[0094] The above describes some implementations of S108. Of course, it should be understood that S108 can also be implemented in other ways, which are not limited in this embodiment of the present application.

[0095] S110: Train the model based on the first conversation text and the third intent to obtain a second model.

[0096] The model here is the same as the model in the above step S106.

[0097] In one implementation, supervised training is used to train a model based on the first conversation text and its corresponding labels, using the third intent as the sample. This allows the model to fully absorb the supervisory signal provided by the third intent, helping the model overcome its inherent recognition biases and limitations and improving its intent recognition capabilities. The trained model is referred to as the second model.

[0098] For example, the model performs intent recognition on the first conversation text to obtain the intent of the first conversation text; then, based on the difference between the first intent and the third intent, the model loss is determined, and the backpropagation algorithm is used to adjust the model parameters based on the loss. This process is repeated multiple times until a preset training stop condition is met, thereby obtaining a second model. The training stop condition can be set according to actual needs, such as when the model loss converges or the number of times the model is adjusted reaches a threshold, and this embodiment of the application is not limited to this.

[0099] In another implementation, some repeated and less frequently occurring third intents and their corresponding first conversation texts may have a negative impact on the training effect of the model. To this end, the distribution of the frequency of occurrence of each third intent in all first conversation texts is first counted, and repeated and less frequently occurring third intents and their corresponding first conversation texts are filtered out. The model is then trained using the remaining third intents and their corresponding first conversation texts, which can significantly improve the training effect of the model.

[0100] Specifically, based on the third intent of each first conversation text, the third intent that appears more than a frequency threshold in multiple first conversation texts is determined as the sixth intent; the model is trained based on the sixth intent and the first conversation texts with the sixth intent to obtain a second model.

[0101] The above describes some implementations of the above S110. Of course, it should be understood that the above S110 can also be implemented in other ways, which are not limited in the present embodiment.

[0102] S112: Compare the recognition accuracy of the first model and the second model on the second conversation text, and determine a third model from the first model and the second model based on the comparison result.

[0103] The second conversation text may be any conversation text used as a test sample. The number of the second conversation texts may be multiple. The second conversation texts may have corresponding intention labels, and the intention labels may represent the true intention of the second conversation texts.

[0104] For example, the first model is used to identify the intent of the second conversation text, and the recognition accuracy of the first model is determined based on the difference between the identified intent and the intent label of the second conversation text. Furthermore, the second model is used to identify the intent of the second conversation text, and the recognition accuracy of the second model is determined based on the difference between the identified intent and the intent label of the second conversation text. Furthermore, the model with the higher recognition accuracy between the first and second models is determined as the third model.

[0105] In the application, the recognition accuracy of the model can be represented by various indicators in the field, such as accuracy, precision, recall, F1 value, etc., which are not limited in the embodiments of the present application. Among them, accuracy refers to the proportion of correctly predicted samples to the total samples. Precision refers to the proportion of samples predicted to be positive that are actually positive. Recall refers to the proportion of samples that are actually positive that are correctly predicted to be positive. F1 score is the harmonic mean of precision and recall, which comprehensively considers precision and recall.

[0106] The model training method provided in the embodiment of the present application identifies a first intention and a first entity information from a first conversation text. The first entity information reflects a specific or abstract object with a clear referential meaning in the first conversation text, and semantics is constructed and expressed through these entity information. On the one hand, the first intention and the first entity information are integrated to obtain the second intention of the first conversation text. Compared with the first intention, the granularity of the second intention is finer, and it can more accurately portray the intention tendency contained in the first conversation text. On this basis, the model is trained based on the first conversation text and its second intention, so that the model can fully absorb the supervision signal provided by the second intention, and deeply learn the precise and comprehensive analysis of the semantics of the first conversation text from multiple dimensions, thereby having the ability to accurately identify the conversation intention. The recognition capability can accurately capture the hidden intention information in various complex conversation texts; on the other hand, the first intention, the first entity information and the second intention are integrated to obtain the third intention of the first conversation text. Compared with the first intention and the second intention, the third intention contains more comprehensive and complete information. On this basis, the model is trained based on the first conversation text and its third intention, which helps the model break through its own recognition preferences and limitations and improve the model's intention recognition capability; finally, the recognition accuracy of the first model and the second model is evaluated based on the second conversation text, and the model with better recognition accuracy is determined from the two as the third model for conversation intention recognition, so as to accurately and comprehensively analyze the conversation semantics during the conversation, and then accurately identify the conversation intention.

[0107] Based on the third model trained by the above model training method, the embodiment of the present application also proposes a response method. Figure 3 , is a flow chart of a response method provided in one embodiment of the present application, the method comprising the following steps: S302: In response to the third conversation text of the first user, obtain the portrait information of the first user.

[0108] The first user's profile information reflects the first user's characteristics and preferences. The profile information may include values ​​for multiple key fields, which can be set based on the application scenario. For example, in a job search session, the key fields may include: gender, age, length of service, professional experience, highest income level, lowest income level, average income level, etc.

[0109] When the third conversation text of the first user is received, the portrait information of the first user can be obtained from the portrait library based on the user identifier of the first user; if the portrait information of the first user does not exist in the portrait library, the portrait information of the first user can be obtained by performing portrait analysis on the previous conversation texts, personal information, etc. of the first user using LLM.

[0110] S304: Perform entity recognition on the third conversation text to obtain third entity information.

[0111] The third entity information reflects a specific or abstract object with a clear referential meaning in the third conversation text. The third entity information may include, but is not limited to: an entity in the third conversation text (herein, the entity is referred to as the third entity), the entity's state, and the relationship between different entities.

[0112] Specifically, the entity recognition model performs entity recognition on the third conversation text to obtain third entity information. The entity recognition model encodes each word in the third conversation text to obtain a high-level semantic representation of each word (i.e., a representation vector for each word). The entity recognition model then uses a CRF to classify each word's representation vector to obtain a classification result for each word, indicating whether the word belongs to an entity. Furthermore, for each entity in the third conversation text, the model extracts features of the entity and features of its contextual information, and identifies the entity's state based on these features. Furthermore, the model extracts features of the different entities in the first and third conversation texts, as well as the distances between them and the grammatical relationships between them. Based on this information, the model predicts the relationships between the different entities.

[0113] In the application, by performing entity recognition on the third conversation text, not only the entities in the third conversation text can be obtained, but also the weights of each entity in the third conversation text can be obtained. The weight of the entity represents the importance of the entity to the third conversation text. The above-mentioned third entity can be an entity in the third conversation text whose weight is greater than the weight threshold. For example, the entities in the third conversation text "I am Xiao Ai Assistant who helps phone owners answer calls. What's the matter?" include "phone owner", "answer the phone", "Xiao Ai Assistant" and "what's the matter". The weights of "Xiao Ai Assistant" and "what's the matter" are higher than the weight threshold, so these two entities are determined as third entities. In this way, when responding to the third conversation text, it is helpful to pay more attention to the important entities in the third conversation text.

[0114] S306: Perform intent recognition on the third conversation text using a third model to obtain the intent of the third conversation text.

[0115] Since the third model already has the ability to map from conversation text to intent, the intent of the third conversation text can be obtained by inputting the third conversation text into the third model.

[0116] S308: Determine the first response text corresponding to the third conversation text based on the third entity information, the intention of the third conversation text, and the portrait information.

[0117] In one implementation, the prompt words corresponding to the response task, the third entity information, the intention of the third conversation text, and the portrait information of the first user are input into the LLM, and the natural language understanding and natural language generation capabilities of the LLM are used to generate the first response text corresponding to the third conversation text.

[0118] In another implementation, based on the third entity information and the intention of the third conversation text, multiple second response texts corresponding to the third conversation text are generated; for each second response text, based on the representation vector of the second response text and the sixth vector corresponding to the portrait information of the first user, the similarity between the second response text and the portrait information is determined; and the second response text among the above multiple second response texts whose similarity with the portrait information is greater than the similarity threshold is determined as the first response text.

[0119] Specifically, the prompt words, third entity information and third conversation text corresponding to the answer task are input into LLM, and the natural language understanding and natural language generation capabilities of LLM are used to generate multiple second answer texts.

[0120] For example, the prompt for the answering task is "You are a professional text student. Please generate diverse answer texts based on the following named entities and text intents." The third conversation text is "I want to find a part-time job close to home with a high salary." The intent of the third conversation text is "find a part-time job." The third entity information includes the third entity [home, salary], the state of the third entity, and the relationship between different third entities [([close] home), ([high] salary)]. Entering this information into the LLM yields the following multiple second answer texts: Response 1: I see. You're looking for a part-time job close to home with a relatively high salary, right? No problem. We can help you keep an eye out for such opportunities.

[0121] Response text 2: OK, you want to find a part-time job that is close to home and has a good salary. We will screen suitable positions for you based on this condition.

[0122] Response 3: Besides being close to home and having a high salary, do you have any other specific requirements for part-time work hours and work content? This will help us make more accurate recommendations for you.

[0123] Furthermore, the similarity between the representation vector of each second response text and the sixth vector corresponding to the portrait information of the first user is calculated. Assuming that the similarity between the response text 2 and the portrait information is greater than the similarity threshold, the response text 2 is determined as the first response text.

[0124] In the embodiment of the present application, the sixth vector corresponding to the portrait information may be a vector obtained after embedding the portrait information, or may be a vector determined by the following steps: Step C1: embed the portrait information to obtain a representation vector of the portrait information.

[0125] Specifically, group embedding is performed according to the type of field value of each key field in the portrait information. For example, for the key field "gender", it has only two values ​​0 and 1, so no embedding is required; for the key fields "age", "maximum income level", "minimum income level" and "average income level", the maximum normalization function can be used. , normalize the field values ​​of such fields to achieve embedding processing of the field values ​​of such fields; for the key field "career experience", the field value of this field can be input into the pre-trained language model to achieve embedding processing of the field value.

[0126] Step C2: embedding the fourth conversation text corresponding to the portrait information to obtain a representation vector of the fourth conversation text.

[0127] The fourth conversation text may be a historical conversation text of the second user with portrait information, or a historical conversation text of the first user before that, etc., which is not limited in the embodiment of the present application. The second user refers to a user other than the first user.

[0128] Step C3: Fusing the representation vector of the portrait information and the representation vector of the fourth conversation text to obtain a sixth vector corresponding to the portrait information.

[0129] Exemplarily, by multiplying the representation vector of the portrait information and the representation vector of the fourth conversation text, the two are fused to obtain the sixth vector corresponding to the portrait information.

[0130] The answering method provided in the embodiment of the present application, during the conversation process, performs intent recognition on the user's conversation text through a third model, and accurately and completely extracts the intent of the conversation text; in addition, entity recognition is also performed on the user's conversation text to obtain corresponding entity information, and the user's portrait information is obtained. The entity information reflects the specific or abstract objects with clear referential meaning in the conversation text, and is the basis for constructing and expressing semantics, while the portrait information reflects the user's characteristics and preferences. These two types of information help to understand the user's real personalized needs more comprehensively and in-depth; further, the user's portrait information, the entity information of the conversation text, and the intent of the conversation text are integrated to determine the answer text corresponding to the conversation text, so that the response to the conversation text is more personalized, targeted, and more in line with user expectations, thereby improving the accuracy of the response.

[0131] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0132] With the above Figure 1 The invention concept of the model training method shown in the figure is the same. The embodiment of the present application also proposes a model training device. Figure 4 , is a structural diagram of a model training device 400 provided in an embodiment of the present application, the device 400 includes: an identification module 410, a first training module 420, a second training module 430 and a first determination module 440.

[0133] The recognition module 410 is configured to recognize a first intention and first entity information of a first conversation text.

[0134] The first training module 420 is used to fuse the first intent and the first entity information to obtain the second intent of the first conversation text, and to train a model based on the first conversation text and the second intent to obtain a first model.

[0135] The second training module 430 is used to fuse the first intent, the first entity information and the second intent to obtain the third intent of the first conversation text, and to train the model based on the first conversation text and the third intent to obtain a second model.

[0136] The first determination module 440 is configured to compare the recognition accuracy of the first model and the second model on the second conversation text, and determine a third model from the first model and the second model based on the comparison result.

[0137] In another embodiment, the identification module is configured to: performing intent recognition on the first conversation text to obtain a fourth intent of the first conversation text; Performing entity recognition on the first conversation text to obtain second entity information of the first conversation text; Performing intent recognition on the first conversation text based on the fourth intent and the second entity information to obtain a first intent of the first conversation text; Entity recognition is performed on the first conversation text based on the fourth intent and the second entity information to obtain first entity information of the first conversation text.

[0138] In another embodiment, when the recognition module performs intent recognition on the first conversation text based on the fourth intent and the second entity information to obtain the first intent of the first conversation text, the recognition module performs the following steps: Encoding and linearly transforming the first conversation text to obtain a first vector; Normalizing the first vector to obtain first probabilities that the first conversation text corresponds to a plurality of preset intents; fusing the first vector and the representation vector of the description text of the first intent to obtain a second vector, and normalizing the second vector to obtain second probabilities that the first conversation text corresponds to the multiple intents respectively; fusing the first vector, the representation vector of the description text of the first intent, and the representation vector of the second entity information to obtain a third vector, and normalizing the third vector to obtain multiple third probabilities that the first conversation text corresponds to the multiple intents respectively; For each intent, determining a fourth probability that the first conversation text corresponds to the intent based on the first probability, the second probability, and the third probability that the first conversation text corresponds to the intent; The fourth intent with the highest probability among the multiple intents is determined as the first intent of the first conversation text.

[0139] In another embodiment, when the recognition module performs entity recognition on the first conversation text based on the fourth intent and the second entity information to obtain the first entity information of the first conversation text, the recognition module performs the following steps: Performing entity recognition on the first conversation text to obtain second entity information of the first conversation text; fusing the representation vector of the second entity information and the representation vector of the fourth intent to obtain a fourth vector; Mapping is performed on the fourth vector to obtain first entity information of the first conversation text.

[0140] In another embodiment, the first training module is used to: Obtaining a template for describing an intent, the template comprising at least one of the following slots: a slot for describing a cause of the intent, a slot for describing a result of the intent, a slot for describing a subject of the intent, and a slot for describing an object of the intent; Embedding the name of each slot in the template to obtain a representation vector for each slot; Determining a first slot in the template that matches the first intent based on a similarity between the representation vector of the first intent and the representation vector of each slot; determining, based on a similarity between a representation vector of the first entity information and a representation vector of each slot, a second slot in the template that matches the first entity information; The first intent is filled into the first slot, and the first entity information is filled into the second slot to obtain the second intent of the first conversation text.

[0141] In another embodiment, the second training module is used to: performing semantic decomposition on the first intent, the first entity information, and the second intent to obtain a plurality of semantic elements, the plurality of semantic elements including the second semantic element of the first intent, the third semantic element of the first entity information, and the fourth semantic element of the second intent; determining a first relevance between the plurality of sememes and a second relevance between each sememe and the first conversation text; determining a first sememe related to the first conversation text from the plurality of sememes based on the first relevance and the second relevance; determining a logical relationship between the first intent, the first entity information, and the second intent based on the first sememe, the second sememe, the third sememe, and the fourth sememe; Based on the logical relationship, the second sememe, the third sememe, and the fourth sememe are combined to obtain a third intention of the first conversation text.

[0142] In another embodiment, when the first training module determines the logical relationship between the first intent, the first entity information, and the second intent based on the first sememe, the second sememe, the third sememe, and the fourth sememe, the first training module performs the following steps: performing operations on the representation vector of the second element, the representation vector of the third element, and the representation vector of the fourth element based on a plurality of operation modes to obtain fifth vectors corresponding to the plurality of operation modes; Map each fifth vector to obtain the fifth element corresponding to each fifth vector; Based on the difference between each fifth sememe and the first sememe and the operation mode corresponding to each fifth sememe, a logical relationship among the first intent, the first entity information, and the second intent is determined.

[0143] In another embodiment, the number of the first conversation texts is multiple; The first training module is used to: Based on the second intent of each first conversation text, determining a second intent that appears in the plurality of first conversation texts at a frequency greater than a frequency threshold as a fifth intent; A model is trained based on the fifth intent and a first conversation text having the fifth intent to obtain a first model.

[0144] In another embodiment, the number of the first conversation texts is multiple; The second training module is used to: Based on the third intent of each first conversation text, determining the third intent whose appearance frequency in the plurality of first conversation texts is greater than a frequency threshold as a sixth intent; The model is trained based on the sixth intent and the first conversation text having the sixth intent to obtain a second model.

[0145] Obviously, the model training device provided in the embodiment of the present application can be used as Figure 1 The execution body of the model training method shown, for example Figure 1 In the model training method shown, step S102 can be performed by Figure 4 The recognition module 410 in the model training device shown in FIG. 1 is executed, and steps S104 and S106 can be performed by Figure 4 The first training module 420 in the model training apparatus shown in FIG. 1 is executed, and steps S108 and S110 can be performed by Figure 4 The second training module 430 in the model training device shown in FIG. 1 is executed, and step S112 can be performed by Figure 4 The first determination module 440 in the model training device shown is executed.

[0146] According to another embodiment of the present application, Figure 4 The various modules in the model training device shown can be individually or completely combined into one or several other modules to form a whole, or one (or more) of the modules can be further divided into multiple smaller modules to form a whole, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of a module can also be implemented by multiple modules, or the functions of multiple modules can be implemented by one module. In an embodiment of the present application, the model training device may also include other modules. In actual applications, these modules can also be implemented with the assistance of other modules, and can be implemented by the collaboration of multiple modules.

[0147] According to another embodiment of the present application, a general computing device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM) and other processing elements and storage elements can be run to execute the following operations: Figure 1A computer program (including program code) for each step involved in the corresponding method shown in FIG. Figure 4 The model training device shown in the figure and the model training method of the embodiment of the present application are implemented. The computer program can be recorded on a computer-readable storage medium, for example, and transferred to an electronic device through the computer-readable storage medium and run therein.

[0148] With the above Figure 3 The present invention also proposes a response device. Figure 5 , is a structural diagram of a response device 500 provided in an embodiment of the present application, the device 500 includes: an acquisition module 510, an identification module 520 and a second determination module 530.

[0149] The acquisition module 510 is used to obtain the portrait information of the first user in response to the third conversation text of the first user.

[0150] The recognition module 520 is configured to perform entity recognition on the third conversation text to obtain third entity information.

[0151] The recognition model is also used to perform intent recognition on the third conversation text through a third model to obtain the intent of the third conversation text. The third model is trained based on the model training method provided in the embodiment of the present application.

[0152] The second determination module 530 is used to determine the first response text corresponding to the third conversation text based on the third entity information, the intention of the third conversation text and the portrait information.

[0153] In another embodiment, the second determining module is configured to: generating a plurality of second response texts corresponding to the third conversation text based on the third entity information and the intention of the third conversation text; For each second response text, determining the similarity between the second response text and the portrait information based on the representation vector of the second response text and the sixth vector corresponding to the portrait information; The second response text among the multiple second response texts whose similarity with the portrait information is greater than a similarity threshold is determined as the first response text.

[0154] In another embodiment, the sixth vector corresponding to the portrait information is determined by: Embedding the portrait information to obtain a representation vector of the portrait information; Performing embedding processing on the fourth conversation text corresponding to the portrait information to obtain a representation vector of the fourth conversation text; The representation vector of the portrait information and the representation vector of the fourth conversation text are fused to obtain a sixth vector corresponding to the portrait information.

[0155] Obviously, the answering device provided in the embodiment of the present application can be used as Figure 3 The execution subject of the response method shown, for example Figure 3 In the response method shown, step S302 can be performed by Figure 5 The acquisition module 310 in the response device shown in FIG. 1 is executed, and steps S304 and S306 can be performed by Figure 5 The identification module 520 in the answering device shown in FIG. 1 is executed, and step S308 can be performed by Figure 5 The second determination module 530 in the response device is shown to execute.

[0156] According to another embodiment of the present application, Figure 5 The various modules in the response device shown can be individually or completely combined into one or more other modules to form a structure, or one (or more) of the modules can be further divided into multiple functionally smaller modules to form a structure, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one module can also be implemented by multiple modules, or the functions of multiple modules can be implemented by one module. In the embodiments of the present application, the response device may also include other modules. In actual applications, these modules can also be assisted by other modules and can be implemented by the collaboration of multiple modules.

[0157] According to another embodiment of the present application, a general computing device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM) and other processing elements and storage elements can be run to execute the following operations: Figure 1 A computer program (including program code) for each step involved in the corresponding method shown in FIG. Figure 5 The computer program can be recorded on a computer-readable storage medium, for example, and transferred to an electronic device through the computer-readable storage medium and run therein.

[0158] Figure 6 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 6At the hardware level, the electronic device includes a processor and, optionally, an internal bus, a network interface, and memory. The memory may include internal memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for its services.

[0159] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 6 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0160] The memory is used to store programs. Specifically, the program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.

[0161] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a model training device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations: identifying a first intent and first entity information of a first conversation text; fusing the first intent and the first entity information to obtain a second intent of the first conversation text, and training a model based on the first conversation text and the second intent to obtain a first model; fusing the first intent, the first entity information, and the second intent to obtain a third intent of the first conversation text, and training the model based on the first conversation text and the third intent to obtain a second model; The recognition accuracy of the first model and the second model on the second conversation text are compared, and based on the comparison result, a third model is determined from the first model and the second model.

[0162] Alternatively, the processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a response device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations: In response to the third conversation text of the first user, obtaining portrait information of the first user; Performing entity recognition on the third conversation text to obtain third entity information; performing intent recognition on the third conversation text using a third model to obtain the intent of the third conversation text, wherein the third model is trained based on the model training method provided in an embodiment of the present application; Based on the third entity information, the intention of the third conversation text and the portrait information, a first response text corresponding to the third conversation text is determined.

[0163] The above application Figure 1 The method performed by the model training device disclosed in the embodiment shown or the above-mentioned application Figure 3 The methods performed by the response device disclosed in the illustrated embodiments can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be performed by hardware integrated logic circuits within the processor or by software instructions. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules within the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0164] The electronic device may also perform Figure 1 Method, and realize the model training device in Figure 1 、 Figure 2 Alternatively, the electronic device may also perform the functions of the embodiment shown. Figure 3 The method and the answering device are implemented in Figure 3 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.

[0165] Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0166] The embodiment of the present application also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by an electronic device including multiple application programs, can enable the electronic device to execute Figure 1 The method of the embodiment shown is specifically used to perform the following operations: identifying a first intent and first entity information of a first conversation text; fusing the first intent and the first entity information to obtain a second intent of the first conversation text, and training a model based on the first conversation text and the second intent to obtain a first model; fusing the first intent, the first entity information, and the second intent to obtain a third intent of the first conversation text, and training the model based on the first conversation text and the third intent to obtain a second model; The recognition accuracy of the first model and the second model on the second conversation text are compared, and based on the comparison result, a third model is determined from the first model and the second model.

[0167] Alternatively, when the instruction is executed by an electronic device including multiple applications, the electronic device can execute Figure 3 The method of the embodiment shown is specifically used to perform the following operations: In response to the third conversation text of the first user, obtaining portrait information of the first user; Performing entity recognition on the third conversation text to obtain third entity information; performing intent recognition on the third conversation text using a third model to obtain the intent of the third conversation text, wherein the third model is trained based on the model training method provided in an embodiment of the present application; Based on the third entity information, the intention of the third conversation text and the portrait information, a first response text corresponding to the third conversation text is determined.

[0168] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute part or all of the steps in the model training method or response method provided in the embodiment of the present application.

[0169] In short, the above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

[0170] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0171] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0172] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0173] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

Claims

1. A model training method, characterized in that: include: identifying a first intent and first entity information of a first conversation text; fusing the first intent and the first entity information to obtain a second intent of the first conversation text, and training a model based on the first conversation text and the second intent to obtain a first model; fusing the first intent, the first entity information, and the second intent to obtain a third intent of the first conversation text, and training the model based on the first conversation text and the third intent to obtain a second model; The recognition accuracy of the first model and the second model on the second conversation text are compared, and based on the comparison result, a third model is determined from the first model and the second model.

2. The method according to claim 1, characterized in that The identifying the first intention and the first entity information of the first conversation text includes: performing intent recognition on the first conversation text to obtain a fourth intent of the first conversation text; Performing entity recognition on the first conversation text to obtain second entity information of the first conversation text; Performing intent recognition on the first conversation text based on the fourth intent and the second entity information to obtain a first intent of the first conversation text; Entity recognition is performed on the first conversation text based on the fourth intent and the second entity information to obtain first entity information of the first conversation text.

3. The method according to claim 2, characterized in that The performing intent recognition on the first conversation text based on the fourth intent and the second entity information to obtain the first intent of the first conversation text includes: Encoding and linearly transforming the first conversation text to obtain a first vector; Normalizing the first vector to obtain first probabilities that the first conversation text corresponds to a plurality of preset intents; fusing the first vector and the representation vector of the description text of the first intent to obtain a second vector, and normalizing the second vector to obtain second probabilities that the first conversation text corresponds to the multiple intents respectively; fusing the first vector, the representation vector of the description text of the first intent, and the representation vector of the second entity information to obtain a third vector, and normalizing the third vector to obtain multiple third probabilities that the first conversation text corresponds to the multiple intents respectively; For each intent, determining a fourth probability that the first conversation text corresponds to the intent based on the first probability, the second probability, and the third probability that the first conversation text corresponds to the intent; The fourth intent with the highest probability among the multiple intents is determined as the first intent of the first conversation text.

4. The method according to claim 2, characterized in that The performing entity recognition on the first conversation text based on the fourth intent and the second entity information to obtain first entity information of the first conversation text includes: Performing entity recognition on the first conversation text to obtain second entity information of the first conversation text; fusing the representation vector of the second entity information and the representation vector of the fourth intent to obtain a fourth vector; Mapping is performed on the fourth vector to obtain first entity information of the first conversation text.

5. The method according to claim 1, wherein The fusing the first intent and the first entity information to obtain the second intent of the first conversation text includes: Obtaining a template for describing an intent, the template comprising at least one of the following slots: a slot for describing a cause of the intent, a slot for describing a result of the intent, a slot for describing a subject of the intent, and a slot for describing an object of the intent; Embedding the name of each slot in the template to obtain a representation vector for each slot; Determining a first slot in the template that matches the first intent based on a similarity between the representation vector of the first intent and the representation vector of each slot; determining, based on a similarity between a representation vector of the first entity information and a representation vector of each slot, a second slot in the template that matches the first entity information; The first intent is filled into the first slot, and the first entity information is filled into the second slot to obtain the second intent of the first conversation text.

6. The method according to claim 1, characterized in that The fusing the first intent, the first entity information, and the second intent to obtain the third intent of the first conversation text includes: performing semantic decomposition on the first intent, the first entity information, and the second intent to obtain a plurality of semantic elements, the plurality of semantic elements including the second semantic element of the first intent, the third semantic element of the first entity information, and the fourth semantic element of the second intent; determining a first relevance between the plurality of sememes and a second relevance between each sememe and the first conversation text; determining a first sememe related to the first conversation text from the plurality of sememes based on the first relevance and the second relevance; determining a logical relationship between the first intent, the first entity information, and the second intent based on the first sememe, the second sememe, the third sememe, and the fourth sememe; Based on the logical relationship, the second sememe, the third sememe, and the fourth sememe are combined to obtain a third intention of the first conversation text.

7. The method according to claim 6, characterized in that The determining the logical relationship between the first intent, the first entity information, and the second intent based on the first sememe, the second sememe, the third sememe, and the fourth sememe includes: performing operations on the representation vector of the second element, the representation vector of the third element, and the representation vector of the fourth element based on a plurality of operation modes to obtain fifth vectors corresponding to the plurality of operation modes; Map each fifth vector to obtain the fifth element corresponding to each fifth vector; Based on the difference between each fifth sememe and the first sememe and the operation mode corresponding to each fifth sememe, a logical relationship among the first intent, the first entity information, and the second intent is determined.

8. The method according to claim 1, characterized in that There are multiple first conversation texts; and training a model based on the first conversation texts and the second intent to obtain a first model includes: Based on the second intent of each first conversation text, determining a second intent that appears in the plurality of first conversation texts at a frequency greater than a frequency threshold as a fifth intent; A model is trained based on the fifth intent and a first conversation text having the fifth intent to obtain a first model.

9. The method according to claim 1, characterized in that There are multiple first conversation texts; and training the model based on the first conversation texts and the third intent to obtain a second model includes: Based on the third intent of each first conversation text, determining the third intent whose appearance frequency in the plurality of first conversation texts is greater than a frequency threshold as a sixth intent; The model is trained based on the sixth intent and the first conversation text having the sixth intent to obtain a second model.

10. A response method, characterized in that: include: In response to the third conversation text of the first user, obtaining portrait information of the first user; Performing entity recognition on the third conversation text to obtain third entity information; performing intent recognition on the third conversation text using a third model to obtain the intent of the third conversation text, wherein the third model is trained based on the model training method according to any one of claims 1 to 9; Based on the third entity information, the intention of the third conversation text and the portrait information, a first response text corresponding to the third conversation text is determined.

11. The method according to claim 10, characterized in that The determining, based on the third entity information, the intent of the third conversation text, and the portrait information, of a first response text corresponding to the third conversation text includes: generating a plurality of second response texts corresponding to the third conversation text based on the third entity information and the intention of the third conversation text; For each second response text, determining the similarity between the second response text and the portrait information based on the representation vector of the second response text and the sixth vector corresponding to the portrait information; The second response text among the multiple second response texts whose similarity with the portrait information is greater than a similarity threshold is determined as the first response text.

12. The method according to claim 11, characterized in that The sixth vector corresponding to the portrait information is determined in the following manner: Embedding the portrait information to obtain a representation vector of the portrait information; Performing embedding processing on the fourth conversation text corresponding to the portrait information to obtain a representation vector of the fourth conversation text; The representation vector of the portrait information and the representation vector of the fourth conversation text are fused to obtain a sixth vector corresponding to the portrait information.

13. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the model training method as described in any one of claims 1 to 9 or the response method as described in any one of claims 10 to 12.

14. A computer program product, characterized in that The computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute part or all of the steps in the model training method as described in any one of claims 1 to 9 or the response method as described in any one of claims 10 to 12.

Citation Information

Cited By

  • Answer determination method and device for intelligent questions and answers

    CN121561050A