Response method and device based on intention recognition, equipment and storage medium
The intent recognition method, which combines multimodal information fusion and deep semantic analysis, solves the problem of low intent recognition accuracy in traditional methods, achieving higher recognition accuracy and lower rule maintenance costs, and improving the adaptability and coherence of intelligent dialogue systems.
Patent Information
- Application Number
- CN202511113975.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-28
AI Technical Summary
Traditional intent recognition methods mainly rely on a single rule matching method or machine learning model, which is difficult to cope with the complex and ever-changing ways users express themselves, resulting in insufficient accuracy in intent recognition and an inability to accurately understand user intent.
We employ a method that combines multimodal information fusion, deep semantic analysis, and contextual association. We first identify the target based on the priority of a pre-defined rule base, then combine deep semantic analysis and entity recognition to generate more accurate intent recognition results. Finally, we dynamically update the rule base through weakly supervised learning.
It significantly improves the accuracy of complex intent recognition, shortens the adaptation cycle to new intents, reduces rule maintenance costs, and enhances the coherence and task completion rate of the agent in multi-turn dialogue scenarios.
Smart Images

Figure CN121031607A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to a response method, apparatus, device, and storage medium based on intent recognition. Background Technology
[0002] Intent recognition systems are widely used in various fields such as dialogue assistants, customer service management, and medical Q&A. These systems can identify a user's intent in response to their input, thereby generating a response to better interact and provide feedback. Traditional intent recognition methods mainly rely on rule-based matching or machine learning models. Rule-based matching alone struggles to handle the complex and ever-changing expressions of users, resulting in poor intent recognition performance. Machine learning models, based on pre-trained content, often fail to quickly adapt to new intents or expressions, leading to inaccurate intent recognition.
[0003] Therefore, how to propose a method that can improve the accuracy of user intent recognition and interact with users has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, this application provides a response method, apparatus, device and storage medium based on intent recognition, the main purpose of which is to solve the problem that the current traditional intent recognition methods mainly rely on a single rule matching method or machine learning model method, the intent recognition accuracy is not accurate enough, and it is impossible to accurately understand the user's intent.
[0005] In a first aspect, this application provides a response method based on intent recognition, comprising:
[0006] The user's multimodal input content is acquired, and the input content is preprocessed to obtain the first text feature;
[0007] Based on the priority of preset rules in the preset rule base, each preset rule is traversed to perform preliminary intent recognition on the first text feature; wherein, the preset rule base includes at least two preset rules, each preset rule includes a second text feature under the current preset rule and an intent tag corresponding to the second text feature, the intent tag being used to determine the intent of the first text feature;
[0008] If the preliminary intent recognition fails to yield a result, deep semantic analysis is performed on the first text features to obtain an intent recognition result; wherein, the intent recognition result includes a comprehensive fusion result of text semantic vectors, multimodal vectors, and context vectors based on the first text features;
[0009] Entity recognition is performed on the first text features, and the entity recognition result is fused with the intent recognition result to obtain a semantic representation;
[0010] Response content is generated based on the semantic representation.
[0011] Optionally, the step of traversing each preset rule to perform preliminary intent recognition on the first text feature according to the priority of preset rules in the preset rule base includes: determining the priority order of preset rules in the preset rule base; matching the second text feature corresponding to the preset rule with the first text feature in descending order of priority until a preliminary intent recognition result is generated or the second text feature corresponding to the preset rule with the lowest priority is matched with the first text feature; wherein, if the second text feature corresponding to the preset rule is successfully matched with the first text feature, the intent label corresponding to the second text feature is output as the preliminary intent recognition result.
[0012] Optionally, after performing preliminary intent recognition on the first text feature by traversing each preset rule according to the priority of preset rules in the preset rule base, the method further includes: obtaining a first text feature that has undergone the preliminary intent recognition within a preset time period but has not obtained the preliminary intent recognition result; generating candidate rules for the first text feature using a weakly supervised learning method; calculating the similarity between the candidate rule and at least one preset rule in the preset rule base; and adding the candidate rule to the preset rule base if the similarity is less than a preset similarity threshold.
[0013] Optionally, if the preliminary intent recognition fails to yield a result, performing deep semantic analysis on the first text features to obtain an intent recognition result includes: mapping the first text features to a text semantic vector using a multi-layer self-attention mechanism and a feedforward neural network; acquiring multimodal information corresponding to the first text features and mapping the multimodal information to the same dimension as the text semantic vector to obtain a multimodal vector; the multimodal information includes at least one of speech, intonation, pitch, speed of sound, and emoticons; concatenating the text semantic vector with the multimodal vector to obtain an enhanced semantic vector; constructing a context vector of the enhanced semantic vector using a weighted summation method based on historical dialogue records; and fusing the context vector with the enhanced semantic vector and using it as input to a deep semantic analysis model to obtain the intent recognition result.
[0014] Optionally, the step of performing entity recognition on the first text feature and fusing the entity recognition result with the intent recognition result to obtain a semantic representation includes: performing forward encoding on the first text feature to obtain a forward hidden state of the first text feature, and performing backward encoding on the first text feature to obtain a backward hidden state of the first text feature; concatenating the forward hidden state and the backward hidden state to obtain a hidden state sequence of the first text feature; performing entity recognition based on the hidden state sequence of the first text feature to obtain the entity recognition result; and fusing the entity recognition result with the intent recognition result to obtain the semantic representation.
[0015] Optionally, the entity recognition result includes the entity and the relationship between entities;
[0016] After obtaining the entity recognition result, the method further includes: performing disambiguation processing on the entity recognition result by combining the context vector; constructing an entity relationship graph by taking the entities in the entity recognition result as nodes and the relationships between the entities in the entity recognition result as edges.
[0017] Optionally, after fusing the entity recognition result and the intent recognition result to obtain the semantic representation, the method further includes: using the entity relationship graph as input to a graph neural network to calculate the semantic correlation degree between the intent recognition result and the entity recognition result; the semantic correlation degree is used to verify the semantic consistency between the intent recognition result and the entity recognition result; if the semantic correlation degree is greater than a preset correlation degree threshold, it is determined that the semantic consistency verification is satisfied; if the semantic correlation degree is less than the preset correlation degree threshold, the entity recognition result and the intent recognition result are re-acquired.
[0018] Secondly, this application provides a response device based on intent recognition, comprising:
[0019] The acquisition unit is configured to acquire multimodal input content from the user and preprocess the input content to obtain a first text feature;
[0020] The matching unit is configured to perform preliminary intent recognition on the first text feature by traversing each preset rule according to the priority of preset rules in the preset rule library; wherein, the preset rule library includes at least two preset rules, each preset rule includes a second text feature under the current preset rule and an intent label corresponding to the second text feature, the intent label being used to determine the intent of the first text feature;
[0021] The analysis unit is configured to perform deep semantic analysis on the first text features to obtain an intent recognition result if the preliminary intent recognition has not been obtained; wherein the intent recognition result includes a comprehensive fusion result of text semantic vectors, multimodal vectors and context vectors based on the first text features;
[0022] The fusion unit is configured to perform entity recognition on the first text features and fuse the entity recognition result with the intent recognition result to obtain a semantic representation;
[0023] The generation unit is configured to generate response content based on the semantic representation.
[0024] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the intent-based response method described in the first aspect.
[0025] Fourthly, this application provides an electronic device including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the intent recognition-based response method described in the first aspect.
[0026] Using the above technical solution, this application provides a response method, apparatus, device, and storage medium based on intent recognition. First, it acquires multimodal input content from a user and preprocesses the input content to obtain a first text feature. Then, based on the priority of preset rules in a preset rule base, it traverses each preset rule to perform preliminary intent recognition on the first text feature. The preset rule base includes at least two preset rules, each including a second text feature under the current preset rule and an intent tag corresponding to the second text feature. The intent tag is used to determine the intent of the first text feature. If the preliminary intent recognition does not yield a result, deep semantic analysis is performed on the first text feature to obtain an intent recognition result. The intent recognition result includes a comprehensive fusion result of text semantic vectors, multimodal vectors, and context vectors based on the first text feature. Entity recognition is performed on the first text feature, and the entity recognition result is fused with the intent recognition result to obtain a semantic representation. Finally, response content is generated based on the semantic representation. This application performs preliminary intent recognition on the preprocessed text one by one according to the priority of the preset rules in the preset rule base, thereby improving the success rate of preliminary recognition; if the preliminary intent recognition result is not obtained, multimodal information and contextual information are then integrated to perform deep semantic analysis, which has higher intent recognition accuracy than the method of analyzing only the text; then the intent recognition result is fused with the entity recognition result to obtain a semantic representation, and response content to the user is generated based on the semantic representation.
[0027] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0028] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 A flowchart illustrating a response method based on intent recognition provided in an embodiment of this application is shown.
[0030] Figure 2 A flowchart illustrating another intent-based response method provided in an embodiment of this application is shown.
[0031] Figure 3 A flowchart illustrating another intent-based response method provided in an embodiment of this application is shown.
[0032] Figure 4 A flowchart illustrating another intent-based response method provided in an embodiment of this application is shown.
[0033] Figure 5 A schematic diagram of the structure of a response device based on intent recognition provided in an embodiment of this application is shown. Detailed Implementation
[0034] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.
[0035] As mentioned in the background, intent recognition systems can identify a user's intent in response to their input, thereby generating response content for better interaction and feedback. Currently, intent recognition systems are typically based on single-rule matching or machine learning models. Single-rule matching involves identifying intent based on pre-defined keywords and expressions, which struggles to handle complex and varied user expressions, resulting in limited ability to recognize complex intents. On the other hand, machine learning models, trained on labeled data, are less effective at understanding unfamiliar data, i.e., new intents (or user expressions).
[0036] To address the aforementioned issues, this embodiment proposes a response method based on intent recognition. Through innovative mechanisms such as multimodal information fusion (including text, speech tone, emoticons, etc.) for deep semantic analysis, context association construction, and deep fusion and verification of intent and entities, it aims to significantly improve the accuracy of complex intent recognition, reduce rule maintenance costs, shorten the adaptation cycle to new intents, and enhance the coherence and task completion rate of the agent in multi-turn dialogue scenarios.
[0037] The intent-based response method proposed in this application can be applied to an intent-based response device or electronic device. This device or electronic device can be installed or integrated into intelligent agents such as dialogue assistants, customer service management systems, and medical Q&A systems. During operation, it can execute any of the intent-based response methods mentioned below. Figure 1 As shown, the method includes:
[0038] S101, Obtain the user's multimodal input content and preprocess the input content to obtain the first text feature;
[0039] First, the user's multimodal input content is obtained. This multimodal form can include, for example, text, speech, or a combination of text and images. The input content is then preprocessed to obtain the first text feature. This first text feature includes keywords and regular expressions corresponding to the preprocessed user input content, and may also include multimodal information such as pitch, speed, intonation, and emoticons. The preprocessing includes modality recognition, format conversion, speech recognition, text cleaning, and word segmentation of the input content to obtain the first text feature.
[0040] Furthermore, the preprocessing process in S101 specifically includes: if the input is in the form of speech or text, the system first identifies the corresponding modality as speech or text, and then converts it into text. If the input is text combined with an image, or an image with text, the image is scanned, its features are extracted, and then converted into text.
[0041] Language recognition calculates the character n-gram features (sequences composed of n consecutive words or characters) through a language model to identify the language type. For example, the probability of the text appearing under each language model is calculated through a 2-gram language model, and the language with the highest probability is selected as the language type. For example, for the text "Hello, how's the weather today?", after extracting its 2-gram features, the probability under the Chinese language model is the highest, so the language type is Chinese.
[0042] Text cleaning is to use regular expressions to remove redundant spaces, line breaks, special characters and other irrelevant information in the text. By defining a series of matching patterns, such as matching consecutive spaces and replacing them with a single space, matching special symbols and directly removing them, etc., the cleaning of the text is achieved. For example, cleaning "Hello! How's the weather #today?" to "Hello! How's the weather today?"
[0043] S102, according to the priority of the preset rules in the preset rule library, traverse each preset rule to perform a preliminary intention recognition on the first text feature;
[0044] Among them, the preset rule library includes at least two preset rules, and each preset rule includes a second text feature under the current preset rule and an intention label corresponding to the second text feature. The intention label is used to determine the intention of the first text feature.
[0045] In S102, the preset rule library refers to a pre-defined rule library. The pre-defined rule library contains at least two preset rules. Each preset rule has a corresponding second text feature and an intention label corresponding to the second text feature. The second text feature includes keywords and regular expressions. Then, in the order of the priority of the preset rules, the keywords and regular expressions of each preset rule are matched with the pre-processed first text feature. If the match is successful, the corresponding intention label is returned to obtain the preliminary intention recognition result.
[0046] S103, in the case where the preliminary intention recognition is performed and no preliminary intention recognition result is obtained, perform a deep semantic analysis on the first text feature to obtain the intention recognition result;
[0047] Among them, the intention recognition result includes the comprehensive fusion result of the text semantic vector, multi-modal vector and context vector based on the first text feature.
[0048] In step S103, if the initial intent identification fails to yield a result, deep semantic analysis is then performed on the first text feature. Firstly, situations where the initial intent identification fails to yield a result include: matching each preset rule with the first text feature in priority order, up to the lowest priority preset rule, without successfully matching the corresponding intent tag; or successfully matching the corresponding intent tag under a certain preset condition, but the confidence level of that intent is less than a preset confidence threshold. The confidence level can be set in advance.
[0049] If initial intent recognition fails, it indicates that rule-based matching is not accurate enough, so deep semantic analysis is used for identification. This involves obtaining a text semantic vector based on first text features, obtaining a multimodal vector based on the multimodal information corresponding to the first text features, and obtaining a context vector from historical dialogue records. These are then integrated to obtain the intent recognition result based on deep semantic analysis. This method is more accurate than rule-based matching or methods that only identify intent based on text features.
[0050] S104, Perform entity recognition on the first text features, and fuse the entity recognition result with the intent recognition result to obtain a semantic representation;
[0051] The intent recognition result was obtained in S103, but intent recognition can only determine the user's abstract goal (such as "book a flight"), lacking the detailed information required for specific execution (such as time, location, etc.). Therefore, entity extraction (such as "next Wednesday", "Beijing", "Shanghai") and fusion are still needed to construct a complete and operable semantic representation for the system.
[0052] S105, Generate response content based on semantic representation.
[0053] The system then generates response content for interaction with the user based on the semantic representation.
[0054] In this embodiment, the user's multimodal input is first acquired and preprocessed to obtain a first text feature. Based on the priority of preset rules in a preset rule base, each preset rule is traversed to perform preliminary intent recognition on the first text feature. The preset rule base includes at least two preset rules, each including a second text feature under the current preset rule and an intent tag corresponding to the second text feature. The intent tag is used to determine the intent of the first text feature. If the preliminary intent recognition does not yield a result, deep semantic analysis is performed on the first text feature to obtain an intent recognition result. This intent recognition result includes a comprehensive fusion result of text semantic vectors, multimodal vectors, and context vectors based on the first text feature. Entity recognition is performed on the first text feature, and the entity recognition result is fused with the intent recognition result to obtain a semantic representation. Response content is generated based on the semantic representation. This embodiment performs preliminary intent recognition on the preprocessed text one by one according to the priority of the preset rules in the preset rule base, thereby improving the success rate of preliminary recognition. If the preliminary intent recognition result is not obtained, multimodal information and contextual information are then fused to perform deep semantic analysis, which has higher intent recognition accuracy than the method of analyzing only the text. Then, the intent recognition result is fused with the entity recognition result to obtain a semantic representation, and response content for the user is generated based on the semantic representation.
[0055] Optionally, based on the priority of preset rules in the preset rule base, each preset rule is traversed to perform preliminary intent recognition on the first text feature, including: determining the priority order of preset rules in the preset rule base; matching the second text feature corresponding to the preset rule with the first text feature in descending order of priority until a preliminary intent recognition result is generated or the second text feature corresponding to the preset rule with the lowest priority is matched with the first text feature; wherein, if the second text feature corresponding to the preset rule is successfully matched with the first text feature, the intent label corresponding to the second text feature is output as the preliminary intent recognition result.
[0056] In this embodiment, the preset rule base first contains at least two preset rules. Each preset rule has a corresponding second text feature and a corresponding intent label, where the second text feature includes keywords and regular expressions. Then, according to the priority order of the preset rules, the second text feature (keywords and regular expressions) of each preset rule is matched with the preprocessed first text feature. If a match is successful, the corresponding intent label is returned, and a preliminary intent recognition result is obtained.
[0057] The priority order of the rules is initially set and can be changed based on factors such as business relevance or frequency of occurrence. For example, in a practical application scenario like customer service, query intents appear frequently and are highly relevant. Therefore, the "query" rule has the highest initial priority, followed by other matching rules such as "pre-sales consultation" and "after-sales assistance." Therefore, according to the query rule, the keywords and regular expressions of the first text feature are searched and matched to see if the corresponding intent tag can be found. For example, there is a rule in the rule base with the pattern "Hello.*|Hi.*", corresponding to the intent tag "Greeting". When the input text is "Hello, how's the weather today?", the match is successful, and the intent "Greeting" is returned directly. By setting at least two matching rules in the preset condition library, the coverage of the initial intent recognition can be improved, thus making intent identification more effective.
[0058] Optionally, after performing preliminary intent recognition on the first text feature by traversing each preset rule according to the priority of preset rules in the preset rule base, the method further includes: obtaining the first text feature that has undergone preliminary intent recognition within a preset time period but has not obtained a preliminary intent recognition result; generating candidate rules for the first text feature through a weakly supervised learning method; calculating the similarity between the candidate rule and at least one preset rule in the preset rule base; and adding the candidate rule to the preset rule base if the similarity is less than a preset similarity threshold.
[0059] In this embodiment, a self-updating mechanism is also added to the preset rule base. This can be combined with... Figure 2 As shown, in cases where no match is found, the first text features of the unmatched text are acquired within a preset time period. These text features represent newly labeled data that has not been encountered before. The purpose of setting a preset time period is to collect newly labeled data over a certain period, thereby increasing the sample size. Then, candidate rules are generated based on the newly labeled data using a weakly supervised learning method. Specifically, this involves performing cluster analysis on the newly labeled data, grouping similar texts into one category, then statistically analyzing the high-frequency words in each category to form new candidate keywords, and generating candidate rules. Finally, the similarity between the candidate rules and existing preset rules is calculated using the Jaccard similarity coefficient:
[0060] Sim(R1,R2)=|K1∩K2| / |K1∪K2| (Formula 1)
[0061] Where R1 represents the candidate rule, K1 represents the set of keywords corresponding to the candidate rule; R2 represents the existing preset rule, and K2 represents the set of keywords corresponding to the existing preset rule.
[0062] If the similarity is lower than the set threshold (e.g., 0.6), it is considered a new rule that is different from the existing rules. The candidate rule is added to the rule base, and the rule priority is reordered. The reordering can be based on the initial priority setting method mentioned above, which will not be repeated here.
[0063] For example, if a new batch of dialogue data is collected and it is found that users frequently use phrases like "Hello, can you help me check this..." to express their query intent, and there are no query-related rules in the pre-matching rules, then cluster analysis and high-frequency word statistics are used to select these as new candidate rules. After calculating their similarity with existing rules, they are added to the rule base and their priority is adjusted.
[0064] It should be noted that in relevant intent recognition systems that rely on rule matching, user input often fails to match existing rules, resulting in situations where "customer service assistants or medical Q&A assistants provide irrelevant answers," indicating low intelligence and a poor user experience. By dynamically updating the aforementioned pre-set rule base, recognition accuracy can be effectively improved, thereby enhancing the user experience.
[0065] Optionally, if preliminary intent recognition is not obtained, deep semantic analysis is performed on the first text features to obtain intent recognition results. This includes: mapping the first text features into text semantic vectors through a multi-layer self-attention mechanism and a feedforward neural network; obtaining multimodal information corresponding to the first text features and mapping the multimodal information to the same dimension as the text semantic vector to obtain a multimodal vector; the multimodal information includes at least one of speech, intonation, pitch, speed of sound, and emoticons; concatenating the text semantic vector and the multimodal vector to obtain an enhanced semantic vector; constructing a context vector of the enhanced semantic vector based on historical dialogue records through a weighted summation method; and fusing the context vector and the enhanced semantic vector as input to the deep semantic analysis model to obtain the intent recognition results.
[0066] In this embodiment, it can be combined with Figure 3As shown, the first text feature is first mapped to a text semantic vector through a multi-layer self-attention mechanism and a feedforward neural network. Specifically, in this embodiment, the deep semantic analysis model adopts the ERNIE model. Texts whose first text features that did not match the rules are input into a text classifier built based on the ERNIE model. The ERNIE model is based on the Transformer architecture and learns the deep semantic representation of the language through pre-training. During the text encoding process, the text is segmented into sub-word units, special tags (such as [CLS], [SEP]) are added, and then each sub-word is mapped to a high-dimensional vector by looking up the word embedding table. The model captures long-distance dependencies and semantic information in the text through a multi-layer self-attention mechanism and a feedforward neural network, and finally outputs the vector corresponding to the [CLS] tag as the semantic vector representation of the text.
[0067] The process then obtains the multimodal information corresponding to the first text feature and maps this information to the same dimension as the text semantic vector to obtain a multimodal vector. The multimodal information includes at least one of speech, intonation, pitch, speed of speech, and emoticons. The text semantic vector and the multimodal vector are then concatenated to obtain an enhanced semantic vector. In one feasible implementation, for speech intonation, features such as pitch and speed of speech are analyzed. The pitch is normalized to the range [0, 1], and the speed of speech can be expressed as the number of words spoken per second. Then, a linear transformation is used to map this into a vector with the same dimension as the text semantic vector. For emoticons, a predefined emoticon semantic vector table is searched. This table assigns a semantic vector to each emoticon through cluster analysis and semantic annotation. The text semantic vector and the multimodal vector are then concatenated to form the enhanced semantic vector.
[0068] For example, speech intonation analysis yields a pitch vector of [0.5, 0.6, ..., 0.8] (128 dimensions) and a speech rate vector of [2.5, 3.0, ..., 2.0] (128 dimensions). After mapping them to the same 768 dimensions as the text semantic vector, they are concatenated with the text semantic vector to form an enhanced semantic vector.
[0069] Finally, based on historical dialogue records, a context vector for the enhanced semantic vector is constructed using a weighted summation method. This context vector is then fused with the enhanced semantic vector and used as input to the deep semantic analysis model to obtain the intent recognition result. Specifically, a context vector is constructed using multi-turn dialogue history. A weighted summation method is used to fuse the enhanced semantic vectors from the previous n turns of dialogue with the current enhanced semantic vector. The weights are set according to the proximity of the current turn to the previous turn, with closer turns having higher weights, as specified in Formula 2:
[0070] V context =α1V pre v1+α2Vpre v2+…+α n V pre v n (Formula 2)
[0071] Here, i represents the reverse sequence number of the dialogue round. For example, if the current dialogue round is the 4th round, and the augmented semantic vectors from the previous 3 rounds are V... pre v1、V pre v2、V pre v3. The corresponding weights are α1 = 3 / 6 = 0.5, α2 = 2 / 6 ≈ 0.333, and α3 = 1 / 6 ≈ 0.167. The context vector and the enhanced semantic vector are then concatenated as the input vector for the final ERNIE model classifier.
[0072] The final input vector is then fed into the ERNIE model classifier, which uses softmax regression and is calculated according to Formula 3:
[0073]
[0074] Where z k This represents the score of the input vector in category k, where K represents the total number of categories. Based on the calculated probability distribution of different intentions, if the maximum probability value is greater than a set threshold (e.g., 0.7), the intention with the highest probability is taken as the intention recognition result; if it is less than the set threshold, it is temporarily defined as an unknown intention, which can be comprehensively judged in conjunction with the subsequent entity extraction results.
[0075] For example, the text "I want to book a flight from Beijing to Shanghai next Wednesday" has a probability of 0.85 for the intent "book a flight," which is greater than the threshold of 0.7. Therefore, the recognition result is "book a flight." By fusing multimodal information and contextual information for deep semantic analysis, intent recognition is more accurate compared to semantic analysis based solely on text features. Furthermore, the multimodal information also considers facial expressions, symbols, and the possibility that the input may be an emoji containing text, thereby improving the accuracy of user intent recognition.
[0076] Optionally, entity recognition is performed on the first text features, and the entity recognition result is fused with the intent recognition result to obtain a semantic representation, including: forward encoding the first text features to obtain the forward hidden state of the first text features, and backward encoding the first text features to obtain the backward hidden state of the first text features; concatenating the forward hidden state and the backward hidden state to obtain the hidden state sequence of the first text features; performing entity recognition based on the hidden state sequence of the first text features to obtain the entity recognition result; and fusing the entity recognition result with the intent recognition result to obtain a semantic representation.
[0077] In this embodiment, the entity extraction process can be combined with Figure 4 As shown, this embodiment uses the MAOE model for entity extraction. First, the entity extraction engine based on the MAOE model is started. This model is based on a BiLSTM-CRF architecture. The BiLSTM layer encodes the first text features using forward LSTM and backward LSTM respectively, obtaining the forward and backward hidden states for each word. These are then concatenated to form the final hidden state. The CRF layer calculates the probability of a conditional random field based on the hidden state sequence and the encoded annotation sequence. A dynamic programming algorithm is then used to decode the sequence and determine the annotation sequence with the highest probability, thus achieving entity extraction. For example, for the text "flight tickets from Beijing to Shanghai next Wednesday", the model extracts the entities "time (next Wednesday)", "departure point (Beijing)", and "destination (Shanghai)". Finally, the entity recognition results are fused with the intent recognition results to obtain the semantic representation.
[0078] Furthermore, the entity recognition results include entities and the relationships between entities; after obtaining the entity recognition results, the process also includes: performing disambiguation processing on the entity recognition results by combining context vectors; and constructing an entity relationship graph by using the entities in the entity recognition results as nodes and the relationships between entities in the entity recognition results as edges.
[0079] In this embodiment, during entity extraction, disambiguation is performed on entities by incorporating contextual information. The cosine similarity between the context vector and candidate entity vectors is calculated, and the most suitable entity type is selected based on this similarity. For example, for the entity "apple," its similarity to the context vector is calculated. If the similarity is greater than a set threshold (e.g., 0.7), the context determines whether it refers to a fruit or an electronics company. Suppose the dialogue mentions "I want to buy an Apple phone." The cosine similarity between the context vector and the semantic vector of "apple" as an electronics company is 0.8, which is greater than the threshold of 0.7. Therefore, the entity "apple" is determined to refer to an electronics company.
[0080] Then, an entity relationship graph is constructed in real time, with extracted entities as nodes and relationships between entities as edges. Formula 4 is used to calculate the weights of relationships between entities:
[0081] W i j = tanh(W) a [h i h j ]+b a (Formula 4)
[0082] Where h i and h jLet represent the vector representations of entities i and j, respectively. Wa and ba are parameters in the attention mechanism. Entity relationships are filtered based on weights, with higher-weighted relationships retained and added to the graph. For example, for the entities "time (next Wednesday)," "departure point (Beijing)," and "destination (Shanghai)," the weights of the relationships between them are calculated. Time has higher weights with departure point and destination, thus establishing "time-departure point" and "time-destination" relationship edges. Through context vector construction and entity relationship graph construction, the semantic associations of the context and the relationships between entities can be understood more accurately. For example, in multi-turn booking dialogues, different entities mentioned by the user (such as location, time, number of people, etc.) can be effectively associated and understood, avoiding intent recognition bias caused by misunderstanding of the context, and improving dialogue coherence and task completion rate.
[0083] Optionally, after fusing the entity recognition result and the intent recognition result to obtain a semantic representation, the method further includes: using the entity relationship graph as input to the graph neural network to calculate the semantic correlation between the intent recognition result and the entity recognition result; the semantic correlation is used to verify the semantic consistency between the intent recognition result and the entity recognition result; if the semantic correlation is greater than a preset correlation threshold, it is determined that the semantic consistency verification is satisfied; if the semantic correlation is less than the preset correlation threshold, the entity recognition result and the intent recognition result are re-acquired.
[0084] In this embodiment, the intent recognition result output by the ERNIE model is first fused with the entity recognition result extracted by the MAOE model to generate a structured semantic representation. For example, the intent "book a flight" is combined with the entities "time (next Wednesday)," "departure point (Beijing)," and "destination (Shanghai)" in JSON format: {"intent": "book a flight", "entity": {"time": "next Wednesday", "departure point": "Beijing", "destination": "Shanghai"}}.
[0085] Then, a graph neural network is used to verify the semantic consistency between intent and entities. The entity relationship graph is input into the graph neural network, and the vector representation of the nodes is updated through graph convolution operations, specifically using Formula 5:
[0086]
[0087] in, Let N(i) be the vector representation of node i at layer l, and let Wg and Wn be the parameters of the graph neural network. The semantic correlation between the intent node and each entity node is calculated. If the correlation is lower than a preset correlation threshold (e.g., 0.6), the intent or entity is corrected.
[0088] For example, after calculation by the graph neural network, the semantic relevance of the intent "book a flight" to the entities "time (next Wednesday)", "departure point (Beijing)" and "destination (Shanghai)" are 0.8, 0.75 and 0.78 respectively, all of which are greater than the threshold of 0.6, indicating that the intent and the entities are reasonably matched; if the relevance of a certain entity is lower than the threshold, the entity extraction or intent recognition results need to be re-examined.
[0089] Furthermore, for cases involving multiple intents, each intent and its corresponding entity are parsed simultaneously, and the multiple intents are sorted according to a preset intent priority strategy. Intent priority P intent The calculation combines user historical preferences and scene popularity, specifically using Formula Six:
[0090] P intent =βP history +(1-β)P scene (Formula Six)
[0091] Where β is the balance coefficient, P history P represents the probability of a user's historical preferences. scene This represents the probability of scenario popularity. For example, if a user's history shows they frequently book flights, then the probability of a "book a flight" intent is high. history The value is 0.8; the current scenario is the peak travel season, and the P value for the intention to "book a flight" is 0.8. scene If β is 0.7, then P = 0.6. intent = 0.6 multiplied by 0.8 + 0.4 multiplied by 0.7 = 0.76. According to P intent The size of the data is used to sort multiple intents, with intents with higher probabilities being processed first.
[0092] Finally, the response content is generated based on the semantic representation. In one feasible implementation, the fused and validated semantic representation is converted into standard JSON-formatted structured data, containing intent tags, various entities and their values, and confidence levels. Then, based on the structured intent result, the corresponding intelligent agent dialogue assistant action is triggered. For example, for the intent to "book a flight," the flight booking interface is called, passing entities such as time, departure point, and destination as parameters. After obtaining the booking result, it is returned to the user. The specific process of the interface call is as follows: first, request parameters are constructed, including entity information and user identifiers; then, an HTTP POST request is sent to the flight booking system, with the request body containing the parameter information; finally, the returned result is received and converted into a user-friendly text reply. For example, after sending the request, a successful booking response is received, returning to the user "Your flight from Beijing to Shanghai next Wednesday has been booked. Please check the relevant information."
[0093] Compared to traditional single-rule matching or machine learning models for intent recognition, this embodiment significantly improves the accuracy of complex intent recognition through multimodal fusion, deep semantic analysis, and context association construction. The introduction of dynamic rule updates and weakly supervised learning optimization mechanisms drastically reduces rule maintenance costs, decreasing manual annotation and rule adjustment workload by 60%-70%. Simultaneously, the adaptation cycle to emerging intents is significantly shortened, enabling automatic identification and adaptation to new user expression patterns in a shorter time, offering greater flexibility and timeliness compared to traditional fixed-rule systems. Furthermore, the construction of context vectors and entity relationship graphs allows for a more accurate understanding of contextual semantic associations and relationships between entities. For example, in multi-turn booking dialogues, different entities mentioned by the user (such as location, time, number of people, etc.) can be effectively associated and understood, avoiding intent recognition deviations caused by contextual misunderstandings and improving dialogue coherence and task completion rates. In addition, the deep utilization of entity information and the intent-entity fusion verification mechanism in this solution address the problem of relatively independent intent recognition and entity extraction in traditional methods.
[0094] Furthermore, as Figures 1 to 4 The specific implementation of the method shown in this embodiment provides a response device based on intent recognition, such as... Figure 5 As shown, the device includes: an acquisition unit 501, a matching unit 502, an analysis unit 503, a fusion unit 504, and a generation unit 505.
[0095] The acquisition unit 501 is configured to acquire multimodal input content from the user and preprocess the input content to obtain a first text feature;
[0096] The matching unit 502 is configured to perform preliminary intent recognition on the first text feature by traversing each preset rule according to the priority of preset rules in the preset rule library; wherein, the preset rule library includes at least two preset rules, each preset rule includes a second text feature under the current preset rule and an intent tag corresponding to the second text feature, the intent tag being used to determine the intent of the first text feature;
[0097] The analysis unit 503 is configured to perform deep semantic analysis on the first text features to obtain an intent recognition result if the preliminary intent recognition has not been obtained; wherein the intent recognition result includes a comprehensive fusion result of text semantic vectors, multimodal vectors and context vectors based on the first text features;
[0098] The fusion unit 504 is configured to perform entity recognition on the first text features and fuse the entity recognition result with the intent recognition result to obtain a semantic representation.
[0099] The generation unit 505 is configured to generate response content based on the semantic representation.
[0100] In a specific application scenario, the matching unit 502 is further configured to determine the priority order of preset rules in the preset rule base; according to the priority order from high to low, the second text feature corresponding to the preset rule is matched with the first text feature in turn until a preliminary intent recognition result is generated or the second text feature corresponding to the preset rule with the lowest priority is matched with the first text feature; wherein, when the second text feature corresponding to the preset rule is successfully matched with the first text feature, the intent label corresponding to the second text feature is output as the preliminary intent recognition result.
[0101] In a specific application scenario, the matching unit 502 is further configured to: acquire a first text feature that has undergone the preliminary intent recognition within a preset time period but has not obtained the preliminary intent recognition result; generate candidate rules for the first text feature using a weakly supervised learning method; calculate the similarity between the candidate rule and at least one preset rule in the preset rule library; and add the candidate rule to the preset rule library if the similarity is less than a preset similarity threshold.
[0102] In a specific application scenario, the analysis unit 503 is further configured to map the first text features into a text semantic vector through a multi-layer self-attention mechanism and a feedforward neural network; obtain multimodal information corresponding to the first text features, and map the multimodal information to the same dimension as the text semantic vector to obtain a multimodal vector; the multimodal information includes at least one of speech, intonation, pitch, speed of sound, and emoticons; concatenate the text semantic vector with the multimodal vector to obtain an enhanced semantic vector; construct a context vector of the enhanced semantic vector based on historical dialogue records through a weighted summation method; fuse the context vector with the enhanced semantic vector and use it as input to a deep semantic analysis model to obtain the intent recognition result.
[0103] In a specific application scenario, the fusion unit 504 is further configured to perform forward encoding on the first text feature to obtain the forward hidden state of the first text feature, and perform backward encoding on the first text feature to obtain the backward hidden state of the first text feature; concatenate the forward hidden state and the backward hidden state to obtain the hidden state sequence of the first text feature; perform entity recognition based on the hidden state sequence of the first text feature to obtain the entity recognition result; and fuse the entity recognition result with the intent recognition result to obtain the semantic representation.
[0104] In specific application scenarios, the fusion unit 504 is further configured to perform disambiguation processing on the entity recognition result by combining the context vector; and to construct an entity relationship graph by using the entities in the entity recognition result as nodes and the relationships between the entities in the entity recognition result as edges.
[0105] In a specific application scenario, the fusion unit 504 is further configured to use the entity relationship graph as input to a graph neural network to calculate the semantic correlation between the intent recognition result and the entity recognition result; the semantic correlation is used to verify the semantic consistency between the intent recognition result and the entity recognition result; if the semantic correlation is greater than a preset correlation threshold, it is determined that the semantic consistency verification is satisfied; if the semantic correlation is less than the preset correlation threshold, the entity recognition result and the intent recognition result are re-acquired.
[0106] It should be noted that other corresponding descriptions of the functional units involved in the intent recognition-based response device provided in this embodiment can be found in [reference]. Figures 1 to 4 The corresponding descriptions in [the document] will not be repeated here.
[0107] Based on the above, Figures 1 to 4 Accordingly, this embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. Figures 1 to 4 The method shown.
[0108] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0109] Based on the above, Figures 1 to 4 The method shown, and Figure 5 To achieve the above objectives, this application also provides an electronic device, which can be configured on a computer side, etc. This device includes a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to achieve the above-described virtual device embodiments. Figures 1 to 4 The method shown.
[0110] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0111] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0112] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0113] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms, or it can be implemented by hardware. By applying the solution of this embodiment, compared with related technologies, the initial intent recognition is performed on the preprocessed text one by one according to the priority of the preset rules in the preset rule base, thereby improving the success rate of the initial recognition; if the initial intent recognition result is not obtained, multimodal information and contextual information are then fused to perform deep semantic analysis, which has higher intent recognition accuracy than the method of analyzing only the text; then the intent recognition result is fused with the entity recognition result to obtain a semantic representation, and response content for the user is generated based on the semantic representation.
[0114] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0115] Similar parts between the embodiments provided in this application can be referred to mutually. The specific implementation methods provided above are only a few examples under the overall concept of this application and do not constitute a limitation on the scope of protection of this application. For those skilled in the art, any other implementation methods extended from the solution of this application without creative effort shall fall within the scope of protection of this application.
Claims
1. A response method based on intent recognition, characterized in that, include: The user's multimodal input content is acquired, and the input content is preprocessed to obtain the first text feature; Based on the priority of preset rules in the preset rule base, each preset rule is traversed to perform preliminary intent recognition on the first text feature; wherein, the preset rule base includes at least two preset rules, each preset rule includes a second text feature under the current preset rule and an intent tag corresponding to the second text feature, the intent tag being used to determine the intent of the first text feature; If the preliminary intent recognition fails to yield a result, deep semantic analysis is performed on the first text features to obtain an intent recognition result; wherein, the intent recognition result includes a comprehensive fusion result of text semantic vectors, multimodal vectors, and context vectors based on the first text features; Entity recognition is performed on the first text features, and the entity recognition result is fused with the intent recognition result to obtain a semantic representation; Response content is generated based on the semantic representation.
2. The method according to claim 1, characterized in that, The step of performing preliminary intent recognition on the first text features by traversing each preset rule according to the priority of preset rules in the preset rule base includes: Determine the priority order of the preset rules in the preset rule base; According to the priority order from high to low, the second text features corresponding to the preset rules are matched with the first text features in turn until a preliminary intent recognition result is generated or the second text feature corresponding to the preset rule with the lowest priority is matched with the first text feature. Specifically, if the second text feature corresponding to the preset rule successfully matches the first text feature, the intent label corresponding to the second text feature is output as the preliminary intent recognition result.
3. The method according to claim 1, characterized in that, After performing preliminary intent recognition on the first text features by traversing each preset rule according to the priority of preset rules in the preset rule base, the method further includes: Obtain the first text feature within a preset time period that has undergone the preliminary intent recognition but has not yielded the preliminary intent recognition result; Candidate rules are generated from the first text features using a weakly supervised learning method; Calculate the similarity between the candidate rule and at least one preset rule in the preset rule base; If the similarity is less than a preset similarity threshold, the candidate rule is added to the preset rule base.
4. The method according to claim 1, characterized in that, In the case where the preliminary intent recognition has not been obtained, the step of performing deep semantic analysis on the first text features to obtain the intent recognition result includes: The first text features are mapped into text semantic vectors through a multi-layer self-attention mechanism and a feedforward neural network. Obtain the multimodal information corresponding to the first text feature, and map the multimodal information to the same dimension as the text semantic vector to obtain a multimodal vector; the multimodal information includes at least one of speech, intonation, pitch, speed of sound, and emoticons; The text semantic vector and the multimodal vector are concatenated to obtain the enhanced semantic vector; Based on historical dialogue records, a context vector for the enhanced semantic vector is constructed using a weighted summation method; The context vector and the enhanced semantic vector are fused and used as input to the deep semantic analysis model to obtain the intent recognition result.
5. The method according to claim 4, characterized in that, The step of performing entity recognition on the first text features and fusing the entity recognition result with the intent recognition result to obtain a semantic representation includes: The first text feature is forward encoded to obtain the forward hidden state of the first text feature, and the first text feature is backward encoded to obtain the backward hidden state of the first text feature. The forward hidden state and the backward hidden state are concatenated to obtain the hidden state sequence of the first text feature; Entity recognition is performed based on the hidden state sequence of the first text features to obtain the entity recognition result; The entity recognition result is fused with the intent recognition result to obtain the semantic representation.
6. The method according to claim 5, characterized in that, The entity recognition results include entities and the relationships between entities; After obtaining the entity recognition result, the method further includes: The entity recognition results are then combined with context vectors for disambiguation processing. The entities in the entity recognition results are used as nodes, and the relationships between the entities in the entity recognition results are used as edges to construct an entity relationship graph.
7. The method according to claim 6, characterized in that, After fusing the entity recognition result with the intent recognition result to obtain the semantic representation, the method further includes: The entity relationship graph is used as input to a graph neural network to calculate the semantic correlation between the intent recognition result and the entity recognition result; the semantic correlation is used to verify the semantic consistency between the intent recognition result and the entity recognition result. If the semantic relevance is greater than a preset relevance threshold, the semantic consistency check is deemed to be satisfied. If the semantic relevance is less than the preset relevance threshold, the entity recognition result and the intent recognition result are re-acquired.
8. A response device based on intent recognition, characterized in that, include: The acquisition unit is configured to acquire multimodal input content from the user and preprocess the input content to obtain a first text feature; The matching unit is configured to perform preliminary intent recognition on the first text feature by traversing each preset rule according to the priority of preset rules in the preset rule library; wherein, the preset rule library includes at least two preset rules, each preset rule includes a second text feature under the current preset rule and an intent label corresponding to the second text feature, the intent label being used to determine the intent of the first text feature; The analysis unit is configured to perform deep semantic analysis on the first text features to obtain an intent recognition result if the preliminary intent recognition has not been obtained; wherein the intent recognition result includes a comprehensive fusion result of text semantic vectors, multimodal vectors and context vectors based on the first text features; The fusion unit is configured to perform entity recognition on the first text features and fuse the entity recognition result with the intent recognition result to obtain a semantic representation; The generation unit is configured to generate response content based on the semantic representation.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the method of any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.