Intention recognition method and related equipment thereof

By determining and processing the idioms of polysynonyms in intention recognition, the problem of intent recognition accuracy caused by polysynonyms is solved, and the accuracy of intent recognition is improved.

CN119988577APending Publication Date: 2025-05-13MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411941909.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

During the intent recognition process, the same word may have multiple meanings, which will affect the accuracy of intent recognition.

Method used

By determining that there are a first word with multiple meanings in the sentence to be identified by the target user, and intent recognition is performed when the word belongs to an idiomatic word. This method can not only determine the first intention of the statement to be identified, but also determine the probability of the word meaning corresponding to the first word in the first intention and the word meaning corresponding to the word meaning, and determine the second intention of the statement to be identified based on this information.

Benefits of technology

By considering the first word of confusing meanings that may exist in the sentence to be identified and its word meaning in the identified first intention, the confusing effect of polysynthetics on the intent recognition results can be avoided, thereby improving the accuracy of intent recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988577A_ABST
    Figure CN119988577A_ABST
Patent Text Reader

Abstract

The invention discloses an intention recognition method and related equipment. The method comprises the steps that it is determined that a first word with multiple meanings exists in a to-be-recognized statement of a target user; determining a first intention of the to-be-identified statement, a word meaning corresponding to the first word in the first intention and a probability corresponding to the word meaning; and determining a second intention of the to-be-identified statement based on the occurrence frequency of the first word in the to-be-identified statement and the probability corresponding to the word meaning of the first word in the first intention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to an intent recognition method and related equipment. Background Art

[0002] Intent recognition is an important branch of natural language processing. Its main task is to determine the intention or purpose of the user during the conversation, so as to provide targeted services to the user based on the identified intention or purpose.

[0003] When performing intent recognition on a sentence, we may encounter a situation where the same word has multiple meanings, that is, polysemous words. How to ensure the accuracy of intent recognition in such a situation is a problem that the existing technology urgently needs to solve. Summary of the invention

[0004] The embodiments of the present application provide an intention recognition method and related devices that overcome the above-mentioned problems or at least partially solve the above-mentioned problems.

[0005] The present application embodiment adopts the following technical solutions: Determining a first word having multiple meanings in a sentence to be recognized by a target user; In the case where the first word is an idiomatic word of the target user, performing intent recognition on the sentence to be recognized, determining a first intent of the sentence to be recognized, and a word meaning corresponding to the first word in the first intent and a probability of the word meaning corresponding to each other; Based on the number of occurrences of the first word in the sentence to be recognized and the probability of the word meaning of the first word in the first intention corresponding to each other, a second intention of the sentence to be recognized is determined.

[0006] In a second aspect, the present application provides an intention recognition device, comprising: A determination module, used to determine a first word having multiple meanings in a sentence to be recognized by a target user; The determination module is also used to perform intent recognition on the sentence to be recognized when the first word is an idiomatic word of the target user, determine the first intent of the sentence to be recognized, and the word meaning corresponding to the first word in the first intent and the probability of the word meaning corresponding; determine the second intent of the sentence to be recognized based on the number of occurrences of the first word in the sentence to be recognized and the probability of the word meaning corresponding to the first word in the first intent.

[0007] In a third aspect, the present application provides an electronic device, including: one or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors, and the one or more programs are configured to execute the steps in the intent recognition method as described in the first aspect.

[0008] In a fourth aspect, the present application provides a computer-readable storage medium, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps in the intent recognition method described in the first aspect above are implemented.

[0009] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps in the intent recognition method described in the first aspect above.

[0010] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: By adopting the intention recognition method provided in the embodiment of the present application, when performing intent recognition on the target user's sentence to be recognized, it is possible to first determine the first word with multiple meanings in the target user's sentence to be recognized, and when it is determined that the first word is an idiomatic word, perform intent recognition on the sentence to be recognized. The intention recognition process can not only determine the first intention of the sentence to be recognized, but also determine the word meaning corresponding to the first word in the first intention that affects the final intention recognition result of the sentence to be recognized and the probability of the word meaning corresponding. Since the word meaning of the first word with confusing meaning in the sentence to be recognized is taken into account in the recognized first intention, and based on the number of occurrences of the first word in the sentence to be recognized and the probability of the word meaning of the first word in the first intention corresponding, the first intention that matches the word meaning of the first word is determined from the first intention as the second intention of the sentence to be recognized. This can avoid the confusing effects of some first words that have multiple meanings on the intention recognition result of the sentence to be recognized during the intention recognition process, thereby improving the accuracy of intent recognition of the sentence to be recognized. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A schematic diagram of an application scenario of the intent recognition method provided in an embodiment of the present application; Figure 2 A schematic diagram of an application environment of the intent recognition method provided in an embodiment of the present application; Figure 3 A schematic diagram of the implementation flow of an intent recognition method is provided for an embodiment of the present application; Figure 4 A schematic diagram of a source conversation being split into multiple words in the intent recognition method provided in an embodiment of the present application; Figure 5 A schematic diagram of the structure of an intention recognition device provided for an exemplary embodiment of the present application; Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0012] In order to make those skilled in the art better understand the present application scheme, the technical scheme in the present application embodiment will be clearly and completely described below in conjunction with the drawings in the present application embodiment. Obviously, the described embodiment is only a part of the present application embodiment, rather than all the embodiments. The components of the present application embodiment usually described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application for protection, but merely represents the selected embodiment of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.

[0013] The following is an explanation of the nouns involved in this application: Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0014] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0015] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.

[0016] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.

[0017] With the rapid development of Internet technology, online intelligent customer service, telephone intelligent customer service, etc., need to use pre-trained AI models to identify the intent of the voice or text sentences output by the user when interacting with the user through text or voice. Since the same word may have multiple meanings, when identifying the intent of the voice or text sentences output by the user, if the meaning of a word with multiple meanings is selected incorrectly, it will affect the intent recognition result of the entire text or voice sentence. Therefore, how to improve the intent recognition accuracy of the AI ​​model for intent recognition of text or voice sentences still requires further solutions.

[0018] To solve the above problems, the inventor of the present application found after careful research that the current natural language processing (NLP) research and model recognition exploration are stagnant at the meaning of the voice or text sentence itself, without considering the individual characteristics of the speaker who issued the voice or text sentence. However, from the actual situation of the conversation sentence, the speaker, as the speaker of the voice or text sentence, will have some personalized speaking habits, such as liking to use some idioms such as catchphrases to express some tone or emotion. For example, the word "not", as an idiom, may not have a practical meaning, but is just a way to start a conversation. If the idiom factor is not considered, the word "not" will be directly used as a negative word. It can be seen from this that by analyzing the personalized speaking habits of different speakers, such as the use of idioms, the true meaning of each word in the voice or text sentence issued by the speaker can be more accurately identified, and then the true intention of the voice or text sentence issued by the speaker can be more accurately identified.

[0019] Based on this, the intention recognition method provided in the embodiment of the present application can first determine the first word with multiple meanings in the target user's sentence to be recognized when performing intention recognition on the target user's sentence to be recognized, and when it is determined that the first word is an idiomatic word, the sentence to be recognized can be recognized. This intention recognition can not only recognize multiple first intentions of the sentence to be recognized, but also determine the word meanings and the probability of word meanings corresponding to the first words in each first intention that affect the final intention recognition result. Because when determining the final intention of the sentence to be recognized, the word meanings of the first words that may have confusing meanings in the sentence to be recognized in the recognized first intention are taken into account, and based on the number of occurrences of the first word in the sentence to be recognized and the probability of the word meanings of the first word in the first intention corresponding, the word meaning of the first word that matches the first intention is determined from the first intention as the second intention of the sentence to be recognized. This can avoid the confusing effects of some first words that have multiple meanings on the intention recognition results of the sentence to be recognized during the intention recognition process, thereby improving the accuracy of intention recognition of the sentence to be recognized.

[0020] The intent recognition method provided in the embodiments of the present application can be applied to any scenario where intent recognition is required, and the method can be applied to products in these scenarios, such as intelligent customer service systems, intelligent marketing systems, intelligent recommendation systems, intelligent navigation systems, intelligent search systems, etc.

[0021] Taking the intelligent search system as an example, the intelligent search system is used to classify the intent of the search text and determine the search intent corresponding to the search text. Among them, intent refers to the purpose you want to achieve, and search intent refers to the content you want to search for. The intelligent search system is used to analyze the search target of the text. The intelligent search system can be applied to a variety of scenarios. For example, in a search scenario, the text can be the text entered by the user. After the user enters the text, the intelligent search system analyzes the corresponding search intent. For another example, in a voice interaction scenario, the text can be the text obtained by voice recognition of the collected voice signal. The user speaks, the device collects the voice signal, recognizes it to obtain text, and determines the voice search command issued by the user through the intelligent search system.

[0022] For example, Figure 1 As shown, a specific voice interaction application scenario is provided. A user 101 sends a voice signal 102, and the device is able to collect the voice signal 102. The voice recognition system 103 performs voice recognition on the voice signal 102 to obtain a text 104 corresponding to the voice signal 102. The intelligent search system 105 determines the search intent corresponding to the text 104 and displays the corresponding search content, thereby determining the target 106 that the voice signal 102 wants to search for.

[0023] In order to better understand the intention recognition method and related equipment provided in the embodiment of the present application, the application environment applicable to the embodiment of the present application is described below.

[0024] See also Figure 2 , Figure 2 The following is a schematic diagram of an application environment of the intent recognition method provided by an embodiment of the present application. For example, the intent recognition method, the training method of the intent recognition model, the intent recognition device, the training device of the intent recognition model, the electronic device, the storage medium and the computer program product provided by the embodiment of the present application can be applied to an electronic device, wherein the electronic device can be such as Figure 2 The server 210 shown in the figure can be linked to the terminal device 220 through a network. The network is used to provide a medium for a communication link between the server 210 and the terminal device 220. The network can include various connection types, such as wired communication links, wireless communication links, etc., which are not limited in the embodiments of the present application. Optionally, in other embodiments, the electronic device can also be a smart phone, a laptop computer, etc.

[0025] It should be understood that Figure 2The server 210, network and terminal device 220 are merely illustrative. Depending on the implementation requirements, there may be any number of servers, networks and terminal devices. For example, the server 210 may be a physical server or a server cluster composed of multiple servers, and the terminal device 220 may be a mobile phone, a tablet, a desktop computer, a laptop computer, and the like. It is understood that the embodiments of the present application may also allow multiple terminal devices 220 to access the server 210 at the same time.

[0026] In some embodiments, the terminal device 220 may obtain a sentence to be recognized input by the user, and then input the sentence to be recognized to the server 210 for intent recognition. In other embodiments, the terminal device 220 may record the user's voice to obtain the user's audio data, and then the audio data may be recognized in the terminal device 220 or the server 210 to obtain text information, and then the text information may be recognized in the server 210 for intent recognition.

[0027] The above application environment is only an example for the convenience of understanding. It can be understood that the embodiments of the present application are not limited to the above application environment.

[0028] The intent recognition method and related equipment provided in the embodiments of the present application will be described in detail below through specific examples.

[0029] See also Figure 3 , which shows a flow chart of the intention recognition method provided by an embodiment of the present application. Figure 3 The process shown in the figure is described in detail, and the intention recognition method may include the following steps: Step 310: Determine a first word having multiple meanings in the target user's sentence to be recognized.

[0030] In some embodiments, the target user may be a user of a terminal device. There is a need for a conversation between the target user and the terminal device, and the intention of an audio or text sentence uttered by the target user may be identified by a server corresponding to the terminal device.

[0031] Due to the different conversation habits of different users, more specifically, the idiom usage habits and commonly used idioms of different users may also be different, and generally speaking, the idioms that users are accustomed to using may be different from the conventional meanings of the words themselves, which results in low accuracy of intent recognition for user conversations according to conventional natural language processing. Based on this, the embodiment of the present application can obtain the user's idioms for different users, and the idioms are the user's commonly used catchphrases, providing personalized recognition basis for the intent recognition of audio or text sentences issued by different users. As an example, the idiom can be displayed through an idiom list. The first word can be any one of a plurality of polysemous words in the target user's sentence to be recognized.

[0032] Optionally, multiple words appearing in the conversation data of the target user in the first time period may be obtained; and the frequency of occurrence of each word in the multiple words is counted to obtain high-frequency words whose frequency of occurrence is greater than or equal to a preset threshold; finally, based on the position flexibility of the high-frequency words in the conversation data in the first time period, the target user's idiomatic word list is determined. In an embodiment of the present application, the historical conversation data of the target user may be divided into multiple words by using a text segmentation method, such as using a dictionary segmentation algorithm such as string matching, or a machine learning algorithm based on statistics, etc., which is not limited in the embodiment of the present application.

[0033] In some embodiments, for the same user, the frequency of use of idiomatic words is usually higher than that of other words. Statistics can be made based on the frequency of occurrence of each word in multiple words, and the frequency of occurrence of multiple words can be sorted in reverse order (in order of frequency of occurrence from high to low), and the first k words are selected as high-frequency words, where k is a positive integer, and its specific value can be set according to actual needs. Optionally, a preset threshold for screening high-frequency words can also be set. When the frequency of occurrence of a word is greater than or equal to the preset threshold, the word is determined to be a high-frequency threshold.

[0034] In some embodiments, the position flexibility of high-frequency words in the source conversation can be determined through the following three dimensions. Specifically, the first position flexibility of the high-frequency words can be determined based on the position of the high-frequency words in the conversation data. The second position flexibility of the high-frequency words can be determined based on the number of categories of syntactic dependencies related to the high-frequency words in the syntactic dependencies of the conversation data. The third position flexibility of the high-frequency words can also be determined based on the degree of semantic influence of the position change of the high-frequency words on the conversation data. Optionally, in some typical embodiments, the position flexibility of each high-frequency word can be obtained by combining the position flexibility of the above three dimensions. That is to say, the position flexibility of the high-frequency words can be determined based on the first position flexibility, the second position flexibility and the third position flexibility of the high-frequency words.

[0035] In some embodiments, optionally, based on the above embodiments, the present embodiment can obtain the list of habitual words of the target user in the following manner. Specifically, the method provided by the embodiments of the present application can determine the habitual words of the target user through the following steps: Step 410: Obtain multiple words that appear in the conversation data of the target user in the first time period.

[0036] In some embodiments, the conversation data of the target user in the first time period (such as the past year, half year, or month) can be obtained. In the embodiments of the present application, the text segmentation method can be used to divide each conversation sentence in the conversation data into multiple words. For example, dictionary segmentation algorithms such as string matching, or machine learning algorithms based on statistics, etc. The embodiments of the present application do not make specific limitations on this.

[0037] In some embodiments, to reduce the interference of some meaningless words, after the text segmentation process of each conversation sentence in the conversation data of the target user in the first time period, further filtering operations can be performed on the multiple words after the segmentation process to remove the function words. Function words generally refer to words without complete meaning, but with grammatical meaning or function. They have the characteristics of having to attach to content words or sentences, expressing grammatical meaning, not being able to form sentences alone, not being able to be used as grammatical components alone, and not being able to be overlapped. As an example, the commonly used stop word list in the NLP field can be used. The words in the stop word list include generally meaningless words such as "de" (的) and "ma" (吗), that is, function words.

[0038] Step 420: For each word, determine the occurrence frequency of the word.

[0039] In some embodiments, the occurrence frequency of each word among the multiple words can be counted to determine the occurrence frequency of each word among the multiple words.

[0040] Step 430: Obtain the high-frequency words whose occurrence frequency is greater than or equal to a preset threshold.

[0041] In some embodiments, the occurrence frequencies of the multiple words can be sorted in reverse order (in the order from high to low occurrence frequency), and the first k words can be selected as the high-frequency words, where k is a positive integer, and its specific value can be set according to actual needs. In some embodiments, a preset threshold for screening high-frequency words can also be set. When the occurrence frequency of a word is greater than or equal to the preset threshold, then the word is determined as the high-frequency threshold. Among them, the occurrence frequency is the number and frequency of the word appearing in the conversation data in the first time period. The more the occurrence frequency, the more frequently the word appears in the conversation data in the first time period.

[0042] Step 440: Determine the target user's idiomatic words based on the position flexibility of the high-frequency words in the conversation data of the first time period.

[0043] In some embodiments, the location flexibility can be used to characterize the variability of the locations where high-frequency words appear in the conversation data of the first time period. The location flexibility of high-frequency words can be determined by: Determining the first position flexibility of the high-frequency words based on the occurrence positions of the high-frequency words in the conversation data; Determining the second position flexibility of the high-frequency words based on the number of categories of syntactic dependency relations related to the high-frequency words in the syntactic dependency relations of the conversation data; Determining the third position flexibility of high-frequency words based on the degree of semantic influence of position changes of high-frequency words on conversation data; Based on the first position flexibility, the second position flexibility and the third position flexibility of the high-frequency words, the position flexibility of the high-frequency words is determined.

[0044] In some embodiments, the position flexibility of each high-frequency word in the conversation data of the first time period may be determined based on the following three dimensions.

[0045] The first dimension: the first position flexibility of high-frequency words can be determined based on the occurrence positions of the high-frequency words in the conversation data.

[0046] As an example, the occurrence position of a high-frequency word in the conversation data may be determined by the following calculation formula: the occurrence position of the high-frequency word in the source conversation = (index of the high-frequency word + 1) / number of words in the source conversation. Figure 4 A schematic diagram showing a source conversation being split into multiple words. Figure 4 In the example, the index of a word is used to represent the position of the word in the source session. The index may start from 0. The word "today" with an index of 0 is the first word in the source session, the word "weather" with an index of 1 is the second word in the source session, and the word "great" with an index of 2 is the third word in the source session. In some embodiments, the number of occurrence positions of each high-frequency word in the source session may be counted, and the number of occurrence positions may be used as the first position flexibility of the high-frequency word in the source session.

[0047] The second dimension: The second position flexibility of high-frequency words can be determined based on the number of syntactic dependencies related to the high-frequency words in the syntactic dependencies of the conversational sentences.

[0048] Among them, syntactic dependency analysis, also known as dependency syntactic parsing, or dependency analysis for short, is used to identify the interdependencies between words in a sentence. Dependency syntax was first proposed by French linguist L. Tesniere. It analyzes a sentence into a dependency syntax tree, describing the dependency relationships between words. That is, it points out the syntactic collocation relationship between words, which is related to semantics. In natural language processing, the framework that uses the dependency relationship between words to describe the language structure is called dependency grammar, also known as dependency grammar. Using dependency syntax for syntactic analysis is one of the important technologies for natural language understanding.

[0049] The following is an example of syntactic dependency: The subject-predicate relationship is the most basic dependency relationship in a sentence, indicating the relationship between the subject and the predicate, for example: Xiao Ming eats. The verb-object relationship is a special case of the subject-predicate relationship, indicating the relationship between the verb and the object, for example: Xiao Ming eats. The attributive-predicate relationship is the relationship between the modifier and the modified, indicating a modification relationship, for example: red flowers. In the syntactic dependency relationships of conversation data, the number of categories of syntactic dependency relationships related to high-frequency words, that is, the number of categories of syntactic dependency relationships formed by high-frequency words with other words in conversation data, such as subject-predicate relationships, verb-object relationships, and attributive-predicate relationships, one of which can be used as a category of syntactic dependency relationships. If the syntactic dependency relationships formed by high-frequency words with other words in conversation data include subject-predicate relationships and verb-object relationships, then the number of categories of syntactic dependency relationships related to high-frequency words in the syntactic dependency relationships of conversation data is 2.

[0050] As an example, a conversation sentence containing high-frequency words can be input into a pre-trained dependency syntax model to output the number of categories of syntactic dependency relationships related to high-frequency words in the syntactic dependency relationships of the conversation sentence. The more categories of syntactic dependency relationships, the higher the second position flexibility of the corresponding high-frequency words. Among them, the dependency syntax model is trained based on conversation sentences with words with multiple different syntactic dependency relationships, and based on the annotation information of the syntactic dependency relationships of conversation sentences with words with multiple different syntactic dependency relationships. In some embodiments, the number of categories of syntactic dependency relationships related to high-frequency words in the syntactic dependency relationships of the conversation data can be determined, and this number can be used as the second position flexibility of the high-frequency words.

[0051] The third dimension: The third position flexibility of high-frequency words can also be determined based on the degree of influence of position changes of high-frequency words on the semantics of conversation data.

[0052] As an example, the high-frequency words in each conversation sentence of the conversation data may be deleted respectively, and it is determined whether the semantics of each conversation sentence changes before and after the deletion of each high-frequency word, and the i high-frequency words with the smallest semantic influence on the conversation sentence before and after the deletion of the high-frequency words are selected. In some embodiments, the value range of the semantic influence of the position change of the high-frequency words on the conversation sentence may be set to [0, 1]. The smaller the value, the smaller the semantic influence of the position change of the high-frequency words on the conversation sentence, that is, the higher the third position flexibility of the high-frequency words in the conversation sentence. As an example, the inverse of the semantic influence of the position change of the high-frequency words on the conversation sentence can be used as the third position flexibility of the high-frequency words in the conversation sentence.

[0053] Optionally, in some typical embodiments, the position flexibility of each high-frequency word can be obtained by combining the position flexibility of the above three dimensions. That is to say, the position flexibility of the high-frequency words in the source session can be determined based on the first position flexibility, the second position flexibility and the third position flexibility of the high-frequency words in the source session. As an example, the weights α, β and γ can be set for the first position flexibility, the second position flexibility and the third position flexibility, respectively, where α+β+γ=1, wherein α, β and γ can be set according to actual needs. Then the position flexibility of the high-frequency words in the source session = first position flexibility × α + second position flexibility × β + third position flexibility × γ.

[0054] In some embodiments, the sentence to be recognized may be text information obtained from a terminal device, or may be text information recognized from audio data collected by the terminal device, and this embodiment of the present application does not limit this. In other embodiments, the sentence to be recognized may also be text information obtained from other servers.

[0055] In the implementation manner of the present application, the sentence to be recognized can be first divided into multiple words by using a text segmentation method, and then the words with at least two meanings in the multiple words can be determined as polysemous words. Alternatively, a pre-trained polysemous word extraction model can be directly used to extract polysemous words with at least two meanings in the sentence to be recognized. Among them, one meaning corresponds to one meaning of the word, some words have one meaning, that is, they only correspond to one meaning, and some words have multiple meanings, that is, they correspond to multiple meanings (i.e., polysemous words).

[0056] If it is determined that a polysemous word in the sentence to be recognized is also a word in the idiomatic word list, since most of the words in the idiomatic word list are virtual words (i.e., words without actual meaning), therefore, when performing intent recognition, the word that is also a word in the idiomatic word list and a polysemous word is a component of the sentence to be recognized, it is likely to be interfered by words that are actually virtual but recognized as real words (i.e., the first word described in the text), thereby affecting the accuracy of intent recognition of the sentence to be recognized. Based on this, the embodiment of the present application can determine the polysemous word in the idiomatic word list as the first word when there is at least one polysemous word in the sentence to be recognized of the target user, and judge whether the first word has actual meaning in the sentence to be recognized through subsequent steps, so as to determine the real second intention corresponding to the sentence to be recognized, thereby effectively avoiding the sentence to be recognized from being interfered by some words without actual meaning during the intent recognition process.

[0057] Step 320, when the first word is a common word of the target user, perform intent recognition on the sentence to be recognized, determine the first intent of the sentence to be recognized, and the word meaning corresponding to the first word in the first intent and the probability of the word meaning corresponding.

[0058] The first word is a common word of the target user, and specifically, the first word is located in a common word list of the target user.

[0059] In some exemplary embodiments, the intent recognition of the sentence to be recognized and the determination of the first intent of the sentence to be recognized can be achieved by performing intent recognition on the sentence to be recognized using a pre-trained target intent recognition model with multi-label classification function, and the target intent recognition model can output the result of multi-dimensional parameters. The training process of the pre-trained target intent recognition model may include: Acquire a training sample, the training sample comprising a plurality of text sentences and annotation information based on the intention of each text sentence and the word role of the first word in each text sentence; Based on multiple text sentences, as well as the intent of each text sentence and the annotation information of the word role of the first word in each text sentence, the intent recognition model is trained to obtain a target intent recognition model.

[0060] As an example, a method for training a target intent recognition model may include the following steps: Step 510: Acquire a training sample, wherein the training sample includes a plurality of text sentences and annotation information based on the intention of each text sentence and the meaning of the first word in each text sentence.

[0061] In some embodiments, the multiple text statements may be text statements containing the habitual words of multiple users, and the text statement of each user corresponds to the list of habitual words of that user. The annotation information of the intention of each text statement and the word role of the first word in each text statement can be obtained based on manual annotation, that is, the possible intentions corresponding to each text statement and the word meaning of the first word in each text statement can be used as the labels of each conversation statement. For example, the intention of a conversation statement includes meaning Figure 1 ~meaning Figure 3 , each intention includes the first word 1, and the word meaning of the first word in each intention includes word meanings 1~2, then meaning Figure 1 ~meaning Figure 3 , and the first word 1 in each intention and the word meanings 1~2 of the first word in each intention are the annotation information of this text statement.

[0062] Step 520, based on the multiple text statements, and based on the annotation information of the intention of each text statement and the word role of the first word in each text statement, train the intention recognition model to obtain the target intention recognition model.

[0063] In some embodiments, based on the multiple text statements, and based on the annotation information of the intention of each text statement and the word role of the first word in each text statement, train the intention recognition model to obtain the target intention recognition model, which may specifically include the following process: S521, load the pre-trained intention recognition model, which may be, for example, a large language model based on Transformer, and this large language model can be trained on a super-large-scale pre-trained data.

[0064] S522, splice the text statement and its corresponding at least two intentions, and the first word and the word meaning of the first word in the text statement as annotation information and then input them into the intention recognition model, and set the task type of the intention recognition model as multi-dimensional sequence annotation. For example: The text statement "God, what's going on?" (where "God" has two meanings, the first meaning is used for the high altitude above the ground, and the second meaning is a function word used to express exclamation) and its two corresponding intentions and the first word are spliced, and the input sequence can be obtained: God, what's going on? + "God" + "what's going on?" + label: ['asking what's going on'] & ['incomplete intention'] + {God: [word meaning: habitual function word]}.

[0065] S523, perform word segmentation on the input sequence "天,这什么事啊 + '天' + '这什么事啊' + label: ['问是什么事'] & ['非完整意图'] + {天: [word meaning: empty habitual word]}", and perform encoding processing on the data after word segmentation to obtain word vector positional embedding, positional encoding positional embedding, and sentence segment encoding segment embedding. Among them, the positional encoding positional embedding is used to write the position information in the input sequence into the sentence vector. The sentence segment vector segment embedding is used to distinguish between two sentences spliced together in the input sequence.

[0066] S524, for the input sequence after word segmentation, the part-of-speech, morphology, and semantic reference of all words in the input sequence after word segmentation can be obtained by retrieving the pre-constructed part-of-speech mapping dictionary and morphology mapping dictionary, and grammatical encoding is performed on the part-of-speech, morphology, and semantic reference of all words in the input sequence after word segmentation to obtain word meaning vectors. Among them, encoding the part of speech can specifically determine the part-of-speech vector part_of_speech_embedding corresponding to the part of speech of each word in the input sequence after word segmentation by querying the pre-constructed mapping relationship table between different parts of speech and vector encodings, and determine the morphology vector morphology_embedding corresponding to the morphology of each word based on the mapping relationship table between different morphologies and vector encodings.

[0067] Among them, semantic reference is the possibility that a certain component of the syntactic structure matches other components (one or several) semantically. Semantic reference studies the semantic collocation relationship between sentence components, and syntactic structure studies the combination relationship of sentence components. It can be understood that semantic reference means which object the semantics of the word explains or discusses (example: 王冕死了父亲, the semantics of '死' refers to the father, not 王冕). Semantic reference can include: ① forward reference and backward reference, where forward reference means the reference object is before and the reference component is after, and backward reference means the reference component is before and the reference object is after; ② internal reference and external reference, where internal reference means the reference component and the reference object are in the same syntactic structure, and external reference means the reference object is outside the syntactic structure on which the reference component depends; ③ single reference and multiple references, where single reference means the reference component has only one reference object, and multiple references means the reference component has several reference objects. Encoding the semantic reference can be specifically obtained by encoding the target object targeted by the semantics of each word.

[0068] Finally, the part-of-speech vector part_of_speech_embedding corresponding to the part of speech of each word in the input sequence after word segmentation, the morphology vector morphology_embedding corresponding to the morphology of each word, and the semantic orientation vector semantic orientation_embedding corresponding to the meaning orientation of each word are merged as the word meaning vector of each word.

[0069] S525, adding sentence vector comparison coding of two intents in the vector coding layer of the intent recognition model, that is, performing vector subtraction on the sentence vectors of the two intents corresponding to the same text sentence, and performing similarity calculation on the vector difference and the word meaning vector of each first word in the text sentence, to obtain the word meaning similarity value of each first word, and using the similarity value as the intent (i.e., the subtracted intent, for example, the first intent Figure 1 - Second intention Figure 2 , then the intention refers to the first intention Figure 1 ), and normalize the probabilities of the word meanings of the same first word in each intention so that the sum of the probabilities of the word meanings of the same first word under the same intention is 1.

[0070] S527, concatenate the intention change vector intention_diff_embedding with the vector in S523, that is, concatenate the word vector encoding positional embedding, positional encoding positional embedding, sentence vector segmentembedding, and word meaning vector to obtain the sentence vector of the text sentence.

[0071] S528, the sentence vector obtained in S527 is used through the softmax method to obtain the probability of each intent, and then the first n intents with larger probability values ​​are selected. The probabilities of these n intents can be set with corresponding probability conditions, such as setting the probabilities of these n intents to be greater than 0.5, and then these n intents are normalized according to the probability so that the sum of the probabilities of these n intents is 1. Finally, the loss function of the intent recognition model is calculated until the calculated loss function value is close to the corresponding label distance, which indicates that the intent recognition model has converged. At this time, the training can be stopped to obtain the target intent recognition model.

[0072] As an example, the input of the target intent recognition model is a text sentence, and the output is multiple intents of the text sentence, the probability corresponding to each intent, and the meaning of the first word in each intent and the corresponding probability. For example, the text sentence input to the target intent recognition model is "No, what are you doing, don't call me anymore", where "what are you doing" has two meanings. The first meaning is a virtual word used to express dissatisfaction, and the second meaning is used to ask what you are doing. The output of the target intent recognition model includes: meaning Figure 1 : Negative + What are you doing + Dislike calling, what should I do? Figure 1 Probability: 0.6. Figure 1 The first words in include: not, what are you doing. Among them, the word meaning 1 of the first word "not" is negation, and the probability of the word meaning 1 is 0.9, and the word meaning 2 of the first word "not" is a vain idiomatic word, and the probability of the word meaning 2 is 0.1; the word meaning 1 of the first word "what are you doing" is what you are doing, and the probability of the word meaning 1 is 0.95, and the word meaning 2 of the first word "what are you doing" is a vain idiomatic word, and the probability of the word meaning 2 is 0.05.

[0073] meaning Figure 2 : Negative + expressing dissatisfaction + disgusted with calling, the meaning Figure 2 The probability of Figure 1 The first words in include: not, what are you doing. Among them, the word meaning 1 of the first word "not" is negation, and the probability of the word meaning 1 is 0.8, and the word meaning 2 of the first word "not" is a virtual idiomatic word, and the probability of the word meaning 2 is 0.2; the word meaning 1 of the first word "what are you doing" is what you are doing, and the probability of the word meaning 1 is 0.9, and the word meaning 2 of the first word "what are you doing" is a virtual idiomatic word, and the probability of the word meaning 2 is 0.1.

[0074] meaning Figure 3 : a cliché with a false meaning + expressing dissatisfaction + dislike of calling. Figure 3 The probability of Figure 1 The first words in include: not, what are you doing. Among them, the word meaning 1 of the first word "not" is negation, and the probability of the word meaning 1 is 0.3, and the word meaning 2 of the first word "not" is a virtual idiomatic word, and the probability of the word meaning 2 is 0.7; the word meaning 1 of the first word "what are you doing" is what you are doing, and the probability of the word meaning 1 is 0.2, and the word meaning 2 of the first word "what are you doing" is a virtual idiomatic word, and the probability of the word meaning 2 is 0.8.

[0075] It can be seen that Figure 1 The probability, meaning Figure 2 The probability and meaning Figure 3The sum of the probabilities of different meanings of the same first word in each intent is 1, and the sum of the probabilities of different meanings of the same first word in each intent is also 1.

[0076] In some embodiments, the method for determining polysemous words may include determining, through a polysemous word extraction model, that the sentence to be recognized input by the target user contains polysemous words with at least two semantics. The training method of the polysemous word extraction model includes: obtaining training samples, the training samples include sample sentences containing polysemous words, and the annotation information of the polysemous words in the sample sentences; splicing the word vectors corresponding to multiple words in the sample sentences, the word attribute vectors, and the vector difference between the word vectors of at least two intentions corresponding to the sample sentences to obtain a spliced ​​sentence vector; based on the spliced ​​sentence vector and the annotation information of the polysemous words in the polysemous sentences, training a preset word extraction model to obtain a polysemous word extraction model.

[0077] In some implementations, a method for obtaining sample sentences containing polysemous words in training samples may include: First, the sample sentences containing polysemous words in the training samples can be defined as text sentences that are identified as having different intentions according to different ways of understanding. For example, in a debt collection dialogue robot, the following two sample sentences containing polysemous words can be identified as having different intentions. (1) "Oh, no, who are you?", its corresponding intentions may include: Figure 1 : Negative (incomplete intent) + unclear caller; intent Figure 2 : Unsure of the caller. (2) “No, what are you doing? Don’t call me anymore”, its corresponding intentions may include: Figure 1 : No, what are you doing + dislike calling; meaning Figure 2 : Denial (incomplete intention) + expression of dissatisfaction + disgust with calling.

[0078] Next, we sort out the keywords or phrases (i.e., polysemous words) that cause confusion from the above-mentioned sample sentences containing polysemous words. For example, the meanings corresponding to the word "not" may include: ① no substantive meaning, mainly used for discourse transitions, for example, some users are accustomed to using this transition when replying to others, meaning that I don't want to talk about this first, I want to talk about the following content; ② negative meaning, used to deny the opinions of others. The meanings corresponding to the word "what are you doing" may include: ① no substantive meaning, mainly used to express dissatisfaction, and can be understood as a modal particle; ② used to ask what you are doing. Finally, the keywords or phrases that cause confusion in the sample sentences containing polysemous words are used as the annotation information of the corresponding sample sentences.

[0079] In some embodiments, based on the sentence vector obtained by concatenating the word vectors corresponding to multiple words in each sample sentence in the training sample, the word attribute vector and the intention change vector corresponding to the sample word and the labeling information of the corresponding confusing word, the pre-trained language model is trained to obtain a polysemous word extraction model, which may specifically include the following process: S611, load a pre-trained language model, where the pre-trained language model can be at least one of BERT (Bidirectional Encoder Representation from Transformers) and RoBERTa (A Robustly Optimized BERT Pretraining Approach).

[0080] S612, the sample sentence and its corresponding two intents and the polysemous words that cause confusion are used as labels (i.e., annotation information) for text concatenation and input into the pre-trained language model, and the task type of the pre-trained language model is set to sequence labeling. For example, the two intents corresponding to the sample sentence "Oh, no, who are you" and the polysemous words that cause confusion are concatenated to obtain the input sequence: Oh, no, who are you & negation (incomplete intent) + unclear caller & unclear caller label: ['no'].

[0081] S613, the input sequence "Oh, no, who are you & negation (incomplete intent) + unclear caller & unclear caller label: ['not']" is segmented, and the data after the word segmentation is encoded into word vectors to obtain word vector positional embedding and segment embedding. Among them, positional embedding is used to write the position information in the input sequence into the word vector. Segment embedding is used to distinguish two sentences spliced ​​together in the input sequence.

[0082] S614, for the input sequence after word segmentation processing, the part of speech and morphology of all words in the input sequence after word segmentation processing can be obtained by calling the pre-built part of speech mapping dictionary and morphology mapping dictionary, and the part of speech and morphology of all words in the input sequence after word segmentation processing are grammatically encoded to obtain word attribute vectors. Among them, the encoding of the part of speech can be specifically achieved by querying the pre-built mapping relationship table between different parts of speech and vector encodings, and the mapping relationship table between different morphologies and vector encodings, and determining the vector part_of_speech_embedding corresponding to the part of speech of each word in the input sequence after word segmentation processing, and the vector morphology_embedding corresponding to the morphology of each word. The vector part_of_speech_embedding corresponding to the part of speech of each word in the input sequence after word segmentation processing, and the vector morphology_embedding corresponding to the morphology of each word are used as the word attribute vectors of each word.

[0083] S615, add two intention comparison encodings, i.e., an encoding module for the intention change vector, to the vector encoding layer of the pre-trained language model, which is used to split the partial sequence after the first "&" in the input sequence according to &, and obtain the corresponding word vectors, i.e., the first word vector and the second word vector, for the two partial sequences after the split, and then perform vector subtraction on the two word vectors (i.e., the first word vector and the second word vector) to obtain the intention change vector intention_diff_embedding.

[0084] S617, concatenate the intention change vector intention_diff_embedding with the vectors in S613 and S614, that is, concatenate the word vector encoding positional embedding and segment embedding, the word attribute vector part_of_speech_embedding and morphology_embedding intention_diff_embedding, and obtain the sentence vector of the confused sentence.

[0085] S618, the sentence vector obtained in S617 is used to calculate the loss function through the softmax method until the calculated loss function value is close to the corresponding label distance, which indicates that the polysemous word extraction model has converged. At this time, the training can be stopped to obtain the polysemous word extraction model.

[0086] In some embodiments, word vectors and word attribute vectors corresponding to multiple words included in the sentence to be recognized can be determined; then, based on the word meaning of the first word in the sentence to be recognized, the word vector and word attribute vector corresponding to the first word among the multiple words are determined; thereafter, based on the word vector and word attribute vector corresponding to the first word among the multiple words, and the word vectors and word attribute vectors corresponding to the words other than the first word among the multiple words, a sentence vector of the sentence to be recognized is constructed; finally, based on the sentence vector of the sentence to be recognized, the intention of the sentence to be recognized is determined.

[0087] In some embodiments, word attributes may include part of speech and morphology. Part of speech refers to the characteristics of a word as the basis for dividing word classes, which may include nouns, verbs, adjectives, numerals, etc. If a word has multiple possible parts of speech, the part of speech of the word can be represented by a part of speech list. Morphology is one of the types at the grammatical level, which is the rules for the formation and use of words in a specific text, and may include simple words, compound words, derived words, etc. As an example, the word attributes of the first word can be obtained based on a pre-constructed part of speech mapping dictionary and morphology mapping dictionary.

[0088] Step 330 , determining the second intent of the sentence to be recognized based on the number of occurrences of the first word in the sentence to be recognized and the probability of the word meaning of the first word in the first intent corresponding to each other.

[0089] As an example, word meaning is used to characterize the degree of influence of the meaning of the first word on the intent of the sentence to be recognized. The degree of influence may include no influence and influence, wherein no influence characterizes that the first word is in a dispensable state in the sentence to be recognized, that is, it has no actual influence on the first intent, and influence characterizes that the first word in the sentence to be recognized plays a role in affecting the recognition result of the intent of the sentence to be recognized, that is, the first word will affect the actual meaning of the first intent. For example, the word meaning of the first word in the first intent characterizes that it is a meaningless word, which is just a modal particle, then the word meaning of the first word can be characterized as having no influence on the second intent.

[0090] In some embodiments, based on the number of occurrences of the first word in the sentence to be recognized and the probability of the word meaning of the first word in the first intention corresponding to each other, determining the second intention of the sentence to be recognized includes: Determining a weighted probability of the first word in the first intention based on the number of occurrences of the first word in the sentence to be recognized and the probability of the word meaning of the first word in the first intention corresponding to the first word; When the weighted probability satisfies the first probability condition, the first intention is determined as the second intention.

[0091] The second intention is a real intention that can characterize the final purpose of the sentence to be recognized. Optionally, the first probability condition can be a probability threshold set according to historical data statistics or experience, and the probability threshold is used to screen out word meanings that match the first intention.

[0092] In some embodiments, since the first word is both a polysemous word and an idiomatic word of the target user, the first word may be identified as a word with actual meaning in actual intent recognition, or a word without actual meaning that is only used as an interjection. That is to say, the word meaning of the first word is likely to affect the intent recognition result of the sentence to be recognized. Based on this, the embodiment of the present application can determine the probability corresponding to the word meaning of the first word in the first intent obtained by the recognition based on the intent recognition result of the sentence to be recognized, and select the word meaning of the first word with the highest probability in the first intent as the first word meaning of the first word in the first intent based on the probability, thereby determining the weighted probability of the first word in the first intent based on the first word meaning and the number of occurrences of the first word in the sentence to be recognized, and thereby determining the true intent of the sentence to be recognized based on the weighted probability. Optionally, determining the weighted probability of the first word in the first intent based on the number of occurrences of the first word in the sentence to be recognized and the probability corresponding to the word meaning of the first word in the first intent includes: Based on the probability corresponding to the word meaning of the first word in the first intention, selecting a word meaning whose probability satisfies a third probability condition from the word meaning of the first word as the first word meaning of the first word in the first intention; Based on the number of occurrences of the first word in the sentence to be recognized and the probability corresponding to the first word meaning of the first word in the first intention, a weighted probability of the first word in the first intention is determined.

[0093] It should be understood that since the first word is both a polysemous word and a common word of the target user, the probability corresponding to the word meaning of the first word in the first intention can be multiple, for example, the probability corresponding to the word meaning 1 of the first word is a, the probability corresponding to the word meaning 2 of the first word is b, and the probability corresponding to the word meaning 3 of the first word is c, where a+b+c=1, assuming a>b>c. As an example, the third probability condition is to select the maximum probability, then, based on the probability corresponding to the word meaning of the first word in the first intention, the word meaning whose probability meets the third probability condition is selected from the word meanings of the first word as the first word meaning of the first word in the first intention, and the word meaning 1 of the first word can be selected as the first word meaning of the first word in the first intention.

[0094] When the number of occurrences of the first word in the sentence to be recognized is not less than 1, the weighted probability of the first word in the first intention can be determined based on the product of the probability of the occurrence of the first word in the sentence to be recognized and the corresponding word meaning of the first word in the first intention. Optionally, when the word meanings of multiple first words in the first intention are the same, the product of the number of occurrences of the first word in the sentence to be recognized and the corresponding probability of the word meaning of the first word in the first intention can be used as the weighted probability of the first word in the first intention. When the word meanings of the same first word in the first intention are different, the sum of the probabilities corresponding to the word meanings of the first word in the first intention can be used as the weighted probability of the first word in the first intention. Among them, the word meaning of each first word should select the word meaning with the highest probability in the first intention.

[0095] In some embodiments, the first intention refers to a third intention among multiple third intentions whose probability satisfies the second probability condition, and the multiple third intentions are obtained by performing intent recognition on the sentence to be recognized. The method provided in the embodiment of the present application also includes: When the weighted probability does not satisfy the first probability condition, the first intention is deleted from the plurality of third intentions to obtain an updated third intention; Based on the updated third intention, the step of determining the first intention of the sentence to be recognized, and the word meaning corresponding to the first word in the first intention and the probability of the word meaning corresponding is performed.

[0096] Among them, multiple third intentions are obtained by performing intent recognition on the sentence to be recognized, each third intention corresponds to an intention probability, and the sum of the intention probabilities corresponding to the multiple third intentions is 1.

[0097] As described above, the first probability condition may be a probability threshold value set based on historical data statistics or experience, and the probability threshold value is used to screen out word meanings that match the first intent. If the weighted probability does not meet the first probability condition, it indicates that the word meaning with the highest probability in the first intent does not match the first intent. At this time, the multiple third intents in the first intent may be deleted to obtain an updated third intent. Optionally, based on the updated third intent, the step of determining the first intent of the sentence to be recognized, as well as the word meaning corresponding to the first word in the first intent and the probability corresponding to the word meaning is performed, including: Determine a fourth intent whose probability satisfies the second probability condition from the updated third intent, and select a word meaning whose probability satisfies the third probability condition from the word meanings of the first word based on the probability corresponding to the word meaning of the first word in the fourth intent as the first word meaning of the first word in the fourth intent; Determining a weighted probability of the first word in the fourth intent based on the number of occurrences of the first word in the sentence to be recognized and the probability of the first word meaning of the first word in the fourth intent corresponding to the first word meaning of the first word; When the weighted probability satisfies the first probability condition, the fourth intention is determined as the second intention.

[0098] By analogy, if the weighted probability of the first word in the fourth intention still does not meet the first probability condition, continue to execute the above steps of updating the third intention and determining the first intention of the sentence to be recognized, as well as the word meaning corresponding to the first word in the first intention and the probability corresponding to the word meaning.

[0099] As an example, the second probability condition is to select the third intention with the highest probability.

[0100] In summary, by using the intent recognition method provided in the embodiment of the present application, when performing intent recognition on the target user's sentence to be recognized, it is possible to first determine the first word with multiple meanings in the target user's sentence to be recognized, and when it is determined that the first word is an idiomatic word, perform intent recognition on the sentence to be recognized. The intent recognition process can not only determine the first intent of the sentence to be recognized, but also determine the word meaning corresponding to the first word in the first intent that affects the final intent recognition result of the sentence to be recognized and the probability of the word meaning corresponding. Since the word meaning of the first word with confusing meanings in the sentence to be recognized is taken into account in the recognized first intent, and based on the number of occurrences of the first word in the sentence to be recognized and the probability of the word meaning of the first word in the first intent corresponding, the first intention that matches the word meaning of the first word is determined from the first intent as the second intention of the sentence to be recognized. This can avoid the confusing effects of some first words that have multiple meanings on the intent recognition result of the sentence to be recognized during the intent recognition process, thereby improving the accuracy of intent recognition of the sentence to be recognized.

[0101] Figure 5 FIG. 5 is a schematic diagram of a structure of an intention recognition device 500 provided in an exemplary embodiment of the present application. Figure 5 As shown, the device 500 includes: a determination module 510, wherein: A determination module 510 is used to determine a first word having multiple meanings in a sentence to be recognized by a target user; The determination module 510 is also used to perform intent recognition on the sentence to be recognized when the first word is an idiomatic word of the target user, determine the first intent of the sentence to be recognized, and the word meaning corresponding to the first word in the first intent and the probability of the word meaning corresponding; determine the second intent of the sentence to be recognized based on the number of occurrences of the first word in the sentence to be recognized and the probability of the word meaning corresponding to the first word in the first intent.

[0102] The intention recognition device 500 can achieve Figure 1~Figure 4 For details, please refer to Figure 1~Figure 4The intention recognition method of the illustrated embodiment will not be described in detail.

[0103] Figure 6 The following is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present application. Figure 6 As shown, the device includes: a memory 61 and a processor 62.

[0104] The memory 61 is used to store computer programs and can be configured to store various other data to support operations on the computing device. Examples of such data include instructions for any application or method operating on the computing device, contact data, phone book data, messages, pictures, videos, etc.

[0105] The processor 62 is coupled to the memory 61 and is configured to execute the computer program in the memory 61 to: Determining a first word having multiple meanings in a sentence to be recognized by a target user; In the case where the first word is an idiomatic word of the target user, performing intent recognition on the sentence to be recognized, determining a first intent of the sentence to be recognized, and a word meaning corresponding to the first word in the first intent and a probability of the word meaning corresponding to each other; Based on the number of occurrences of the first word in the sentence to be recognized and the probability of the word meaning of the first word in the first intention corresponding to each other, a second intention of the sentence to be recognized is determined.

[0106] By using the electronic device provided by the embodiment of the present application, when performing intent recognition on the target user's sentence to be recognized, it is possible to first determine a first word with multiple meanings in the target user's sentence to be recognized, and when it is determined that the first word is an idiomatic word, perform intent recognition on the sentence to be recognized. The intent recognition process can not only determine the first intent of the sentence to be recognized, but also determine the word meaning corresponding to the first word in the first intent that affects the final intent recognition result of the sentence to be recognized and the probability of the word meaning corresponding. Since the word meaning of the first word that may have confusing meanings in the sentence to be recognized is taken into account in the recognized first intent, and based on the number of occurrences of the first word in the sentence to be recognized and the probability of the word meaning of the first word in the first intent corresponding, a first intention that matches the word meaning of the first word is determined from the first intent as the second intent of the sentence to be recognized. This can avoid confusing effects of some first words that have multiple meanings on the intent recognition result of the sentence to be recognized during the intent recognition process, thereby improving the accuracy of intent recognition of the sentence to be recognized.

[0107] Further, if Figure 6 As shown, the electronic device also includes: a communication component 63, a display 64, a power component 65, an audio component 66 and other components. Figure 6Only some components are shown schematically, which does not mean that the electronic device only includes Figure 6 In addition, depending on the implementation form of the traffic playback device, Figure 6 The components in the dashed box are optional components, not mandatory components. For example, when the electronic device is implemented as a terminal device such as a smartphone, tablet computer or desktop computer, it may include Figure 6 Components in the dashed box; when the electronic device is implemented as a server-side device such as a conventional server, cloud server, data center or server array, it may not include Figure 6 Components within the dashed box.

[0108] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor is enabled to implement the steps in the above method embodiment.

[0109] Accordingly, the embodiment of the present application further provides a computer program product, including a computer program / instruction, which can implement each step that can be performed by the electronic device in the above method embodiment when the computer program / instruction is executed. Optionally, the computer program product, in addition to executing each step in the above method embodiment.

[0110] Above Figure 6 The communication component in the communication component is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component may also include a near field communication (NFC) module, a radio frequency identification (RFID) technology, an infrared data association (IrDA) technology, an ultra-wideband (UWB) technology, a Bluetooth (BT) technology, etc.

[0111] Above Figure 6 The memory in the invention can be implemented by any type of volatile or non-volatile memory device or a combination of them, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0112] Above Figure 6The display in the embodiment includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundary of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.

[0113] Above Figure 6 The power supply component in the device provides power to various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device where the power supply component is located.

[0114] Above Figure 6 The audio component in can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in a memory or sent via a communication component. In some embodiments, the audio component also includes a speaker for outputting an audio signal.

[0115] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0116] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0117] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0118] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0120] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0121] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0122] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0123] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0124] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A method for intention recognition, characterized in that: include: Determining a first word having multiple meanings in a sentence to be recognized by a target user; In the case where the first word is an idiomatic word of the target user, performing intent recognition on the sentence to be recognized, determining a first intent of the sentence to be recognized, and a word meaning corresponding to the first word in the first intent and a probability of the word meaning corresponding to each other; Based on the number of occurrences of the first word in the sentence to be recognized and the probability of the word meaning of the first word in the first intention corresponding to each other, a second intention of the sentence to be recognized is determined.

2. The method according to claim 1, characterized in that The determining the second intention of the sentence to be recognized based on the number of occurrences of the first word in the sentence to be recognized and the probability of the word meaning of the first word in the first intention corresponding to each other includes: Determining a weighted probability of the first word in the first intent based on the number of occurrences of the first word in the sentence to be recognized and the probability of the word meaning of the first word in the first intent corresponding to the first word; When the weighted probability satisfies a first probability condition, the first intention is determined as a second intention.

3. The method according to claim 2, characterized in that The first intention refers to a third intention among multiple third intentions whose probability satisfies a second probability condition, and the multiple third intentions are obtained by performing intent recognition on the sentence to be recognized, and the method further includes: When the weighted probability does not satisfy the first probability condition, deleting the first intention from the multiple third intentions to obtain an updated third intention; Based on the updated third intention, a step of determining a first intention of the sentence to be recognized, and a word meaning corresponding to the first word in the first intention and a probability of the word meaning corresponding is performed.

4. The method according to claim 2, characterized in that The determining, based on the number of occurrences of the first word in the sentence to be recognized and the probability corresponding to the word meaning of the first word in the first intention, a weighted probability of the first word in the first intention includes: Based on the probabilities corresponding to the word meanings of the first word in the first intention, selecting a word meaning whose probability satisfies a third probability condition from the word meanings of the first word as the first word meaning of the first word in the first intention; Based on the number of occurrences of the first word in the sentence to be recognized and the probability corresponding to the first word meaning of the first word in the first intention, a weighted probability of the first word in the first intention is determined.

5. The method according to claim 1, characterized in that The method further comprises: Acquire multiple words that appear in the session data of the target user in the first time period; For each word, determining the frequency of occurrence of the word; Obtain high-frequency words whose occurrence frequency is greater than or equal to a preset threshold; Based on the position flexibility of the high-frequency words in the conversation data of the first time period, the idiomatic words of the target user are determined.

6. The method according to claim 5, characterized in that The method for determining the position flexibility of the high-frequency words includes: Determining a first position flexibility of the high-frequency words based on the occurrence positions of the high-frequency words in the conversation data; Determining a second position flexibility of the high-frequency words based on the number of categories of syntactic dependency relationships related to the high-frequency words in the syntactic dependency relationships of the conversation data; Determining a third position flexibility of the high-frequency words based on the degree of influence of the position change of the high-frequency words on the semantics of the conversation data; The position flexibility of the high-frequency words is determined based on the first position flexibility, the second position flexibility and the third position flexibility of the high-frequency words.

7. An intention recognition device, characterized in that: include: A determination module, used to determine a first word having multiple meanings in a sentence to be recognized by a target user; The determination module is also used to perform intent recognition on the sentence to be recognized when the first word is an idiomatic word of the target user, determine the first intent of the sentence to be recognized, and the word meaning corresponding to the first word in the first intent and the probability of the word meaning corresponding; determine the second intent of the sentence to be recognized based on the number of occurrences of the first word in the sentence to be recognized and the probability of the word meaning corresponding to the first word in the first intent.

8. An electronic device, characterized in that: include: one or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors, and the one or more programs are configured to perform the steps in the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The method comprises a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 6.