Intent recognition method, electronic apparatus, and vehicle

By using deep learning models to recognize and guide information interaction based on voice commands, the problem of vehicles recognizing ambiguous voice commands has been solved, improving the accuracy of intent recognition and user experience.

WO2026026334A1PCT designated stage Publication Date: 2026-02-05BYD CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/103886
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-02
Filing Date
2025-06-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

When vehicles encounter voice commands that are not properly worded or have unclear intent, they usually refuse to recognize or perform generalized processing, resulting in a poor user experience.

Method used

The system uses a deep learning model to recognize voice commands, generates voice recognition results, and issues guidance information when preset guidance conditions are met. It also interacts with the user to confirm their intent and generates intent recognition results.

Benefits of technology

This improves the accuracy of vehicle intent recognition and user experience, avoiding situations where incorrect intents are refused to be recognized or executed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025103886_05022026_PF_FP_ABST
    Figure CN2025103886_05022026_PF_FP_ABST
Patent Text Reader

Abstract

An intent recognition method, an electronic apparatus, and a vehicle (100). The method comprises: when a speech instruction has been received, recognizing the speech instruction, so as to obtain a speech recognition result (011); when the speech recognition result meets a preset guidance condition, sending guidance information, wherein the guidance information is generated on the basis of the speech recognition result (012); and when feedback speech for the guidance information has been received, recognizing the feedback speech, so as to obtain an intent recognition result (013).
Need to check novelty before this filing date? Find Prior Art

Description

Intention recognition method, electronic device and vehicle

[0001] Priority information

[0002] The present application claims priority to and the benefit of the filing date of the Chinese Patent Application No. 2024110544725 filed on August 2, 2024 in the China National Intellectual Property Office, and incorporates by reference the entirety of the aforementioned application. TECHNICAL FIELD

[0003] The present application relates to the technical field of vehicles, and more particularly, to an intention recognition method, an electronic device and a vehicle. BACKGROUND

[0004] After the vehicle performs speech recognition and intention understanding on a sound signal, for a sentence with an explicit intention, the vehicle can accurately understand and execute the corresponding intention, thereby bringing a good experience to the user. The user does not need to repeatedly explain or clarify his / her own needs, and the vehicle can make a satisfactory response based on the speech recognition and intention understanding.

[0005] However, in daily life and applications, not all the speech received by the vehicle contains an explicit intention. When the vehicle encounters speech with an unstandardized expression or an unclear intention, the vehicle usually chooses to refuse to recognize or sends a prompt of “unable to recognize” to the user, or directly executes an intention obtained by generalizing the speech with an unclear intention (the intention is usually obtained by generalizing the unclear intention when the vehicle fails to correctly recognize, and the correctness is usually low), thereby causing a poor user experience. SUMMARY

[0006] Therefore, the present application provides an intention recognition method, an electronic device and a vehicle, which can improve the user experience.

[0007] The intention recognition method of the present application comprises: in the case of receiving a speech instruction, recognizing the speech instruction to obtain a speech recognition result; in the case that the speech recognition result meets a preset guidance condition, issuing guidance information, the guidance information being generated based on the speech recognition result; and in the case of receiving feedback speech for the guidance information, recognizing the feedback speech to obtain an intention recognition result.

[0008] In some embodiments, the speech recognition result includes a classification result and extracted words, the classification result includes whether there is an intent, and the extracted words include at least one of an entity word, a verb, and a locative word, and the speech recognition result satisfies a preset guiding condition, including at least one of: the classification result is that there is an intent, the extracted words only include an entity word, and a confidence of the entity word is greater than a first confidence threshold; the classification result is that there is an intent, the extracted words only include an entity word and a verb, and a confidence of the entity word is between a second confidence threshold and the first confidence threshold, the second confidence threshold being less than the first confidence threshold; the classification result is that there is an intent, the extracted words only include an entity word, and a confidence of the entity word is between a second confidence threshold and the first confidence threshold, the second confidence threshold being less than the first confidence threshold; the classification result is that there is an intent, and the extracted words only include a locative word; the classification result is that there is an intent, and the extracted words only include a verb; the classification result is that there is an intent, and the extracted words only include a locative word and a verb.

[0009] In some embodiments, the speech recognition result includes extracted words, and the extracted words include at least one of an entity word, a verb, and a locative word, and the method further includes: generating the guiding information based on at least one of a target word missing in the extracted words and a confidence of the entity word, the target word including at least one of a verb and a locative word.

[0010] In some embodiments, the generating the guiding information based on at least one of a target word missing in the extracted words and a confidence of the entity word includes: in a case where the extracted words include an entity word, generating the guiding information based on at least one of a target word missing in the extracted words and the confidence of the entity word; and in a case where the extracted words do not include an entity word, generating the guiding information based on a target word missing in the extracted words.

[0011] In some embodiments, the guidance information includes orientation guidance information, action guidance information, action and orientation guidance information, entity guidance information, action and entity guidance information, orientation and entity guidance information; the generating the guidance information based on the missing target word in the extracted word and the confidence of the entity word in the case that the extracted word contains an entity word includes: in the case that there is only an entity word in the extracted word and the confidence of the entity word is greater than a first confidence threshold, generating the action guidance information or the action and orientation guidance information; in the case that there is only an entity word and a verb in the extracted word and the confidence of the entity word is between the second confidence threshold and the first confidence threshold, generating the entity guidance information or the orientation and entity guidance information; in the case that there is only an entity word in the extracted word and the confidence of the entity word is between the second confidence threshold and the first confidence threshold, generating the action and entity guidance information; the generating the guidance information based on the missing target word in the extracted word in the case that the extracted word does not contain an entity word includes: in the case that there is only an orientation word in the extracted word, generating the action and entity guidance information; in the case that there is only a verb in the extracted word, generating the orientation and entity guidance information; in the case that there is only an orientation word and a verb in the extracted word, generating the entity guidance information.

[0012] In some embodiments, the speech recognition result includes a classification result and an extracted word, the classification result includes whether there is an intent, and the extracted word includes at least one of an entity word, a verb, and an orientation word, and the method further includes: in the case that the recognition result meets a preset intent generation condition, generating the intent recognition result based on the speech recognition result; wherein the intent generation condition includes that the classification result is that there is an intent, the extracted word contains a verb and an entity word, and the confidence of the entity word is greater than a first confidence threshold, or the classification result is that there is an intent, the extracted word contains a verb, an orientation word, and an entity word, and the confidence of the entity word is greater than the first confidence threshold.

[0013] In some embodiments, the recognition result includes a classification result and an extracted word, the classification result includes whether there is an intent, and the extracted word includes at least one of an entity word, a verb, and an orientation word, and the method further includes: in the case that the recognition result meets a preset stop condition, stopping intent recognition; wherein the recognition result meeting the preset stop condition includes at least one of the following: the classification result is that there is no intent; the extracted word has no orientation word; the confidence of the entity word in the extracted word is less than a third confidence threshold.

[0014] In some embodiments, the confidence of the entity word is determined based on a first confidence and a second confidence, the first confidence is determined based on a maximum similarity between the entity word and each matched entity word in the preset intent configuration table that matches the entity word, and the second confidence is determined based on a similarity between a pinyin of the entity word and a pinyin of the matched entity word corresponding to the maximum similarity.

[0015] In some embodiments, the speech instruction is recognized to obtain a speech recognition result based on a preset deep learning model.

[0016] The electronic device of the embodiments of the present application includes a processor connected with a memory; the memory stores a computer program, and the processor executes the computer program to implement the instructions of the intent recognition method of any of the above embodiments.

[0017] The vehicle of the embodiments of the present application includes the electronic device of any of the above embodiments.

[0018] The non-volatile computer readable storage medium of the embodiments of the present application includes a computer program, which, when executed by a processor, causes the processor to execute the intent recognition method of any of the above embodiments.

[0019] The intent recognition method, electronic device and vehicle of the embodiments of the present application, in the case of receiving a speech instruction, recognize the speech instruction to obtain a speech recognition result; in the case that the speech recognition result meets a preset guidance condition, issue a guidance information, the guidance information is generated based on the speech recognition result; in the case of receiving a feedback speech for the guidance information, recognize the feedback speech to obtain an intent recognition result, that is, based on the speech recognition result obtained after recognizing the speech instruction, issue the guidance information to the user, and then receive and recognize the feedback speech of the user to obtain the intent recognition result, that is, even if a speech instruction that is difficult for the vehicle to determine the intent or the vehicle cannot guarantee the accuracy of the recognition is received, the guidance information can be issued to interact with the user to obtain the intent recognition result, instead of rejecting the recognition, replying to the user that the intent cannot be recognized or executing the wrong intent, etc., thereby improving the accuracy of the executed intent and improving the user experience.

[0020] Additional aspects and advantages of the embodiments of the present application will be in part apparent and in part pointed out hereinafter in the description of the embodiments of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0021] The above and / or additional aspects and advantages of the present application will become apparent and more readily appreciated from the following description of the embodiments, taken in conjunction with the accompanying drawings in which:

[0022] FIG. 1 is a schematic diagram of an application scenario of an intention recognition method according to some embodiments of the present application;

[0023] FIG. 2 is a schematic diagram of a flow of an intention recognition method according to some embodiments of the present application;

[0024] FIG. 3 is a schematic diagram of a flow of an intention recognition method according to some embodiments of the present application;

[0025] FIG. 4 is a schematic diagram of a flow of an intention recognition method according to some embodiments of the present application;

[0026] FIG. 5 is a schematic diagram of a flow of an intention recognition method according to some embodiments of the present application;

[0027] FIG. 6 is a schematic diagram of an application scenario of an intention recognition method according to some embodiments of the present application;

[0028] FIG. 7 is a schematic diagram of an application scenario of an intention recognition method according to some embodiments of the present application;

[0029] FIG. 8 is a schematic diagram of an application scenario of an intention recognition method according to some embodiments of the present application;

[0030] FIG. 9 is a schematic diagram of an application scenario of an intention recognition method according to some embodiments of the present application;

[0031] FIG. 10 is a schematic diagram of a flow of an intention recognition method according to some embodiments of the present application;

[0032] FIG. 11 is a schematic diagram of a flow of an intention recognition method according to some embodiments of the present application;

[0033] FIG. 12 is a schematic diagram of modules of an intention recognition apparatus according to some embodiments of the present application;

[0034] FIG. 13 is a schematic diagram of a structure of a vehicle according to some embodiments of the present application;

[0035] FIG. 14 is a schematic diagram of a connection state of a non-volatile computer readable storage medium and a processor according to some embodiments of the present application. DETAILED DESCRIPTION

[0036] Embodiments of the present application are described in detail below with reference to the accompanying drawings, in which the same or similar components have the same or similar reference numerals throughout the drawings and a detailed description of the embodiments is given below by way of example only, which is intended to explain the embodiments of the present application and cannot be understood as a limitation of the embodiments of the present application.

[0037] For the convenience of understanding the present application, the terms appearing in the present application are explained as follows:

[0038] 1. Artificial Intelligence (AI) is the theory, methods, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making functions. AI technology is a comprehensive discipline involving a wide range of fields, including both hardware and software technologies. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operating / interactive systems, and mechatronics. AI software technologies mainly include computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. The technical solutions provided in this application mainly involve natural language processing technology and machine learning / deep learning in artificial intelligence.

[0039] 2. Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.

[0040] 3. Deep Learning (DL): A branch of machine learning, it's an algorithm that attempts to perform high-level abstraction of data using multiple processing layers containing complex structures or multiple nonlinear transformations. Deep learning learns the inherent patterns and hierarchical representations of training sample data. The information gained during this learning process greatly aids in interpreting data such as text, images, and sound. The ultimate goal of deep learning is to enable machines to possess analytical and learning capabilities like humans, capable of recognizing data such as text, images, and sound. Deep learning is a complex machine learning algorithm, and its performance in speech and image recognition far surpasses previous related technologies.

[0041] 4. Natural Language Understanding (NLU): is a general term for all methods and models or tasks that support machine understanding of text content. NLU plays a very important role in text information processing systems and is a necessary module for recommendation, question answering, search, and other systems. In the embodiments of the present application, NLU is used to learn and understand user intent, so as to better respond to instructions and requirements and improve interaction efficiency and accuracy.

[0042] 5. Automatic Speech Recognition (ASR): aims to convert the lexical content in human speech into computer-readable input, such as keystrokes, binary codes, or character sequences.

[0043] Vehicle speech recognition technology refers to the automatic recognition and understanding of vehicle status, operation, or environment by analyzing sound signals inside or around the vehicle, which can be used in vehicle safety, driving assistance systems, and vehicle entertainment.

[0044] In vehicle speech data processing, the explicit intent included in the sound signal is the key to the system making correct decisions and actions. After the vehicle performs speech recognition and intent understanding on the sound signal, for sentences with explicit intent, the corresponding intent can be accurately understood and executed, providing a good experience for the user. The user does not need to repeatedly explain or clarify their needs, and the vehicle can respond satisfactorily based on speech recognition and intent understanding.

[0045] However, in daily life and applications, not all the speech received by the vehicle contains explicit intent. When the vehicle encounters speech that is not standardized or has unclear intent, it usually chooses to refuse recognition or prompts the user with "unable to recognize", or directly executes the intent obtained by generalizing the speech with unclear intent (the intent is obtained by generalizing the unclear intent when the vehicle cannot correctly recognize it, and the correctness is usually low), resulting in a poor user experience.

[0046] To solve the above technical problems, the embodiments of the present application provide an intent recognition method.

[0047] First, an application scenario of the technical solution of the present application will be introduced. As shown in FIG. 1, the intent recognition method provided by the present application can be applied to the application scenario as shown in FIG. 1. The intent recognition method can be applied to a vehicle 100.

[0048] The vehicle 100 is any vehicle 100 that can process speech data, such as a car, a truck, etc.

[0049] The vehicle 100 comprises a vehicle body 50.

[0050] In an embodiment, the vehicle 100 further comprises a microphone 20 arranged inside the vehicle body 50, which is used to collect audio information inside or around the vehicle, and through which the user can issue voice instructions to the vehicle, etc.

[0051] The intention recognition method of the present application will be described in detail as follows:

[0052] Referring to FIGS. 1 and 2, the intention recognition method provided by the present application comprises:

[0053] Step 011: In the case of receiving a voice instruction, the voice instruction is recognized to obtain a voice recognition result.

[0054] The voice instruction can include a request or query sent by the user to the system through voice, representing the function or action that the user wants the vehicle to perform. In the field of voice recognition and voice control, the terms "voice request", "voice instruction", "voice command", "voice control instruction", etc. are often used to describe the voice phrase used by the user to trigger a specific operation or task. These voice instructions can be a single word, a phrase, or a complete sentence, which are referred to as "queries". That is, the voice instruction can be a query used by the user in voice data processing. For example, the voice instruction can be "please open the window", "raise the driver's seat", and "turn off the vehicle air conditioner", etc.

[0055] Specifically, the user can issue a voice instruction through the microphone on the vehicle. The voice instruction can be a voice instruction issued by the user that is collected by the microphone of the vehicle; or the vehicle is associated with a terminal (such as connected through a network, etc.), and the user issues a voice instruction through the terminal, and the terminal sends the received voice instruction to the vehicle after receiving the voice instruction. In the case of receiving the voice instruction, the vehicle can first preprocess the voice instruction, for example, through preprocessing methods such as noise reduction, audio enhancement, audio segmentation, etc., to improve the audio quality; then based on ASR voice recognition technology, etc., through voice feature extraction, etc., the voice instruction is recognized, and the corresponding voice recognition result is generated.

[0056] Referring to FIG. 3, optionally, step 011: In the case of receiving a voice instruction, the voice instruction is recognized to obtain a voice recognition result, comprising:

[0057] Step 0111: The voice instruction is recognized based on a pre-set deep learning model to obtain a voice recognition result.

[0058] The preset deep learning model can be a neural network model based on deep learning. The deep learning model can recognize the voice instruction to output a corresponding voice recognition result (for example, a voice recognition result in the form of text).

[0059] The preset deep learning model can be a named entity recognition (NER) model, which can include a binary classification output (i.e., whether the voice instruction includes an intent, etc.). The model structure can include Bert (an open-source pre-trained model), CRF (a conditional random field model commonly used in named entity recognition tasks), and Dense (a fully connected layer in deep learning). Each character is tokenized (a simplified term for character mapping to a numerical sequence in deep learning), and then input into the Bert pre-trained encode layer (input layer) to obtain a feature vector representation of each character. The feature vector is then input into the model to output the results of the CRF layer and the Dense layer, i.e., the classification result of the voice recognition result and the words (such as verbs, entity words, etc.) included in the voice recognition result.

[0060] Specifically, the deep learning model can be obtained by training. The training process is as follows: by obtaining a sample of a voice instruction and label information corresponding to the sample of the voice instruction, the label information including a voice recognition result corresponding to the voice instruction; inputting the voice instruction into the deep learning model for training to output a training voice recognition result; adjusting the deep learning model based on the loss value between the training voice recognition result and the voice recognition result included in the label information until the deep learning model converges.

[0061] Upon receiving the voice instruction, the corresponding voice recognition result can be output by inputting the voice instruction into the deep learning model.

[0062] Step 012: In the case where the voice recognition result meets a preset guidance condition, issuing guidance information, the guidance information being generated based on the voice recognition result.

[0063] The preset guidance condition can be a condition for distinguishing whether to issue guidance information and what kind of guidance information to issue. For example, the preset guidance condition can be a certain keyword. In the case where the voice recognition result includes the keyword, it is considered that the voice recognition result meets the preset guidance condition, and the vehicle is controlled to issue the guidance information. The guidance information can be information for guiding the user to confirm the intent included in the voice instruction, and the guidance information is generated based on the voice recognition result.

[0064] Specifically, in a case where the voice recognition result meets a preset guidance condition (i.e., in a case where the vehicle receives a voice instruction with an ambiguous intention), guidance information can be sent to the user through a speaker, a display screen, or the like of the vehicle to guide the user to confirm the intention included in the voice instruction. For a received voice instruction with an ambiguous intention, the voice recognition result generated based on the voice instruction can be judged, and corresponding guidance information can be generated based on the voice recognition result to guide the user to confirm.

[0065] Optionally, the voice recognition result includes a classification result and extracted words, the classification result includes whether there is an intention, and the extracted words include at least one of an entity word, a verb, and a position word, the entity word is used to represent an object of an operation of the voice instruction, for example, in a vehicle, the entity word can be a window, a seat, or the like of the vehicle.

[0066] Optionally, the classification result includes whether there is an intention, and the voice received by the vehicle is not all voice including an intention. In a case where the vehicle receives voice, the voice can be classified to determine whether the voice received at this time contains an intention.

[0067] The entity word can be used to refer to an object or device existing in the vehicle, and the entity word can include an object in the intention that the user expects to perform a certain action. For example, the entity word can be a window, a sunroof, or the like. The verb can be an action in the intention that the user expects to perform. For example, open, close, or the like. The position word can be used to represent the position of the entity word, for example, the driver, the front passenger, or the like.

[0068] It can be understood that the vehicle understands the intention of the voice instruction of the user through the entity word, the verb, and the position word existing in the intention, and wants to perform an action on what (position word) device (entity word).

[0069] Optionally, the voice recognition result meets a preset guidance condition, including at least one of the following:

[0070] (1) the classification result is that there is an intention, there is only an entity word in the extracted words, and the confidence of the entity word is greater than a first confidence threshold.

[0071] Optionally, the confidence of the entity word is determined based on a first confidence and a second confidence, the first confidence is determined based on a maximum similarity between the entity word and each matching entity word matched with the entity word in a preset intention configuration table, and the second confidence is determined based on a similarity between a pinyin of the entity word and a pinyin of the matching entity word corresponding to the maximum similarity.

[0072] Optionally, the preset intention configuration table can be generated by experiment, table lookup, or the like, and the preset intention configuration table includes an entity word (such as a window) existing in the vehicle and pinyin information corresponding to each entity word.

[0073] By obtaining a preset intention configuration table, the intention configuration table includes entity words existing in the vehicle and pinyin corresponding to the entity words. According to the intention configuration table, a matching entity word with the maximum similarity to the entity word is determined, and a first confidence is determined based on the entity word and the matching entity word by obtaining the matching entity word. Then, the pinyin information of the entity word is obtained (such as by a pinyin function), and a second confidence is determined based on the pinyin information of the entity word and the pinyin information of the matching entity word, which can avoid the problem of misinterpretation of homophonic words caused by different tones of the same word in different regions during speech recognition. For example, by using a preset similarity comparison model, the similarity comparison model can be used to determine the matching entity word with the maximum similarity to the entity word and the similarity between the entity word and the matching entity word (i.e., the first confidence). The entity word is input into the similarity comparison model to output the first confidence. Then, based on the pinyin information of the entity word and the pinyin information of the matching entity word, the second confidence is obtained, for example, by using a cosine function to calculate the second confidence.

[0074] wherein the confidence of the entity word can be determined based on the first confidence and the second confidence. For example, the confidence S = 0.5 * the first confidence S1 + 0.5 * the second confidence S2; or the confidence S = 0.4 * the first confidence S1 + 0.6 * the second confidence S2, etc. By combining the first confidence and the second confidence to determine the confidence, the misrecognition of homophonic words and synonyms can be reduced, and the recognition accuracy can be improved.

[0075] wherein the first confidence threshold can be determined based on an empirical value, for example, the first confidence threshold can be 0.5, 0.6, 0.7, 1, etc.

[0076] In the case where the classification result is that there is an intention and only the entity word exists in the extracted words, since only the entity word exists in the extracted words of the voice instruction, it cannot be determined that the user wants the vehicle to perform which intention based on the entity word. However, in the case where the confidence of the entity word is greater than the first confidence threshold, it can be considered that the matching entity word is included in the voice instruction with a high probability. Therefore, the guiding information is sent to the user, and the intention of the user is confirmed through interaction with the user to improve the user experience. That is, in the case where the classification result is that there is an intention, only the entity word exists in the extracted words, and the confidence of the entity word is greater than the first confidence threshold, it is considered that the preset guiding condition is met, and the user is guided to confirm the intention by sending the guiding information.

[0077] (2) the classification result is that there is an intention, only the entity word and the verb exist in the extracted words, and the confidence of the entity word is between the second confidence threshold and the first confidence threshold, the second confidence threshold is less than the first confidence threshold.

[0078] The second confidence threshold is less than the first confidence threshold, for example, the first confidence threshold can be 1, and the second confidence threshold is 0.6; or the first confidence can be 0.8, and the second confidence threshold is 0.4, etc. In the case where the confidence of the entity word is between the second confidence threshold and the first confidence threshold, it can be considered that the probability of including the matching entity word in the voice instruction is moderate.

[0079] In the case where the classification result is that there is an intention, and only entity words exist in the extracted words, and the confidence of the entity word is between the second confidence threshold and the first confidence threshold, the vehicle can determine that the user's voice instruction contains an intention, but cannot determine the direction of the intention, so the user's intention can be confirmed by issuing guidance information to the user (such as confirming the entity word, confirming the verb and the entity word, etc.), so as to improve the user's experience.

[0080] (3) The classification result is that there is an intention, only entity words exist in the extracted words, and the confidence of the entity word is between the second confidence threshold and the first confidence threshold, and the second confidence threshold is less than the first confidence threshold.

[0081] In the case where the classification result is that there is an intention, only entity words exist in the extracted words, and the confidence of the entity word is between the second confidence threshold and the first confidence threshold, the vehicle can determine that the user's voice instruction contains an intention, but cannot determine the direction of the intention, so the user's intention can be confirmed by issuing guidance information to the user (such as confirming the entity word, confirming the verb and the entity word, etc.), so as to improve the user's experience.

[0082] (4) The classification result is that there is an intention, and only orientation words exist in the extracted words.

[0083] In the case where the classification result is that there is an intention, and only orientation words exist in the extracted words, the vehicle can only determine the orientation, and cannot determine the specific device based on the orientation, etc. The user's intention can be confirmed by issuing guidance information to the user (such as confirming the entity word, confirming the verb and the entity word, etc.).

[0084] (5) The classification result is that there is an intention, and only verbs exist in the extracted words.

[0085] In the case where the classification result is that there is an intention, and only orientation words exist in the extracted words, the vehicle can only confirm the action expected to be performed by the user, and cannot determine the entity of the performed action included in the intention, etc. Therefore, the user's intention can be confirmed by issuing guidance information to the user (such as confirming the entity word, confirming the orientation word, etc.).

[0086] (6) The classification result is that there is an intention, and only orientation words and verbs exist in the extracted words.

[0087] In the case that the classification result is that the intention exists, the extracted words only include the position words and the verbs, the vehicle cannot determine the entity of the performed action included in the intention, etc., and thus the intention of the user can be confirmed (e.g., the entity word is confirmed) by issuing the guidance information to the user.

[0088] Step 013: In the case that the feedback speech for the guidance information is received, the feedback speech is recognized to obtain an intention recognition result.

[0089] The feedback speech can be a response or feedback made by the user to the guidance information.

[0090] Specifically, the vehicle issues the guidance information to guide the user to confirm the intention, and in the case that the feedback speech for the guidance information is received, the feedback speech is recognized to obtain the intention recognition result, the accuracy of the recognition of the intention is improved through the interaction with the user, and thus the use experience of the user is improved.

[0091] In this way, in the case that the speech instruction is received, the speech instruction is recognized to obtain a speech recognition result, in the case that the speech recognition result meets a preset guidance condition, the guidance information is issued, the guidance information is generated based on the speech recognition result, in the case that the feedback speech for the guidance information is received, the feedback speech is recognized to obtain an intention recognition result, that is, the guidance information is issued to the user based on the speech recognition result obtained after the speech instruction is recognized, and then the feedback speech of the user is received and recognized to obtain the intention recognition result, even if the speech instruction in which the intention is difficult for the vehicle to determine or the vehicle cannot guarantee the accuracy of the recognition is received, the guidance information can be issued to interact with the user, and thus the intention recognition result is obtained, rather than rejecting the recognition, replying that the intention cannot be recognized or the intention is wrong, etc., the accuracy of the performed intention is improved, and the use experience of the user is improved.

[0092] Referring to FIG. 4, in some embodiments, the speech recognition result includes the extracted words, the extracted words include at least one of the entity words, the verbs and the position words, and the intention recognition method further includes:

[0093] Step 014: The guidance information is generated based on at least one of the target words missing in the extracted words and the confidence of the entity words, the target words including at least one of the verbs and the position words.

[0094] The guidance information can be generated based on the target words missing in the extracted words, or the guidance information can be generated based on the confidence of the entity words, or the guidance information can be generated based on the target words missing in the extracted words and the confidence of the entity words, and the present application does not limit this.

[0095] The target word can include a verb, or the target word can include a direction word, or the target word can include a verb and a direction word, and the application does not limit this.

[0096] Specifically, the guide information is generated according to at least one of the confidence of the target word and the entity word missing in the extracted word. The vehicle determines the user's intention through the understanding of the extracted word, and therefore, the corresponding guide information can be generated based on the target word missing in the extracted word. In the case where the confidence of the entity word is high, the vehicle can determine the entity included in the intention, and in the case where the confidence of the entity word is low, the vehicle cannot determine the entity included in the intention. In order to ensure the user's experience, the corresponding guide information can also be generated to guide the user to confirm.

[0097] Referring to FIG. 5, optionally, step 014: generating guide information based on at least one of the confidence of the target word and the entity word missing in the extracted word, the target word including at least one of a verb and a direction word, including:

[0098] Step 0141: in the case where the extracted word contains the entity word, generating guide information based on the confidence of the target word and the entity word missing in the extracted word;

[0099] Step 0142: in the case where the extracted word does not contain the entity word, generating guide information based on the target word missing in the extracted word.

[0100] Specifically, the corresponding guide information can be generated by determining whether the extracted word contains the entity word. In the case where the extracted word contains the entity word, the guide information is generated based on the confidence of the target word and the entity word missing in the extracted word. Based on the missing target word, the operation expected by the user's voice instruction to be performed by the vehicle is determined, and based on the confidence of the entity word, it is judged whether the entity word needs to be confirmed, thereby generating the corresponding guide information; in the case where the extracted word does not contain the entity word, the vehicle cannot determine the intention, and therefore, the guide information can be generated based on the target word missing in the extracted word.

[0101] For example, in the case that the voice instruction of the user received is "open the main driver's window", the voice instruction is recognized to obtain a voice recognition result, the voice recognition result includes the verb "open", the direction word "main driver", and the entity word "window", the extracted words include the entity word, and there is no missing target word, and based on the confidence of the entity word, corresponding guide information can be generated: according to the intent configuration table, it is determined that the matching entity word with the maximum similarity to "window" is "window", based on "window" and "window", a first confidence is determined, and based on the pinyin "chechuang" of "window" and the pinyin "chechuang" of "window", a second confidence is determined, according to the first confidence and the second confidence, the confidence is determined, and according to the confidence of the entity word, the corresponding guide information is generated, such as "Is it to open the main driver's window? You can try to tell me to open the main driver's window", and in the case that the feedback voice for the guide information is received, such as the user's answer "yes" or "open the main driver's window", the feedback voice is recognized, and an intent recognition result is obtained to execute the corresponding intent.

[0102] For another example, in the case that the voice instruction of the user received is "Tianchuang", the voice instruction is recognized to obtain a voice recognition result, and the voice recognition result only includes the entity word "Tianchuang", and lacks target words (verbs and direction words), and then the confidence of "Tianchuang" and "skylight" is calculated, and based on the missing target words (verbs and direction words) and the confidence of "Tianchuang" and "skylight", corresponding guide information is generated, such as "Is it to open or close the skylight? You can try to tell me to open the skylight", and in the case that the feedback voice for the guide information is received, such as the user's answer "yes" or "open the skylight", the feedback voice is recognized, and an intent recognition result is obtained to execute the corresponding intent.

[0103] For another example, in the case that the voice instruction of the user received is "close the co-driver", the voice instruction is recognized to obtain a voice recognition result, and the voice recognition result includes the verb "close" and the direction word "co-driver", and does not include the entity word, therefore, the guide information generated based on the missing target words (i.e. entity words) in the extracted words can be: such as "Is it to close the co-driver? You can try to tell me to close the co-driver's window", in the case that the feedback voice for the guide information is received, such as the user's answer "yes" or "close the co-driver XX", the feedback voice is recognized, and an intent recognition result is obtained to execute the corresponding intent.

[0104] Optionally, the guide information includes direction guide information, action guide information, action and direction guide information, entity guide information, action and entity guide information, and direction and entity guide information.

[0105] The guidance information includes position guidance information, action guidance information, action and position guidance information, entity guidance information, action and entity guidance information, and position and entity guidance information. The position guidance information is used to guide the user to confirm the position in the voice instruction; the action guidance information is used to guide the user to confirm the action expected to be performed by the vehicle in the voice instruction; the action and position guidance information is used to guide the user to confirm the action and the position in the voice instruction; the entity guidance information is used to guide the user to confirm the entity of the action performed in the voice instruction; the action and entity guidance information is used to guide the user to confirm the entity of the action performed and the specific action to be performed; and the position and entity guidance information is used to guide the user to confirm the entity of the action performed and the specific position of the entity.

[0106] In a case where the extracted words include entity words, the guidance information is generated based on the target words missing in the extracted words and the confidence of the entity words, including:

[0107] (1) In a case where only entity words exist in the extracted words and the confidence of the entity words is greater than a first confidence threshold, action guidance information or action and position guidance information is generated.

[0108] In a case where only entity words exist in the extracted words and the confidence of the entity words is greater than a first confidence threshold, action guidance information or action and position guidance information is generated to guide the user to confirm the action performed or the action and the position performed. For example, referring to FIG. 6, in a case where the received voice instruction is “please open the window”, that is, the extracted words include only the entity word “window” and the confidence of the entity word (0.9) is greater than the first confidence threshold (0.8), since there is no verb and position word in the voice instruction, the action guidance information such as “do you want to open or close the window? You can try to tell me to open the window.” or the action and position guidance information such as “do you want to open or close the window of the driver? You can try to tell me to open the window of the driver” and the like can be generated.

[0109] (2) In a case where only entity words and verbs exist in the extracted words and the confidence of the entity words is between a second confidence threshold and the first confidence threshold, entity guidance information or position and entity guidance information is generated.

[0110] In the case that only entity words and verbs exist in the extracted words, and the confidence of the entity words is between the second confidence threshold and the first confidence threshold, the entity guiding information or the position and entity guiding information is generated to guide the user to confirm the entity of the performed action or the entity of the performed action and the position where the entity is located. For example, in the case that the received voice instruction is "please open the Tianchuang for me", that is, only entity words ("Tianchuang") and verbs ("open") exist in the extracted words, and the confidence of the entity words (0.7) is between the second confidence threshold (0.6) and the first confidence threshold (0.8), since the confidence of the entity word "Tianchuang" and "sunroof" in the intent configuration table is 0.7, the entity guiding information such as "is it to open the sunroof" can be generated to confirm the entity word "sunroof"; or the position and entity guiding information such as "is it to open the sunroof in the front of the vehicle? You can try to tell me to open the sunroof in the front of the vehicle" and the like is generated to guide the user to confirm.

[0111] (3) In the case that only entity words exist in the extracted words, and the confidence of the entity words is between the second confidence threshold and the first confidence threshold, the action and entity guiding information is generated;

[0112] In the case that only entity words exist in the extracted words, and the confidence of the entity words is between the second confidence threshold and the first confidence threshold, the action and entity guiding information is generated to guide the user to confirm the performed action and the entity of the performed action. For example, in the case that the received voice instruction is "please open the vehicle", that is, only entity words ("vehicle") exist in the extracted words, and the confidence of the entity words (0.7) is between the second confidence threshold (0.6) and the first confidence threshold (1), since the confidence of the entity word "vehicle" and "vehicle window" in the intent configuration table is 0.7, the action and entity guiding information such as "is it to open or close the vehicle window? You can try to tell me to open the vehicle window" can be generated to guide the user to confirm the entity ("vehicle window") and the action ("open" or "close") performed on the "vehicle window".

[0113] Optionally, in the case that the extracted words do not contain entity words, the guiding information is generated based on the missing target words in the extracted words, including:

[0114] (1) In the case that only position words exist in the extracted words, the action and entity guiding information is generated;

[0115] In the case that only the position word exists in the extracted words, the action and entity guide information is generated to guide the user to confirm the action performed and the entity on which the action is performed. For example, referring to FIG. 7, in the case that the received voice instruction is "main driver", i.e., only the position word ("main driver") exists in the extracted words, the action and entity guide information such as "what do you want to open the main driver? You can try to tell me to open the main driver's window" is generated to guide the user to confirm the instruction.

[0116] (2) In the case that only the verb exists in the extracted words, the position and entity guide information is generated.

[0117] In the case that only the verb exists in the extracted words, the position and entity guide information is generated to guide the user to confirm the entity on which the action is performed and the position of the entity. For example, referring to FIG. 8, in the case that the received voice instruction is "close", i.e., only the verb ("close") exists in the extracted words, the position and entity guide information such as "what do you want to close? You can try to tell me to close the main driver's window" is generated to guide the user to confirm the instruction.

[0118] (3) In the case that only the position word and the verb exist in the extracted words, the entity guide information is generated.

[0119] In the case that only the position word and the verb exist in the extracted words, the entity guide information is generated to guide the user to confirm the entity on which the action is performed. For example, referring to FIG. 9, in the case that the received voice instruction is "open the main driver for me", i.e., only the position word ("main driver") and the verb ("open") exist in the extracted words, the entity guide information such as "what do you want to open the main driver? You can try to tell me to open the main driver's window" is generated to guide the user to confirm the instruction.

[0120] Referring to FIG. 10, in some embodiments, the voice recognition result includes a classification result and extracted words, the classification result includes whether an intent exists, and the extracted words include at least one of an entity word, a verb, and a position word, and the method further includes:

[0121] Step 015: In the case that the recognition result meets a preset intent generation condition, generating an intent recognition result based on the voice recognition result;

[0122] The preset intent generation condition includes that the classification result is that the intent exists, the extracted words include the verb and the entity word, and the confidence of the entity word is greater than a first confidence threshold, or the classification result is that the intent exists, the extracted words include the verb, the position word, and the entity word, and the confidence of the entity word is greater than the first confidence threshold.

[0123] Specifically, in a case where the classification result is that there is an intent, the extracted words contain a verb and an entity word, and the confidence of the entity word is greater than the first confidence threshold (for example, in a case where the extracted words include "open sunroof", and the confidence of "sunroof" is greater than the first confidence threshold (0.9)); or in a case where the classification result is that there is an intent, the extracted words contain a verb, a direction word, and an entity word, and the confidence of the entity word is greater than the first confidence threshold (for example, in a case where the extracted words include "open driver's window", and the confidence of "window" is greater than the first confidence threshold (0.9)), it can be considered that the vehicle can explicitly express the intent, and therefore, in a case where the recognition result satisfies a preset intent generation condition, an intent recognition result can be generated based on the speech recognition result.

[0124] Referring to FIG. 11, in some embodiments, the recognition result includes a classification result and extracted words, the classification result includes whether there is an intent, and the extracted words include at least one of an entity word, a verb, and a direction word, and the method further includes:

[0125] Step 016: In a case where the recognition result satisfies a preset stopping condition, stopping intent recognition.

[0126] In a case where the recognition result satisfies the preset stopping condition, the recognition result satisfies at least one of the following:

[0127] (1) The classification result is that there is no intent.

[0128] In a case where the vehicle identifies the received speech instruction, the vehicle can classify whether the speech instruction has an intent. For example, in a scenario where the user is chatting on the vehicle, the user does not expect to control the vehicle. At this time, the vehicle classifies the received speech, and confirms that the classification result is that there is no intent, and in a case where the recognition result satisfies a preset stopping condition that the classification result is that there is no intent, the vehicle stops intent recognition.

[0129] (2) The extracted words do not contain an entity word, a verb, and a direction word.

[0130] In a case where the extracted words do not contain an entity word, a verb, and a direction word, it can be considered that the user does not need the vehicle to execute the corresponding instruction, and the vehicle can stop intent recognition.

[0131] (3) The confidence of the entity word in the extracted words is less than a third confidence threshold.

[0132] In a case where the confidence of the entity word in the extracted words is less than the third confidence threshold, it can be considered that there is a misjudgment of whether the speech instruction has an intent, or it can be considered that the user does not need the vehicle to execute the intent at this time, and the vehicle can stop intent recognition.

[0133] Referring to FIG. 12, to better implement the intent recognition method of the embodiments of the present application, the embodiments of the present application further provide an intent recognition device 300. The intent recognition device 300 can include a speech recognition module 301, a guiding module 302, and an intent recognition module 303. The speech recognition module 301 is configured to, in a case where a speech instruction is received, recognize the speech instruction to obtain a speech recognition result; the guiding module 302 is configured to, in a case where the speech recognition result satisfies a preset guiding condition, issue guiding information, the guiding information being generated based on the speech recognition result; and the intent recognition module 303 is configured to, in a case where feedback speech for the guiding information is received, recognize the feedback speech to obtain an intent recognition result.

[0134] In one embodiment, the speech recognition result includes extracted words, the extracted words including at least one of an entity word, a verb, and a position word, and the intent recognition device 300 can further include a first generation module 304, the first generation module 304 being configured to generate the guiding information based on at least one of a target word missing in the extracted words and a confidence degree of the entity word, the target word including at least one of the verb and the position word.

[0135] In one embodiment, the first generation module 304 is specifically configured to, in a case where the extracted words include the entity word, generate the guiding information based on the target word missing in the extracted words and the confidence degree of the entity word; and in a case where the extracted words do not include the entity word, generate the guiding information based on the target word missing in the extracted words.

[0136] In one embodiment, the guiding information includes position guiding information, action guiding information, action and position guiding information, entity guiding information, action and entity guiding information, and position and entity guiding information; and the first generation module 304 is specifically configured to, in a case where only the entity word exists in the extracted words and the confidence degree of the entity word is greater than a first confidence degree threshold, generate the action guiding information or the action and position guiding information; in a case where only the entity word and the verb exist in the extracted words and the confidence degree of the entity word is between a second confidence degree threshold and the first confidence degree threshold, generate the entity guiding information or the position and entity guiding information; in a case where only the entity word exists in the extracted words and the confidence degree of the entity word is between the second confidence degree threshold and the first confidence degree threshold, generate the action and entity guiding information; in a case where only the position word exists in the extracted words, generate the action and entity guiding information; in a case where only the verb exists in the extracted words, generate the position and entity guiding information; and in a case where only the position word and the verb exist in the extracted words, generate the entity guiding information.

[0137] In an embodiment, the speech recognition result comprises a classification result and extracted words, the classification result comprises whether an intent exists, and the extracted words comprise at least one of an entity word, a verb, and a directional word. The intent recognition apparatus 300 can further comprise a second generation module 305 configured to generate an intent recognition result based on the speech recognition result when the recognition result satisfies a preset intent generation condition. The intent generation condition comprises that the classification result is that an intent exists, the extracted words comprise a verb and an entity word, and the confidence of the entity word is greater than a first confidence threshold, or the classification result is that an intent exists, the extracted words comprise a verb, a directional word, and an entity word, and the confidence of the entity word is greater than the first confidence threshold.

[0138] In an embodiment, the recognition result comprises a classification result and extracted words, the classification result comprises whether an intent exists, and the extracted words comprise at least one of an entity word, a verb, and a directional word. The intent recognition apparatus 300 can further comprise a stopping module 306 configured to stop intent recognition when the recognition result satisfies a preset stopping condition. The recognition result satisfying the preset stopping condition comprises at least one of the following: the classification result is that an intent does not exist; the extracted words have no entity word, verb, or directional word; and the confidence of the entity word in the extracted words is less than a third confidence threshold.

[0139] In an embodiment, the speech recognition module 301 is specifically configured to recognize the speech instruction based on a preset deep learning model to obtain the speech recognition result.

[0140] The above describes the intent recognition apparatus 300 from the perspective of functional modules, which can be implemented in the form of hardware, implemented by instructions in the form of software, or implemented by a combination of hardware and software modules. Specifically, each step of the method embodiment in the embodiment of the present application can be completed by integrated logic circuits of hardware in a processor and / or instructions in the form of software. The steps of the method disclosed in the embodiment of the present application can be directly embodied as hardware coding processor execution completion, or executed by a combination of hardware and software modules in the coding processor. Alternatively, the software module can be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps in the above method embodiment.

[0141] The electronic device of the embodiment of the present application comprises a processor connected with a memory, the memory stores a computer program, and the processor executes the computer program to implement the intent recognition method described in any one of the above embodiments. For brevity, details are not repeated here.

[0142] Referring to FIG. 13, the electronic device can be a processor of the vehicle, and the electronic device can be installed in the vehicle to enable the vehicle to implement the intention recognition method of any one of the above embodiments.

[0143] The vehicle of the embodiments of the present application includes the intention recognition device or the electronic device of the above embodiments, and the intention recognition device or the electronic device is a processor of the vehicle. The vehicle implements the intention recognition method of any one of the above embodiments through the intention recognition device or the electronic device.

[0144] In one embodiment, the internal structure of the vehicle can be as shown in FIG. 13, including a processor 402, a memory 403, a network interface 404, a display screen 401 and an input device 405 connected through a system bus.

[0145] The processor 402 of the vehicle is configured to provide computing and control capabilities. The memory 403 of the vehicle includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface 404 of the vehicle is configured to connect with external devices through a network to recognize intentions. The computer program is executed by the processor to implement the intention recognition method of any one of the above embodiments. The display screen 401 of the vehicle can be a liquid crystal display screen or an electronic ink display screen, and the input device 405 of the vehicle can be a touch layer overlaid on the display screen 401, or a button, trackball or touchpad provided on the vehicle, or an external keyboard, touchpad or mouse, etc.

[0146] Those skilled in the art can understand that the structure shown in FIG. 13 is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0147] Referring to FIG. 14, the embodiments of the present application further provide a computer readable storage medium 600 having a computer program 610 stored thereon. When the computer program 610 is executed by a processor 620, the steps of the intention recognition method of any one of the above embodiments are implemented. For brevity, details are not repeated here.

[0148] In the description of the specification, the description of the terms "certain embodiments", "in an example", "exemplarily" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In the specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Also, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the specification and the features of the different embodiments or examples without contradiction.

[0149] Any process or method descriptions or descriptions of the flow diagrams in the flow charts described herein and elsewhere can be understood as representing code modules, segments, or portions of code which include one or more executable instructions for performing specific logic functions or steps in the process, and that the various systems described herein can include one or more circuits, circuitry, or other hardware for implementing the described functions or steps. The various systems described herein can form part of a machine in the form of a computer, embedded computer, arithmetical logic unit, or other device for example.

[0150] Although the embodiments of the present application have been shown and described above, it is understood that the above-described embodiments are exemplary and are not to be construed as limiting the present application, and the person skilled in the art can make changes, modifications, replacements and variations to the above-described embodiments within the scope of the present application.

Claims

1. An intention recognition method, wherein, Comprise: In the case of receiving a voice instruction, recognizing the voice instruction to obtain a voice recognition result; In the case that the voice recognition result meets a preset guidance condition, issuing guidance information, the guidance information being generated based on the voice recognition result; In the case of receiving feedback voice for the guidance information, recognizing the feedback voice to obtain an intent recognition result.

2. The intention recognition method according to claim 1, wherein The voice recognition result includes a classification result and an extracted word, the classification result includes whether there is an intent, and the extracted word includes at least one of an entity word, a verb, and a position word, the entity word is used to represent an object operated by the voice instruction, and the voice recognition result meets a preset guidance condition, including at least one of: The classification result is that there is an intent, there is only an entity word in the extracted word, and the confidence of the entity word is greater than a first confidence threshold; The classification result is that there is an intent, there is only an entity word and a verb in the extracted word, and the confidence of the entity word is between a second confidence threshold and the first confidence threshold, the second confidence threshold being less than the first confidence threshold; The classification result is that there is an intent, there is only an entity word in the extracted word, and the confidence of the entity word is between a second confidence threshold and the first confidence threshold, the second confidence threshold being less than the first confidence threshold; The classification result is that there is an intent, and there is only a position word in the extracted word; The classification result is that there is an intent, and there is only a verb in the extracted word; The classification result is that there is an intent, and there is only a position word and a verb in the extracted word.

3. The intention recognition method according to claim 1 or 2, wherein The voice recognition result includes an extracted word, the extracted word including at least one of an entity word, a verb, and a position word, and the method further comprises: Generating the guidance information based on at least one of a target word missing in the extracted word and a confidence of the entity word, the target word including at least one of a verb and a position word.

4. The intention recognition method according to claim 3, wherein The generating the guidance information based on at least one of a target word missing in the extracted word and a confidence of the entity word, comprises: In the case that the extracted word contains an entity word, generating the guidance information based on the target word missing in the extracted word and the confidence of the entity word; In the case that the extracted word does not contain an entity word, generating the guidance information based on the target word missing in the extracted word.

5. The intention recognition method according to claim 4, wherein The guidance information includes at least one of action guidance information, action and position guidance information, entity guidance information, action and entity guidance information, and position and entity guidance information; The generating the guidance information based on at least one of a target word missing in the extracted word and a confidence of the entity word in the case that the extracted word contains an entity word, comprises: In the case that there is only an entity word in the extracted word and the confidence of the entity word is greater than a first confidence threshold, generating the action guidance information or the action and position guidance information; In the case that there is only an entity word and a verb in the extracted word and the confidence of the entity word is between a second confidence threshold and the first confidence threshold, generating the entity guidance information or the position and entity guidance information; generate the action and entity guide information in a case where only entity words exist in the extracted words and a confidence of the entity words is between a second confidence threshold and the first confidence threshold; the generating the guide information based on a target word missing in the extracted words in a case where the extracted words do not contain entity words, comprises: generate the action and entity guide information in a case where only location words exist in the extracted words; generate the location and entity guide information in a case where only verbs exist in the extracted words; generate the entity guide information in a case where only location words and verbs exist in the extracted words.

6. The intent recognition method according to any one of claims 1-5, wherein, The speech recognition result comprises a classification result and extracted words, the classification result comprises whether an intent exists, and the extracted words comprise at least one of entity words, verbs, and location words, and the method further comprises: generate the intent recognition result based on the speech recognition result in a case where the recognition result meets a preset intent generation condition; The intent generation condition comprises that the classification result is that an intent exists, the extracted words contain verbs and entity words, and a confidence of the entity words is greater than a first confidence threshold, or the classification result is that an intent exists, the extracted words contain verbs, location words, and entity words, and the confidence of the entity words is greater than the first confidence threshold.

7. The intent recognition method according to any one of claims 1-6, wherein, The recognition result comprises a classification result and extracted words, the classification result comprises whether an intent exists, and the extracted words comprise at least one of entity words, verbs, and location words, and the method further comprises: stop intent recognition in a case where the recognition result meets a preset stop condition; The recognition result meets the preset stop condition comprises at least one of the following: the classification result is that an intent does not exist; the extracted words do not contain entity words, verbs, and location words; a confidence of the entity words in the extracted words is less than a third confidence threshold.

8. The intent recognition method of any one of claims 2-7, wherein, The confidence of the entity words is determined based on a first confidence and a second confidence, the first confidence is determined based on a maximum similarity between each matching entity word matched with the entity word and the entity word in a preset intent configuration table, and the second confidence is determined based on a similarity between a pinyin of the entity word and a pinyin of the matching entity word corresponding to the maximum similarity.

9. The intent recognition method according to any one of claims 1-8, wherein, The speech recognition result is obtained by recognizing the speech instruction based on a preset deep learning model. The speech recognition result is obtained by recognizing the speech instruction based on a preset deep learning model.

10. An electronic device, wherein, comprise: a processor connected with a memory; the memory stores a computer program, and the processor executes the computer program to implement instructions of the intent recognition method in any one of claims 1 to 9.

11. A vehicle, wherein, comprise: the electronic device in claim 10.

12. The non-transitory computer readable storage medium of claim 1, having stored thereon a computer program, wherein, The computer program is executed by the processor to implement the intent recognition method in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Airport service guiding method, system and device based on intention recognition

    CN115168563A

  • Non-instruction voice rejection method, vehicle-mounted voice recognition system and automobile

    CN115331656A

  • Information interaction method and system, electronic equipment, vehicle and storage medium

    CN116486803A

  • Voice interaction method and device, equipment and storage medium

    CN117524223A

  • Intention recognition method, electronic device, vehicle and medium

    CN118588065A