User intention recognition method, user intention recognition device and storage medium

By performing multi-dimensional recognition and correction of audio and video data, the problem that the video automatic response system fails to effectively utilize video data is solved, and the accuracy and interactive experience of user intention recognition are improved.

CN113903335BActive Publication Date: 2025-08-22IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111138281.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-27
Publication Date
2025-08-22
Estimated Expiration
2041-09-27

AI Technical Summary

Technical Problem

The existing video automatic response system fails to effectively use video data for user intention recognition, resulting in poor user interaction experience and unable to provide business value better than audio automatic response system.

Method used

By processing audio and video data, multiple sub-recognition results are generated, and the first and second sub-recognition results are selected using preset rules for correction, and user intention information is finally obtained, and multi-dimensional recognition and correction are combined with video and audio data.

Benefits of technology

It improves the accuracy of user intention recognition, provides services that are more in line with user needs, enhances the user interaction dimension, and provides additional service channels for people with language barriers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113903335B_ABST
    Figure CN113903335B_ABST
Patent Text Reader

Abstract

This application discloses a user intent recognition method, a user intent recognition device, and a computer-readable storage medium. The user intent recognition method includes: processing acquired audio and video data to generate a recognition result, the recognition result including multiple sub-recognition results; determining a first sub-recognition result and a second sub-recognition result from the multiple sub-recognition results; using the second sub-recognition result to correct the first sub-recognition result to generate a corrected recognition result; and obtaining user intent information based on the corrected recognition result. Through the above methods, the application can accurately identify user intent in multiple dimensions, improving the accuracy of intent recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method for identifying user intention, a device for identifying user intention, and a computer-readable storage medium. Background Art

[0002] The audio automatic answering system can recognize and semantically analyze the user's voice and obtain the user's intention. The video automatic answering system (Interactive Voice and Video Response, IVVR) in related technologies only adds video interaction to the audio automatic answering system to respond to user requests. The video data collected on the user side is idle and is not involved in user data analysis and response logic. It does not substantially improve the audio automatic answering system, and the user's interactive experience is not improved. How to use video resources and computer technology to give the video automatic answering system business value and user experience that are superior to the audio automatic answering system needs to be solved urgently. Summary of the Invention

[0003] The present application provides a user intention recognition method, a user intention recognition device and a computer-readable storage medium, which can accurately recognize user intention in multiple dimensions and improve the accuracy of intention recognition.

[0004] In order to solve the above technical problems, the technical solution adopted in this application is: to provide a user intention recognition method, which includes: processing the acquired audio and video data to generate a recognition result, and the recognition result includes multiple sub-recognition results; determining a first sub-recognition result and a second sub-recognition result from the multiple sub-recognition results; using the second sub-recognition result to correct the first sub-recognition result to generate a corrected recognition result; and obtaining user intention information based on the corrected recognition result.

[0005] In order to solve the above technical problems, another technical solution adopted in this application is: to provide a user intention recognition device, which includes a memory and a processor connected to each other, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the user intention recognition method in the above technical solution.

[0006] To solve the above technical problems, another technical solution adopted in this application is: providing a computer-readable storage medium, which is used to store a computer program. When the computer program is executed by a processor, it is used to implement the user intention recognition method in the above technical solution.

[0007] Through the above scheme, the beneficial effect of the present application is: first obtain audio and video data, then identify and process the audio and video data to generate an identification result including multiple sub-identification results; then determine the first sub-identification result and the second sub-identification result from the multiple sub-identification results, the first sub-identification result is the sub-identification result that needs to be corrected, and the second sub-identification result is the sub-identification result used to correct the first sub-identification result; then use the second sub-identification result to correct the first sub-identification result to generate a corrected identification result; then use the corrected identification result to know the user's intention; because the user's intention is multi-facetedly identified using video data and audio data and is corrected after identification, more accurate user intention information can be obtained based on the corrected identification result, which can improve the accuracy of intention recognition so as to provide users with services that better meet their own needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. Among them:

[0009] Figure 1 This is a flowchart of an embodiment of a method for identifying user intent provided by this application;

[0010] Figure 2 This is a flowchart of another embodiment of the user intention recognition method provided by the present application;

[0011] Figure 3 This is a flow chart of an embodiment of a correction method provided by the present application;

[0012] Figure 4 is a flow chart of another embodiment of the correction method provided by the present application;

[0013] Figure 5 This is a structural diagram of an embodiment of a user intention recognition device provided by the present application;

[0014] Figure 6 It is a structural diagram of an embodiment of a computer-readable storage medium provided by this application. DETAILED DESCRIPTION

[0015] The present application will be further described in detail below in conjunction with the accompanying drawings and examples. It is particularly noted that the following examples are only intended to illustrate the present application and are not intended to limit the scope of the present application. Similarly, the following examples are only some examples of the present application and not all examples. All other examples obtained by those of ordinary skill in the art without creative work are intended to fall within the scope of protection of this application.

[0016] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0017] It should be noted that the terms "first", "second" and "third" in this application are only used for descriptive purposes and should not be understood as indicating or suggesting relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first", "second" and "third" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units that are inherent to these processes, methods, products or devices.

[0018] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of a method for identifying user intent provided by the present application, which includes:

[0019] Step 11: Process the acquired audio and video data to generate recognition results.

[0020] Audio and video data may include video data and audio data. A camera can be used to shoot the target in the current monitoring scene to obtain video data, such as the user's face or motion image; an audio acquisition device can be used to collect sound to obtain audio data, such as the user's interactive voice; and then the obtained video data and audio data are processed to generate recognition results.

[0021] Furthermore, the recognition result includes multiple sub-recognition results, which may include video recognition results and audio recognition results. Video data can be processed to obtain a video recognition result, for example, the collected user's motion information can be recognized to obtain a corresponding motion recognition result, in which case the motion recognition result is a sub-recognition result. Audio data can be processed to obtain an audio recognition result, for example, the collected user's voice information can be recognized to obtain a corresponding voice recognition result, in which case the voice recognition result is also a sub-recognition result. It is understandable that the above description only uses motion information as an example. In other embodiments, the user's facial information or gesture information can also be recognized based on video data.

[0022] Step 12: Determine a first sub-recognition result and a second sub-recognition result from the multiple sub-recognition results.

[0023] The first sub-recognition result and the corresponding second sub-recognition result can be determined according to preset rules. The preset rules may be related to the interaction mode selected by the user. The interaction mode can be selected by the user independently, and the preset rules can be set according to actual conditions; for example: assuming that the user selects the voice interaction mode, then the recognized voice recognition result can be selected as the first sub-recognition result, and other sub-recognition results other than the voice recognition result can be used as the second sub-recognition result, such as: using the action recognition result as the second sub-recognition result, so as to use the voice recognition result and the action recognition result to analyze the user intention, so that the analyzed user intention is more accurate.

[0024] In a specific embodiment, the user can select the desired interaction mode by inputting interaction selection instructions, and thus express the needs in the selected interaction mode in a corresponding manner; specifically, the interaction selection instructions may include voice interaction instructions and action interaction instructions, and the interaction mode may accordingly include voice interaction mode and action interaction mode. The user can be guided to input the interaction selection instructions by playing guidance audio / guidance video to select the interaction mode, and then after the user selects the corresponding interaction mode, the user is guided to input the corresponding voice or perform the corresponding action according to the specification, wherein the voice interaction mode is a mode in which the user interacts by inputting voice, and the action interaction mode is a mode in which the user interacts by performing gestures / actions; it can be understood that when selecting the interaction mode, the user can also make a choice in a variety of interaction methods such as voice / gestures / actions according to the content of the guidance audio / guidance video, thereby increasing the interaction dimension with the user in all directions, and also providing a comprehensive service channel for people with language barriers.

[0025] Furthermore, when the user selects the voice interaction mode, the user can be guided to input the corresponding voice according to the guide word format. For example, the guide word for searching for a location can be "I want to go to XX", and the user can express their needs according to the guide word format. For example, the user can say "I want to go to the parking lot" to find the location information of a nearby parking lot. When the user selects the motion interaction mode, the user can be guided in the standardization of body movements, and the user can be guided to make corresponding gestures, actions, or express their needs through sign language. This embodiment is mainly described in terms of voice interaction mode and motion interaction mode. The interaction mode is not limited to the above two interaction modes. For example, it can also include a hybrid interaction mode that realizes interaction by inputting voice and performing gestures / actions at the same time. The specific setting can be made according to the actual situation.

[0026] Step 13: Use the second sub-recognition result to correct the first sub-recognition result to generate a corrected recognition result.

[0027] Using the second sub-recognition result to correct the first sub-recognition result can generate a more accurate corrected recognition result, which can improve the accuracy of the recognition result; for example, the first sub-recognition result is selected as the voice recognition result, and the second sub-recognition result is selected as the action recognition result. At this time, the action recognition result can be used to correct the voice recognition result. When the voice recognition result is inaccurate, a more accurate corrected recognition result can be obtained by correcting the action recognition result.

[0028] It is understandable that in other specific embodiments, two or more second sub-recognition results may be selected to correct the first sub-result, for example, using lip recognition results and gesture recognition results to correct the speech recognition result.

[0029] Step 14: Based on the corrected recognition results, obtain user intent information.

[0030] The user intention information is obtained from the corrected recognition result. For example, the corrected recognition result obtained based on the voice recognition result and the action recognition result is "I want to go to the parking lot". The user's intention information is obtained based on the corrected recognition result, and it is known that the user wants to obtain the location information of nearby parking lots.

[0031] In this embodiment, audio and video data are first obtained, and then the audio and video data are recognized and processed to generate multiple sub-recognition results including audio recognition results and video recognition results; then a first sub-recognition result and a second sub-recognition result are determined from the multiple sub-recognition results, the first sub-recognition result is the sub-recognition result that needs to be corrected, and the second sub-recognition result is the sub-recognition result used to correct the first sub-recognition result; the second sub-recognition result is then used to correct the first sub-recognition result to generate a corrected recognition result; then the corrected recognition result can be used to know the user's intention; since video data and audio data are used for multi-faceted recognition and correction is performed after recognition, more accurate user intention information can be obtained based on the corrected recognition result, which can improve the accuracy of intention recognition so as to provide users with services that better meet their needs.

[0032] See also Figure 2 , Figure 2 : is a flow chart of another embodiment of the user intention recognition method provided by the present application, the method comprising:

[0033] Step 21: Process the acquired audio and video data to generate recognition results.

[0034] Audio and video data may include video data and audio data, and the recognition results include multiple sub-recognition results. Specifically, the multiple sub-recognition results include lip recognition results, voice recognition results or motion recognition results. The voice information in the voice data can be recognized and processed to obtain a voice recognition result, and the motion information in the video data can be recognized and processed to obtain a motion recognition result; and / or, the lip information can be recognized and processed to obtain a lip recognition result.

[0035] Step 22: Determine whether an interactive selection instruction is received.

[0036] After identifying multiple sub-recognition results, it is necessary to determine the first sub-recognition result and the second sub-recognition result from the multiple sub-recognition results, and then use the second sub-recognition result to correct the first sub-recognition result to generate a corrected recognition result; specifically, it is possible to first determine whether an interactive selection instruction input by the user is received, and then select the corresponding first sub-recognition result and the second sub-recognition result based on the interactive selection instruction.

[0037] In a specific embodiment, the received interactive selection instruction is parsed to generate selection information; based on the selection information, a first sub-recognition result is determined, and the first sub-recognition result is the sub-recognition result corresponding to the interactive selection instruction; the second sub-recognition result is determined according to a preset rule, and the second sub-recognition result is a sub-recognition result other than the first sub-recognition result among multiple sub-recognition results; for example, when a voice interaction instruction input by a user is received, the voice interaction instruction is parsed to generate corresponding selection information, and it is determined that the user has selected a voice interaction mode. At this time, based on the selection result of the voice interaction mode, the first sub-recognition result is determined as a voice recognition result, and the second sub-recognition result is a sub-recognition result other than the voice recognition result. At this time, the second sub-recognition result can be determined according to the preset rules, and then the first sub-recognition result is corrected using the second sub-recognition result to generate a corrected recognition result; further, the preset rules and the correction method are related to the interaction mode selected by the user. The following steps 23 to 24 are the correction methods used in the action interaction mode / voice interaction mode / no interaction mode.

[0038] Furthermore, if an interactive selection instruction is received, the first sub-recognition result is an action recognition result / speech recognition result, and the second sub-recognition result is a lip shape recognition result; or the first sub-recognition result is a speech recognition result, and the second sub-recognition result is an action recognition result; or the first sub-recognition result is an action recognition result, and the second sub-recognition result is a speech recognition result. Specifically, the specific information of the first sub-recognition result and the second sub-recognition result can be determined by analyzing the interactive selection instruction; for example, four operation buttons can be set on the user interaction interface: a first recognition button, a second recognition button, a third recognition button, and a fourth recognition button. The voice interaction mode is related to the first recognition button and the second recognition button. When the first recognition button is triggered, it indicates that the user wants to realize intention recognition through the "voice-based + lip shape-assisted" method. At this time, the first sub-recognition result is a speech recognition result, and the second sub-recognition result is a lip shape recognition result; when the second recognition button is triggered, it indicates that the user wants to realize intention recognition through the "voice-based + lip shape-assisted" method. As a supplementary method, the first sub-recognition result is the voice recognition result, and the second sub-recognition result is the action recognition result; the action interaction mode is related to the third recognition button and the fourth recognition button. When the third recognition button is triggered, it indicates that the user wants to realize the intention recognition through the "action as the main + lip shape as the supplementary" method. At this time, the first sub-recognition result is the action recognition result, and the second sub-recognition result is the lip shape recognition result; when the fourth recognition button is triggered, it indicates that the user wants to realize the intention recognition through the "action as the main + voice as the supplementary" method. At this time, the first sub-recognition result is the action recognition result, and the second sub-recognition result is the voice recognition result.

[0039] Step 23: When the interactive selection instruction is received, the action recognition result / speech recognition result is corrected based on the lip recognition result to obtain a corrected recognition result.

[0040] When the user selects the motion interaction mode / voice interaction mode, the first sub-recognition result is the motion recognition result / voice recognition result, and the second sub-recognition result is the lip recognition result. At this time, the motion recognition result / voice recognition result is corrected based on the lip recognition result, such as Figure 3 The specific steps are shown in steps 231a to 234a below:

[0041] Step 231a: Match the lip recognition result with the preset knowledge base to obtain a matching degree.

[0042] The lip shape recognition result can be an ordered set of words. By recognizing the facial lip shape in the video data, words such as people or places can be identified. For example, if the user says "I want to go to the basketball hall", the corresponding keyword "basketball hall" can be identified based on the lip shape data. The speech recognition results and action recognition results are both sentence results, that is, they can recognize complete sentences, for example: "I want to go to the basketball hall."

[0043] Furthermore, the lip recognition results are matched with a preset knowledge base. The preset knowledge base may be a knowledge base containing social hot words, names or professional terms. The words in the lip recognition results may be matched with the words in the preset knowledge base for similarity, and the matching degree corresponding to the words in the lip recognition results is obtained according to the matching algorithm in the relevant technology. For example, if the lip recognition result is "Jong Kook", the name "Kim Jong Kook" can be matched in the preset knowledge base, and the cosine similarity between the two is calculated as the matching degree.

[0044] Step 232a: Replace the corresponding words in the speech recognition result / action recognition result with the lip recognition result to obtain a new recognition result, and calculate the fit of the new recognition result.

[0045] The words in the lip recognition results can be replaced with the corresponding word positions in the complete sentence of the speech recognition result / action recognition result based on the grammatical attributes of the words in the lip recognition results. For example, the lip recognition result is "cola", which is a noun, and the sentence produced by the speech recognition result / action recognition result is "I want to drink milk tea". Therefore, "cola" can be replaced with the name "milk tea" in the sentence to obtain a new recognition result "I want to drink cola". Then, the fitting algorithm in the relevant technology can be used to calculate the fitting of the new recognition result. It can be understood that this embodiment only uses the replacement based on the grammatical attributes of words as an example. In other embodiments, the replacement operation can also be performed by comprehensively analyzing the order or semantic attributes of the words in the lip recognition results.

[0046] Step 233a: Determine whether the product of the matching degree and the fitting degree is greater than a preset threshold.

[0047] After obtaining the matching degree of the lip recognition result and the fitting degree of the new recognition result, the matching degree and the fitting degree are multiplied and then compared with the preset threshold. Specifically, the preset threshold is a pre-set value, which can be a value set based on experience or the optimal value debugged by the deep learning model. The accuracy of the new recognition result can be judged by comparing the product of the matching degree and the fitting degree with .

[0048] Step 234a: If the product of the matching degree and the fitting degree is greater than a preset threshold, the new recognition result is used as the revised recognition result.

[0049] Taking the preset threshold of 0.7 as an example, if the product of the matching and fitting degrees is greater than 0.7, it means that the new recognition result obtained by correcting the lip recognition result is relatively accurate and can be used as the corrected recognition result. If the product of the matching and fitting degrees is less than or equal to 0.7, it means that the new recognition result is inaccurate. In this case, the original speech recognition result / action recognition result will not be corrected and will still be used as the final recognition result.

[0050] In other specific embodiments, the correction method may also be: when the first sub-recognition result is a speech recognition result and the second sub-recognition result is a motion recognition result, the speech recognition result is corrected based on the motion recognition result; or when the first sub-recognition result is a motion recognition result and the second sub-recognition result is a speech recognition result, the motion recognition result is corrected based on the speech recognition result. At this time, the speech recognition result and the motion recognition result can be directly compared, and the more accurate one can be selected as the corrected recognition result. Specifically, Figure 4 As shown in steps 231b to 234b:

[0051] Step 231b: Score the action recognition result and the speech recognition result respectively to obtain a first score value and a second score value.

[0052] A scoring model can be used to score the action recognition results / speech recognition results. The complete sentence of the action recognition result / speech recognition result is input into the scoring model to obtain a first scoring value corresponding to the action recognition result and a second scoring value corresponding to the speech recognition result. Specifically, the scoring model can perform a weighted operation on the location, person, behavior, or transaction in the complete sentence to obtain a separate scoring value, and then add up each scoring value to obtain a final scoring value for the sentence.

[0053] Step 232b: Determine whether the first score value is greater than the second score value.

[0054] The first score value is compared with the second score value to determine which of the action recognition result and the speech recognition result is more accurate.

[0055] Step 233b: If the first score value is greater than the second score value, the recognition result is modified to be an action recognition result.

[0056] If the first score is greater than the second score, it indicates that the action recognition result corresponding to the first score is more accurate than the speech recognition result corresponding to the second score. In this case, the modified recognition result is determined as the action recognition result. For example, if the speech recognition result is "I want to go to the Netherlands" and the action recognition result is "I want to go to Henan," and the action recognition result has a higher score, the action recognition result can be output as "I want to go to Henan" as the modified recognition result.

[0057] Step 234b: If the first score value is less than or equal to the second score value, the recognition result is modified to be a speech recognition result.

[0058] If the first score is less than or equal to the second score, the speech recognition result corresponding to the second score is more accurate than the action recognition result corresponding to the first score. In this case, the revised recognition result is determined as the speech recognition result. For example, if the action recognition result is "I want to go to Henan" and the speech recognition result is "I want to go to the Netherlands," and the speech recognition result has a higher score, the speech recognition result may be output as "I want to go to the Netherlands" as the revised recognition result.

[0059] Step 24: If no interactive selection instruction is received, the action recognition result is corrected based on the lip recognition result and the voice recognition result to obtain a corrected recognition result.

[0060] When it is determined that no interactive selection instruction is received, that is, the user has not selected the interactive mode, since the action recognition result is based on the user's action / gesture / sign language, its accuracy is generally higher than the voice recognition result / lip recognition result. At this time, the first sub-recognition result is the action recognition result, and the second sub-recognition result is the lip recognition result and the voice recognition result. The action recognition result is corrected using the lip recognition result and the voice recognition result to obtain a corrected recognition result. It can be understood that in other embodiments, the voice recognition result can also be corrected based on the lip recognition result and the action recognition result, which is not limited here.

[0061] Specifically, the action recognition result is corrected using the speech recognition result to obtain an intermediate corrected result, wherein the method of correcting the action recognition result using the speech recognition result is the same as the above-mentioned method of correcting the action recognition result based on the speech recognition result, and will not be repeated here. The intermediate corrected result is the one with the higher score between the action recognition result and the speech recognition result; then the intermediate corrected result is corrected using the lip recognition result to obtain a corrected recognition result, wherein the method of correcting the intermediate corrected result (i.e., speech recognition result / action recognition result) using the lip recognition result is the same as the above-mentioned method of correcting the action recognition result / speech recognition result based on the lip recognition result, and will not be repeated here.

[0062] The corrected recognition results obtained according to the above correction method are all complete sentences. It can be seen that they only contain the literal meaning expressed by the sentence structure and do not involve semantic expression. At this time, the user's emotions can be identified based on the acquired video data, and then the emotion recognition results and the corrected recognition results can be semantically analyzed to finally analyze the user's intention. Since analyzing the user's intention based only on the sentence structure without combining semantic expression may be biased and the user's true intention cannot be obtained, the method of involving emotions in semantic analysis can make the user intention analysis more accurate.

[0063] Step 25: Perform emotion recognition processing on the video data to obtain emotion recognition results.

[0064] Emotion recognition processing is performed based on the acquired video data to obtain an emotion recognition result. Specifically, emotion recognition can be performed based on the user's face / tone of voice to identify the user's emotions such as anger, calmness, or joy.

[0065] Step 26: Input the corrected recognition result and the emotion recognition result into the semantic analysis engine to obtain user intent information.

[0066] After obtaining the emotion recognition result, the corrected recognition result and the emotion recognition result are input into the semantic analysis engine together to analyze and obtain the user intention information. For example, the corrected recognition result is "I want to drink a drink". The sentence structure simply expresses that the user wants to drink a drink, that is, to query information related to the drink. However, after identifying the user's emotion, it is found that the user is in an angry mood, which means that the user's true intention at this time may not be to query information related to the drink. By using the semantic analysis engine to analyze the user intention information together with the corrected recognition result and the emotion recognition result, it is finally analyzed that the user actually wants to complain. The deviation between the user's true intention and the corrected recognition result is corrected by the participation of the emotion recognition result in the intention analysis. Specifically, the semantic analysis engine can use the corrected recognition result as input, add the vector of the emotion recognition result and assign weights to it, and then perform similarity matching on the corrected recognition result with the added emotion recognition result vector and the original word model to obtain the user intention information.

[0067] Step 27: Based on the user intention information, query the preset database to obtain query data; or based on the user intention information, generate response information, and play and / or display the response information.

[0068] The preset database can be a database containing information about life, society, film and television. When the user has a query demand, that is, when the user intention information is the query data type, for example: querying hotels or restaurants, etc., the user intention information can be queried in the preset database to obtain the corresponding query data; for example, when the user queries a certain movie, the corresponding movie information can be queried in the preset database, and the movie can be played to the user through the display interface. When the user queries a certain song, the corresponding song information can be queried in the preset database, and the song can be played to the user through the display interface. When the user queries a hotel, the corresponding hotel information can be queried in the preset database, and the nearby hotel information can be displayed to the user through the display interface.

[0069] The above-mentioned user intention recognition method can be applied to scenarios such as intelligent robots, video customer service or live broadcast. In a specific embodiment, in addition to the intention to query data, the user also has the need to chat. At this time, a chat service can be provided to the user based on the chat intention, and the built-in knowledge base can be used to select appropriate response information according to the context or semantics. For example, if the user says "Hello, the weather is very good today", the corresponding response information "Hello, today is a good day for traveling" can be generated and played and / or displayed; it can be understood that in other embodiments, when it is identified that the user is in a negative emotion (for example: angry or sad, etc.), corresponding soothing response information can also be generated according to the user's emotions to soothe the user's emotions.

[0070] In this embodiment, when an interactive selection instruction is received, the action recognition result / speech recognition result is corrected based on the lip recognition result; if no interactive selection instruction is received, the action recognition result is corrected based on the lip recognition result and the speech recognition result, thereby obtaining a corrected recognition result; then, the corrected recognition result and the user's emotion recognition result are used to obtain user intent information, thereby responding to the user intent information to meet the user's needs. By performing multi-dimensional analysis on audio data and video data, the utilization of user video resources is greatly improved, breaking through the upper limit of the accuracy of identifying user intent through language and acoustic analysis alone, and increasing the user's interactive dimension. At the same time, the added action interaction method also provides a service channel for people with language barriers, making it easier for these users to use, and increasing the reference source of user intent analysis in the visual dimension. At the same time, the emotional content generated by the user's expression and tone of voice is used to assist in user intent recognition, thereby improving the accuracy of intent analysis and providing users with better services.

[0071] See also Figure 5 , Figure 5 It is a structural diagram of an embodiment of a user intention recognition device provided in the present application. The user intention recognition device 50 includes a memory 51 and a processor 52 connected to each other. The memory 51 is used to store a computer program. When the computer program is executed by the processor 52, it is used to implement the user intention recognition method in the above embodiment.

[0072] See also Figure 6 , Figure 6 This is a structural diagram of an embodiment of a computer-readable storage medium provided in the present application. The computer-readable storage medium 60 is used to store a computer program 61. When the computer program 61 is executed by the processor, it is used to implement the user intention recognition method in the above embodiment.

[0073] The computer-readable storage medium 60 can be a server, a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which can store program codes.

[0074] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is merely a logical functional division. In actual implementation, other division methods may be used. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not implemented.

[0075] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0076] In addition, each functional unit in each embodiment of the present application may be integrated into a processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The above-mentioned integrated units may be implemented in the form of hardware or software functional units.

[0077] The above description is merely an embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method for identifying user intention, characterized in that: include: Processing the acquired audio and video data to generate a recognition result, wherein the recognition result includes multiple sub-recognition results; determining a first sub-recognition result and a second sub-recognition result from the plurality of sub-recognition results; Correcting the first sub-recognition result using the second sub-recognition result to generate a corrected recognition result; Obtaining user intention information based on the corrected recognition result; The multiple sub-recognition results include lip recognition results, speech recognition results, or action recognition results. The step of using the second sub-recognition result to correct the first sub-recognition result to generate a corrected recognition result includes: Determining whether an interactive selection instruction is received; If not, the first sub-recognition result is the action recognition result, and the second sub-recognition result is the lip recognition result and the voice recognition result. Based on the lip recognition result and the voice recognition result, the action recognition result is corrected to obtain the corrected recognition result.

2. The user intention recognition method according to claim 1, characterized in that: The step of correcting the action recognition result based on the lip recognition result and the speech recognition result includes: Correcting the action recognition result using the speech recognition result to obtain an intermediate correction result; The intermediate correction result is corrected using the lip recognition result to obtain the corrected recognition result.

3. The user intention recognition method according to claim 1, characterized in that: The method further comprises: If the interactive selection instruction is received, the first sub-recognition result is the action recognition result / the voice recognition result, and the second sub-recognition result is the lip recognition result; or the first sub-recognition result is the voice recognition result, and the second sub-recognition result is the action recognition result; or the first sub-recognition result is the action recognition result, and the second sub-recognition result is the voice recognition result.

4. The method for identifying user intention according to claim 3, wherein: The step of correcting the action recognition result / the speech recognition result based on the lip recognition result includes: Matching the lip recognition result with a preset knowledge base to obtain a matching degree; Replacing corresponding words in the speech recognition result / action recognition result with the lip recognition result to obtain a new recognition result, and calculating the fitting degree of the new recognition result; Determining whether the product of the matching degree and the fitting degree is greater than a preset threshold; If so, the new recognition result is used as the revised recognition result.

5. The method for identifying user intention according to claim 3, wherein: The method further comprises: Scoring the action recognition result and the speech recognition result respectively to obtain a first scoring value and a second scoring value; Determining whether the first score value is greater than the second score value; If so, the modified recognition result is the action recognition result; If not, the modified recognition result is the speech recognition result.

6. The method for identifying user intention according to claim 1, wherein: The audio and video data includes video data and audio data, the multiple sub-recognition results include video recognition results and voice recognition results, and the method further includes: Using a camera to shoot a target in a current monitoring scene to obtain the video data, and performing recognition processing on the video data to obtain the video recognition result; The audio data is collected by an audio collection device, and the audio data is recognized and processed to obtain the speech recognition result.

7. The method for identifying user intention according to claim 1, wherein: The method further comprises: Performing emotion recognition processing on the video data to obtain an emotion recognition result; Inputting the modified recognition result and the emotion recognition result into a semantic analysis engine to obtain the user intention information; Based on the user intention information, query a preset database to obtain query data; or Based on the user intention information, response information is generated, and the response information is played and / or displayed.

8. A user intention recognition device, characterized in that: The invention comprises a memory and a processor connected to each other, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, it is used to implement the user intention recognition method according to any one of claims 1 to 7.

9. A computer-readable storage medium for storing a computer program, characterized in that: When the computer program is executed by a processor, it is used to implement the user intention recognition method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice emotion interaction method, computer equipment and computer readable storage medium

    CN110085221A

  • Expression judgment voice recognition method based on face recognition, server and air conditioner

    CN112687260A