User intention recognition method and device based on multi-modal judgment and storage medium

By obtaining user behavior status information and spatial information, combined with the functions of intelligent interactive devices, the target user intention is generated, and the problem of low user intention recognition efficiency in smart homes is solved, and efficient and accurate intention recognition is achieved.

CN120234643APending Publication Date: 2025-07-01QINGDAO HAIER TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311851328.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In smart home scenarios, it is difficult for current technology to efficiently and accurately identify the user's true intentions, and usually requires multiple rounds of interaction, which reduces the recognition efficiency and affects the user experience.

Method used

By calling the target intelligent interactive device, the user's behavior status information is obtained, the initial intention is determined, and the matching intention type is determined based on the spatial information in which the user is located, and the target user's intention is finally generated that matches the behavior status information.

Benefits of technology

It realizes efficient and accurate identification of users' true intentions while reducing interaction rounds, improves recognition efficiency and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234643A_ABST
    Figure CN120234643A_ABST
Patent Text Reader

Abstract

The invention discloses a user intention recognition method and device based on multi-modal judgment and a storage medium, and relates to the technical field of smart home, and the user intention recognition method based on multi-modal judgment comprises the steps: calling target intelligent interaction equipment, obtaining the behavior state information of a user, determining an initial intention of the user according to the behavior state information; determining information of a space where the user is located; based on the spatial information, determining an intention type matched with the spatial information; and based on the intention type and the initial intention, generating a target user intention matched with the behavior state information. The real intention of the user can be efficiently and accurately recognized on the premise that the interaction turns are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of smart home, and particularly to a method, device and storage medium for user intention recognition based on multi-modal decision-making. Background Art

[0002] In the smart home scenario, it is often necessary to obtain a conversation interaction with the user to obtain the user's intention.

[0003] However, the same sentence spoken by the user may contain different intentions. Currently, in order to clarify the user's true intention, multiple rounds of interaction with the user are often required, which reduces the recognition efficiency of the user's true intention, and the multiple rounds of interaction also affect the user experience.

[0004] Therefore, finding an efficient and accurate method for recognizing user intentions has become a current research hotspot. Summary of the Invention

[0005] This application provides a method, device and storage medium for user intention recognition based on multi-modal decision-making, which can efficiently and accurately recognize the user's true intention.

[0006] This application provides a method for user intention recognition based on multi-modal decision-making. The method includes: calling a target intelligent interaction device to obtain the user's behavior state information, and determining the user's initial intention according to the behavior state information; determining the spatial information where the user is located; based on the spatial information, determining an intention type matching the spatial information; and generating a target user intention matching the behavior state information based on the intention type and the initial intention.

[0007] According to the method for user intention recognition based on multi-modal decision-making provided by this application, specifically, determining the user's initial intention according to the behavior state information includes: parsing the behavior state information to obtain a parsing result, and determining multiple initial intentions of the user according to the parsing result; specifically, generating a target user intention matching the behavior state information based on the intention type and the initial intention includes: screening a target initial intention matching the intention type from multiple initial intentions according to the intention type; and using the screened target initial intention matching the intention type as the target user intention matching the behavior state information.

[0008] A user intention recognition method based on multimodal decision provided by the present application, wherein the behavior state information includes preset round interaction dialogue information initiated by the user to the target intelligent interaction device, and the number of the preset rounds is less than or equal to a number threshold; after using the filtered target initial intention that matches the intention type as the target user intention that matches the behavior state information, the method further includes: generating a running instruction of the target intelligent interaction device that matches the target user intention based on the target user intention; calling the target intelligent interaction device to run according to the running instruction to respond to the preset round interaction dialogue information initiated by the user to the target intelligent interaction device.

[0009] A user intention recognition method based on multimodal decision provided by the present application, wherein the behavior state information includes a sequence of behavior actions executed by the user; after using the filtered target initial intention that matches the intention type as the target user intention that matches the behavior state information, the method further includes: generating a running instruction of the target intelligent interaction device that matches the target user intention based on the target user intention; calling the target intelligent interaction device to respond to the running instruction to actively respond to the sequence of behavior actions executed by the user.

[0010] A user intention recognition method based on multimodal decision provided by the present application, determining the spatial information where the user is located is achieved by the following method: determining the spatial information where the user is located based on the location information of the target intelligent interaction device, where the target intelligent interaction device and the user are in the same space, and / or calling the target intelligent interaction device to wake up other intelligent interaction devices under the same local area network as the target intelligent interaction device, where the other intelligent interaction devices are equipped with image acquisition devices; collecting an environmental image of the space where the user is located based on the image acquisition devices of the other intelligent interaction devices; identifying the spatial information where the user is located according to the environmental image.

[0011] A user intention recognition method based on multimodal decision provided by the present application, determining the location information of the target intelligent interaction device is achieved by the following method: determining the device type of the target intelligent interaction device; determining the location information of the target intelligent interaction device according to the device type.

[0012] A user intention recognition method based on multimodal decision provided by the present application. Before determining the intention type that matches the spatial information based on the spatial information, the method further includes: obtaining multiple sets of historical interaction intentions of the user in multiple spaces, where the historical interaction intentions are interaction intentions determined during the prior interaction process between the user and the intelligent interaction device; fitting and constructing a spatial interaction intention type mapping table according to the multiple sets of historical interaction intentions, where the spatial interaction intention type mapping table includes the corresponding relationships formed between different spaces of the user and different interaction intention types; the determining the intention type that matches the spatial information based on the spatial information specifically includes: determining the intention type that matches the spatial information based on the corresponding relationship and the spatial information.

[0013] The present application also provides a user intention recognition device based on multimodal decision. The device includes: an acquisition module, configured to call a target intelligent interaction device, obtain the behavior state information of the user, and determine the initial intention of the user according to the behavior state information; a determination module, configured to determine the spatial information where the user is located; a processing module, configured to determine the intention type that matches the spatial information based on the spatial information; a generation module, configured to generate a target user intention that matches the behavior state information based on the intention type and the initial intention.

[0014] The present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute to implement the user intention recognition method based on multimodal decision as described in any one of the above.

[0015] The present application also provides a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it executes to implement the user intention recognition method based on multimodal decision as described in any one of the above.

[0016] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the user intention recognition method based on multimodal decision as described in any one of the above.

[0017] The user intention recognition method, device, and storage medium based on multimodal decision provided by this application call a target intelligent interaction device to obtain the user's behavior status information, and determine the user's initial intention based on the behavior status information; then determine the spatial information where the user is located, and based on the spatial information, determine the intention type that matches the spatial information, and generate a target user intention that matches the behavior status information based on the intention type and the initial intention. In this application, by sensing the spatial information where the user is located, the ambiguity in the user's initial intention can be eliminated, and the target user intention can be efficiently and accurately recognized on the premise of reducing the interaction rounds. Description of the Drawings

[0018] The drawings here are incorporated into the specification and form a part of this specification, showing the embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a schematic diagram of the hardware environment of a user intention recognition method based on multimodal decision according to an embodiment of this application;

[0021] Figure 2 It is one of the schematic flowcharts of the user intention recognition method based on multimodal decision provided by this application;

[0022] Figure 3 It is a schematic flowchart of generating a target user intention that matches the behavior status information based on the intention type and the initial intention provided by this application;

[0023] Figure 4 It is another schematic flowchart of the user intention recognition method based on multimodal decision provided by this application;

[0024] Figure 5 It is a schematic diagram of the application scenario of the user intention recognition method based on multimodal decision provided by this application;

[0025] Figure 6 It is a schematic diagram of the structure of the user intention recognition device based on multimodal decision provided by this application;

[0026] Figure 7 It is a schematic diagram of the structure of the electronic device provided by this application. Detailed Embodiments

[0027] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0029] According to one aspect of the embodiments of this application, a user intention recognition method based on multimodal decision is provided. The user intention recognition method based on multimodal decision is widely applied to whole-house intelligent digital control application scenarios such as Smart Home, smart home, intelligent home device ecosystem, and Intelligence House ecosystem. Optionally, in this embodiment, the above-mentioned user intention recognition method based on multimodal decision can be applied to, for example, Figure 1 the hardware environment composed of the terminal device 102 and the server 104 as shown. As Figure 1 shown, the server 104 is connected to the terminal device 102 through the network, and can be used to provide services (such as application services, etc.) for the terminal or the client installed on the terminal. A database can be set on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data operation services for the server 104.

[0030] The above network may include, but is not limited to, at least one of the following: a wired network, a wireless network. The above wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, a local area network. The above wireless network may include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 is not limited to a PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projection device, smart TV, smart drying rack, smart curtain, smart audio and video, smart socket, smart speaker, smart sound box, smart fresh air device, smart kitchen and bathroom equipment, smart bathroom equipment, smart floor sweeping robot, smart window cleaning robot, smart mopping robot, smart air purification device, smart steam box, smart microwave oven, smart kitchen water heater, smart purifier, smart water dispenser, smart door lock, etc.

[0031] In yet another embodiment, the user intention recognition method based on multi-modal decision provided by the present application can be applied to smart home appliances. Among them, smart home appliances refer to home appliance products formed after introducing microprocessors, sensor technologies, and network communication technologies into home appliance devices, which have the ability to automatically sense the state of the residential space, the state of the home appliances themselves, and the service state of the home appliances, and can automatically control and receive control instructions from residential users inside or remotely from the residence. It can be understood that smart home appliances are components of smart homes.

[0032] For the user intention recognition method based on multi-modal decision provided by the present application, after a single smart device (corresponding to the target smart interaction device) receives a user's ambiguous intention, if it cannot obtain more information to determine the current spatial information by itself, it can cooperate with other smart devices (corresponding to other smart interaction devices) in the current space to obtain the corresponding missing information. After obtaining the information, the current device makes a decision on the user's true intention and gives a corresponding reply.

[0033] Figure 2 It is one of the flow diagrams of the user intention recognition method based on multi-modal decision provided by the present application.

[0034] Next, in conjunction with Figure 2 The process of the user intention recognition method based on multi-modal decision provided by the present application will be described.

[0035] In another exemplary embodiment of the present application, in conjunction with Figure 2 It can be seen that the user intention recognition method based on multi-modal decision may include steps 210 to 240, and each step will be introduced separately below.

[0036] In step 210, the target smart interaction device is called to obtain the user's behavior state information, and the initial intention of the user is determined according to the behavior state information.

[0037] In step 220, the spatial information where the user is located is determined.

[0038] In one embodiment, the target intelligent interaction device can be considered as an intelligent device that can establish an interaction behavior with the user and perform corresponding functions according to the user's interaction instructions. Among them, the target intelligent interaction device can be an intelligent speaker, a smart phone, etc. In this embodiment, no specific limitation is imposed on the target intelligent interaction device.

[0039] In another embodiment, the target intelligent interaction device can be called to obtain the user's behavior state information, and the user's initial intention can be determined according to the user's behavior state information.

[0040] In an example, the user's behavior state information can be the preset round interaction dialogue information actively initiated by the user to the target intelligent interaction device, where the number of preset rounds is less than or equal to the number threshold. It can be understood that the preset round interaction dialogue information can be considered as dialogue information with fewer interaction rounds. In this application, no specific limitation is imposed on the number of interaction rounds of the preset round interaction dialogue information. For the convenience of description, in this application, the single-round interaction dialogue information will be used as an example of the preset round interaction dialogue information for description.

[0041] Taking the target intelligent interaction device as an intelligent speaker as an example for description, the user's behavior state information can be the voice actively sent by the user to the target intelligent interaction device. For example, the user says "Kung Pao Chicken" to the intelligent speaker. Since "Kung Pao Chicken" may be a recipe or a song name, the user's initial intention can be determined according to the behavior state information. Among them, the initial intention is a vague intention and needs to be confirmed again.

[0042] In another example, the user's behavior state information can also be the sequence of behavioral actions performed by the user. The user's behavior state information can be the action sequence of the user washing tomatoes and eggs. Since the action sequence of the user washing tomatoes and eggs will have different intentions in different scenarios, the user's initial intention can be determined according to the behavior state information. Among them, the initial intention is a vague intention and needs to be confirmed again.

[0043] In step 230, based on the spatial information, the intention type matching the spatial information is determined.

[0044] In step 240, based on the intention type and the initial intention, the target user intention matching the behavior state information is generated.

[0045] In another embodiment, after obtaining the user's ambiguous intention, i.e., the initial intention, it is necessary to reprocess the initial intention to obtain the user's true intention that matches the behavior state information, that is, the target user intention that can match the behavior state information can be obtained.

[0046] In one embodiment, the microphone, camera, radar and other sensing components of other intelligent interaction devices can be used to intelligently sense the space where the user is currently located, obtain the intention type corresponding to the space information, so as to eliminate the ambiguity in the user's dialogue intention.

[0047] In an example, continue to take the user's utterance of "Kung Pao Chicken" in the previous text as an example. If it is sensed that the user is in the kitchen space, it can be determined that the user's true intention is to obtain the recipe for Kung Pao Chicken. In this implementation, by sensing the space information where the user is located, the ambiguity in the user's initial intention can be eliminated, and the user's true intention can be efficiently and accurately recognized on the premise of reducing the number of interaction rounds.

[0048] Figure 3 It is a schematic flow chart of generating a target user intention that matches the behavior state information based on the intention type and the initial intention provided by the present application.

[0049] Next, it will be combined with Figure 3 to illustrate the process of generating a target user intention that matches the behavior state information based on the intention type and the initial intention.

[0050] In an exemplary embodiment of the present application, combined with Figure 3 it can be known that generating a target user intention that matches the behavior state information based on the intention type and the initial intention may include steps 310 to 330, and each step will be introduced separately below.

[0051] In step 310, the behavior state information is parsed to obtain a parsing result, and multiple initial intentions of the user are determined according to the parsing result.

[0052] In step 320, according to the intention type, a target initial intention that matches the intention type is screened out from the multiple initial intentions.

[0053] In step 330, the target initial intention that is screened out and matches the intention type is used as the target user intention that matches the behavior state information.

[0054] In one embodiment, the behavior state information of the user can be parsed to obtain a parsing result corresponding to the behavior state information. Among them, the parsing result can be the key action or keyword regarding the behavior state information. Further, multiple initial intentions of the user are determined according to the parsing result.

[0055] Since behavior state information often represents multiple intentions without auxiliary judgment. For example, when a user says "Kung Pao Chicken" to a smart speaker, since "Kung Pao Chicken" may be a recipe or a song name, the user's initial intention can be determined based on the behavior state information. Among them, the initial intention is a vague intention that needs to be confirmed. In the application process, in order to accurately determine the user's true intention, also known as the target user intention, the intention type determined according to the spatial information where the user is located is used to filter out the target initial intention that matches the intention type from multiple initial intentions.

[0056] In one embodiment, if it is sensed that the user is in the kitchen space, the intention type corresponding to the kitchen space can be determined as a recipe. Therefore, based on the recipe intention type, the user's true intention of wanting to obtain a Kung Pao Chicken recipe can be filtered out from multiple initial intentions. In this implementation, by sensing the spatial information where the user is located, the intention type can be determined, and then based on the intention type, the user's true intention can be determined, thereby eliminating the ambiguity in the user's initial intention and achieving efficient and accurate recognition of the user's true intention on the premise of reducing the number of interaction rounds.

[0057] Figure 4 It is the second flow diagram of the user intention recognition method based on multi-modal decision provided by this application.

[0058] To further introduce the user intention recognition method based on multi-modal decision provided by this application, the following will be combined with Figure 4 for description.

[0059] In another exemplary embodiment of this application, the behavior state information may include the preset round interaction dialogue information initiated by the user to the target intelligent interaction device, for example, it may be single-round interaction dialogue information; the behavior state information may also include the sequence of behavior actions executed by the user. The following will be combined with Figure 4 to describe the user intention recognition method based on multi-modal decision under two types of behavior state information.

[0060] Combined with Figure 4 it can be seen that the user intention recognition method based on multi-modal decision may include steps 410 to 480. Among them, steps 410 to 440 are the same as or similar to steps 210 to 240 described above respectively. For their specific implementation manners and beneficial effects, please refer to the previous description and will not be specifically limited in this embodiment. The following will separately introduce steps 450 to 480.

[0061] In step 450, based on the target user intention, a running instruction for the target intelligent interaction device that matches the target user intention is generated.

[0062] In step 460, the target intelligent interaction device is called to run according to the operation instruction to respond to the single-round interaction dialogue information initiated by the user to the target intelligent interaction device.

[0063] In one embodiment, when the behavior state information is the single-round interaction dialogue information initiated by the user to the target intelligent interaction device, after determining the true intention of the user corresponding to the single-round interaction dialogue information initiated by the user to the target intelligent interaction device, an operation instruction for the target intelligent interaction device corresponding to the true intention of the user can be generated.

[0064] Continuing with the previous "Kung Pao Chicken" embodiment as an example, when the true intention of the user is parsed as obtaining the Kung Pao Chicken recipe, an operation instruction for the target intelligent interaction device to query and display the Kung Pao Chicken recipe to the user can be generated. Further, the target intelligent interaction device, such as a smart speaker, is called to run according to the operation instruction, that is, to query and display the Kung Pao Chicken recipe to the user, so as to respond to the single-round interaction dialogue information initiated by the user to the target intelligent interaction device.

[0065] In step 470, based on the target user intention, an operation instruction for the target intelligent interaction device that matches the target user intention is generated.

[0066] In step 480, the target intelligent interaction device is called to respond to the operation instruction to actively respond to the sequence of behavioral actions performed by the user.

[0067] In another embodiment, when the behavior state information is the sequence of behavioral actions performed by the user, after determining the true intention of the user corresponding to the sequence of behavioral actions performed by the user, also known as the target user intention, an operation instruction for the target intelligent interaction device corresponding to the true intention of the user can be generated.

[0068] Continuing with the previous embodiment of the user's action sequence of washing tomatoes and eggs as an example, when the true intention of the user is parsed as making scrambled eggs with tomatoes, an operation instruction for the target intelligent interaction device to query and display the scrambled eggs with tomatoes recipe to the user can be generated. Further, the target intelligent interaction device, such as a smart speaker, is called to run according to the operation instruction, that is, to actively display the scrambled eggs with tomatoes recipe to the user, so as to actively respond to the sequence of behavioral actions performed by the user. In this embodiment, when the user does not interact with the intelligent device but has potential needs, the user's needs can also be actively met to improve the satisfaction rate of human-computer interaction.

[0069] In another embodiment, when a smart device, such as a smart speaker, is placed in the kitchen and obtains the current spatial information as the kitchen by collaborating with other smart devices, and it is also recognized that there are ingredients - tomatoes and eggs placed on the chopping board, it can be analyzed that the user's most urgent current need is to search for a recipe. At this time, it actively voice broadcasts "Guess you want to make scrambled eggs with tomatoes. I've found a selected recipe video for you. Do you want to play it?" If the user says "Okay", then the smart device plays the retrieved recipe video.

[0070] Through the foregoing embodiments, it is possible to predict the most likely needs of the user in the current space, and determine the interaction mode and feedback content with the user through analysis, so as to actively respond to the sequence of behavioral actions performed by the user.

[0071] In another exemplary embodiment of the present application, the spatial information where the user is located can be determined in the following manner:

[0072] Based on the location information of the target intelligent interaction device, determine the spatial information of the user, where the target intelligent interaction device and the user are in the same space, and / or

[0073] Call the target intelligent interaction device to wake up other intelligent interaction devices under the same local area network as the target intelligent interaction device, where the other intelligent interaction devices are equipped with image acquisition devices;

[0074] Based on the image acquisition device of the other intelligent interaction devices, acquire the environmental image of the space where the user is located;

[0075] According to the environmental image, identify the spatial information where the user is located.

[0076] In one embodiment, the location information of the target intelligent interaction device can be determined by using sensors, cameras, radars, etc. in the target intelligent interaction device. Further, since the target intelligent interaction device and the user are in the same space, the spatial information of the user can be determined according to the location information of the target intelligent interaction device.

[0077] In another embodiment, other smart devices can also be linked to obtain multi-dimensional information and intelligently perceive the current spatial information. During the application process, the target intelligent interaction device can be called to wake up other intelligent interaction devices under the same local area network as the target intelligent interaction device, and the environmental image of the space where the user is located can be acquired based on the image acquisition device of the other intelligent interaction devices. Since the environmental image can reflect the spatial information, the spatial information where the user is located can be determined. For example, if it is acquired that the environment where the user is located includes a refrigerator, it can be determined that the space where the user is located is the kitchen.

[0078] In another exemplary embodiment of the present application, taking the embodiment described above as an example for illustration, the location information of the target intelligent interaction device can be determined in the following manner:

[0079] Determine the device type of the target intelligent interaction device;

[0080] Based on the device type, determine the location information of the target intelligent interaction device.

[0081] In one embodiment, since the device type of the target intelligent interaction device can determine what the target intelligent interaction device specifically is, and different intelligent interaction devices are in different spaces, therefore, based on the device type, the location information of the target intelligent interaction device can be determined. In an example, when the device type of the target intelligent interaction device is a refrigerator, it can be stated that the location information of the target intelligent interaction device is the kitchen.

[0082] In another exemplary embodiment of the present application, taking the embodiment described above as an example for illustration, before obtaining the intention type corresponding to the spatial information based on the spatial information (corresponding to step 230), the user intention recognition method based on multimodal decision-making may further include the following steps:

[0083] Obtain multiple groups of historical interaction intentions of the user in multiple spaces, where the historical interaction intention is the interaction intention determined during the prior interaction process between the user and the intelligent interaction device;

[0084] According to the multiple groups of historical interaction intentions, fit and construct a spatial interaction intention type mapping table, where the spatial interaction intention type mapping table includes the corresponding relationship formed between different spaces of the user and different interaction intention types;

[0085] Among them, obtaining the intention type corresponding to the spatial information based on the spatial information can be implemented in the following manner:

[0086] Based on the corresponding relationship and the spatial information, determine the intention type that matches the spatial information.

[0087] In one embodiment, the interaction habits of a specific user in a specific space of the family can be recorded. During the application process, multiple groups of historical interaction intentions of the user in multiple spaces can be obtained, and according to the multiple groups of historical interaction intentions, a spatial interaction intention type mapping table can be fit and constructed, where the spatial interaction intention type mapping table includes the corresponding relationship formed between different spaces of the user and different interaction intention types. In other words, based on the spatial interaction intention type mapping table, the intention types of the user in different spaces can be obtained. For example, for user A, the kitchen space corresponds to the recipe intention type; the living room space corresponds to the video or music intention type; the study space corresponds to the book intention type, etc.

[0088] Further, based on the corresponding relationship in the spatial interaction intention type mapping table and the spatial information where the user is currently located, the intention type corresponding to the spatial information can be obtained.

[0089] In this embodiment, when the user's intention is a vague intention, based on the corresponding relationship of the constructed spatial interaction intention type mapping table and the current spatial information, the clear intention can be comprehensively judged without multiple rounds of interaction and confirmation with the user.

[0090] Figure 5 It is a schematic diagram of the application scenario of the user intention recognition method based on multi-modal decision provided by this application.

[0091] To further introduce the user intention recognition method based on multi-modal decision provided by this application, the following will be combined with Figure 5 for description.

[0092] In another exemplary embodiment of this application, when the user says "Kung Pao Chicken" to the smart speaker by voice, since the user's intention is a vague intention, if the smart speaker itself cannot obtain more information to judge the current spatial information, it needs to cooperate with other smart devices in the current space to obtain the corresponding missing information. After obtaining the information, the current device decides the user's true intention and gives a corresponding reply.

[0093] In another embodiment, other smart devices can be combined to determine the home space X where the user is located through images or sensors. If it is determined that the home space X where the user is located is the kitchen, the smart device analyzes that the most likely intention of the user is to check the recipe, so that the ambiguity of the user-initiated "Kung Pao Chicken" can be resolved, and the user's intention is clarified as the recipe. In this scenario, the smart speaker can reply by voice "Guess you want to know how to make Kung Pao Chicken, and I've found a cooking video for you", and then play the specified video recipe.

[0094] In another embodiment, if it is determined that the home space X where the user is located is the bedroom, the smart device analyzes by voice that the most likely intention of the user is to listen to music, and replies "Guess you want to listen to Kung Pao Chicken by XX", and then plays the specified music.

[0095] In this application, through a distributed collaborative information acquisition method, the most urgent needs of the user in the current scenario can be analyzed, so as to actively meet the user's needs when the user's speech is ambiguous or the user does not interact with the smart device but has potential needs, so as to improve the satisfaction rate of human-computer interaction.

[0096] As described above, the user intention recognition method based on multi-modal decision provided by this application obtains the behavior state information of the user by calling the target intelligent interaction device, and determines the initial intention of the user according to the behavior state information; then determines the spatial information where the user is located, and based on the spatial information, determines the intention type that matches the spatial information, and based on the intention type and the initial intention, generates the target user intention that matches the behavior state information. In this application, by sensing the spatial information where the user is located, the ambiguity in the user's initial intention can be eliminated, and the target user intention can be efficiently and accurately recognized on the premise of reducing the number of interaction rounds.

[0097] Figure 6 It is a schematic structural diagram of the user intention recognition device based on multi-modal decision provided by this application.

[0098] Next, the user intention recognition device based on multi-modal decision provided by this application will be described in conjunction with Figure 6 The user intention recognition device based on multi-modal decision described below can be correspondingly referred to the user intention recognition method based on multi-modal decision described above.

[0099] In another exemplary embodiment of this application, the user intention recognition device based on multi-modal decision may include an acquisition module 610, a determination module 620, a processing module 630, and a generation module 640. Each module will be introduced separately below.

[0100] The acquisition module 610 can be configured to call the target intelligent interaction device, obtain the behavior state information of the user, and determine the initial intention of the user according to the behavior state information;

[0101] The determination module 620 can be configured to determine the spatial information where the user is located;

[0102] The processing module 630 can be configured to determine the intention type that matches the spatial information based on the spatial information;

[0103] The generation module 640 can be configured to generate the target user intention that matches the behavior state information based on the intention type and the initial intention.

[0104] In another exemplary embodiment of this application, the acquisition module 610 can adopt the following method to determine the initial intention of the user according to the behavior state information:

[0105] Analyze the behavior state information to obtain an analysis result, and determine multiple initial intentions of the user according to the analysis result;

[0106] The generation module 640 may generate a target user intention that matches the behavior status information based on the intention type and the initial intention in the following manner:

[0107] Filter out a target initial intention that matches the intention type from the multiple initial intentions according to the intention type;

[0108] Use the filtered target initial intention that matches the intention type as the target user intention that matches the behavior status information.

[0109] In another exemplary embodiment of the present application, the behavior status information includes preset round interaction dialogue information initiated by the user to the target intelligent interaction device, where the number of times of the preset round is less than or equal to a threshold number of times;

[0110] The generation module 640 may also be configured to:

[0111] Generate an operation instruction of the target intelligent interaction device that matches the target user intention based on the target user intention;

[0112] Call the target intelligent interaction device to operate according to the operation instruction to respond to the preset round interaction dialogue information initiated by the user to the target intelligent interaction device.

[0113] In another exemplary embodiment of the present application, the behavior status information includes a sequence of behavioral actions performed by the user;

[0114] The generation module 640 may also be configured to:

[0115] Generate an operation instruction of the target intelligent interaction device that matches the target user intention based on the target user intention;

[0116] Call the target intelligent interaction device to respond to the operation instruction to actively respond to the sequence of behavioral actions performed by the user.

[0117] In another exemplary embodiment of the present application, the determination module 620 may determine the spatial information where the user is located in the following manner:

[0118] Determine the spatial information of the user based on the location information of the target intelligent interaction device, where the target intelligent interaction device and the user are in the same space, and / or

[0119] Call the target intelligent interaction device to wake up other intelligent interaction devices in the same local area network as the target intelligent interaction device, where the other intelligent interaction devices are equipped with image acquisition devices;

[0120] The image acquisition device based on the other intelligent interaction device acquires an environmental image of the space where the user is located;

[0121] Based on the environmental image, the space information where the user is located is recognized.

[0122] In another exemplary embodiment of the present application, the determination module 620 may implement the determination of the position information of the target intelligent interaction device in the following manner:

[0123] Determine the device type of the target intelligent interaction device;

[0124] Based on the device type, determine the position information of the target intelligent interaction device.

[0125] In another exemplary embodiment of the present application, the processing module 630 may further be configured to:

[0126] Obtain multiple groups of historical interaction intentions of the user in multiple spaces, where the historical interaction intention is the interaction intention determined during the prior interaction process between the user and the intelligent interaction device;

[0127] Based on the multiple groups of historical interaction intentions, fit and construct a spatial interaction intention type mapping table, where the spatial interaction intention type mapping table includes the corresponding relationship formed between different spaces of the user and different interaction intention types;

[0128] The processing module 630 may implement the determination of the intention type matching the space information based on the space information in the following manner:

[0129] Based on the corresponding relationship and the space information, determine the intention type matching the space information.

[0130] Figure 7 Illustrates a schematic physical structure diagram of an electronic device, such as Figure 7As shown in the figure, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communications interface 720, and the memory 730 complete their mutual communication through the communication bus 740. The processor 710 may call the logical instructions in the memory 730 to execute a user intention recognition method based on multimodal decision-making. The method includes: calling a target intelligent interaction device, obtaining the behavior state information of the user, and determining the initial intention of the user according to the behavior state information; determining the spatial information where the user is located; based on the spatial information, determining an intention type that matches the spatial information; based on the intention type and the initial intention, generating a target user intention that matches the behavior state information.

[0131] In addition, when the logical instructions in the above-mentioned memory 730 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0132] On the other hand, this application also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the user intention recognition method based on multimodal decision-making provided by the above-mentioned various methods. The method includes: calling a target intelligent interaction device, obtaining the behavior state information of the user, and determining the initial intention of the user according to the behavior state information; determining the spatial information where the user is located; based on the spatial information, determining an intention type that matches the spatial information; based on the intention type and the initial intention, generating a target user intention that matches the behavior state information.

[0133] In another aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium includes a stored program. When the program runs, it executes the user intention recognition method based on multimodal judgment provided by the above-mentioned various methods. The method includes: calling a target intelligent interaction device, obtaining the behavior state information of the user, and determining the initial intention of the user according to the behavior state information; determining the spatial information where the user is located; based on the spatial information, determining an intention type that matches the spatial information; and generating a target user intention that matches the behavior state information based on the intention type and the initial intention.

[0134] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0135] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A user intention recognition method based on multimodal decision, characterized in that, The method includes: Invoking a target intelligent interaction device, obtaining the behavior status information of the user, and determining the initial intention of the user according to the behavior status information; Determining the spatial information where the user is located; Based on the spatial information, determining an intention type that matches the spatial information; Based on the intention type and the initial intention, generating a target user intention that matches the behavior status information.

2. The user intention recognition method based on multimodal decision according to claim 1, characterized in that The determining the initial intention of the user according to the behavior status information specifically includes: Parsing the behavior status information to obtain a parsing result, and determining multiple initial intentions of the user according to the parsing result; The generating a target user intention that matches the behavior status information based on the intention type and the initial intention specifically includes: According to the intention type, screening out a target initial intention that matches the intention type from multiple initial intentions; Taking the screened target initial intention that matches the intention type as the target user intention that matches the behavior status information.

3. The method for identifying user intention based on multimodal decision according to claim 2, wherein, The behavior status information includes preset round interaction dialogue information initiated by the user to the target intelligent interaction device, where the number of the preset rounds is less than or equal to a threshold number; After taking the screened target initial intention that matches the intention type as the target user intention that matches the behavior status information, the method further includes: Generating an operation instruction for the target intelligent interaction device that matches the target user intention based on the target user intention; Invoking the target intelligent interaction device to operate according to the operation instruction to respond to the preset round interaction dialogue information initiated by the user to the target intelligent interaction device.

4. The user intention recognition method based on multimodal decision according to claim 2, wherein, The behavior status information includes a sequence of behavior actions executed by the user; After taking the screened target initial intention that matches the intention type as the target user intention that matches the behavior status information, the method further includes: Generating an operation instruction for the target intelligent interaction device that matches the target user intention based on the target user intention; Invoking the target intelligent interaction device to respond to the operation instruction to actively respond to the sequence of behavior actions executed by the user.

5. The method for identifying user intention based on multimodal decision according to claim 1, characterized in that, The determining the spatial information where the user is located is implemented in the following manner: Based on the location information of the target intelligent interaction device, determining the spatial information of the user, where the target intelligent interaction device and the user are in the same space, and / or Invoking the target intelligent interaction device to wake up other intelligent interaction devices under the same local area network as the target intelligent interaction device, where the other intelligent interaction devices are equipped with image acquisition devices; Based on the image acquisition devices of the other intelligent interaction devices, collecting an environmental image of the space where the user is located; According to the environmental image, identifying the spatial information where the user is located.

6. The method for identifying user intention based on multi-modal decision according to claim 5, characterized in that, The determining the location information of the target intelligent interaction device is implemented in the following manner: Determining the device type of the target intelligent interaction device; Determine the location information of the target intelligent interaction device according to the device type.

7. The user intention recognition method based on multi-modal decision according to claim 1, wherein Before determining the intent type matching the spatial information based on the spatial information, the method further includes: Obtain multiple groups of historical interaction intents of the user in multiple spaces, where the historical interaction intent is the interaction intent determined during the prior interaction process between the user and the intelligent interaction device; Fit and construct a spatial interaction intent type mapping table according to the multiple groups of historical interaction intents, where the spatial interaction intent type mapping table includes the corresponding relationships formed between different spaces of the user and different interaction intent types; The determining, based on the spatial information, of the intent type matching the spatial information specifically includes: Based on the corresponding relationship and the spatial information, determine the intent type matching the spatial information.

8. A user intention recognition device based on multi-modal decision, characterized in that, The apparatus includes: An acquisition module, configured to call a target intelligent interaction device, acquire the behavior state information of the user, and determine the initial intent of the user according to the behavior state information; A determination module, configured to determine the spatial information where the user is located; A processing module, configured to determine the intent type matching the spatial information based on the spatial information; A generation module, configured to generate a target user intent matching the behavior state information based on the intent type and the initial intent.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, where the program, when running, executes the multi-modal decision-based user intent recognition method according to any one of claims 1 to 7.

10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the multi-modal decision-based user intent recognition method according to any one of claims 1 to 7 through the computer program.