Multi-modal man-machine interaction method of intelligent assistant robot
The intelligent assistant robot automatically processes message cards through multimodal human-computer interaction method, solving the inefficiency and anxiety caused by information overload, and achieving efficient information processing and management.
Patent Information
- Application Number
- CN202510553818.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Anxiety and inefficiency caused by information overload, especially when dealing with message cards from multiple information sources, users find it difficult to concentrate and often miss important information.
Through the intelligent assistant robot, it provides multimodal human-computer interaction method to automatically identify the context of the message card or determine whether to activate the assistant robot through semantic analysis, realize the binding of the message card and application functions, and automatically process information.
It improves the efficiency and user experience of information processing, reduces information anxiety, and ensures that important information is not missed.
Smart Images

Figure CN120492070A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and human-computer interaction technology, and in particular relates to a multimodal human-computer interaction method for an intelligent assistant robot, a computer-readable storage medium / computer program product for implementing the method, an electronic device, and a related cloud server cluster system. Background Art
[0002] With the rapid development of technologies like mobile internet and the Internet of Things, smartphones, smart wearables, and smart home devices are proliferating, enabling users to create and share information anytime, anywhere. From the quick morning greetings sent upon waking to the frequent emails and work group messages received at work; from daily life shared by friends on social media platforms to the massive amounts of entertainment content pushed by short video platforms, information in various forms, including text, images, and videos, is flooding in like a tide. To promote their products and services, businesses are also constantly producing marketing information, exacerbating the information flood. This information overload is seriously impacting people's lives and work efficiency.
[0003] In related technologies, Chinese invention patent application CN202310714614 proposes a method for processing information content, which can quickly obtain key content from a large amount of information, thereby increasing information density and improving information processing efficiency. CN202011399447 proposes an AI robot system for collaborative systems, allowing users to interact with the functions of third-party applications in a portable manner using commands, text, or voice semantics. However, long-term information overload can reduce human capabilities, make decision-making and cognition more difficult, and even lead to physical and mental health issues such as insomnia and anxiety.
[0004] At the individual level, people often suffer from "information anxiety," fearing they'll miss important messages. They frequently refresh social media apps and news apps, but struggle to focus deeply on any given piece of information. At work, employees spend significant time sifting through and processing work messages from various communication channels. This constant interruption slows work progress and puts important tasks on hold. Summary of the Invention
[0005] In response to the above technical problems, the present invention proposes a multimodal human-computer interaction method for an intelligent assistant robot, a computer-readable storage medium / computer program product for implementing the method, an electronic device, and a related cloud server cluster system.
[0006] In a first aspect of the present invention, a multimodal human-computer interaction method for an intelligent assistant robot is proposed, the method comprising the following steps:
[0007] Displaying a first message card on a first page of a first application;
[0008] Determining whether a condition for activating at least one intelligent assistant robot is met based on the conversation context of the first message card on the first page;
[0009] When it is determined that the conditions for activating at least one intelligent assistant robot are met, an activation button is displayed at a preset position of the first message card;
[0010] In response to the first user clicking the activation button, activating at least one intelligent assistant robot, so that the intelligent assistant robot uses the adapted functional modality to bind the first message card to the first target function of the first application;
[0011] The activated intelligent assistant robot provides multiple functional modes, and different functional modes correspond to different processing modes of the first message card.
[0012] The first message card comes from a second application; the first application and the second application are different types of applications, or the first application and the second application are two applications of the same type.
[0013] During actual execution, in the method, when it is determined that the conditions for activating at least one intelligent assistant robot are not met, it is detected whether an instruction to activate at least one intelligent assistant robot is received from the first user; if so, an activation button is displayed at a preset position of the first message card.
[0014] During actual execution, in the method, when it is determined that the conditions for activating at least one intelligent assistant robot are met based on the conversation context of the first message card on the first page, the activated intelligent assistant robot determines the adapted functional mode based on the conversation context of the first message card on the first page.
[0015] During actual execution, in the method, when an instruction from the first user to activate at least one intelligent assistant robot is received, the activated intelligent assistant robot determines the adapted functional mode based on the instruction from the first user to activate at least one intelligent assistant robot.
[0016] In one case, the activated intelligent assistant robot provides a first functional mode. After binding the first message card with the first target function of the first application under the first functional mode, if the first message card generates an update message, the intelligent assistant robot creates a conversation group with the first user and displays the first message card and the update message in the conversation group.
[0017] In one embodiment, the first message card is sent by a second user to the first user;
[0018] The activated intelligent assistant robot provides a second functional mode. After binding the first message card with the first target function of the first application under the second functional mode, if the first message card generates an update message, the intelligent assistant robot creates a conversation group including the first user and the second user, and displays the first message card and the update message in the conversation group.
[0019] During actual execution, in the method, if it is not possible to determine whether the condition for activating at least one intelligent assistant robot is met based on the conversation context of the first message card on the first page, the method further includes:
[0020] In response to detecting that the first user has opened the first message card, performing semantic analysis on message content contained in the first message card, and determining whether a condition for activating at least one intelligent assistant robot is met based on a result of the semantic analysis;
[0021] If it is determined based on the result of the semantic analysis that the conditions for activating at least one intelligent assistant robot are met, in response to detecting that the first user closes the first message card, an activation button is displayed at a preset position of the first message card.
[0022] In the second aspect of the present invention, an electronic device is also proposed, which includes one or more processors; one or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device executes the multimodal human-computer interaction method of the intelligent assistant robot mentioned above.
[0023] In the third aspect of the present invention, a readable storage medium / computer program product is also proposed, wherein the readable storage medium and computer program product store / include a computer program. When the computer program is executed, all or part of the steps of the multimodal human-computer interaction method of the intelligent assistant robot mentioned above are implemented.
[0024] In the fourth aspect of the present invention, a cloud server cluster system is also proposed, which includes multiple cloud servers, each cloud server provides a cloud-based intelligent assistant robot database, and the cloud-based intelligent assistant robot database stores multiple digital assistant robots with different functional modes for the aforementioned electronic devices to call when executing the multimodal human-computer interaction method.
[0025] In actual applications, the first user usually needs to process multiple message cards from multiple information sources in turn in multiple applications. Therefore, in the fifth aspect of the present invention, it is also proposed that the multimodal human-computer interaction method automatically executes a loop processing flow scheme when processing multiple target file cards on multiple pages, including displaying the first message card as the target message card on the first page of the first application and then switching to the next target message card to continue execution. The next target message card can be the second message card displayed on the first page of the first application, the third message card displayed on the second page of the first application, the fourth message card displayed on the first page of the second application, etc., thereby realizing the automatic processing of multiple message cards from multiple information sources in turn in multiple applications.
[0026] The technical solution of the present invention can automatically call an intelligent assistant robot to realize multimodal human-computer interaction, thereby improving the efficiency of human-computer interaction and information processing. Its specific advantages and implementation principles will be further reflected in detail in the specific embodiment part in combination with the drawings of the specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 This is a schematic diagram of the main flow of a multimodal human-computer interaction method for an intelligent assistant robot according to one embodiment of the present invention;
[0029] Figure 2 It is executed Figure 1 A schematic diagram of a portion of the display interface of an electronic device according to the method;
[0030] Figure 3 It is executed Figure 1 Schematic diagram of the human-computer interaction steps of the method;
[0031] Figure 4 It is a main flow chart of a multimodal human-computer interaction method of an intelligent assistant robot according to another preferred embodiment of the present invention. DETAILED DESCRIPTION
[0032] In the specific implementation of this application, if the embodiments of the relevant technical solutions involve user-related data, when the embodiments of this application are applied to specific products or technologies, user permission or consent must be obtained, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0033] See also Figure 1 , Figure 1 It is a main flow diagram of a multimodal human-computer interaction method of an intelligent assistant robot according to an embodiment of the present invention.
[0034] Figure 1 The method shows four flow blocks, which are respectively labeled as steps S100-S400 for simplicity of description. The details of each step are as follows (numbers are omitted in the drawings):
[0035] S100: Displaying a first message card on a first page of a first application;
[0036] S200: Determine whether a condition for activating at least one intelligent assistant robot is met based on the conversation context of the first message card on the first page;
[0037] S300: When it is determined that the conditions for activating at least one intelligent assistant robot are met, displaying an activation button at a preset position of the first message card;
[0038] S400: In response to the first user clicking the activation button, at least one intelligent assistant robot is activated, so that the intelligent assistant robot uses an adapted functional mode to bind the first message card with the first target function of the first application.
[0039] Next, we will combine specific examples and Figure 2-Figure 3 , which expands on each of the above steps.
[0040] First, step S100: displaying a first message card on a first page of a first application.
[0041] The method can be specifically executed by an electronic terminal with information processing capabilities, such as various smart terminals, preferably smart office terminals, including smart phone terminals, etc.
[0042] The first application can be any application (APP) with interactive information function installed on the smart terminal device, including social applications, office applications, video applications, etc. The first application has a first display page, and various corresponding message cards can be displayed on the first display page.
[0043] Message cards, also known as card messages, are a type of message with interactive functionality on smart devices. Card messages present information in a card format, offering a richer display format and interactive functionality. Common types include:
[0044] 5G Card Messages: A 5G messaging product, 5G Card Messages leverages native mobile messaging portals to precisely reach users, delivering interactive services in vertical scenarios using 5G cards. They support over ten template types, including single image and text, multiple images and text, horizontal scrolling, carousels, red envelopes, videos, e-commerce, notifications, and long images and text. Unlike traditional SMS, 5G Card Messages present text messages in a more intuitive and visually impactful format. They also offer one-click access to WeChat mini-programs and apps, effectively shortening the user journey from reading to purchasing.
[0045] Card SMS: Seamlessly upgrade traditional text messages into interactive rich media messages or file links. Through the SMS gateway, a text message with a link to intelligent SMS parsing is sent to the mobile user. Upon receiving the message, SMS enhancement technology parses the message content into registered rich media messages such as images, videos, and red envelopes. Card SMS offers richer and more diverse content on mobile devices, making it suitable for use in scenarios such as product launches, promotions, and event marketing.
[0046] In-app card messaging: In Android development, card messaging is a practical UI element, commonly used in scenarios such as notifications, chat apps, or social platforms. Cards can contain text, images, buttons, file names, and other elements, presenting content in a concise and clear manner, making user interactions more enjoyable and efficient.
[0047] See also Figure 2 , Figure 2 A schematic diagram of a first page of a first application is shown. The first page includes a (command) input area at the bottom of the page and a display area at the top of the page, wherein a first message card B is displayed in the display area.
[0048] As a preference, the first application can be an office application, such as an email processing APP. In this case, the first page includes an email reply area located at the bottom of the page and an email content display area located at the top of the page. The email content that can be displayed in the display area includes the email body, email title, and email message card.
[0049] Specifically, the email message card can be an email attachment sent along with the email body, including an attachment file, attachment picture, attachment audio, attachment video, etc., and can also be a hyperlink URL, file link, image, etc. embedded in the email body.
[0050] As another preference, the first application can be a social application, such as a conversation chat APP. In this case, the first page includes a chat input box located at the bottom of the page and a conversation content display area located at the upper part of the page. The conversation content that can be displayed in the display area includes text, pictures, videos and other files presented as message cards. When the user clicks on the message card, the corresponding file can be opened or jumped to the corresponding file.
[0051] exist Figure 2 In the schematic diagram in , a session context A, a first message card B and a session context C are shown.
[0052] As an example, Figure 2 The conversation context A or the conversation context C shown can be the email body, email title, conversation text, etc. mentioned above; the first message card B can be the email attachment mentioned above (attachment file, attachment picture, attachment audio, attachment video), a hyperlink URL embedded in the email body, a file link, a picture, and other files displayed in the message card in the social application APP.
[0053] Understandable, yes, although Figure 2 The arrangement order of session context A, first message card B, and session context C is shown in the figure. However, in actual applications, the display positions of session context A, first message card B, and session context C are not limited to this. Session context A and session context C may not appear at the same time, that is, only session context A or only session context C.
[0054] In most cases, when the first page displays the first message card, it is usually accompanied by a context related to the first message card.
[0055] For example:
[0056] Session context A: A file is received, see the card below.
[0057] First message card B (only one message card indication button is displayed):
[0058] Session context C: This file is updated every Monday. Please check it carefully.
[0059] At this point, step S200 may be entered: based on the conversation context of the first message card on the first page, determining whether the conditions for activating at least one intelligent assistant robot are met.
[0060] In the above example, based on the conversation context of the first message card on the first page ("A file has been sent to you, please check it. The file is updated every Monday, please check it carefully"), it can be seen that the first application may expect the user to pay attention to the updated data of the file every Monday. That is, based on the conversation context of the first message card on the first page, it is determined that an information processing operation with a corresponding purpose exists. In this case, it is considered that the condition for activating at least one intelligent assistant robot is met;
[0061] Then, proceed to step S300: when it is determined that the conditions for activating at least one intelligent assistant robot are met, an activation button is displayed at a preset position of the first message card. Figure 2 An activation button is shown below the first message card;
[0062] Then, it is up to the user to determine whether it is really necessary to perform the information processing operation for the corresponding purpose (this file updates data every Monday, please pay attention to check). If so, the user can click the activation button, that is, execute step S400: in response to the first user clicking the activation button, activate at least one intelligent assistant robot, so that the intelligent assistant robot uses an adapted functional mode to bind the first message card to the first target function of the first application.
[0063] In the above example, after the user determines that an information processing operation for the corresponding purpose needs to be performed (the file updates data every Monday, please pay attention to check) and clicks the activation button, an intelligent assistant robot will be activated. After the intelligent assistant robot uses an adapted functional mode to bind the first message card with the first target function of the first application, it will execute the corresponding information processing mode according to the information processing operation for the corresponding purpose.
[0064] The activated intelligent assistant robot provides multiple functional modes, and different functional modes correspond to different processing modes of the first message card.
[0065] The multiple functional modes include at least a first functional mode and a second functional mode.
[0066] As an example, after binding the first message card with the first target function of the first application in the first functional mode, if the first message card generates an update message, the intelligent assistant robot creates a conversation group with the first user and displays the first message card and the update message in the conversation group.
[0067] As an example, after binding the first message card with the first target function of the first application in the second functional mode, if the first message card generates an update message, the intelligent assistant robot creates a conversation group including the first user and the second user, and displays the first message card and the update message in the conversation group.
[0068] Of course, the intelligent assistant robot can also provide more other functional modes to perform more other functions.
[0069] In the above embodiment, the activated intelligent assistant robot determines the adapted functional mode based on the conversation context of the first message card on the first page.
[0070] Schematically, after the activated intelligent assistant robot analyzes the conversation context of the first message card on the first page (this file updates data every Monday, please pay attention to check), it can determine that the functional mode currently adapted by the intelligent assistant robot is the first functional mode.
[0071] When the first application is an email processing app, the first target function of the first application includes creating emails, replying to emails, and subscribing to emails;
[0072] In the above embodiment, in response to the first user clicking the activation button, at least one intelligent assistant robot is activated, so that the intelligent assistant robot uses an adapted functional mode to bind the first message card with the email subscription function of the email processing APP, so that at the scheduled time every Monday, the message update data of the first message card B is provided, and after reminding the user to open the email APP, the email conversation group between the robot and the first user is displayed, and the first message card and the updated message are displayed in the email conversation group.
[0073] When the first application is a social application APP, the first target function of the first application includes resuming a conversation, creating a conversation group, entering a group conversation, etc.;
[0074] In the above embodiment, in response to the first user clicking the activation button, at least one intelligent assistant robot is activated, so that the intelligent assistant robot uses an adapted functional mode to bind the first message card with the group conversation function of the social application APP, so that at the scheduled time every Monday, after providing the message update data of the first message card B, a target group is created or entered, and the target group is composed of the first user and the intelligent assistant robot.
[0075] In the above example, the first message card displayed on the first page of the first application may come from different message sources, such as being sent automatically by the system or by a second user. The second user may send the message through the same type of first application or a second application of a different type. However, this message card only needs to interact with the first user and does not involve other users.
[0076] As a further preference, in a wider range of application scenarios, the first message card comes from a second application; the first application and the second application are different types of applications.
[0077] Preferably, the first message card is sent by the second user to the first user, and an interactive operation needs to be performed between the first user and the second user.
[0078] For example:
[0079] Session context A: Send you a file.
[0080] First message card B (only one message card indication button is displayed):
[0081] Session Context C: Please update this document every Monday and send it to me.
[0082] In the above example, based on the conversation context of the first message card on the first page (send you a file, please update the file every Monday and send it to me), it can be known that the first user needs to update the data of the file every Monday and send it to the second user. That is, based on the conversation context of the first message card on the first page, it is determined that there is an information processing operation with a corresponding purpose (the first user updates the file every Monday and sends it to the second user). At this time, it is considered that the condition for activating at least one intelligent assistant robot is met;
[0083] Then, an activation button is displayed at a preset position of the first message card. Figure 2 An activation button is shown below the first message card;
[0084] Then, the first user determines whether it is really necessary to perform the information processing operation for the corresponding purpose (the first user updates the file every Monday and sends it to the second user). If so, the user can click the activation button to execute step S400: in response to the first user clicking the activation button, at least one intelligent assistant robot is activated, so that the intelligent assistant robot uses an adapted functional mode to bind the first message card to the first target function of the first application.
[0085] In the above example, after the user determines that an information processing operation for the corresponding purpose needs to be performed (the first user updates the file every Monday and sends it to the second user) and clicks the activation button, an intelligent assistant robot will be activated. The intelligent assistant robot uses an adapted functional mode to bind the first message card with the first target function of the first application, and then executes the corresponding information processing mode according to the information processing operation for the corresponding purpose.
[0086] In the above embodiment, the activated intelligent assistant robot determines the adapted functional mode based on the conversation context of the first message card on the first page.
[0087] Schematically, after the activated intelligent assistant robot analyzes the conversation context of the first message card on the first page, it can determine that the functional mode currently adapted by the intelligent assistant robot is the second functional mode.
[0088] When the first application is an email processing app, the first target function of the first application includes creating emails, replying to emails, and subscribing to emails;
[0089] In the above embodiment, in response to the first user clicking the activation button, at least one intelligent assistant robot is activated, so that the intelligent assistant robot uses an adapted functional mode to bind the first message card with the reply email function of the email processing APP, so that at the scheduled time every Monday, the message update data of the first message card B is provided, and the user is reminded to open the email APP, and the email conversation group between the robot and the first user is displayed, and the first message card and the updated message are displayed in the email conversation group, and the email is replied to the second user.
[0090] When the first application is a social application APP, the first target function of the first application includes resuming a conversation, creating a conversation group, entering a group conversation, etc.;
[0091] In the above embodiment, in response to the first user clicking the activation button, at least one intelligent assistant robot is activated, so that the intelligent assistant robot uses an adapted functional mode to bind the first message card with the group conversation function of the social application APP, so that at the scheduled time every Monday, after providing the message update data of the first message card B, a target group is created or entered, and the target group users are composed of the first user and the second user, excluding the intelligent assistant robot.
[0092] The above embodiment is directed to the situation where there is a conversation context for the first message and whether the conditions for activating at least one intelligent assistant robot are met can be determined by analyzing the conversation context.
[0093] However, in some cases, the first user may only receive a first message card, such as an email processing app that only receives an email attachment file, or a social application app that only receives a file hyperlink card, without any additional instructions. In other words, no conversation context can be obtained at this time.
[0094] Alternatively, although a conversation context exists, the conversation context cannot provide any valid information, that is, no information processing operation with corresponding purpose can be derived, that is, based on the conversation context of the first message card on the first page, it is impossible to determine whether the conditions for activating at least one intelligent assistant robot are met.
[0095] For this situation, a more preferred embodiment can be found in Figure 3 .
[0096] based on Figure 3 , a multimodal human-computer interaction method of an intelligent assistant robot can be implemented as follows:
[0097] Displaying a first message card on a first page of a first application;
[0098] If it cannot be determined whether a condition for activating at least one intelligent assistant robot is met based on the conversation context of the first message card on the first page, in response to detecting that the first user has opened the first message card, performing semantic analysis on the message content contained in the first message card, and determining whether a condition for activating at least one intelligent assistant robot is met based on a result of the semantic analysis;
[0099] If it is determined based on the result of the semantic analysis that the conditions for activating at least one intelligent assistant robot are met, in response to detecting that the first user closes the first message card, an activation button is displayed at a preset position of the first message card.
[0100] At this time, since there is no conversation context, or the conversation context is invalid for determining whether to activate at least one intelligent assistant robot, it is necessary to track the user's interactive operations on the first message card.
[0101] Generally speaking, if a user is interested in the first message card, he or she will usually open the message card to browse the text. At this time, in response to detecting that the first user has opened the first message card, a semantic analysis is performed on the message content contained in the first message card, and based on the results of the semantic analysis, it is determined whether the conditions for activating at least one intelligent assistant robot are met.
[0102] Figure 3It is shown that the message body contains the semantics "This file is updated every Wednesday, click the original link to view the updated content", which means that there is an information processing operation that needs to be performed for the corresponding purpose (this file is updated every Wednesday, click the original link to view the updated content), that is, the conditions for activating at least one intelligent assistant robot are met, and an activation button can be displayed at the preset position of the first message card to continue executing the method.
[0103] In another case, in addition to being automatically executed, steps S200-S400 of the method may also include user interaction. In this case, the method includes:
[0104] Displaying a first message card on a first page of a first application;
[0105] If it is determined based on the conversation context of the first message card on the first page that the condition for activating at least one intelligent assistant robot is not met, detecting whether an instruction from the first user to activate at least one intelligent assistant robot is received;
[0106] If yes, an activation button is displayed at a preset position of the first message card.
[0107] Preferably, the method further comprises:
[0108] Displaying a first message card on a first page of a first application;
[0109] If it cannot be determined whether a condition for activating at least one intelligent assistant robot is met based on the conversation context of the first message card on the first page, in response to detecting that the first user has opened the first message card, performing semantic analysis on the message content contained in the first message card, and determining whether a condition for activating at least one intelligent assistant robot is met based on a result of the semantic analysis;
[0110] If it is determined based on the result of the semantic analysis that the condition for activating at least one intelligent assistant robot is not met, detecting whether an instruction from the first user to activate at least one intelligent assistant robot is received;
[0111] If yes, an activation button is displayed at a preset position of the first message card.
[0112] It can be seen that when the multimodal human-computer interaction method of the intelligent assistant robot of the present invention determines whether the conditions for activating at least one intelligent assistant robot are met, the judgment order (conditions) has a priority:
[0113] First priority: If it can be determined based on the conversation context of the first message card on the first page that the conditions for activating at least one intelligent assistant robot are met, an activation button is displayed at a preset position of the first message card;
[0114] Otherwise, proceed to the second priority level: if it cannot be determined based on the conversation context of the first message card on the first page that the conditions for activating at least one intelligent assistant robot are met (or it is determined based on the conversation context of the first message card on the first page that the conditions for activating at least one intelligent assistant robot are not met), then in response to detecting that the first user has opened the first message card, perform semantic analysis on the message content contained in the first message card, and determine whether the conditions for activating at least one intelligent assistant robot are met based on the result of the semantic analysis;
[0115] If it is determined based on the result of the semantic analysis that the condition for activating at least one intelligent assistant robot is met, when the user closes the first message card, an activation button is displayed at a preset position of the first message card;
[0116] If the result of the semantic analysis cannot determine that the conditions for activating at least one intelligent assistant robot are met (or in other words, the result of the semantic analysis determines that the conditions for activating at least one intelligent assistant robot are not met), then enter the third priority level;
[0117] Third priority: detecting whether an instruction to activate at least one intelligent assistant robot is received from the first user; if so, displaying an activation button at a preset position of the first message card.
[0118] The above-mentioned progressive priorities can avoid users from manually inputting commands as much as possible, thereby maximizing the intelligence of human-computer interaction and improving processing efficiency.
[0119] In one embodiment, the first user's instruction to activate at least one intelligent assistant robot may be a text instruction input with specific characters, such as "%% robot A", "@ robot B", etc.
[0120] When the user actively intervenes, the user can also send instructions to specify the processing mode, such as "%% robot A reminds file updates every Wednesday", "@ robot B sends an email containing the file update data to XX every Monday", etc.
[0121] At this time, when the instruction of activating at least one intelligent assistant robot is received from the first user, the activated intelligent assistant robot can determine the adapted functional mode based on the instruction of activating at least one intelligent assistant robot from the first user.
[0122] When the method of the present invention is actually executed, in most cases it can be automatically executed based on a pre-trained intelligent assistant robot without the need for human intervention. Preferably, the first application is implemented based on the support of a cloud server cluster system, the cluster system includes multiple cloud servers, each of which provides a cloud-based intelligent assistant robot database, and the cloud-based intelligent assistant robot database stores multiple digital assistant robots with different functional modes for the first application of the aforementioned electronic device to call when executing the multimodal human-computer interaction method.
[0123] In actual implementation, the program instruction flow of the more preferred automated multimodal human-computer interaction method can be found in Figure 4 Flowchart of the process. Figure 1-Figure 3 The embodiment only provides an embodiment for processing a certain "first message card". In actual applications, the first user usually needs to process multiple message cards from multiple information sources in multiple applications in turn.
[0124] to this end, Figure 4 The flowchart shows the automatic execution cycle process of the multimodal human-computer interaction method in actual application when processing multiple target file cards on multiple pages.
[0125] Described in the form of computer pseudocode Figure 4 The examples are as follows:
[0126] Step 1: Display the target message card on the target page of the target application;
[0127] Step 2: Determine whether the target message card has a conversation context;
[0128] If it does not exist, go to step 3; otherwise go to step 4;
[0129] Step 3: In response to detecting that the first user opens the target message card, the message content contained in the target message card is used as the conversation context, and step 4 is entered;
[0130] Step 4: Parse the session context;
[0131] Step 5: Based on the analysis result of step 4, determine whether the activation conditions are met. If yes, proceed to step 6; otherwise, proceed to step 11;
[0132] Step 6: Display the activation button;
[0133] Step 7: Determine whether the user clicks the activation button. If yes, go to step 8; otherwise, go to step 11;
[0134] Step 8: Display the adapted functional mode;
[0135] Step 9: Determine whether the user confirms the function mode. If yes, proceed to step 10; otherwise, update the matching function mode and return to step 8;
[0136] Step 10: Bind the target message card to the target function of the target application and go to step 11;
[0137] Step 11: Go to the next target message card and return to step 1.
[0138] In the above process, step 1 may be to display the first message card as the target message card on the first page of the first application; and in step 11, when switching to the next target message card, the next target message card may be the second message card displayed on the first page of the first application, the third message card displayed on the second page of the first application, the fourth message card displayed on the first page of the second application, and so on, thereby automatically processing multiple message cards from multiple information sources in multiple applications in turn. The entire process can automatically call the intelligent assistant robot to realize multimodal human-computer interaction, thereby improving the efficiency of human-computer interaction and information processing.
[0139] Although not shown in the accompanying drawings, preferably, more product embodiments can also be a cloud server cluster system, which includes multiple cloud servers, each cloud server provides a cloud-based intelligent assistant robot database, and the cloud-based intelligent assistant robot database stores multiple digital assistant robots with different functional modes for the aforementioned electronic devices to call when executing the multimodal human-computer interaction method.
[0140] Although not shown in the accompanying drawings, preferably, further product embodiments may also include an electronic device comprising a memory and one or more processors. The memory stores one or more application programs, which are suitable for being executed by the one or more processors to implement the aforementioned multimodal human-computer interaction method for the intelligent assistant robot.
[0141] Although not shown in the accompanying drawings, further embodiments further include a computer medium storing a computer program. When the computer program is executed, all or part of the steps of the multimodal human-computer interaction method of the intelligent assistant robot mentioned above are implemented, for example, Figure 4 The computer pseudocode corresponds to a computer program.
[0142] It can be understood that the system, product, device, medium embodiments and method implementations correspond to each other and can reference each other. Their principles are similar or the same, so they will not be repeated.
[0143] For other technologies, principles, algorithms or models not elaborated in detail in this application, please refer to the existing technology.
[0144] The foregoing has shown and described the method embodiments and system of the present invention, but it is understood by those skilled in the art that various changes, modifications, substitutions and variations may be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multimodal human-computer interaction method for an intelligent assistant robot, characterized in that: The method comprises the following steps: Displaying a first message card on a first page of a first application; Determining whether a condition for activating at least one intelligent assistant robot is met based on the conversation context of the first message card on the first page; When it is determined that the conditions for activating at least one intelligent assistant robot are met, an activation button is displayed at a preset position of the first message card; In response to the first user clicking the activation button, activating at least one intelligent assistant robot, so that the intelligent assistant robot uses the adapted functional modality to bind the first message card to the first target function of the first application; The activated intelligent assistant robot provides multiple functional modes, and different functional modes correspond to different processing modes of the first message card.
2. The multimodal human-computer interaction method of an intelligent assistant robot according to claim 1, characterized in that: The first message card comes from a second application; the first application and the second application are different types of applications.
3. The multimodal human-computer interaction method of an intelligent assistant robot according to claim 1, characterized in that: In the method, when it is determined that the condition for activating at least one intelligent assistant robot is not met, detecting whether an instruction to activate at least one intelligent assistant robot is received from the first user; If yes, an activation button is displayed at a preset position of the first message card.
4. The multimodal human-computer interaction method of an intelligent assistant robot according to claim 1, wherein: When it is determined that the conditions for activating at least one intelligent assistant robot are met based on the conversation context of the first message card on the first page, the activated intelligent assistant robot determines the adapted functional mode based on the conversation context of the first message card on the first page.
5. The multimodal human-computer interaction method of an intelligent assistant robot according to claim 3, characterized in that: When receiving the instruction of the first user to activate at least one intelligent assistant robot, the activated intelligent assistant robot determines the adapted functional modality based on the instruction of the first user to activate at least one intelligent assistant robot.
6. The multimodal human-computer interaction method of an intelligent assistant robot according to claim 1, characterized in that: The activated intelligent assistant robot provides a first functional mode. After binding the first message card to the first target function of the first application under the first functional mode, if the first message card generates an update message, the intelligent assistant robot creates a conversation group with the first user and displays the first message card and the update message in the conversation group.
7. The multimodal human-computer interaction method of an intelligent assistant robot according to claim 1, characterized in that: The first message card is sent by the second user to the first user; The activated intelligent assistant robot provides a second functional mode. After binding the first message card with the first target function of the first application under the second functional mode, if the first message card generates an update message, the intelligent assistant robot creates a conversation group including the first user and the second user, and displays the first message card and the update message in the conversation group.
8. The multimodal human-computer interaction method of an intelligent assistant robot according to claim 1, characterized in that: If it cannot be determined whether a condition for activating at least one intelligent assistant robot is met based on the conversation context of the first message card on the first page, the method further includes: In response to detecting that the first user has opened the first message card, performing semantic analysis on message content contained in the first message card, and determining whether a condition for activating at least one intelligent assistant robot is met based on a result of the semantic analysis; If it is determined based on the result of the semantic analysis that the conditions for activating at least one intelligent assistant robot are met, in response to detecting that the first user closes the first message card, an activation button is displayed at a preset position of the first message card.
9. An electronic device comprising a memory and a processor, wherein the memory comprises an executable computer program code, and when the computer program code is executed by the processor, a multimodal human-computer interaction method of an intelligent assistant robot as described in any one of claims 1 to 8 is implemented on a visual interface of the electronic device.
10. A computer-readable storage medium storing computer program instructions. When the computer program instructions are executed, a multimodal human-computer interaction method of an intelligent assistant robot as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
AI robot system applied to cooperative system
CN112565061A
Information content processing method and device, equipment and storage medium
CN119155274A