Task execution method and device and electronic equipment
By entering dialogue information in the AI dialogue window and combining the application interface content, the task intention is automatically determined and the task is executed, which solves the problems of user cumbersome operations and reduced equipment heating and battery life, and improves the task execution efficiency.
Patent Information
- Application Number
- CN202510532497.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
When users perform tasks through AI assistants, they need to enter a large amount of text description information in the AI dialogue window, which leads to cumbersome and time-consuming operation steps, and frequent screenshots will lead to the electronic device's heating and battery life reduction.
By entering dialogue information in the AI dialogue window, combining the interface content of the first application interface, the task intention information is automatically determined, and the corresponding application is controlled to perform tasks through the AI assistant, avoiding manual screenshots and circle screen content.
It reduces user operation steps, improves the efficiency of electronic equipment to perform tasks, and avoids the problems of equipment heating and reduced battery life.
Smart Images

Figure CN120448014A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to a task execution method, device and electronic device. Background Art
[0002] With the development of electronic devices, users can use artificial intelligence (AI) assistants in electronic devices to perform tasks such as information query, image processing, audio conversion, etc. When a user wants to use an AI assistant to perform a task, the user can enter information in the AI dialogue window, and the AI assistant will perform the corresponding task based on the information entered by the user.
[0003] However, in order to ensure that the AI assistant can accurately determine the user's task intention, the user generally needs to enter a large amount of text description information in the AI dialogue window, or enter images and text in the AI dialogue window so that the AI assistant can determine the user's task intention as accurately as possible. For example, if a user wants to use the AI assistant to query the species of a butterfly on the lawn, the user needs to enter a large amount of butterfly description information in the AI dialogue window, such as "blue butterfly, wings and body with spots, slender abdomen, wide wings, body shape of about 5 cm, what kind of butterfly is this", so that the AI assistant can identify the type of butterfly as accurately as possible. This makes the user's operation steps cumbersome and time-consuming. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a task execution method, device and electronic device, which can avoid problems such as heating and reduced battery life caused by frequent screenshots of electronic devices, while reducing the number of operating steps for users to trigger electronic devices to perform required tasks, thereby improving the efficiency of electronic devices in performing tasks.
[0005] In a first aspect, an embodiment of the present application provides a task execution method, the method comprising: receiving dialogue information input by the user in the AI dialogue window when an AI dialogue window between the user and the AI assistant is displayed; the dialogue information includes description information of a first task to be executed, and the dialogue information lacks at least one task element for executing the first task; determining first task intention information based on the dialogue information and the interface content of the first application interface, and the AI dialogue window is displayed on the first application interface; the first task intention information is used to describe the task content of the first task; controlling the first application through the AI assistant to execute the first task according to the first task intention information; wherein the first application is an application that can execute the first task.
[0006] In the second aspect, an embodiment of the present application provides a task execution device, which includes: a receiving module, a determining module and a processing module. The receiving module is used to receive the dialogue information input by the user in the AI dialogue window when the AI dialogue window between the user and the AI assistant is displayed; the dialogue information includes description information of the first task to be executed, and the dialogue information lacks at least one task element for executing the first task. The determining module is used to determine the first task intention information based on the dialogue information received by the receiving module and the interface content of the first application interface, and the AI dialogue window is displayed on the first application interface; the first task intention information is used to describe the task content of the first task. The processing module is used to control the first application through the AI assistant to execute the first task according to the first task intention information determined by the determining module; wherein the first application is an application that can execute the first task.
[0007] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
[0008] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0009] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method described in the first aspect.
[0010] In a sixth aspect, an embodiment of the present application provides a computer program / program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.
[0011] In an embodiment of the present application, when an AI dialogue window between a user and an AI assistant is displayed, dialogue information input by the user in the AI dialogue window is received; the dialogue information includes description information of a first task to be performed, and the dialogue information lacks at least one task element for performing the first task; based on the dialogue information and the interface content of the first application interface, the first task intention information is determined, and the AI dialogue window is displayed on the first application interface; the first task intention information is used to describe the task content of the first task; the first application is controlled by the AI assistant to perform the first task according to the first task intention information; wherein, the first application is an application that can execute the first task. In this solution, after the user enters the conversation information in the AI conversation window, the electronic device can automatically combine the conversation information entered by the user in the AI conversation window and the interface content of the first application interface displayed in the lower layer of the AI conversation window to determine the user's task intention information, and then control the first application through the AI assistant to perform the task corresponding to the task intention information. That is, the user does not need to manually take screenshots or manually circle the screen content. The user only needs to input voice or text input instructions in the AI conversation window, and the AI assistant can understand the screen content and automatically execute the user's instructions. Therefore, while avoiding problems such as heat and reduced battery life caused by frequent screenshots of electronic devices, it also reduces the user's operating steps to trigger the electronic device to perform the required tasks, thereby improving the efficiency of the electronic device in performing tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a flowchart of a task execution method provided by some embodiments of the present application;
[0013] Figure 2A is a schematic diagram of an example of a display conversation interface provided by some embodiments of the present application;
[0014] Figure 2B This is a schematic diagram of an example of displaying an AI dialogue window provided by some embodiments of the present application;
[0015] Figure 2C This is a schematic diagram of an example of inputting dialogue information in an AI dialogue window provided by some embodiments of the present application;
[0016] Figure 3A is a schematic diagram of an example of displaying a shooting preview interface provided by some embodiments of the present application;
[0017] Figure 3B This is a schematic diagram of an example of displaying an AI dialogue window provided by some embodiments of the present application;
[0018] Figure 3C This is a schematic diagram of an example of inputting dialogue information in an AI dialogue window provided by some embodiments of the present application;
[0019] Figure 4 is a flowchart of a task execution method provided by some embodiments of the present application;
[0020] Figure 5 is a schematic diagram of an example of performing a task through an AI assistant provided in some embodiments of the present application;
[0021] Figure 6A is a schematic diagram of an example of a display conversation interface provided by some embodiments of the present application;
[0022] Figure 6B This is a schematic diagram of an example of displaying an AI dialogue window provided by some embodiments of the present application;
[0023] Figure 6C This is a schematic diagram of an example of inputting dialogue information in an AI dialogue window provided by some embodiments of the present application;
[0024] Figure 6D This is a schematic diagram of an example of prompting a user to switch interfaces provided by some embodiments of the present application;
[0025] Figure 7 is a schematic diagram of an example of performing a task through an AI assistant provided in some embodiments of the present application;
[0026] Figure 8 is a flowchart of a task execution method provided by some embodiments of the present application;
[0027] Figure 9 is a schematic diagram of an example of performing a task through an AI assistant provided in some embodiments of the present application;
[0028] Figure 10A This is a schematic diagram of an example of inputting dialogue information in an AI dialogue window provided by some embodiments of the present application;
[0029] Figure 10B This is a schematic diagram of an example of performing a task through an AI assistant provided by some embodiments of the present application;
[0030] Figure 11A This is a schematic diagram of an example of inputting dialogue information in an AI dialogue window provided by some embodiments of the present application;
[0031] Figure 11B This is a schematic diagram of an example of displaying an inquiry message in an AI dialogue window provided by some embodiments of the present application;
[0032] Figure 11C This is a schematic diagram of an example of inputting a reply message in an AI dialogue window provided by some embodiments of the present application;
[0033] Figure 12 is a structural diagram of a task execution device provided in some embodiments of the present application;
[0034] Figure 13 is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of the present application;
[0035] Figure 14 This is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of the present application. DETAILED DESCRIPTION
[0036] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0037] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0038] The terms "at least one" and "at least one of" in the specification and claims of this application refer to any one, any two, or a combination of more than two of the objects included. For example, at least one of a, b, and c can be represented by: "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple. Similarly, "at least two" means two or more, and its meaning is similar to "at least one".
[0039] The terms used in the embodiments of this application are only used to explain the specific embodiments of this application and are not intended to limit this application. The following is an explanation of the terms involved in the embodiments of this application.
[0040] Interface: The medium through which people interact with electronic devices. An interface allows users to send commands to the system through input devices and receive feedback through output devices. Input devices can include keyboards, mice, touch screens, and other devices; output devices can also include displays, speakers, and other devices.
[0041] Shooting preview interface: It is the area used to display the shooting scene in real time before taking a photo or recording a video.
[0042] AI Assistant: An embedded system service or functional module with natural language processing and machine learning capabilities to facilitate human-computer interaction through voice or text interfaces. This assistant is an intelligent assistance tool based on artificial intelligence technology that can automate tasks, provide search and information services, personalize responses, and optimize the user experience. It aims to help users efficiently complete various tasks, obtain information, or create content. The AI Assistant has natural language processing capabilities and can understand and process natural language commands entered by users through voice or text. This includes recognizing, parsing, and generating appropriate responses.
[0043] Optical Character Recognition (OCR) technology: is a technology that converts text on an image into machine-readable text.
[0044] BlueLM-V-3B: is a multimodal large language model (MLLM) optimized for mobile devices that can process both images and text content.
[0045] Soft fine-tuning (SFT) is a technique used in natural language processing and machine learning. SFT fine-tuning typically involves performing relatively small and gentle parameter adjustments on a large, pre-trained language model. Its primary purpose is to better adapt the model to a specific task or domain.
[0046] Control: A control (also called part, component, widget or control) is a graphical user interface element and the basic building block of the user interface, such as a window or text box, displayed in the program interface of any application. A control can be a button, text box, label, etc., used to control all data processed by each application and the interactive operations on this data.
[0047] The following describes in detail the task execution method provided by the embodiment of the present application through specific embodiments and their application scenarios in conjunction with the accompanying drawings.
[0048] The task execution method in the embodiments of the present application can be applied to scenarios where tasks are performed by an AI assistant.
[0049] For example, in scenario 1 where an AI assistant is used to navigate a user, the user receives a chat message from contact A in the conversation interface with contact A, saying, "Let's play badminton tonight at Fandou Garden on Fandou Street in the East District. See you there." The user wants to navigate to the address "Fandou Garden" on the screen.
[0050] Another example is scenario 2, where an AI assistant is used to identify butterfly species for a user. The user sees a butterfly outdoors and wants to know what species it is.
[0051] The execution subject of the task execution method provided in the embodiment of the present application may be a task execution device, which may be an electronic device, or a functional module or entity in the electronic device. The technical solution provided in the embodiment of the present application is described below using an electronic device as an example.
[0052] The present application embodiment provides a task execution method, Figure 1 FIG1 shows a flow chart of a task execution method provided by an embodiment of the present application, which can be applied to electronic devices. Figure 1 As shown, the task execution method provided in the embodiment of the present application may include the following steps 201 to 203.
[0053] Step 201: When an AI dialogue window between a user and an AI assistant is displayed, the electronic device receives dialogue information input by the user in the AI dialogue window.
[0054] In some embodiments of the present application, the electronic device can receive the user's AI assistant wake-up input and display the AI dialogue window on the current display interface.
[0055] In some embodiments of the present application, the AI assistant wake-up input includes but is not limited to: a user's finger touch input on the power button, or a voice command input by the user. The specific input can be determined based on actual usage needs and is not limited in the present application.
[0056] For example: the above voice command can be "Xiao V Xiao V".
[0057] In some embodiments of the present application, the user can perform editing input or voice input in the AI dialogue window, so that the electronic device can receive the dialogue information input by the user in the AI dialogue window.
[0058] It is understandable that users can communicate with the AI assistant in the AI dialogue window by inputting voice or text, so that the AI assistant can perform corresponding tasks according to the user's needs.
[0059] In an embodiment of the present application, the above-mentioned dialogue information includes description information of the first task to be executed.
[0060] In some embodiments of the present application, the first task may include any of the following: a search task, a parameter setting task, an image processing task, an audio conversion task, a text extraction task, a navigation task, etc. The specific task may be determined according to actual use requirements and is not limited in the embodiments of the present application.
[0061] In some embodiments of the present application, the above-mentioned dialogue information lacks at least one task element for performing the first task.
[0062] For example, the above dialogue message may be “Help me navigate to this address”, but the dialogue message lacks the task element “navigation destination”.
[0063] In some examples, when the above-mentioned first task is a search task, the task elements for executing the first task may include: search content; when the above-mentioned first task is a parameter setting task, the task elements for executing the first task may include: parameter setting items, parameter setting values; when the above-mentioned first task is an image processing task, the task elements for executing the first task may include: the image to be processed, and the image processing requirements; when the above-mentioned first task is an audio conversion task, the task elements for executing the first task may include: the audio to be converted; when the above-mentioned first task is a text extraction task, the task elements for executing the first task may include: text content; when the above-mentioned first task is a navigation task, the task elements for executing the first task may include: navigation destination.
[0064] Step 202: The electronic device determines first task intention information based on the dialogue information and the interface content of the first application interface.
[0065] In some embodiments of the present application, the above-mentioned AI dialogue window is displayed on the first application interface.
[0066] In some embodiments of the present application, the first application interface may be any of the following: a conversation page, a camera preview page, a shopping page, a game page, an information page, a desktop page, a document page, etc. The specific one may be determined according to actual usage requirements and is not limited in the present application.
[0067] In some embodiments of the present application, when the first application interface is displayed, the electronic device can receive the user's AI assistant wake-up input to display the AI dialogue window in a floating manner on the first application interface.
[0068] In some embodiments of the present application, the electronic device may display the AI dialogue window in the top area of the first application interface, or in the bottom area of the first application interface, or at any location on the first application interface. The specific location may be determined based on actual usage requirements and is not limited in the embodiments of the present application.
[0069] In the embodiment of the present application, the first task intention information is used to describe the task content of the first task.
[0070] For example, when the first task is to navigate to the dumpling garden, the first task intention information may be "navigate to the dumpling garden".
[0071] In some embodiments of the present application, when the determination of the missing task elements of the first task requires associating with the screen content, the electronic device can perform graphic and text analysis on the dialogue information input by the user in the AI dialogue window and the interface content of the first application interface to determine the first task intention information.
[0072] In some embodiments of the present application, the electronic device may process the first application interface through OCR to obtain the text content in the first application interface, and then the electronic device may perform semantic analysis on the text content and the dialogue information entered by the user in the AI dialogue window to obtain the above-mentioned first task intention information.
[0073] It should be noted that for the description of the electronic device determining the first task intention information based on the dialogue information input by the user in the AI dialogue window and the interface content of the first application interface, please refer to the description of graphic analysis of text information and image content in the relevant technology, which will not be repeated here.
[0074] Step 203: The electronic device controls the first application through the AI assistant to perform the first task according to the first task intention information.
[0075] In some embodiments of the present application, the first application is an application that can execute the first task, that is, an application that has the ability to execute the first task.
[0076] In some embodiments of the present application, the electronic device can determine the first application to perform the first task through the first task intention information, and then control the first application to perform the first task according to the first task intention information through the AI assistant.
[0077] Exemplarily, when the first task is a search task, the above-mentioned first application may be a browser; when the first task is an image processing task, the above-mentioned first application may be an application with an image processing function; when the first task is a parameter setting task, the above-mentioned first application may be a setting application; when the first task is an audio conversion task, the above-mentioned first application may be an application with an audio conversion function; when the first task is a text extraction task, the above-mentioned first application may be an application with a text extraction function; when the first task is a navigation task, the above-mentioned first application may be an application with a navigation function.
[0078] In some embodiments of the present application, the above-mentioned first task can be understood as a task corresponding to the first task intention information.
[0079] Example 1, combined with scenario 1, such as Figure 2A As shown, the user receives a chat message from contact A in the conversation interface 11 with contact A, that is, the first application interface mentioned above, "Let's play badminton tonight at Fandou Garden on Fandou Street in the East District. See you there." If the user wants to quickly navigate to the address "Fandou Garden" on the screen, he can long press the power button to wake up the AI assistant, and Figure 2B As shown, the AI dialogue window 12 is displayed floating on the conversation interface 11, and the user can long press the voice control 13 in the AI dialogue window 12, as shown in FIG. Figure 2C As shown, by entering the dialogue information 14 "Help me navigate to the address on the screen" in the AI dialogue window 12, the electronic device can use the AI assistant to determine the first task intention information as "Navigate to the Fandou Garden on Fandou Street in the East District" based on the dialogue information 14 "Help me navigate to the address on the screen" entered by the user in the AI dialogue window 12 and the interface content of the conversation interface 11. Then, the electronic device can use the AI assistant to control the navigation application to navigate to the Fandou Garden on Fandou Street in the East District according to the first task intention information, that is, to execute the first task.
[0080] Example 2, combined with scenario 2, if the user sees a butterfly outdoors and wants to know what kind of butterfly it is, the user can operate the camera icon on the electronic device to control the electronic device to run the camera application, display the shooting preview interface on the display screen of the electronic device, and then point the camera of the electronic device at the butterfly, such as Figure 3A As shown, the shooting preview interface 21, that is, the first application interface mentioned above contains a butterfly, and then the user can long press the power button to wake up the AI assistant, and as shown in FIG. Figure 3B As shown, an AI dialogue window 22 is displayed floating on the shooting preview interface 21, and the user can long press the voice control 23 in the AI dialogue window 22, such as Figure 3CAs shown, by inputting the dialogue information 24 "What kind of butterfly is this" in the AI dialogue window 22, the electronic device can use the AI assistant to determine that the first task intention information is "identify the type of butterfly on the screen" based on the dialogue information 24 "What kind of butterfly is this" input by the user in the AI dialogue window 22 and the interface content of the shooting preview interface 21. Then, the electronic device can use the AI assistant to control the browser application to identify the type of butterfly on the screen according to the first task intention information, that is, to execute the first task.
[0081] In the task execution method provided in the embodiment of the present application, after the user enters the conversation information in the AI conversation window, the electronic device can automatically combine the conversation information entered by the user in the AI conversation window and the interface content of the first application interface displayed in the lower layer of the AI conversation window to determine the user's task intention information, and then control the first application through the AI assistant to execute the task corresponding to the task intention information. That is, the user does not need to manually take screenshots or manually circle the screen content. The user only needs to input voice or text input instructions in the AI conversation window, and the AI assistant can understand the screen content and automatically execute the user instructions. Therefore, while avoiding the problems of heating and reduced battery life caused by frequent screenshots of the electronic device, it also reduces the user's operating steps to trigger the electronic device to perform the required tasks, thereby improving the efficiency of the electronic device in performing tasks.
[0082] In some embodiments of the present application, Figure 1 ,like Figure 4 As shown, before the above step 202, the task execution method provided by the embodiment of the present application further includes the following steps 301 and 302. In addition, the above step 202 can be specifically implemented by the following step 202a.
[0083] Step 301: The electronic device inputs the display status information and dialogue information of the AI dialogue window into the intent recognition model through the AI assistant, and outputs the first intent type when the display status information indicates that the AI dialogue window is in a floating display state and the dialogue information includes screen query keywords.
[0084] In some embodiments of the present application, the display status information of the above-mentioned AI dialogue window is used to indicate the display status of the AI dialogue window.
[0085] In some embodiments of the present application, the display state of the AI dialogue window may include any one of the following: a floating display state and a full-screen display state.
[0086] It can be understood that when the AI dialogue window is displayed floating on the application interface, the AI dialogue window is in a floating display state. When the AI dialogue window is in the floating display state, the AI dialogue window blocks part of the page content of the application interface.
[0087] In some embodiments of the present application, the above-mentioned screen query keywords may include at least one of the following: screen, page, interface, display screen, screen, display, image, picture, photo, etc. The specific keywords can be determined according to actual usage requirements and are not limited in the embodiments of the present application.
[0088] In some embodiments of the present application, the above-mentioned first intention type may include any one of the following: first type, second type.
[0089] In an embodiment of the present application, the first type indicates that the determination of the task elements missing from the first task needs to be associated with the screen content. The second type indicates that the determination of the task elements missing from the first task needs to be associated with the screen content.
[0090] In some embodiments of the present application, the determination of missing task elements of the first task requires associating with screen content, which can be understood as: the electronic device needs to associate with screen content to determine the complete task elements of the first task.
[0091] It should be noted that the first intent type is only used to indicate whether the determination of the task element needs to be associated with the screen content when the first task lacks a task element. However, it does not determine whether the dialogue information input by the user lacks at least one task element for executing the first task.
[0092] In some embodiments of the present application, the electronic device can generate a second instruction based on the display status information of the AI dialogue window, the dialogue information input by the user in the AI dialogue window, and the first instruction (prompt), and then input the second instruction into the intent recognition model to obtain the first intent type output by the intent recognition model.
[0093] For example, the first instruction is as follows:
[0094] Now you are an expert at determining whether a user command contains a screen question. You can combine the [Current State] and [User Input] to determine whether the text corresponding to the [User Input] falls into the category of a "screen question." Please output the final [Answer] as required below.
[0095] <Output requirements>
[0096] 1. Output in a fixed JSON format and do not output anything else.
[0097] 2. [User Input] Output: {"Type":"Question Screen"}
[0098] [User input] Output when it does not belong to the "question screen" type: {"type":"other"}
[0099] 3. [Current Status] includes two enumeration values: "Full Screen Display Status" and "Floating Display Status"
[0100] From now on, please strictly follow the above <output requirements> to judge the intention of the [user input] and output the final [answer].
[0101] [Current Status]: [Status Placeholder]
[0102]
User input
Command placeholder
[0103]
answer
[0104] Specifically, the electronic device can replace the [status placeholder] in the first instruction with the display status of the AI dialogue window, and replace the [instruction placeholder] in the first instruction with the dialogue information entered by the user in the AI dialogue window to generate a second instruction, and then input the second instruction into the intent recognition model to obtain the first intent type output by the intent recognition model.
[0105] In some embodiments of the present application, when [user input] belongs to the "question screen" type, the first intent type output by the intent recognition model can be the first type, for example: the intent recognition model outputs: {"type":"question screen"}.
[0106] In some embodiments of the present application, when [user input] belongs to the "other" type, the first intent type output by the intent recognition model may be the second type, for example: the intent recognition model outputs: {"type":"other"}.
[0107] For example, assuming that the display state of the AI dialogue window is a floating display state, and the dialogue message entered by the user in the AI dialogue window is "Help me navigate to the address on the screen", the electronic device may replace the [state placeholder] in the first instruction with the display state of the AI dialogue window, and replace the [instruction placeholder] in the first instruction with the dialogue message entered by the user in the AI dialogue window, to generate a second instruction, which is as follows:
[0108] Now you are an expert at determining whether a user command contains a screen question. You can combine the [Current State] and [User Input] to determine whether the text corresponding to the [User Input] falls into the category of a "screen question." Please output the final [Answer] as required below.
[0109] <Output requirements>
[0110] 1. Output in a fixed JSON format and do not output anything else.
[0111] 2. [User Input] Output: {"Type":"Question Screen"}
[0112] [User input] Output when it does not belong to the "question screen" type: {"type":"other"}
[0113] 3. [Current Status] includes two enumeration values: "Full Screen Display Status" and "Floating Display Status"
[0114] From now on, please strictly follow the above <output requirements> to judge the intention of the [user input] and output the final [answer].
[0115] [Current status]: [Floating display status]
[0116] [User input]: [Help me navigate to the address on the screen]
[0117]
answer
[0118] Then, the electronic device can input the second instruction into the intent recognition model. Since the display state of the AI dialogue window is a floating display state, the dialogue information "Help me navigate to the address on the screen" entered by the user in the AI dialogue window includes the screen-asking keyword "screen", so the intent recognition model can output [Answer]: {"Type":"Ask Screen"}", to indicate that the determination of the missing task elements of the first task requires associating the screen content.
[0119] For another example: assuming that the display state of the AI dialogue window is a floating display state, and the user enters the dialogue information in the AI dialogue window as "What kind of butterfly is this?", the electronic device may replace the "state placeholder" in the first instruction with the display state of the AI dialogue window, and replace the "command placeholder" in the first instruction with the dialogue information entered by the user in the AI dialogue window, to generate a second instruction, which is as follows:
[0120] Now you are an expert at determining whether a user command contains a screen question. You can combine the [Current State] and [User Input] to determine whether the text corresponding to the [User Input] falls into the category of a "screen question." Please output the final [Answer] as required below.
[0121] <Output requirements>
[0122] 1. Output in a fixed JSON format and do not output anything else.
[0123] 2. [User Input] Output: {"Type":"Question Screen"}
[0124] [User input] Output when it does not belong to the "question screen" type: {"type":"other"}
[0125] 3. [Current Status] includes two enumeration values: "Full Screen Display Status" and "Floating Display Status"
[0126] From now on, please strictly follow the above <output requirements> to judge the intention of the [user input] and output the final [answer].
[0127] [Current status]: [Floating display status]
[0128]
User input
What kind of butterfly is this
[0129]
answer
[0130] Then, the electronic device can input the second instruction into the intent recognition model. Since the display state of the AI dialogue window is a floating display state, the dialogue information "What kind of butterfly is this" entered by the user in the AI dialogue window does not include the screen inquiry keyword. Therefore, the intent recognition model can output [Answer]: {"Type": "Other"}", to indicate that the determination of the missing task elements of the first task does not require the association of screen content.
[0131] Step 302: When the first intent type is the first type, the electronic device determines whether the interface content of the first application interface includes the task elements missing from the first task.
[0132] In some embodiments of the present application, the "electronic device determines whether the interface content of the first application interface includes the task elements missing from the first task" in the above step 302 can be specifically implemented through the following steps 302a to 302d.
[0133] Step 302a: The electronic device takes a screenshot of the first application interface through the AI assistant to obtain a screenshot of the first interface.
[0134] In some embodiments of the present application, before the electronic device takes a screenshot of the first application interface via an AI assistant to obtain a screenshot of the first interface, the electronic device may display a permission confirmation interface to inform the user that the current operation requires access to the user's screen content. Upon receiving the user's confirmation input on the permission confirmation interface, the electronic device may take a screenshot of the first application interface via the AI assistant to obtain a screenshot of the first interface.
[0135] In some embodiments of the present application, the above-mentioned confirmation input is used to allow the electronic device to obtain the user's screen content.
[0136] Step 302b: The electronic device sends a screenshot of the first interface and conversation information to the server through the AI assistant.
[0137] In some embodiments of the present application, when the electronic device sends a screenshot of the first interface to the server through the AI assistant, the electronic device can display a dynamic effect of uploading the image.
[0138] Step 302c: The electronic device receives the first message sent by the server.
[0139] In an embodiment of the present application, the above-mentioned first message includes first judgment result information.
[0140] In an embodiment of the present application, the first judgment result information is obtained by the server inputting the first interface screenshot and the conversation information into the first graphic analysis model on the server side.
[0141] In some embodiments of the present application, the above-mentioned first judgment result information may indicate any one of the following: the interface content of the first application interface includes task element information of the task elements that are missing from the first task, and the interface content of the first application interface does not include task element information of the task elements that are missing from the first task.
[0142] In some embodiments of the present application, after the server receives the first interface screenshot and the conversation information entered by the user in the AI conversation window, the first interface screenshot and the conversation information can be input into the first graphic analysis model on the server side, so that the first graphic analysis model analyzes the first interface screenshot and the conversation information to obtain the first judgment result information output by the first graphic analysis model, and then the server can send a first message including the first judgment result information to the electronic device.
[0143] In some embodiments of the present application, the server can perform SFT fine-tuning on the second image and text analysis model to obtain the above-mentioned first image and text analysis model, and analyze the first interface screenshot and dialogue information through the first image and text analysis model to obtain the first judgment result information output by the first image and text analysis model.
[0144] In some embodiments of the present application, the second graphic analysis model may be BlueLM-V-3B.
[0145] In some embodiments of the present application, the server performs SFT fine-tuning on the second image-text analysis model to obtain the first image-text analysis model through the following steps A1-A5.
[0146] A1. Based on the conversation data after users upload images, a large amount of conversation data containing images is collected.
[0147] In some embodiments of the present application, the diversity of conversation scenarios must be ensured during collection, covering the user's common operations as much as possible.
[0148] A2. Use OCR technology to identify text content in the collected images.
[0149] A3. Design a third instruction that can determine the relevance between the OCR content and the user input content. The third instruction is as follows:
[0150] "Now you are an expert in determining whether OCR and user input are related. You can combine the [OCR content] and [user input] to determine whether they are related. Please output the final [answer] according to the following requirements.
[0151] <Judgment Rules>
[0152] 1. There is a semantic association between the [OCR content] and the [user input], for example, the [user input] is navigation, and the [OCR content] contains an address;
[0153] 2. The [OCR content] and [user input] contain the same keywords and entity words, such as [user input] "What kind of insect is this?", while the [OCR content] contains an insect introduction;
[0154] 3. The words "picture," "photo," "screen," etc. referring to the screen are present in [User Input];
[0155] <Output requirements>
[0156] 1. Output in a fixed JSON format and do not output anything else.
[0157] 2. The output result is: {"Type":"Relevant / Not relevant","Confidence":"0-100%"}
[0158] From now on, please strictly follow the above <Output Requirements> to judge the relevance between [OCR content] and [user input], and output the final [answer]
[0159] [OCR content]: [OCR placeholder]
[0160]
User input
Command placeholder
[0161]
answer
[0162] A4. The collected OCR content and user input are combined with the third instruction to generate a fourth instruction. This fourth instruction is then fed into the second image-text analysis model to obtain relevant pseudo-labels. After obtaining a large number of pseudo-labels, manual verification is performed to obtain a large amount of high-quality SFT annotated data.
[0163] For example, assuming that the extracted OCR content is: "Let's have lunch on the second floor of Joy City at noon today, see you there\nOkay", and the user input content is: "Help me navigate to the address on the screen", then replace the [OCR placeholder] and [instruction placeholder] in the third instruction respectively to obtain the fourth instruction, and input the fourth instruction into the second image and text analysis model to determine the relevance between the OCR content and the user input content. For example: the result output by the second image and text analysis model is: {"Type":"Related","Confidence":"90%"}.
[0164] For example, we can collect 100,000 pairs of OCR content and user input, covering 500 common user scenarios. This data is then cleaned to remove non-text symbols. The cleaned OCR content and the corresponding user input are then dynamically concatenated into the third instruction, resulting in the fourth instruction. The second graph-text analysis model is then used for batch inference to determine the correlation and corresponding confidence levels of the data pairs. Finally, the confidence levels are divided into three levels: confidence levels >85% are assumed to be true labels; confidence levels between 60% and 85% are fully manually annotated; and confidence levels below 60% are sampled and manually annotated for a 20% sample. This approach improves annotation efficiency by approximately 50%, resulting in approximately 20,000 pieces of manually annotated high-quality SFT data in a short period of time.
[0165] A5. Use the fine-tuning data constructed in step A4 to train the second image-text analysis model to obtain the first image-text analysis model.
[0166] In an embodiment of the present application, the first image-text analysis model can further filter out screen content in the first application interface that is irrelevant to the conversation information input by the user. Therefore, compared with the second image-text analysis model, the first image-text analysis model has better and more accurate judgment results when judging the correlation between the interface content of the first application interface and the conversation information input by the user.
[0167] In some embodiments of the present application, determining the correlation between the interface content of the first application interface and the dialogue information input by the user can also be understood as: when the dialogue information input by the user includes description information of the first task to be executed, and the dialogue information lacks at least one task element for executing the first task, determining whether the interface content of the first application interface includes the task element missing from the first task.
[0168] In some embodiments of the present application, when the determination of the missing task elements of the first task requires associating with the screen content, but the interface content of the first application interface does not include the missing task elements of the first task, the server can delete the above-mentioned first interface screenshot, that is, delete the data after it is used up, thereby maximizing the protection of user privacy and security.
[0169] Step 302d: When the first judgment result information indicates that the interface content of the first application interface includes task element information of the task element missing from the first task, the electronic device determines that the interface content of the first application interface includes the task element missing from the first task.
[0170] In some embodiments of the present application, when the above-mentioned first task is a search task, the task element information of the task element for executing the first task may include: search content; when the above-mentioned first task is a parameter setting task, the task element information of the task element for executing the first task may include: parameter setting items, parameter setting values; when the above-mentioned first task is an image processing task, the task element information of the task element for executing the first task may include: the image to be processed, the image processing requirements; when the above-mentioned first task is an audio conversion task, the task element information of the task element for executing the first task may include: the audio to be converted; when the above-mentioned first task is a text extraction task, the task element information of the task element for executing the first task may include: text content; when the above-mentioned first task is a navigation task, the task element information of the task element for executing the first task may include: navigation destination information.
[0171] Step 202a: When the first intention type is the first type and the interface content of the first application interface includes task elements missing from the first task, the electronic device determines the first task intention information based on the dialogue information and the interface content of the first application interface.
[0172] It can be understood that after determining that the task elements missing from the first task need to be associated with the screen content, the electronic device can further determine whether the interface content of the first application interface displayed in the lower layer of the AI dialogue window includes the task elements missing from the first task. When the interface content of the first application interface includes the task elements missing from the first task, the electronic device can determine the first task intention information in combination with the dialogue information input by the user in the AI dialogue window and the interface content of the first application interface.
[0173] Example 3, combined with Example 1, after the user long presses the voice control 13 in the AI dialogue window 12 and enters the dialogue information 14 "Help me navigate to the address on the screen" in the AI dialogue window 12, the electronic device can use the AI assistant to determine whether the determination of the missing task elements of the first task needs to be associated with the screen content according to the display state of the AI dialogue window 12 and the dialogue information 14 entered by the user in the AI dialogue window 12; when the AI dialogue window 12 is displayed in a floating manner on the conversation interface 11 and the dialogue information 14 entered by the user in the AI dialogue window 12 contains the screen query keyword "screen", the electronic device can use the AI assistant to determine that the determination of the missing task elements of the first task needs to be associated with the screen content; then the electronic device can The AI assistant further determines whether the interface content of the conversation interface 11 includes the task element "navigation destination" that is missing in the first task; when the conversation interface 11 includes the address "Fandou Garden on Fandou Street in the East District", the electronic device can determine through the AI assistant that the interface content of the conversation interface 11 includes the task element that is missing in the first task; therefore, the electronic device can determine through the AI assistant that the first task intention information is "navigate to Fandou Garden on Fandou Street in the East District" based on the conversation information 14 input by the user in the AI conversation window 12 and the interface content of the conversation interface 11. Then, the electronic device can control the navigation application through the AI assistant to navigate to Fandou Garden on Fandou Street in the East District according to the first task intention information, and, combined with Figure 2C ,like Figure 5 As shown, the execution result 15 "The navigation application has been opened for you, navigating to the Fandou Garden on Fandou Street in the East District" is displayed in the AI dialogue window 12.
[0174] In this way, since the electronic device can first determine the user's intention by combining the screen content and the question entered by the user in the AI dialogue window, the AI assistant will only identify the screen content when the user's question is an "ask screen" intention, thus avoiding infringement of user privacy to the greatest extent. Then, the electronic device can use the AI assistant to determine on the server side whether the interface content of the first application interface displayed in the lower layer of the AI dialogue window includes the task elements missing from the first task. When the interface content of the first application interface includes the task elements missing from the first task, the task intention information is determined by combining the question entered by the user in the AI dialogue window and the interface content of the first application interface. That is, the user does not need to take a screenshot or circle the screen content, but can directly ask the content on the screen of the electronic device through voice input or text input in the AI dialogue window in a simple, natural and humane communication language. The way to trigger the electronic device to perform tasks is more convenient, and the interaction between the screen and text is smoother. Therefore, while avoiding infringement of user privacy to the greatest extent, it also reduces problems such as heating and reduced battery life caused by frequent screenshots of electronic devices.
[0175] In some embodiments of the present application, after the above step 302, the task execution method provided by the embodiment of the present application further includes the following steps 401 to 405.
[0176] Step 401: When the interface content of the first application interface does not include the task elements missing from the first task, the electronic device displays a first prompt message.
[0177] In an embodiment of the present application, the first prompt information is used to prompt the user to switch to an application interface that includes task elements that are missing from the first task.
[0178] It can be understood that since the determination of the missing task elements of the first task requires the association of screen content, but the interface content of the first application interface does not include the missing task elements of the first task, it is judged that the interface currently opened by the user may be an incorrect interface. The electronic device can prompt the user to switch to the application interface including the missing task elements of the first task through the first prompt information.
[0179] In some embodiments of the present application, the electronic device may display the above-mentioned first prompt information in the AI dialogue window.
[0180] In some embodiments of the present application, the electronic device may display a floating window on the first application interface, where the floating window includes the first prompt information.
[0181] For example, the first prompt message may be “The current screen content is not associated with your instruction, please switch to the correct interface.”
[0182] Example 4, combined with scenario 1, such as Figure 2A As shown, the user receives a chat message from contact A in the conversation interface 11 with contact A, "Let's play badminton tonight at Fandou Garden on Fandou Street in the East District. See you there." When the user wants to quickly navigate to the address "Fandou Garden" on the screen, the user can wake up the AI assistant. If the user presses and holds the power button to wake up the AI assistant, as shown in FIG. Figure 6A As shown, the user accidentally touches the screen, causing the electronic device to display the conversation interface 31 with contact B, that is, the first application interface mentioned above, as shown in FIG. Figure 6B As shown, the electronic device displays the AI dialogue window 32 on the conversation interface 31. If the user does not find that the wrong interface has been opened by mistake, and the user long presses the voice control 33 in the AI dialogue window 32, as shown in FIG. Figure 6CAs shown, the dialogue message 34 "Help me navigate to the address on the screen" is entered in the AI dialogue window 32. The electronic device can use the AI assistant to determine whether the determination of the missing task elements of the first task needs to be associated with the screen content based on the display state of the AI dialogue window 32 and the dialogue message 34 entered by the user in the AI dialogue window 32. When the AI dialogue window 32 is displayed in a floating manner on the conversation interface 31 and the dialogue message 34 entered by the user in the AI dialogue window 32 contains the screen query keyword "screen", the electronic device can use the AI assistant to determine that the determination of the missing task elements of the first task needs to be associated with the screen content. Then, the electronic device can use the AI assistant to further determine whether the interface content of the conversation interface 31 includes the missing task element "navigation destination" of the first task. When the conversation interface 31 does not include the address information, the electronic device can use the AI assistant to determine that the interface content of the conversation interface 31 does not include the missing task elements of the first task. Since it is determined that the determination of the missing task elements of the first task needs to be associated with the screen content, but the interface content of the conversation interface 31 does not include the missing task elements of the first task, it is determined that the interface currently opened by the user may be an incorrect interface. Figure 6D As shown, the electronic device can display a first prompt message 35 "The current screen content is not related to your instructions, please switch to the page containing the address" in the AI dialogue window 32 through the AI assistant to prompt the user to switch to the correct application interface.
[0183] Step 402: The electronic device receives an interface switching input from the user.
[0184] In some embodiments of the present application, the above-mentioned interface switching input can be a touch input of the user to the screen through a touch device such as a finger or a stylus.
[0185] In some embodiments of the present application, the above-mentioned interface switching input is used to switch the interface displayed by the electronic device.
[0186] Step 403: The electronic device switches the first application interface to the second application interface in response to the interface switching input.
[0187] In an embodiment of the present application, the interface content of the second application interface includes task elements that are missing from the first task.
[0188] In some embodiments of the present application, the second application interface may include any of the following: a conversation page, a camera preview page, a shopping page, a game page, an information page, a desktop page, a document page, etc. The specific one may be determined based on actual usage requirements and is not limited in the present application.
[0189] In some embodiments of the present application, the first application interface and the second application interface may be application interfaces of the same application or application interfaces of different applications, which is not limited in the present application.
[0190] In an embodiment of the present application, during the process of switching the first application interface to the second application interface, the AI dialogue window remains displayed.
[0191] It can be understood that in the process of switching the first application interface to the second application interface, since the AI dialogue window remains displayed, the AI assistant can detect whether the application interface is switched to the task elements missing from the first task.
[0192] Example 5, combined with Example 4, the electronic device displays the first prompt message 35 "The current screen content is not related to your instruction, please switch to the page containing the address" in the AI dialogue window 32 through the AI assistant. The user can click the return control, and after the electronic device displays the message list interface, click the conversation identifier of contact A. Figure 6D ,like Figure 7 As shown, the electronic device displays a conversation interface 41 with contact A, that is, the second application interface mentioned above. In the process of switching from the conversation interface 31 to the conversation interface 41, the AI conversation window 32 remains displayed; then, the electronic device can use the AI assistant to determine that the second task intention information is "Navigate to the Fandou Garden on Fandou Street in the East District" based on the conversation information 34 "Help me navigate to the address on the screen" entered by the user in the AI conversation window 32 and the interface content of the conversation interface 41. Then, the electronic device can control the navigation application through the AI assistant to navigate to the Fandou Garden on Fandou Street in the East District according to the second task intention information, and display the execution result 42 "The navigation application has been opened for you, navigate to the Fandou Garden on Fandou Street in the East District" in the AI conversation window 32.
[0193] Step 404: When the AI assistant detects that the second application interface is displayed, the electronic device determines the second task intention information based on the conversation information and the interface content of the second application interface.
[0194] In the embodiment of the present application, the second task intention information is used to describe the task content of the second task to be executed.
[0195] It should be noted that the electronic device determines the description of the second task intent information based on the conversation information and the interface content of the second application interface. Reference can be made to the electronic device determining the description of the first task intent information based on the conversation information and the interface content of the first application interface in step 202 above. This will not be further elaborated here.
[0196] Step 405: The electronic device controls the second application through the AI assistant to perform the second task according to the second task intention information.
[0197] In an embodiment of the present application, the second application is an application that can execute a second task.
[0198] It should be noted that the description of the electronic device controlling the second application program to perform the second task according to the second task intent information through the AI assistant can refer to the description of the electronic device controlling the first application program to perform the first task according to the first task intent information through the AI assistant in step 203 above. Detailed description is not repeated here.
[0199] In this way, since it is determined that the determination of the missing task elements of the first task requires associating the screen content, but the interface content of the first application interface does not include the missing task elements of the first task, the electronic device determines that the interface currently opened by the user may be the wrong interface, and then can prompt the user to switch to the application interface including the missing task elements of the first task through a prompt message, so as to combine the switched application interface and the dialogue information entered by the user in the AI dialogue window to determine the accurate task intention information, that is, after identifying that the user's intention is the "ask screen" intention, if the screen content and the question raised by the user are not related, the electronic device will prompt the user to switch to the correct screen content. After the user switches the screen, the electronic device will re-identify the information on the screen and automatically reply to the user's question in combination with the screen content. The user no longer needs to manually enter the missing information, but only needs to switch the screen, making it more convenient and efficient for the user to trigger the electronic device to perform the required task.
[0200] In some embodiments of the present application, when the determination of the missing task elements of the first task requires associating with the screen content, but the interface content of the first application interface does not include the missing task elements of the first task, the electronic device can display a second prompt message.
[0201] In some embodiments of the present application, the second prompt information is used to prompt the user to select a first object to be processed.
[0202] In some embodiments of the present application, the first object may be any of the following: a file, an address, an audio, an image, a text, etc. The specific one may be determined according to actual use requirements and is not limited in the embodiments of the present application.
[0203] It can be understood that since the determination of the missing task elements of the first task requires the association of the screen content, but the interface content of the first application interface does not include the missing task elements of the first task, it is speculated that the electronic device may not accurately identify the screen content. The electronic device can prompt the user to select the first object to be processed through the second prompt information.
[0204] For example, the second prompt message may be “The current screen content is not associated with your instruction, please select the object you want to process.”
[0205] In this way, since the determination of the missing task elements of the first task requires the association of the screen content, but the interface content of the first application interface does not include the missing task elements of the first task, the electronic device can prompt the user to select the object to be processed through prompt information, so as to combine the object selected by the user and the dialogue information entered by the user in the AI dialogue window to determine the accurate task intention information, thereby improving the accuracy and flexibility of the electronic device in performing tasks.
[0206] In some embodiments of the present application, Figure 1 ,like Figure 8 As shown, before the above step 202, the task execution method provided by the embodiment of the present application further includes the following steps 501 to 503. In addition, the above step 202 can be specifically implemented by the following step 202b.
[0207] Step 501: The electronic device inputs the display status information and dialogue information of the AI dialogue window into the intent recognition model through the AI assistant, and outputs the second intent type when the display status information indicates that the AI dialogue window is in a floating display state and the dialogue information does not include screen query keywords.
[0208] In some embodiments of the present application, the second intent type may include any one of the following: the first type, the second type.
[0209] In an embodiment of the present application, the first type indicates that the determination of the task elements missing from the first task needs to be associated with the screen content. The second type indicates that the determination of the task elements missing from the first task needs to be associated with the screen content.
[0210] It should be noted that for the detailed description of the above step 501, reference may be made to the description of step 301 in the above embodiment, which will not be repeated here.
[0211] Step 502: When the second intention type is the second type, the electronic device determines whether the dialogue information lacks a task element.
[0212] In an embodiment of the present application, the second type of determination indicating that the first task lacks a task element does not need to be associated with screen content.
[0213] In some embodiments of the present application, the electronic device may perform semantic analysis on the dialogue information to determine whether the dialogue information lacks task elements.
[0214] In some examples, when the dialog information is “Help me navigate to this address”, the dialog information lacks the task element “navigation destination”.
[0215] In some examples, when the dialogue message is “help me find the price of this mobile phone”, the dialogue message lacks the task element “relevant information about the mobile phone”, such as: a picture of the mobile phone, or the model of the mobile phone.
[0216] Step 503: When at least one task element is missing from the dialog information, the electronic device determines whether the interface content of the first application interface includes the missing task element of the first task.
[0217] In some embodiments of the present application, the above step 503 of "the electronic device determines whether the interface content of the first application interface includes the task elements missing from the first task" can be specifically implemented through the following steps 503a and 503b.
[0218] Step 503a: The electronic device takes a screenshot of the first application interface through the AI assistant to obtain a screenshot of the second interface.
[0219] Step 503b: The electronic device inputs the second interface screenshot and the conversation information into the second graphic analysis model on the electronic device side through the AI assistant, outputs second judgment result information, and when the second judgment result information indicates that the interface content of the first application interface includes task element information of the task elements that are missing from the first task, the electronic device determines that the interface content of the first application interface includes the task elements that are missing from the first task.
[0220] In some embodiments of the present application, the above-mentioned second judgment result information may indicate any one of the following: the interface content of the first application interface includes task element information of the task elements that are missing from the first task, and the interface content of the first application interface does not include task element information of the task elements that are missing from the first task.
[0221] It can be understood that when the second judgment result information indicates that the interface content of the first application interface does not include the task element information of the task element missing from the first task, the electronic device determines that the interface content of the first application interface does not include the task element missing from the first task.
[0222] In some embodiments of the present application, the electronic device can generate a sixth instruction based on the dialogue information and the fifth instruction input by the user in the AI dialogue window, and then input the sixth instruction and the second interface screenshot into the second graphic analysis model to obtain the second judgment result information output by the second graphic analysis model.
[0223] For example, the fifth instruction is as follows:
[0224] "Now you are an expert at determining whether the user input and the image content are related. You can combine the image content and the [user input] to determine whether they are related. Please output the final [answer] according to the following requirements.
[0225] <Output requirements>
[0226] 1. Output in a fixed JSON format and do not output anything else.
[0227] 2. The output result is: {"type":"relevant / irrelevant"}
[0228] From now on, please strictly follow the above <Output Requirements> to judge the relevance between the image content and the [user input], and output the final [answer]
[0229]
User input
Command placeholder
[0230]
answer
[0231] Specifically, the electronic device can replace the [instruction placeholder] in the fifth instruction with the dialogue information entered by the user in the AI dialogue window to generate a sixth instruction, and then input the sixth instruction and the second interface screenshot into the second graphic analysis model to obtain the second judgment result information output by the second graphic analysis model.
[0232] In some embodiments of the present application, when the second judgment result information output by the second graphic analysis model is {"type":"related"}, the second judgment result information indicates that the interface content of the first application interface includes task element information of the task elements that are missing from the first task.
[0233] In some embodiments of the present application, when the second judgment result information output by the second graphic analysis model is {"type":"irrelevant"}, the second judgment result information indicates that the interface content of the first application interface does not include task element information of the task elements missing from the first task.
[0234] For example, assuming that an AI conversation window is displayed on a conversation interface between a user and contact A, and the conversation interface includes a chat message sent by contact A, "Let's play badminton tonight at Fandou Garden on Fandou Street in the East District. See you there," and the user enters the conversation message "Help me navigate to the address on the screen" in the AI conversation window, the electronic device may replace the [command placeholder] in the fifth instruction with the conversation message entered by the user in the AI conversation window to generate a sixth instruction, which is as follows:
[0235] "Now you are an expert at determining whether the user input and the image content are related. You can combine the image content and the [user input] to determine whether they are related. Please output the final [answer] according to the following requirements.
[0236] <Output requirements>
[0237] 1. Output in a fixed JSON format and do not output anything else.
[0238] 2. The output result is: {"type":"relevant / irrelevant"}
[0239] From now on, please strictly follow the above <Output Requirements> to judge the relevance between the image content and the [user input], and output the final [answer]
[0240]
User input
Help me navigate to the address on the screen
[0241]
answer
[0242] Then, the electronic device can input the sixth instruction and the interface screenshot of the conversation interface between the user and contact A into the second graphic analysis model. Since the interface screenshot of the conversation interface includes the task element information "navigation destination information" of the task elements that are missing for the task to be executed, the second graphic analysis model can output [Answer]: {"Type":"Related"} to indicate that the interface content of the conversation interface includes the task element information of the task elements that are missing for the task to be executed.
[0243] For another example, suppose an AI conversation window is displayed on a conversation interface between a user and contact B. The conversation interface includes a chat message from contact B, "Are you free at noon? Let's have lunch together." The user enters the conversation message "Help me navigate to the address on the screen" in the AI conversation window. The electronic device may replace the "command placeholder" in the fifth instruction with the conversation message entered by the user in the AI conversation window to generate a sixth instruction. The sixth instruction is as follows:
[0244] "Now you are an expert at determining whether the user input and the image content are related. You can combine the image content and the [user input] to determine whether they are related. Please output the final [answer] according to the following requirements.
[0245] <Output requirements>
[0246] 1. Output in a fixed JSON format and do not output anything else.
[0247] 2. The output result is: {"type":"relevant / irrelevant"}
[0248] From now on, please strictly follow the above <Output Requirements> to judge the relevance between the image content and the [user input], and output the final [answer]
[0249] [User input]: [Are you free at noon? Let's have lunch together]
[0250]
answer
[0251] Then, the electronic device can input the sixth instruction and the interface screenshot of the conversation interface between the user and contact B into the second graphic analysis model. Since the interface screenshot of the conversation interface does not include the task element information "navigation destination information" of the task elements that are missing for the task to be performed, the second graphic analysis model can output [Answer]: {"Type":"Irrelevant"} to indicate that the interface content of the conversation interface does not include the task element information of the task elements that are missing for the task to be performed.
[0252] In this way, since the determination of the missing task elements of the first task needs to be associated with the screen content, the electronic device can use the AI assistant to determine on the electronic device side whether the interface content of the first application interface displayed in the lower layer of the AI dialogue window includes the missing task elements of the first task, so as to determine the task intention information in combination with the dialogue information input by the user in the AI dialogue window and the interface content of the first application interface when the interface content of the first application interface includes the missing task elements of the first task, thereby improving the accuracy and efficiency of determining the task intention information of the task to be executed.
[0253] In this way, since the electronic device can first determine the user's intent by combining the screen content and the question entered by the user in the AI dialogue window, when the user's question is "other", the AI assistant does not first recognize the screen content, and further determines whether the question entered by the user in the AI dialogue window is missing a task element. If at least one task element is missing in the entered question, the AI assistant then recognizes the screen content, thereby minimizing the violation of user privacy. The electronic device can then use the AI assistant to determine whether the interface content of the first application interface displayed below the AI dialogue window includes the missing task element of the first task. When the interface content of the first application interface includes the missing task element of the first task, the task intent information is determined by combining the question entered by the user in the AI dialogue window and the interface content of the first application interface. That is, the user does not need to take a screenshot or circle the screen content, but can directly use voice input or text input in the AI dialogue window to inquire about the content on the electronic device screen in simple and natural communication language. The way to trigger the electronic device to perform tasks is more convenient, and the interaction between the screen and text is also smoother. Therefore, while minimizing the violation of user privacy, it also reduces the problems of heating and reduced battery life caused by frequent screenshots of electronic devices.
[0254] In some embodiments of the present application, when the interface content of the first application interface does not include the task elements missing from the first task, the electronic device may delete the above-mentioned second interface screenshot.
[0255] In some embodiments of the present application, when the electronic device determines whether the interface content of the first application interface includes the task elements missing from the first task, the electronic device may display a first identifier on the first application interface.
[0256] In some examples, the first identifier is used to indicate that the AI assistant is analyzing the second interface screenshot and conversation information, which can also be understood as the AI assistant is understanding the screen content.
[0257] Exemplarily, the first identifier may include any one of the following: a progress bar, a light wave identifier, etc.
[0258] For example: When the electronic device is determining whether the interface content of the first application interface includes the task elements missing from the first task, the electronic device can display a light wave on the first application interface that sweeps across the screen from top to bottom to prompt the user that the AI assistant is understanding the screen content.
[0259] In some embodiments of the present application, the above-mentioned step 503b of "when the second judgment result information indicates that the interface content of the first application interface includes task element information of the task elements that are missing from the first task, the electronic device determines that the interface content of the first application interface includes task elements that are missing from the first task" can be specifically implemented through the following steps 503b1 to 503b3.
[0260] Step 503b1: When the second judgment result information indicates that the interface content of the first application interface includes task element information of the task element missing from the first task, the electronic device sends a screenshot of the second interface and dialogue information to the server through the AI assistant.
[0261] In some embodiments of the present application, when the electronic device sends a screenshot of the second interface to the server through the AI assistant, the electronic device can display a dynamic effect of uploading the image.
[0262] Step 503b2: The electronic device receives the second message sent by the server.
[0263] In an embodiment of the present application, the above-mentioned second message includes third judgment result information.
[0264] In an embodiment of the present application, the third judgment result information is obtained by the server inputting the second interface screenshot and the conversation information into the first graphic analysis model on the server side.
[0265] In some embodiments of the present application, the above-mentioned third judgment result information may indicate any one of the following: the interface content of the first application interface includes task element information of the task elements that are missing from the first task, and the interface content of the first application interface does not include task element information of the task elements that are missing from the first task.
[0266] In some embodiments of the present application, after the server receives the second interface screenshot and conversation information, it can input the second interface screenshot and conversation information into the first graphic analysis model on the server side, so that the first graphic analysis model analyzes the second interface screenshot and conversation information to obtain the third judgment result information output by the first graphic analysis model, and then the server can send a second message including the third judgment result information to the electronic device.
[0267] Step 503b3: When both the third judgment result information and the second judgment result information indicate that the interface content of the first application interface includes task element information of the task element missing from the first task, the electronic device determines that the interface content of the first application interface includes task elements missing from the first task.
[0268] It can be understood that when the second judgment result information indicates that the interface content of the first application interface includes task element information of the task elements that are missing from the first task, the electronic device can send a second interface screenshot and dialogue information to the server, so that the server can further determine whether the interface content of the first application interface includes task element information of the task elements that are missing from the first task. When the third judgment result information output by the first graphic analysis model on the server side also indicates that the interface content of the first application interface includes task element information of the task elements that are missing from the first task, the electronic device can determine that the interface content of the first application interface includes task elements that are missing from the first task.
[0269] In this way, the electronic device determines that the interface content of the first application interface includes the task elements missing from the first task only when the judgment result information output by the first graphic analysis model and the second graphic analysis model both indicate that the interface content of the first application interface includes the task element information that is missing from the first task, and then determines the task intention information by combining the interface content of the first application interface and the dialogue information input by the user in the AI dialogue window, so as to control the application program through the AI assistant to perform the task according to the task intention information, thereby improving the accuracy and efficiency of the electronic device in performing tasks. In addition, since the server side performs a second judgment only after the judgment result information output by the first graphic analysis model on the electronic device side indicates that the interface content of the first application interface includes the task element information that is missing from the first task, the screenshot content can be "not uploaded to the cloud unless necessary", thereby maximizing the protection of user privacy and security.
[0270] Step 202b: When at least one task element is missing from the dialogue information and the interface content of the first application interface includes at least one task element, the electronic device determines the first task intention information based on the dialogue information and the interface content of the first application interface.
[0271] It can be understood that after the determination of the missing task elements of the first task does not need to be associated with the screen content, it means that there is no need to associate the dialogue information input by the user in the AI dialogue window with the screen content. The electronic device can determine the task intention information only based on the dialogue information input by the user in the AI dialogue window; however, when the dialogue information input by the user in the AI dialogue window lacks task elements, the electronic device cannot determine the task intention information. Therefore, it is speculated that the dialogue information input by the user in the AI dialogue window may need to be associated with the screen content, that is, the determination of the missing task elements of the first task needs to be associated with the screen content. Therefore, the electronic device can further determine whether the interface content of the first application interface displayed in the lower layer of the AI dialogue window includes the missing task elements of the first task. When the interface content of the first application interface includes the missing task elements of the first task, the electronic device can determine the first task intention information based on the dialogue information input by the user in the AI dialogue window and the interface content of the first application interface.
[0272] Example 6. In combination with Example 2, after the user long presses the voice control 23 in the preview interface 21 and enters the dialogue information 24 "What kind of butterfly is this" in the AI dialogue window 22, the electronic device can use the AI assistant to determine whether the determination of the missing task elements of the first task needs to be associated with the screen content based on the display status of the AI dialogue window 22 and the dialogue information 24 entered by the user in the AI dialogue window 22; when the AI dialogue window 22 is displayed floating on the shooting preview interface 21 and the dialogue information 24 entered by the user in the AI dialogue window 22 does not contain the question-screen keyword, the electronic device can use the AI assistant to determine that the determination of the missing task elements of the first task does not need to be associated with the screen content; then the electronic device can use the AI assistant to further determine whether the dialogue information 24 entered by the user in the AI dialogue window 22 lacks task elements, and when the user enters When the task element "relevant information about the butterfly" is missing in the dialogue information 24 of the shooting preview interface 21, the electronic device cannot determine the task intention information; therefore, it is speculated that the previous judgment is wrong, and the determination of the missing task element of the first task actually needs to be associated with the screen content; therefore, the electronic device can further determine through the AI assistant whether the interface content of the shooting preview interface 21 includes the missing task element "relevant information about the butterfly". When the interface content of the shooting preview interface 21 includes a butterfly image, that is, includes the missing task element "relevant information about the butterfly", the electronic device can determine the first task intention information through the AI assistant in combination with the dialogue information 24 "What kind of butterfly is this" input by the user in the AI dialogue window 22 and the butterfly image in the shooting preview interface 21. Then, the electronic device can control the browser application through the AI assistant to identify the type of butterfly on the screen according to the first task intention information, and, as shown in FIG. Figure 9 As shown, the AI dialogue window 22 displays the recognition result 25 “the species of this butterfly is Papilio palinurus”.
[0273] In this way, after the determination of the missing task elements of the first task does not need to be associated with the screen content, but the task elements are missing in the dialogue information input by the user in the AI dialogue window, the electronic device can further determine whether the interface content of the first application interface displayed in the lower layer of the AI dialogue window includes the task elements missing from the first task. When the interface content of the first application interface includes the task elements missing from the first task, the electronic device can determine the first task intention information in combination with the dialogue information input by the user in the AI dialogue window and the interface content of the first application interface, and then control the application through the AI assistant to perform the task corresponding to the task intention information, thereby improving the accuracy and efficiency of the electronic device in performing tasks.
[0274] In some embodiments of the present application, after the above step 503, the task execution method provided by the embodiment of the present application further includes the following steps 601 and 602.
[0275] Step 601: When the task elements in the dialogue information are complete, the electronic device determines third task intention information based on the dialogue information.
[0276] In the embodiment of the present application, the third task intention information is determined without associating it with the screen content, and the third task intention information is used to describe the task content of the third task to be executed.
[0277] For example, when the user inputs the dialogue information "What family does butterfly belong to?" in the AI dialogue window, the task elements in the dialogue information are complete, so the electronic device can determine the third task intention information as "Search what family does butterfly belong to" based on the dialogue information.
[0278] In some embodiments of the present application, when the task elements in the dialogue information are complete, the electronic device may perform semantic analysis on the dialogue information to determine the third task intention information.
[0279] It can be understood that when the task elements in the dialogue information are complete, the execution of the task does not require associated screen content. Therefore, the electronic device can determine the third task intention information based only on the dialogue information entered by the user in the AI dialogue window, thereby controlling the third application through the AI assistant to execute the third task according to the third task intention information.
[0280] It can be understood that the third task is different from the first task mentioned above. The execution of the third task does not need to be associated with screen content, while the execution of the first task needs to be associated with screen content.
[0281] Step 602: The electronic device controls the third application through the AI assistant to perform the third task according to the third task intention information.
[0282] In an embodiment of the present application, the third application is an application that can execute a third task.
[0283] It should be noted that the description of the electronic device controlling the application program to perform the task corresponding to the third task intent information through the AI assistant can be referred to the description of the electronic device controlling the first application program to perform the first task according to the first task intent information through the AI assistant in step 203 above. Detailed description is omitted here.
[0284] Example 7: Figure 3A As shown, in the shooting preview interface 21, that is, when the first application interface mentioned above contains a butterfly, if the user wants to know what family the butterfly belongs to, he can long press the power button to wake up the AI assistant, and Figure 3B As shown, an AI dialogue window 22 is displayed floating on the shooting preview interface 21, and the user can long press the voice control 23 in the AI dialogue window 22, such as Figure 10A As shown, by inputting the dialogue information 51 "What family does the butterfly belong to?" in the AI dialogue window 22, the electronic device can determine through the AI assistant whether the determination of the missing task elements of the third task needs to be associated with the screen content according to the display state of the AI dialogue window 22 and the dialogue information 51 input by the user in the AI dialogue window 22; when the AI dialogue window 22 is displayed floating on the shooting preview interface 21 and the dialogue information 51 input by the user in the AI dialogue window 22 does not contain the screen query keyword, the electronic device can determine through the AI assistant that the determination of the missing task elements of the third task does not need to be associated with the screen content; then the electronic device can further determine through the AI assistant whether the dialogue information 51 input by the user in the AI dialogue window 23 lacks the task element; when the dialogue information 51 "What family does the butterfly belong to?" input by the user in the AI dialogue window 23 has complete task elements, the electronic device can determine through the AI assistant that the third task intention information is "search what family the butterfly belongs to" according to the dialogue information 51 input by the user in the AI dialogue window 23, and then the electronic device can control the browser application through the AI assistant to search what family the butterfly belongs to according to the third task intention information, that is, execute the third task, and as shown in FIG. Figure 10B As shown, the AI dialogue window 22 displays the recognition result 52 "butterfly in biological taxonomy refers to a superfamily-level evolutionary branch called Papilioidea in the Lepidoptera, also known as True Butterflyoidea."
[0285] In some embodiments of the present application, after the above step 503, the task execution method provided by the embodiment of the present application further includes the following steps 701 to 704.
[0286] Step 701: When at least one task element is missing from the dialogue information and the interface content of the first application interface does not include the missing task element of the first task, the electronic device displays an inquiry message in the AI dialogue window.
[0287] In an embodiment of the present application, the inquiry message is used to inquire about missing task elements of the first task.
[0288] For example: Suppose the dialogue information entered by the user in the AI dialogue window is "Help me navigate to the address on the screen", the dialogue information lacks the task element "navigation destination", and the interface content of the first application interface displayed in the lower layer of the AI dialogue window does not include the task element, then the electronic device can display the inquiry message "Where is the navigation destination?" in the AI dialogue window.
[0289] It can be understood that since it is first determined that the determination of the missing task elements of the first task does not need to be associated with the screen content, the electronic device can determine the task intention information only based on the dialogue information entered by the user in the AI dialogue window; however, the dialogue information entered by the user in the AI dialogue window lacks at least one task element, and the electronic device cannot determine the task intention information. Therefore, it is speculated that the dialogue information entered by the user in the AI dialogue window may still need to be associated with the screen content. Therefore, the electronic device can further determine whether the interface content of the first application interface includes the task elements missing from the first task. When the interface content of the first application interface also does not include the task elements missing from the first task, it means that the dialogue information entered by the user in the AI dialogue window is incomplete. Therefore, the electronic device can display an inquiry message in the AI dialogue window to obtain at least one missing task element for performing the first task, so that the electronic device can determine the fourth task intention information based on the dialogue information and reply information entered by the user in the AI dialogue window, thereby controlling the fourth application to perform the fourth task according to the fourth task intention information through the AI assistant.
[0290] Step 702: The electronic device receives reply information input by the user in the AI dialogue window based on the inquiry information.
[0291] In some embodiments of the present application, the reply information includes at least one missing task element for executing the first task.
[0292] In some examples, the reply information may include at least one of the following: an image, text, a file, audio, etc.
[0293] For example, when the dialogue message is “Help me navigate to this address”, the dialogue message lacks the task element “navigation destination”, and the reply message may include the navigation destination.
[0294] For example, when the dialogue message is "Help me check the price of this mobile phone", the dialogue message lacks the task element "relevant information about the mobile phone", and the reply message may include a picture of the mobile phone.
[0295] Step 703: The electronic device determines fourth task intention information based on the conversation information and the reply information.
[0296] In the embodiment of the present application, the fourth task intention information is used to describe the task content of the fourth task to be executed.
[0297] For example: when the dialogue information entered by the user in the AI dialogue window is "Help me navigate to the address on the screen", and the user enters the reply information "Fandou Garden on Fandou Street in the East District" in the AI dialogue window based on the inquiry information "Where is the navigation destination?", the electronic device can determine that the fourth task intention information is "Navigate to Fandou Garden on Fandou Street in the East District" based on the dialogue information and the reply information.
[0298] In some embodiments of the present application, when at least one task element is missing in the conversation information and the interface content of the first application interface does not include the missing task element of the first task, the electronic device can perform semantic analysis on the conversation information and the reply information to obtain the above-mentioned fourth task intention information.
[0299] Step 704: The electronic device controls the fourth application through the AI assistant to perform the fourth task according to the fourth task intention information.
[0300] In the embodiment of the present application, the fourth application is an application that can execute the fourth task.
[0301] Example 8, such as Figure 6A As shown, when the conversation interface 31 with contact B is displayed, if the user wants to navigate to the Fandou Garden on Fandou Street in the East District, the user can long press the power button to wake up the AI assistant; Figure 6B As shown, the electronic device displays an AI dialogue window 32 on the conversation interface 31 in a floating manner. The user can long press the voice control 33 in the AI dialogue window 32 and enter the dialogue message "Help me navigate to the Doudou Garden on Doudou Street in the East District" in the AI dialogue window 32. Figure 11AAs shown, assuming that during the process of inputting the dialogue information, the user's finger leaves the screen in advance, resulting in the electronic device only inputting the dialogue information 61 "Help me navigate" in the AI dialogue window 32, the electronic device can use the AI assistant to determine whether the determination of the missing task elements of the fourth task needs to be associated with the screen content according to the display state of the AI dialogue window 32 and the dialogue information 61 input by the user in the AI dialogue window 32; when the AI dialogue window 32 is displayed floating on the conversation interface 31 and the dialogue information 31 input by the user in the AI dialogue window 32 does not contain the screen query keyword, the electronic device can use the AI assistant to determine that the determination of the missing task elements of the fourth task does not need to be associated with the screen content, so the electronic device can use the AI assistant only based on the display state of the AI dialogue window 32 and the dialogue information 61 input by the user in the AI dialogue window 32. The electronic device can determine the task intention information by using the dialogue information 61 input by the user in the dialogue window 32; then the electronic device can further determine whether at least one task element is missing in the dialogue information 61 through the AI assistant. When the task element "navigation destination" is missing in the dialogue information 61, the electronic device cannot determine the task intention information. Therefore, it is speculated that the dialogue information 61 input by the user in the AI dialogue window 32 may still need to be associated with the screen content. Therefore, the electronic device can further determine whether the interface content of the conversation interface 31 includes the task element missing for the fourth task through the AI assistant. When the interface content of the conversation interface 31 does not include the task element "navigation destination" missing for the fourth task, it means that the dialogue information 61 input by the user in the AI dialogue window 32 itself is incomplete. Figure 11B As shown, the electronic device can display a query message 62 "Where is the navigation destination?" in the AI dialogue window 32 through the AI assistant to obtain the missing task elements; then the user can long press the voice control 33 in the AI dialogue window 32, as shown in FIG. Figure 11C As shown, the reply information 63 "Fandou Garden on Fandou Street in the East District" is entered in the AI dialogue window 32, and then the electronic device can use the AI assistant to determine that the fourth task intention information is "Navigate to the Fandou Garden on Fandou Street in the East District" based on the dialogue information 61 and the reply information 63 entered by the user in the AI dialogue window 32. The electronic device can then control the navigation application through the AI assistant to navigate to the Fandou Garden on Fandou Street in the East District according to the fourth task intention information, that is, to execute the fourth task.
[0302] It should be noted that the description of the electronic device controlling the fourth application through the AI assistant to perform the fourth task according to the fourth task intention information can be referred to the description of the electronic device controlling the first application through the AI assistant to perform the first task according to the first task intention information in the above step 203, which will not be repeated here.
[0303] In this way, when it is determined that the dialogue information input by the user in the AI dialogue window is incomplete and the task elements for task execution are missing, the electronic device can display an inquiry message in the AI dialogue window to obtain the missing task elements for task execution. Therefore, the electronic device can determine the accurate task intention information, and then control the application through the AI assistant to execute the task corresponding to the task intention information, thereby improving the reliability and flexibility of the electronic device in performing tasks.
[0304] It should be noted that the task execution method provided in the embodiment of the present application can be performed by a task execution device. In the embodiment of the present application, the task execution device provided in the embodiment of the present application is described by taking the task execution method performed by the task execution device as an example.
[0305] It should be noted that the above-mentioned method embodiments, or various possible implementation methods in each method embodiment, can be executed separately, or, under the premise that there is no contradiction, can also be executed in combination with each other. The specific implementation can be determined according to actual usage requirements, and the embodiments of this application do not limit this.
[0306] Figure 12 A possible structural diagram of the task execution device involved in the embodiment of the present application is shown. Figure 12 As shown, the task execution device 70 may include: a receiving module 71, a determining module 72 and a processing module 73;
[0307] The receiving module 71 is configured to receive, when displaying an AI dialogue window between the user and the AI assistant, dialogue information input by the user in the AI dialogue window; the dialogue information includes description information of a first task to be performed, and the dialogue information lacks at least one task element for performing the first task;
[0308] Determining module 73, configured to determine first task intent information based on the dialogue information received by receiving module 71 and the interface content of the first application interface, and display the AI dialogue window on the first application interface; the first task intent information is used to describe the task content of the first task;
[0309] The processing module 73 is used to control the first application through the AI assistant to perform the first task according to the first task intention information determined by the determination module 72; wherein the first application is an application that can execute the first task.
[0310] An embodiment of the present application provides a task execution device. After the user inputs conversation information in the AI conversation window, the electronic device can automatically combine the conversation information input by the user in the AI conversation window and the interface content of the first application interface displayed in the lower layer of the AI conversation window to determine the user's task intention information, and then control the first application through the AI assistant to execute the task corresponding to the task intention information. That is, the user does not need to manually take screenshots or manually circle the screen content. The user only needs to input voice or text input instructions in the AI conversation window, and the AI assistant can understand the screen content and automatically execute the user's instructions. Therefore, while avoiding problems such as heat and reduced battery life caused by frequent screenshots of the electronic device, it also reduces the user's operating steps to trigger the electronic device to perform the required tasks, thereby improving the efficiency of the electronic device in performing tasks.
[0311] In one possible implementation, the processing module 73 is also used to input the display status information and dialogue information of the AI dialogue window into the intention recognition model through the AI assistant before the determination module 72 determines the first task intention information based on the dialogue information and the interface content of the first application interface, and output the first intention type when the display status information indicates that the AI dialogue window is in a floating display state and the dialogue information includes screen inquiry keywords. The determination module 72 is also used to determine whether the interface content of the first application interface includes task elements that are missing from the first task when the first intention type is the first type; wherein the determination of the task elements that are missing from the first task indicated by the first type requires associating the screen content. The determination module 72 is specifically used to determine the first task intention information based on the dialogue information and the interface content of the first application interface when the first intention type is the first type and the interface content of the first application interface includes task elements that are missing from the first task.
[0312] In a possible implementation, the task execution device provided by the embodiment of the present application also includes: a sending module. The processing module 73 is also used to take a screenshot of the first application interface through the AI assistant to obtain a first interface screenshot. The sending module is used to send the first interface screenshot and conversation information obtained by the processing module 73 to the server through the AI assistant. The receiving module 71 is also used to receive a first message sent by the server, and the first message includes first judgment result information. The determination module 72 is specifically used to determine that the interface content of the first application interface includes the task elements missing from the first task when the first judgment result information received by the receiving module 71 indicates that the interface content of the first application interface includes the task element information of the task elements missing from the first task; wherein the first judgment result information is obtained by the server inputting the first interface screenshot and conversation information into the first graphic analysis model on the server side.
[0313] In one possible implementation, the task execution device provided in an embodiment of the present application further includes a display module. The display module is configured to, after the determination module 72 determines whether the interface content of the first application interface includes the task elements missing from the first task, display a first prompt message if the interface content of the first application interface does not include the task elements missing from the first task. The first prompt message prompts the user to switch to the application interface that includes the task elements missing from the first task. The receiving module 71 is further configured to receive user interface switching input. The processing module 73 is further configured to, in response to the interface switching input received by the receiving module 71, switch the first application interface to a second application interface, wherein the interface content of the second application interface includes the task elements missing from the first task. The AI dialogue window remains displayed during the switching process. The determination module 72 is further configured to, upon detecting the display of the second application interface through the AI assistant, determine second task intent information based on the dialogue information and the interface content of the second application interface. The second task intent information describes the task content of the second task to be executed. The processing module 73 is further configured to, through the AI assistant, control a second application to execute the second task in accordance with the second task intent information determined by the determination module 72; wherein the second application is an application capable of executing the second task.
[0314] In one possible implementation, the processing module 73 is further configured to input the display status information and the dialogue information of the AI dialogue window into the intent recognition model through the AI assistant before the determination module 72 determines the first task intent information based on the dialogue information and the interface content of the first application interface, and output a second intent type when the display status information indicates that the AI dialogue window is in a suspended display state and the dialogue information does not include screen query keywords. The determination module 72 is further configured to determine whether the dialogue information lacks task elements when the second intent type is the second type; the second type indicates that the determination of the missing task elements of the first task does not require associating with the screen content; and when at least one task element is missing from the dialogue information, determine whether the interface content of the first application interface includes the missing task elements of the first task. The determination module 72 is specifically configured to determine the first task intent information based on the dialogue information and the interface content of the first application interface when at least one task element is missing from the dialogue information and the interface content of the first application interface includes at least one task element.
[0315] In one possible implementation, the processing module 73 is specifically used to take a screenshot of the first application interface through the AI assistant to obtain a second interface screenshot; and input the second interface screenshot and conversation information into the second graphic analysis model on the electronic device side through the AI assistant to output second judgment result information; and when the second judgment result information indicates that the interface content of the first application interface includes task element information of the task elements that are missing from the first task, determine that the interface content of the first application interface includes the task elements that are missing from the first task.
[0316] In one possible implementation, the task execution device provided by the embodiment of the present application also includes: a sending module. The sending module is used to send the second interface screenshot and dialogue information to the server through the AI assistant when the second judgment result information indicates that the interface content of the first application interface includes task element information of the task elements that are missing from the first task. The receiving module 71 is also used to receive a second message sent by the server, and the second message includes third judgment result information. The determination module 72 is specifically used to determine that the interface content of the first application interface includes the task elements that are missing from the first task when the third judgment result information and the second judgment result information received by the receiving module 71 both indicate that the interface content of the first application interface includes task element information of the task elements that are missing from the first task; wherein the third judgment result information is obtained by the server inputting the second interface screenshot and dialogue information into the first graphic analysis model on the server side.
[0317] In one possible implementation, determination module 72 is further configured to, after determining whether the conversation information lacks task elements, determine third task intent information based on the conversation information if the task elements are complete. The third task intent information is determined without associating it with screen content and is used to describe the task content of the third task to be executed. Processing module 73 is further configured to control, through the AI assistant, a third application to execute the third task in accordance with the third task intent information determined by determination module 72; the third application is an application capable of executing the third task.
[0318] In one possible implementation, the task execution device provided by the embodiment of the present application further includes: a display module. The display module is used to display an inquiry message in the AI dialogue window after the determination module 72 determines whether the interface content of the first application interface includes the task elements missing from the first task, if at least one task element is missing from the dialogue information and the interface content of the first application interface does not include the task elements missing from the first task. The inquiry message is used to inquire about the task elements missing from the first task. The receiving module 71 is also used to receive the reply information input by the user in the AI dialogue window based on the inquiry information. The determination module 72 is also used to determine the fourth task intention information based on the dialogue information and the reply information received by the receiving module 71; the fourth task intention information is used to describe the task content of the fourth task to be executed. The processing module 73 is also used to control the fourth application through the AI assistant to perform the fourth task according to the fourth task intention information determined by the determination module 72; wherein the fourth application is an application that can execute the fourth task.
[0319] The task execution device in the embodiment of the present application can be an electronic device, or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or a device other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., and the embodiment of the present application does not specifically limit it.
[0320] The task execution device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0321] The task execution device provided in the embodiment of the present application can implement each process implemented in the above method embodiment. To avoid repetition, it will not be described here.
[0322] Alternatively, as Figure 13 As shown, an embodiment of the present application also provides an electronic device 900, including a processor 901 and a memory 902, wherein the memory 902 stores a program or instruction that can be run on the processor 901, and when the program or instruction is executed by the processor 901, the various steps of the above-mentioned method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0323] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.
[0324] Figure 14 A schematic diagram of the hardware structure of an electronic device implementing an embodiment of the present application.
[0325] The electronic device 100 includes but is not limited to components such as a radio frequency unit 101 , a network module 102 , an audio output unit 103 , an input unit 104 , a sensor 105 , a display unit 106 , a user input unit 107 , an interface unit 108 , a memory 109 , and a processor 110 .
[0326] Those skilled in the art will understand that the electronic device 100 may also include a power source (such as a battery) to power each component, and the power source may be logically connected to the processor 110 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system. Figure 14 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be repeated here.
[0327] The user input unit 107 is configured to receive, when displaying an AI dialogue window between the user and the AI assistant, dialogue information input by the user in the AI dialogue window; the dialogue information includes description information of a first task to be performed, and the dialogue information lacks at least one task element for performing the first task;
[0328] The processor 110 is used to determine the first task intention information based on the dialogue information received by the user input unit 107 and the interface content of the first application interface, and the AI dialogue window is displayed on the first application interface; the first task intention information is used to describe the task content of the first task; and the AI assistant is used to control the first application to perform the first task according to the first task intention information determined by the determination module; wherein the first application is an application that can execute the first task.
[0329] An embodiment of the present application provides an electronic device. After a user inputs conversation information in an AI conversation window, the electronic device can automatically combine the conversation information input by the user in the AI conversation window and the interface content of the first application interface displayed in the lower layer of the AI conversation window to determine the user's task intention information, and then control the first application through the AI assistant to perform the task corresponding to the task intention information. That is, the user does not need to manually take screenshots or manually circle the screen content. The user only needs to input voice or text input instructions in the AI conversation window, and the AI assistant can understand the screen content and automatically execute the user's instructions. Therefore, while avoiding problems such as heat and reduced battery life caused by frequent screenshots of the electronic device, it also reduces the user's operating steps to trigger the electronic device to perform the required tasks, thereby improving the efficiency of the electronic device in performing tasks.
[0330] In some embodiments of the present application, the processor 110 is further used to input the display status information and dialogue information of the AI dialogue window into the intention recognition model through the AI assistant before determining the first task intention information based on the dialogue information and the interface content of the first application interface, and output the first intention type when the display status information indicates that the AI dialogue window is in a floating display state and the dialogue information includes screen inquiry keywords; and when the first intention type is the first type, determine whether the interface content of the first application interface includes task elements that are missing from the first task; wherein, the first type indicates that the determination of the task elements that are missing from the first task requires associating with the screen content.
[0331] The processor 110 is specifically configured to determine the first task intention information based on the dialogue information and the interface content of the first application interface when the first intent type is the first type and the interface content of the first application interface includes task elements missing from the first task.
[0332] In some embodiments of the present application, the processor 110 is further configured to take a screenshot of the first application interface through an AI assistant to obtain a screenshot of the first interface.
[0333] The radio frequency unit 101 is used to send the first interface screenshot and dialogue information obtained by the processor 110 to the server through the AI assistant; and receive the first message sent by the server, where the first message includes the first judgment result information.
[0334] The processor 110 is specifically configured to determine that the interface content of the first application interface includes task elements that are missing from the first task when the first judgment result information received by the radio frequency unit 101 indicates that the interface content of the first application interface includes task element information of the task elements that are missing from the first task; wherein the first judgment result information is obtained by the server inputting the first interface screenshot and conversation information into the first graphic and text analysis model on the server side.
[0335] In some embodiments of the present application, the display unit 106 is used to display a first prompt message after the processor 110 determines whether the interface content of the first application interface includes the task elements that are missing from the first task, if the interface content of the first application interface does not include the task elements that are missing from the first task. The first prompt message is used to prompt the user to switch to the application interface that includes the task elements that are missing from the first task.
[0336] The user input unit 107 is further configured to receive user interface switching input.
[0337] The processor 110 is also used to switch the first application interface to the second application interface in response to the interface switching input received by the user input unit 107, where the interface content of the second application interface includes task elements that are missing from the first task. During the process of switching the first application interface to the second application interface, the AI dialogue window remains displayed; and when the AI assistant detects that the second application interface is displayed, the second task intention information is determined based on the dialogue information and the interface content of the second application interface, where the second task intention information is used to describe the task content of the second task to be executed; and the AI assistant is used to control the second application to execute the second task according to the second task intention information determined by the processor 110; wherein the second application is an application that can execute the second task.
[0338] In some embodiments of the present application, the processor 110 is specifically used to input the display status information and dialogue information of the AI dialogue window into the intent recognition model through the AI assistant before determining the first task intention information based on the dialogue information and the interface content of the first application interface, and output a second intention type when the display status information indicates that the AI dialogue window is in a floating display state and the dialogue information does not include screen inquiry keywords; and when the second intention type is the second type, determine whether the dialogue information lacks task elements; the second type indicates that the determination of the task elements missing from the first task does not require associating screen content; and when at least one task element is missing in the dialogue information, determine whether the interface content of the first application interface includes the task elements missing from the first task; and when at least one task element is missing in the dialogue information and the interface content of the first application interface includes at least one task element, determine the first task intention information based on the dialogue information and the interface content of the first application interface.
[0339] In some embodiments of the present application, the processor 110 is specifically used to take a screenshot of the first application interface through an AI assistant to obtain a second interface screenshot; and input the second interface screenshot and conversation information into a second graphic analysis model on the electronic device side through the AI assistant to output second judgment result information; and when the second judgment result information indicates that the interface content of the first application interface includes task element information of the task elements that are missing from the first task, determine that the interface content of the first application interface includes task elements that are missing from the first task.
[0340] In some embodiments of the present application, the radio frequency unit 101 is also used to send a second interface screenshot and dialogue information to the server through the AI assistant when the second judgment result information indicates that the interface content of the first application interface includes task element information of the task element missing from the first task; and receive a second message sent by the server, the second message including third judgment result information.
[0341] The processor 110 is specifically used to determine that the interface content of the first application interface includes task elements that are missing from the first task when the third judgment result information and the second judgment result information received by the user input unit 10771 both indicate that the interface content of the first application interface includes task element information that is missing from the first task; wherein the third judgment result information is obtained by the server inputting the second interface screenshot and dialogue information into the first graphic analysis model on the server side.
[0342] In some embodiments of the present application, the processor 110 is further used to, after determining whether the task elements are missing in the dialogue information, determine third task intention information based on the dialogue information when the task elements in the dialogue information are complete, where the third task intention information is determined without associating the screen content, and the third task intention information is used to describe the task content of the third task to be executed; and control the third application through the AI assistant to execute the third task according to the third task intention information; wherein the third application is an application that can execute the third task.
[0343] In some embodiments of the present application, the display unit 106 is also used to display an inquiry message in the AI dialogue window after the processor 110 determines whether the interface content of the first application interface includes the task elements missing from the first task, if at least one task element is missing in the dialogue information and the interface content of the first application interface does not include the task elements missing from the first task, and the inquiry message is used to inquire about the task elements missing from the first task.
[0344] The user input unit 107 is further configured to receive reply information input by the user in the AI dialogue window based on the query information.
[0345] The processor 110 is further used to determine fourth task intention information based on the dialogue information and the reply information received by the user input unit 107; the fourth task intention information is used to describe the task content of the fourth task to be executed; and to control the fourth application through the AI assistant to execute the fourth task according to the fourth task intention information; wherein the fourth application is an application that can execute the fourth task.
[0346] The electronic device provided in the embodiment of the present application can implement each process implemented in the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described here.
[0347] The beneficial effects of various implementations in this embodiment can be specifically referred to the beneficial effects of the corresponding implementations in the above method embodiment. To avoid repetition, they will not be described here.
[0348] It should be understood that in an embodiment of the present application, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042, and the graphics processor 1041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 107 includes a touch panel 1071 and at least one of other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include two parts: a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an operating stick, which will not be repeated here.
[0349] The memory 109 can be used to store software programs and various data. The memory 109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 109 may include a volatile memory or a non-volatile memory, or the memory 109 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct memory bus random access memory (DRRAM). The memory 109 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0350] Processor 110 may include one or more processing units. Optionally, processor 110 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 110.
[0351] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0352] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0353] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned method embodiment and achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0354] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0355] An embodiment of the present application provides a computer program / program product, which is stored in a storage medium. The program / program product is executed by at least one processor to implement the various processes of the above-mentioned method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0356] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0357] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0358] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A task execution method, characterized in that: The method comprises: In a case where an AI dialogue window between a user and an artificial intelligence (AI) assistant is displayed, receiving dialogue information input by the user in the AI dialogue window; the dialogue information includes description information of a first task to be performed, and the dialogue information lacks at least one task element for performing the first task; Determining first task intent information based on the dialogue information and the interface content of the first application interface, and displaying the AI dialogue window on the first application interface; the first task intent information is used to describe the task content of the first task; The AI assistant controls the first application to execute the first task according to the first task intention information; wherein the first application is an application that can execute the first task.
2. The method according to claim 1, characterized in that Before determining the first task intention information based on the dialogue information and the interface content of the first application interface, the method further includes: Inputting the display state information of the AI dialogue window and the dialogue information into an intent recognition model through the AI assistant, and outputting a first intent type when the display state information indicates that the AI dialogue window is in a suspended display state and the dialogue information includes a screen query keyword; In a case where the first intent type is a first type, determining whether the interface content of the first application interface includes a task element missing from the first task; wherein the first type indicates that determination of the task element missing from the first task requires associating with screen content; The determining of the first task intention information according to the dialogue information and the interface content of the first application interface includes: When the first intent type is the first type and the interface content of the first application interface includes task elements missing from the first task, the first task intent information is determined based on the dialogue information and the interface content of the first application interface.
3. The method according to claim 2, characterized in that The determining whether the interface content of the first application interface includes task elements missing from the first task includes: Taking a screenshot of the first application interface by the AI assistant to obtain a first interface screenshot; Sending the first interface screenshot and the conversation information to the server through the AI assistant; receiving a first message sent by the server, where the first message includes first judgment result information; In a case where the first judgment result information indicates that the interface content of the first application interface includes task element information of the task element missing from the first task, determining that the interface content of the first application interface includes the task element missing from the first task; The first judgment result information is obtained by the server inputting the first interface screenshot and the conversation information into a first graphic analysis model on the server side.
4. The method according to claim 2, characterized in that After determining whether the interface content of the first application interface includes task elements missing from the first task, the method further includes: If the interface content of the first application interface does not include the task elements missing from the first task, displaying first prompt information, where the first prompt information is used to prompt the user to switch to the application interface that includes the task elements missing from the first task; Receive user interface switching input; In response to the interface switching input, the first application interface is switched to a second application interface, where the interface content of the second application interface includes task elements that are missing from the first task, and the AI dialogue window remains displayed during the process of switching from the first application interface to the second application interface; When the AI assistant detects that the second application interface is displayed, determining second task intent information based on the conversation information and the interface content of the second application interface, where the second task intent information is used to describe the task content of the second task to be performed; The AI assistant controls the second application to execute the second task according to the second task intention information; wherein the second application is an application that can execute the second task.
5. The method according to claim 1, wherein Before determining the first task intention information based on the conversation information and the interface content of the first application interface, the method further includes: Inputting the display state information of the AI dialogue window and the dialogue information into an intent recognition model through the AI assistant, and outputting a second intent type when the display state information indicates that the AI dialogue window is in a suspended display state and the dialogue information does not include a screen query keyword; In the case where the second intention type is the second type, determining whether the dialog information lacks a task element; the second type indicates that the determination of the task element missing from the first task does not need to be associated with screen content; In a case where at least one task element is missing from the conversation information, determining whether the interface content of the first application interface includes the task element missing from the first task; The determining of the first task intention information according to the dialogue information and the interface content of the first application interface includes: When the dialog information lacks the at least one task element and the interface content of the first application interface includes the at least one task element, the first task intention information is determined according to the dialog information and the interface content of the first application interface.
6. The method according to claim 5, characterized in that The determining whether the interface content of the first application interface includes task elements missing from the first task includes: Taking a screenshot of the first application interface through the AI assistant to obtain a screenshot of the second interface; The AI assistant inputs the second interface screenshot and the conversation information into the second graphic analysis model on the electronic device side, outputs second judgment result information, and when the second judgment result information indicates that the interface content of the first application interface includes task element information of the task elements that are missing from the first task, determines that the interface content of the first application interface includes the task elements that are missing from the first task.
7. The method according to claim 6, characterized in that When the second judgment result information indicates that the interface content of the first application interface includes task element information of the task element missing from the first task, determining that the interface content of the first application interface includes the task element missing from the first task includes: If the second judgment result information indicates that the interface content of the first application interface includes task element information of the task element missing from the first task, sending the second interface screenshot and the conversation information to the server through the AI assistant; receiving a second message sent by the server, where the second message includes third judgment result information; When both the third judgment result information and the second judgment result information indicate that the interface content of the first application interface includes task element information of the task element missing from the first task, determining that the interface content of the first application interface includes the task element missing from the first task; The third judgment result information is obtained by the server inputting the second interface screenshot and the conversation information into the first graphic analysis model on the server side.
8. The method according to claim 5, characterized in that After determining whether the dialogue information lacks task elements, the method further includes: When the task elements in the dialogue information are complete, determining third task intention information based on the dialogue information, wherein the third task intention information is determined without associating it with screen content, and the third task intention information is used to describe the task content of the third task to be executed; The AI assistant controls a third application to execute the third task according to the third task intention information; wherein the third application is an application that can execute the third task.
9. The method according to claim 5, characterized in that After determining whether the interface content of the first application interface includes task elements missing from the first task, the method further includes: If at least one task element is missing from the dialogue information and the interface content of the first application interface does not include the missing task element of the first task, displaying an inquiry message in the AI dialogue window, wherein the inquiry message is used to inquire about the missing task element of the first task; Receiving a reply message input by the user in the AI dialogue window based on the query message; Determining fourth task intention information according to the conversation information and the reply information; the fourth task intention information is used to describe the task content of the fourth task to be performed; The fourth application is controlled by the AI assistant to execute the fourth task according to the fourth task intention information; wherein the fourth application is an application that can execute the fourth task.
10. A task execution device, characterized in that: include: A receiving module, configured to receive dialogue information input by the user in the AI dialogue window when the AI dialogue window between the user and the AI assistant is displayed; The dialogue information includes description information of a first task to be performed, and the dialogue information lacks at least one task element for performing the first task; a determining module, configured to determine first task intent information based on the dialogue information received by the receiving module and the interface content of the first application interface, wherein the AI dialogue window is displayed on the first application interface; the first task intent information is used to describe the task content of the first task; A processing module is used to control the first application through the AI assistant to perform the first task according to the first task intention information determined by the determination module; wherein the first application is an application that can execute the first task.
11. The device according to claim 10, characterized in that The processing module is further configured to input the display state information of the AI dialogue window and the dialogue information into an intent recognition model through the AI assistant before the determination module determines the first task intention information based on the dialogue information and the interface content of the first application interface, and output a first intent type when the display state information indicates that the AI dialogue window is in a suspended display state and the dialogue information includes a screen query keyword; The determining module 72 is further configured to, when the first intent type is a first type, determine whether the interface content of the first application interface includes a task element missing from the first task; wherein the first type indicates that the determination of the task element missing from the first task requires association with screen content; The determination module is specifically used to determine the first task intention information based on the dialogue information and the interface content of the first application interface when the first intention type is the first type and the interface content of the first application interface includes task elements missing from the first task.
12. The device according to claim 11, characterized in that The device further includes: a sending module; The processing module is further configured to take a screenshot of the first application interface through the AI assistant to obtain a first interface screenshot; The sending module is configured to send the first interface screenshot and the conversation information obtained by the processing module through the AI assistant to the server; The receiving module is further configured to receive a first message sent by the server, where the first message includes first judgment result information; The determining module is specifically configured to determine that the interface content of the first application interface includes the task element missing from the first task, if the first judgment result information received by the receiving module indicates that the interface content of the first application interface includes the task element information of the task element missing from the first task; The first judgment result information is obtained by the server inputting the first interface screenshot and the conversation information into a first graphic analysis model on the server side.
13. The device according to claim 11, characterized in that The device further includes: a display module, the display module is configured to, after the determination module determines whether the interface content of the first application interface includes the task elements missing from the first task, display a first prompt message if the interface content of the first application interface does not include the task elements missing from the first task, the first prompt message being configured to prompt a user to switch to an application interface that includes the task elements missing from the first task; The receiving module is further configured to receive user interface switching input; The processing module is further configured to, in response to the interface switching input received by the receiving module, switch the first application interface to a second application interface, wherein the interface content of the second application interface includes task elements missing from the first task, and the AI dialogue window remains displayed during the process of switching from the first application interface to the second application interface; The determining module is further configured to, when the AI assistant detects that the second application interface is displayed, determine second task intent information based on the conversation information and the interface content of the second application interface, where the second task intent information is used to describe the task content of the second task to be executed; The processing module is further used to control the second application through the AI assistant to perform the second task according to the second task intention information determined by the determination module; wherein the second application is an application that can execute the second task.
14. The device according to claim 10, characterized in that The processing module is further configured to input the display state information of the AI dialogue window and the dialogue information into an intent recognition model through the AI assistant before the determination module determines the first task intent information based on the dialogue information and the interface content of the first application interface, and output a second intent type when the display state information indicates that the AI dialogue window is in a suspended display state and the dialogue information does not include a screen query keyword; The determining module is further configured to determine whether the dialog information lacks a task element when the second intent type is a second type; the second type indicating that the determination of the task element missing from the first task does not require association with screen content; and, in the case where at least one task element is missing from the conversation information, determining whether the interface content of the first application interface includes the task element missing from the first task; The determination module is specifically used to determine the first task intention information based on the dialogue information and the interface content of the first application interface when the at least one task element is missing in the dialogue information and the interface content of the first application interface includes the at least one task element.
15. The device according to claim 14, characterized in that The processing module is further configured to take a screenshot of the first application interface through the AI assistant to obtain a second interface screenshot; and input the second interface screenshot and the conversation information into a second graphic analysis model on the electronic device side through the AI assistant to output second judgment result information; The determination module is specifically used to determine that the interface content of the first application interface includes the task elements missing from the first task when the second judgment result information indicates that the interface content of the first application interface includes task element information of the task elements missing from the first task.
16. The device according to claim 15, characterized in that The device further includes: a sending module; The sending module is configured to send the second interface screenshot and the conversation information to the server through the AI assistant when the second judgment result information indicates that the interface content of the first application interface includes task element information of the task element missing from the first task; The receiving module is further configured to receive a second message sent by the server, where the second message includes third judgment result information; The determining module is specifically configured to determine that the interface content of the first application interface includes the task element missing from the first task, if the third judgment result information and the second judgment result information received by the receiving module both indicate that the interface content of the first application interface includes task element information of the task element missing from the first task; The third judgment result information is obtained by the server inputting the second interface screenshot and the conversation information into the first graphic analysis model on the server side.
17. The device according to claim 14, characterized in that The determining module is further configured to, after determining whether the dialog information lacks task elements, determine third task intent information based on the dialog information if the task elements in the dialog information are complete, wherein the third task intent information is determined without associating it with screen content, and the third task intent information is used to describe the task content of the third task to be executed; The processing module is further used to control a third application through the AI assistant to perform the third task according to the third task intention information determined by the determination module; wherein the third application is an application that can execute the third task.
18. The device according to claim 14, characterized in that The device further includes: a display module; The display module is configured to, after the determination module determines whether the interface content of the first application interface includes the task elements missing from the first task, display a query message in the AI dialogue window if at least one task element is missing from the dialogue information and the interface content of the first application interface does not include the task elements missing from the first task, wherein the query message is used to inquire about the task elements missing from the first task; The receiving module is further configured to receive a reply message input by the user in the AI dialogue window based on the query information; The determining module is further configured to determine fourth task intention information based on the conversation information and the reply information received by the receiving module; the fourth task intention information is used to describe the task content of the fourth task to be executed; The processing module is further used to control the fourth application through the AI assistant to perform the fourth task according to the fourth task intention information determined by the determination module; wherein the fourth application is an application that can execute the fourth task.
19. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the task execution method according to any one of claims 1 to 9 are implemented.