Intelligent interaction method, intelligent interaction application deployment method and related equipment
By identifying and assigning operation sequence numbers, generating reference images and prompts, and inputting them into multimodal large model, the problem of unnatural user interaction methods in the prior art is solved, and an intelligent interaction method for users to operate intelligent devices through task descriptions is realized.
Patent Information
- Application Number
- CN202510180024.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-10
AI Technical Summary
In the prior art, the interaction mode between users and electronic devices mainly relies on the graphical user interface, lacking natural and intuitive interaction modes, and cannot meet the users' growing intelligent interaction needs.
By receiving intelligent interaction tasks input by users, you can obtain screenshots of the user's interaction interface, identify operable page elements, assign operation numbers, generate reference images and prompts, and input them into the multimodal large model to output the next operation, and perform the operation through the system debugging tool.
It realizes that users use multimodal large models to operate intelligent electronic devices through task descriptions to complete intelligent interactive tasks, improve user operation experience, and achieve more natural and intuitive device interaction.
Smart Images

Figure CN120122862A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular, to an intelligent interaction method, an intelligent interaction English deployment method, and related devices. Background Art
[0002] Currently, the interaction method with electronic devices mainly relies on the graphical user interface (GUI). For example, users currently mainly operate applications by touching icons and buttons on the screen. However, with the popularization of intelligent electronic devices and the rapid development of the mobile Internet, users' demand for intelligent interaction is increasing day by day. Users expect to be able to directly interact with applications on electronic devices in a more natural and intuitive way. Summary of the Invention
[0003] In view of this, embodiments of the present disclosure provide an intelligent interaction method, enabling users to operate intelligent electronic devices through task descriptions and utilize multimodal large models to complete corresponding intelligent interaction tasks.
[0004] The intelligent interaction method described in the embodiments of the present disclosure includes: receiving an intelligent interaction task input by a user; obtaining a first screenshot of the user interaction interface; identifying at least one operable page element on the first screenshot; respectively assigning operation numbers to the at least one operable page element; generating a reference image based on the first screenshot and the operation numbers of the at least one operable page element; generating a prompt based on the intelligent interaction task and the at least one operable page element; inputting the prompt and the reference image into a multimodal large model, and outputting the next operation by the multimodal large model; and calling a system debugging tool to execute the operation, and returning to the step of obtaining the first screenshot of the user interaction interface.
[0005] In the embodiments of the present disclosure, obtaining a first screenshot of the user interaction interface includes: calling the system debugging tool to obtain a screenshot of the current user interaction interface as the first screenshot.
[0006] In the embodiments of the present disclosure, identifying at least one operable page element on the first screenshot includes: obtaining a page layout file corresponding to the first screenshot; and parsing the page layout file to obtain at least one operable page element on the first screenshot and its position on the first screenshot.
[0007] In an embodiment of the present disclosure, parsing the page layout file to obtain at least one operable page element on the first screenshot includes: parsing the page layout file to obtain at least one page element on the first screenshot; determining whether the at least one page element is an operable page element based on relevant information of the at least one page element; wherein, the operable page elements include: clickable elements, slidable elements, long-pressable elements, focusable elements, and text elements.
[0008] In an embodiment of the present disclosure, generating a reference image based on the first screenshot and the operation sequence numbers of the at least one operable page element includes: adding the operation sequence numbers of the at least one operable page element to corresponding positions on the first screenshot according to the positions of the at least one operable page element on the first screenshot to obtain the reference image.
[0009] In an embodiment of the present disclosure, generating a prompt based on the intelligent interaction task and the at least one operable page element includes: generating an operation manual part of the prompt based on the functions and operation sequence numbers of the at least one operable page element; generating a role setting part of the prompt based on a preset role template; generating a task design part of the prompt based on a task description text corresponding to the intelligent interaction task; generating a memory part of the prompt based on a summary of the previous operation output by the multimodal large model; and generating an output specification part of the prompt based on a preset output specification template; wherein, the output of the multimodal large model defined by the output specification includes: the next operation and a summary of the previous operation.
[0010] In an embodiment of the present disclosure, the output of the multimodal large model further includes: an operation target. The above intelligent interaction method further includes: after invoking the system debugging tool to execute the operation, obtaining a second screenshot of the user interaction interface after the operation; inputting the first screenshot, the second screenshot, and the operation target into the multimodal large model, and determining by the multimodal large model whether the execution result of the operation has achieved the operation target; in response to determining that the execution result of the operation has achieved the operation target, recording the function of the operable page element corresponding to the operation and adding it to the operation manual part of the prompt; and in response to determining that the execution result of the operation has not achieved the operation target, invoking the system debugging tool to return to the previous user interaction interface, and adding a record that the execution result of the operation has not achieved the operation target to the memory part of the prompt.
[0011] The intelligent interaction method according to the embodiments of the present disclosure further includes: initializing the operation count parameter to 0; after calling the system debugging tool to execute the operation, incrementing the operation count parameter by 1; determining whether the operation count parameter reaches a preset target operation count; in response to determining that the operation count parameter does not reach the target operation count, returning to the step of obtaining the first screenshot of the user interaction interface; and in response to determining that the operation count parameter has reached the target operation count, ending the intelligent interaction method.
[0012] The intelligent interaction method according to the embodiments of the present disclosure further includes: calling the system debugging tool to execute a screen recording operation to obtain a target video; obtaining a target text corresponding to the target video through speech recognition; extracting key frames of the target video; and inputting the key frames and the target text into the multimodal large model while inputting the prompt and the reference image into the multimodal large model.
[0013] Corresponding to the above intelligent interaction method, the embodiments of the present disclosure also disclose a deployment method for an intelligent interaction application, including: deploying a distributed virtual machine cluster; marking the states of each virtual machine in the virtual machine cluster; wherein, the states include: an idle state and a busy state; after receiving a deployment request for the intelligent interaction application, determining a target virtual machine in the idle state; establishing a connection with the target virtual machine; and operating the target virtual machine to execute the above intelligent interaction method.
[0014] The deployment method for the intelligent interaction application according to the embodiments of the present disclosure further includes: obtaining a pre - script corresponding to the intelligent interaction application; and before obtaining the first screenshot, operating the target virtual machine to call the system debugging tool to execute the pre - script.
[0015] Corresponding to the above intelligent interaction method, the embodiments of the present disclosure also disclose an intelligent interaction device, including:
[0016] An intelligent interaction task receiving module, configured to receive an intelligent interaction task input by a user;
[0017] A screenshot module, configured to obtain a first screenshot of the user interaction interface;
[0018] A page element recognition module, configured to recognize at least one operable page element on the first screenshot;
[0019] A reference image generation module, configured to respectively assign operation serial numbers to the at least one operable page element and generate a reference image based on the first screenshot and the operation serial numbers of the at least one operable page element;
[0020] A prompt generation module, configured to generate a prompt based on the intelligent interaction task and the at least one operable page element;
[0021] An analysis module, configured to input the prompt and the reference image into a multi-modal large model, and output the next operation by the multi-modal large model; and
[0022] An execution module, configured to call a system debugging tool to execute the operation; and
[0023] A control module, configured to trigger the screenshot module to execute the step of obtaining the first screenshot of the user interface again after executing the operation.
[0024] In addition, an embodiment of the present disclosure further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the above-mentioned intelligent interaction method is implemented.
[0025] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, where the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the above-mentioned intelligent interaction method.
[0026] An embodiment of the present disclosure further provides a computer program product, including computer program instructions, where when the computer program instructions run on a computer, the computer is caused to execute the above-mentioned intelligent interaction method.
[0027] It can be seen from this that the intelligent interaction method, the intelligent interaction application deployment method, and related devices provided by some embodiments of the present disclosure enable users to operate an intelligent electronic device to complete corresponding intelligent interaction tasks by using a multi-modal large model in the manner of inputting an intelligent interaction task in the form of natural language, thereby achieving the goal that users expect to directly interact with applications on the electronic device in a more natural and intuitive manner. Description of the Drawings
[0028] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only the embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0029] Figure 1 Shows the implementation process of the intelligent interaction method described in some embodiments of the present disclosure.
[0030] Figure 2 Shows an example of a first screenshot described in an embodiment of the present disclosure.
[0031] Figure 3 Shows an example of a reference image described in an embodiment of the present disclosure.
[0032] Figure 4 Shows the implementation process of the operation check method described in an embodiment of the present disclosure.
[0033] Figure 5 Shows the implementation process of the method for extracting and using auxiliary information described in an embodiment of the present disclosure.
[0034] Figure 6 Shows a schematic diagram of a system for deploying intelligent interaction applications through a distributed virtual machine cluster described in an embodiment of the present disclosure.
[0035] Figure 7 Shows the internal structure of an intelligent interaction device described in some embodiments of the present disclosure.
[0036] Figure 8 Shows a more specific schematic diagram of the hardware structure of an electronic device described in some embodiments of the present disclosure. Detailed implementation manners
[0037] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the present disclosure will be further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0038] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the field to which the present disclosure belongs. The "first", "second", and similar terms used in the embodiments of the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0039] It can be understood that before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner and the user's authorization will be obtained.
[0040] For example, when receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the present disclosure's technical solution based on the prompt message.
[0041] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0042] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0043] As mentioned above, with the popularization of intelligent electronic devices and the rapid development of the mobile Internet, users' demand for intelligent interaction is increasing day by day. Users expect to be able to directly interact with applications on electronic devices in a more natural and intuitive way. At the same time, in recent years, significant progress has been made in the research of large language models (LLMs) and multimodal large models (MLLMs). Through their powerful natural language processing capabilities, these models enable intelligent agents to understand complex tasks. It can be imagined that applying multimodal large models to the intelligent interaction of intelligent electronic devices will help improve the user's operation experience and help users better use intelligent electronic devices in barrier-free scenarios. For example, users can use multimodal large models to operate intelligent electronic devices and complete corresponding tasks in the way of task description.
[0044] In view of this, some embodiments of the present disclosure provide an intelligent interaction method that enables users to operate an intelligent electronic device by means of task description and use a multimodal large model to complete corresponding intelligent interaction tasks. In the embodiments of the present disclosure, the above intelligent interaction method may be executed by an intermediate server, and the intermediate server may act as a medium between the multimodal large model and the intelligent electronic device and cooperate with the multimodal large model and the intelligent electronic device to complete intelligent interaction tasks.
[0045] Figure 1 Shows the implementation process of the intelligent interaction method described in some embodiments of the present disclosure. As Figure 1 shown, the above intelligent interaction method may include the following multiple steps:
[0046] In step 110, receive an intelligent interaction task input by the user.
[0047] In step 120, obtain the first screenshot of the user interaction interface.
[0048] In step 130, identify at least one operable page element on the first screenshot.
[0049] In step 140, assign operation numbers to at least one operable page element respectively.
[0050] In step 150, generate a reference image based on the above first screenshot and the operation numbers of the above at least one operable page element.
[0051] In step 160, generate a prompt based on the above intelligent interaction task and the above at least one operable page element.
[0052] In step 170, input the prompt and the reference image into the multimodal large model, and the multimodal large model outputs the next operation.
[0053] In step 180, call the system debugging tool to execute the above operation.
[0054] After executing the above step 180, it is possible to return to the step of obtaining the first screenshot of the user interaction interface until the multimodal large model determines that the intelligent interaction task is completed.
[0055] In addition, in order to avoid excessive number of operations output by the multimodal large model, as an alternative to the above solution, a target number of operations can be further set, that is, the maximum number of operations required to execute the task. When the total number of operations output by the multimodal large model has reached the above target number of operations, even if the intelligent interaction task has not been completed, the above intelligent task interaction method will be terminated. Specifically, to achieve the above goal, the intelligent interaction method described in the embodiments of the present disclosure may further include the following steps: First, initialize the operation number parameter to 0; after calling the system debugging tool to execute the corresponding operation, increment the operation number parameter by 1; further, determine whether the operation number parameter has reached the pre-set target number of operations; in response to determining that the operation number parameter has not reached the target number of operations, return to step 120; and in response to determining that the operation number parameter has reached the target number of operations, end the above intelligent interaction method.
[0056] Next, the specific execution methods of each step in the above intelligent interaction method will be further described in detail with specific examples.
[0057] Regarding the above step 110, formally, the intelligent interaction task input by the user can be implemented by the user through various input methods. For example, it can be a text task input by the user through the text input method, or it can be a voice task input by the user through the voice input method, etc. For the text input method, the input text itself can be used as the text description of the intelligent interaction task; for the voice input method, the input voice task can be first subjected to automatic speech recognition (ASR), and the text obtained after recognition can be used as the text description of the intelligent interaction task. That is, in the embodiments of the present disclosure, for an intelligent interaction task input through an input method other than the text method, it needs to be first converted into a text description of the intelligent interaction task, and the text description is used as the input to interact with the multimodal large model.
[0058] In addition, functionally, the intelligent interaction task can be a natural language description of a task for the user to interact with an application on the intelligent electronic device. It can be understood that the goal of the intelligent interaction method described in the embodiments of the present disclosure is that after the intelligent electronic device receives an intelligent interaction task from the user, the intelligent electronic device automatically completes a series of operations required to implement the intelligent interaction task, without the user having to complete the operations of each step personally. For example, if the user wishes to send a message containing content B to user C through instant messaging software A, the user can directly submit an intelligent interaction task of "sending a message with content B to C through A" to the intelligent electronic device. The goal of the intelligent interaction method described in the embodiments of the present disclosure is that after the intelligent electronic device receives the above intelligent interaction task, it can automatically complete a series of operations required to implement the intelligent interaction task, such as including the following series of operations: opening instant messaging software A; searching for and finding user C; opening the conversation page of user C; entering content B in the input box; and clicking send. It can be seen that if the intelligent interaction method described in the embodiments of the present disclosure can achieve the above goal, the user can operate the intelligent electronic device by inputting a natural language task description, so as to achieve the goal of directly interacting with the application on the electronic device in a more natural and intuitive way.
[0059] Regarding the above-mentioned step 120, in an embodiment of the present disclosure, the specific method for obtaining the first screenshot of the user interface may include: calling a system debugging tool to obtain a screenshot of the interactive interface of the intelligent electronic device used by the current user as the first screenshot. Among them, the above-mentioned system debugging tool will be related to the operating system of the intelligent electronic device used by the user, that is, different operating systems will correspond to different system debugging tools. For example, for the Android system, the above-mentioned system debugging tool may specifically be appium-uiautomator2-driver. By calling the screenshot function provided by the above-mentioned system debugging tool, a screenshot of the interactive interface of the intelligent electronic device used by the current user can be captured. Figure 2 Shows an example of a first screenshot described in an embodiment of the present disclosure. As Figure 2 shown, the above-mentioned first screenshot may include multiple page elements, where each page element has specific attributes and functions.
[0060] Regarding the above-mentioned step 130, in an embodiment of the present disclosure, the specific method for identifying at least one operable page element on the first screenshot may specifically include the following steps: First, obtain a page layout file corresponding to the first screenshot; and parse the above-mentioned page layout file to obtain at least one operable page element on the first screenshot and its position on the first screenshot.
[0061] Specifically, in an embodiment of the present disclosure, the page layout file corresponding to the first screenshot may also be obtained by calling a system debugging tool. It can be understood that in the above-mentioned page layout file, all relevant information required to display page elements on the page, such as multiple page elements on the page and the corresponding positions, attributes, and functions of each page element, is defined.
[0062] Thus, in an embodiment of the present disclosure, the method for parsing the page layout file to obtain at least one operable page element on the first screenshot may specifically include: parsing the above-mentioned page layout file to obtain at least one page element; and determining whether the above-mentioned at least one page element is an operable page element based on the relevant information of the above-mentioned at least one page element. Usually, it can be directly determined whether a page element is an operable page element according to the attributes of the page element.
[0063] Specifically, in the embodiments of the present disclosure, the above-mentioned operable page elements generally may include: clickable elements, scrollable elements, long-clickable elements, focusable elements, and text elements. Among them, the above-mentioned clickable elements generally refer to the elements on the page that can be executed with a click (Tap) operation; the above-mentioned scrollable elements generally refer to the elements on the page that can be executed with a swipe (Swipe) operation; the above-mentioned long-clickable elements generally refer to the elements on the page that can be executed with a long-press (Long_press) operation; and the above-mentioned text elements generally refer to the elements on the page that can be executed with a tap on the text (Tap_phrase) operation (that is, in combination with optical character recognition (OCR), for some places where the element positions cannot be located, the operation is completed by clicking on the specific text position). In Figure 2 In the first screenshot example shown, the current user interface includes a clickable element 210, a scrollable element 220, and a long-clickable element 230.
[0064] Next, in the above step 140, an operation serial number is assigned to each of the at least one operable page element. The above operation serial number is mainly used to refer to each operable page element on the first screenshot when interacting with the multimodal large model. Among them, the operation serial numbers assigned to each operable page element should be unique on the current user interface, that is, the operation serial numbers corresponding to different operable page elements on the same user interface should be different.
[0065] Regarding the above step 150, in the embodiments of the present disclosure, the method for generating a reference image based on the first screenshot and the operation serial numbers of the at least one operable page element may include: adding the operation serial numbers of the at least one operable page element to the corresponding positions on the first screenshot according to the positions of the at least one operable page element on the first screenshot of the user interface, so as to obtain a reference image. Specifically, in the above step, the operation of adding the operation serial numbers to the first screenshot can be realized by any image editing tool. The embodiments of the present disclosure do not limit the image editing tool used. Thus, it can be understood that compared with the first screenshot, the page elements included in the reference image remain unchanged, but the operation serial numbers corresponding to each operable page element are added to the reference image.
[0066] Figure 3 Shows an example of a reference image described in the embodiments of the present disclosure. Figure 3 The reference image shown is based on Figure 2 The first screenshot shown is the generated reference image. As Figure 3 Shown, in the reference image, in Figure 2In the first screenshot shown, operation numbers 1, 2, and 3 corresponding to the above-mentioned operable page elements are added at the positions of the clickable element 210, the slidable element 220, and the long-pressable element 230 respectively.
[0067] Regarding the above step 160, in some embodiments of the present disclosure, the generated prompt may include the following multiple parts.
[0068] 1. Operation manual part: used to describe the functions corresponding to each operable page element, and to establish the correspondence between the operable page element and the page element on the reference image through the operation number. The purpose of the above operation manual part is to enable the multimodal large model to clearly understand the roles played by each operable page element, so as to better complete the intelligent interaction task.
[0069] 2. Role setting part: used to set the user role, enable the multimodal large model to simulate human behavior, and perform more preference-based operations according to the personality characteristics set for the user role.
[0070] 3. Task design part: used to describe the intelligent interaction task to be completed.
[0071] 4. Memory part: used to record the summary of a series of previous operations output by the multimodal large model.
[0072] 5. Output specification part: used to standardize the output of the multimodal large model, so that the multimodal large model outputs text in a fixed format.
[0073] In some specific examples, the above output specification part may specifically include the following multiple sub-parts:
[0074] a. Operation: used to indicate the next operation to be performed on a certain operable page element;
[0075] b. Summary of previous operations: describes the reason for performing this operation during the task execution process, and summarizes the behavior chain from the start of the task execution to the present. In the embodiments of the present disclosure, the above summary of previous operations will be passed as memory to the next round of conversation with the multimodal large model.
[0076] More specifically, in the embodiments of the present disclosure, the above operation may include one of the following specific multiple operations:
[0077] 1. Tap(n): Click on the page element with the operation number n;
[0078] 2. Swipe(n, distance, direction): Slide the page element with the operation number n, and specifically adjust the sliding method according to the sliding distance (distance) and the sliding direction (direction);
[0079] 3. Long_press(n): Long press the page element with the operation number n.
[0080] 4. Back(): Return to the previous-level user interaction interface.
[0081] 5. Wait(n): Wait for n seconds.
[0082] 6. Tap_phrase(string): In combination with optical character recognition (OCR), for some places where the positions of page elements cannot be located, complete the operation by clicking on the specific text position.
[0083] In some other specific examples, in addition to the operations and the summary of the previous operations, the above output specification part may further include: c. Operation objective: used to explain the objective of performing the above operations.
[0084] Based on the above setting of the prompt structure, in the embodiments of the present disclosure, generating a prompt based on the intelligent interaction task and the at least one operable page element may specifically include the following multiple steps: generating the operation manual part of the prompt based on the functions and operation numbers of the at least one operable page element; generating the role setting part of the prompt based on a pre-set role template; generating the task design part of the prompt based on the task description text corresponding to the intelligent interaction task; generating the memory part of the prompt based on the summary of the previous operations output by the multimodal large model; and generating the output specification part of the prompt based on a pre-set output specification template. For example, in a specific example, a record in the generated operation manual may include the following information: "The page element with the operation number 1 is a clickable page element. After performing a click operation on the page element with the operation number 1, the previous-level user interaction interface can be returned."
[0085] Next, in the above step 170, by inputting the prompt and the reference image into the multimodal large model, the multimodal large model can predict the next operation to be performed according to the input prompt and reference image. At this time, the multimodal large model will output according to the format defined in the output specification part of the prompt. It can be seen that the output of the multimodal large model includes the next operation. For example, "Click the page element with the operation number 1". Further, the output of the above multimodal large model may further include a summary of the previous operations and even the operation objective.
[0086] After receiving the next operation output by the multi-modal large model, for example, "click on the page element with operation number 1", at step 180, the system debugging tool can be called to execute the above operation. For example, by calling the corresponding method of appium-uiautomator2-driver to click on the page element with operation number 1.
[0087] As described above, after executing the above step 180, it is possible to return to step 120 and repeat the above steps until the multi-modal large model determines that the intelligent interaction task has been completed or the target number of operations has been reached. Usually, when the multi-modal large model determines that the intelligent interaction task of the user input has been completed based on the input information, the next operation output can be empty or end.
[0088] From the definition of the prompt described above, it can be seen that during the above loop, the summary of the previous operation output by the multi-modal large model can be added to the memory part of the current prompt, so as to achieve continuous update of the prompt, thereby assisting the multi-modal large model in predicting the next operation and achieving the goal of improving the execution efficiency of the multi-modal large model.
[0089] To ensure the execution effect of the multi-modal large model and ensure that each operation can achieve its operation goal, the embodiments of the present disclosure further provide a method for checking each operation. Figure 4 Shows the implementation process of the operation checking method described in the embodiments of the present disclosure. As Figure 4 shown, the operation checking method described in the embodiments of the present disclosure includes the following multiple steps.
[0090] At step 410, after calling the system debugging tool to execute the next operation output by the multi-modal large model, obtain the second screenshot of the user interaction interface after the operation.
[0091] It can be understood that in the embodiments of the present disclosure, the above first screenshot is a screenshot of the user interaction interface before the operation; and the above second screenshot is a screenshot of the user interaction interface after the operation.
[0092] At step 420, input the above first screenshot, second screenshot, and the operation goal output by the multi-modal large model into the multi-modal large model again, and let the multi-modal large model determine whether the execution result of this operation has achieved the above operation goal.
[0093] At step 430, in response to determining that the execution result of the operation has achieved the operation goal, add the function of the operable page element corresponding to this operation to the operation manual part of the prompt.
[0094] In step 440, in response to determining that the execution result of the operation has not achieved the operation goal, the system debugging tool is called to return to the previous user interface, that is, the user interface before the execution of this operation, and a record that the execution result of the operation has not achieved the operation goal is added to the memory part of the prompt.
[0095] Through the above operation inspection method, after each step output by the multi-modal large model is executed, it can be verified in a timely manner whether the operation goal has been achieved. For the operations that have achieved the operation goal, the functions of the page elements involved in this operation are strengthened by adding the operation records to the operation manual part of the prompt; while for the operations that have not achieved the operation goal, by adding the operation records to the memory part of the prompt, the operation records that have achieved the operation goal are recorded and passed down, so as to effectively avoid the multi-modal large model repeating previous mistakes.
[0096] Furthermore, in some other embodiments of the present disclosure, to parse video-type inputs, these embodiments add a screen recording function and an automatic speech recognition (ASR) function, and use the key frames in the video file obtained by screen recording and the recognized ASR text as auxiliary information, which is used as one of the inputs of the multi-modal large model when needed, thereby enhancing the perception ability of the multi-modal large model. Figure 5 Shows the implementation process of the method for extracting and using the auxiliary information described in the embodiments of the present disclosure. As Figure 5 shown, in the embodiments of the present disclosure, the method for extracting and using the auxiliary information may include the following multiple steps.
[0097] In step 510, the system debugging tool is called to perform screen recording to obtain the target video.
[0098] In step 520, the target text corresponding to the target video is obtained through speech recognition.
[0099] In step 530, the key frames of the target video are extracted.
[0100] In step 540, while inputting the prompt and the reference image into the multi-modal large model, the above key frames and the target text are input into the multi-modal large model.
[0101] Through the above method, when the user is using the smart electronic device to browse the video, the multi-modal large model can perceive the content of the video being browsed by the user, thereby assisting the multi-modal large model in analysis and prediction, and thus more efficiently completing the intelligent interaction task.
[0102] It can be seen that the intelligent interaction method provided by some embodiments of the present disclosure enables users to use a multimodal large model to operate an intelligent electronic device to complete corresponding intelligent interaction tasks by inputting intelligent interaction tasks in the form of natural language, thereby achieving the goal of directly interacting with applications on the electronic device in a more natural and intuitive manner as expected by the user.
[0103] Corresponding to the above intelligent interaction method, embodiments of the present disclosure further provide a deployment method for an intelligent interaction application. Specifically, the deployment method for the intelligent interaction application may include the following multiple steps: First, deploy a distributed virtual machine cluster; then, mark the status of each virtual machine in the virtual machine cluster; where the above status may include: an idle state and a busy state; next, after receiving a deployment request for the intelligent interaction application, determine a target virtual machine in the idle state and establish a connection with the target virtual machine; finally, operate the target virtual machine to execute the above intelligent interaction method. Figure 6 shows a schematic diagram of a system for deploying an intelligent interaction application through a distributed virtual machine cluster in an embodiment of the present disclosure. In Figure 6 the shown system, the intermediate server 610 is mainly used to execute the above deployment method for the intelligent interaction application; the distributed virtual machine cluster 620 is used to represent the intelligent electronic device; the multimodal large model 630 is used to complete the prediction of the next operation based on the input prompt and reference image; the database 640 is mainly used to mark the status of each virtual machine in the distributed virtual machine cluster 620.
[0104] To reduce the time required to complete the deployment of the intelligent interaction application, for the intelligent interaction application, the number of steps required to complete the intelligent interaction task can be reduced by means of a fixed pre - script, thereby improving the task execution efficiency. Specifically, the corresponding pre - script for the intelligent interaction application can be obtained first; then, before obtaining the above first screenshot, operate the target virtual machine to call the system debugging tool to execute the above pre - script, so that the virtual machine completes the preparation for executing the intelligent interaction task in advance, thereby reducing the number of steps required to complete the intelligent interaction task and improving the efficiency of deploying the intelligent interaction application.
[0105] Corresponding to the above intelligent interaction method, embodiments of the present disclosure also disclose an intelligent interaction device. Figure 7 shows the internal structure of the intelligent interaction device described in an embodiment of the present disclosure. As Figure 7 shown, the above intelligent interaction device may include:
[0106] An intelligent interaction task receiving module 710, configured to receive an intelligent interaction task input by a user;
[0107] A screenshot module 720, configured to obtain a first screenshot of the user interaction interface;
[0108] A page element recognition module 730, configured to recognize at least one operable page element on the first screenshot;
[0109] A reference image generation module 740, configured to respectively assign operation numbers to the at least one operable page element and generate a reference image based on the first screenshot and the operation numbers of the at least one operable page element;
[0110] A prompt generation module 750, configured to generate a prompt based on the intelligent interaction task and the at least one operable page element;
[0111] An analysis module 760, configured to input the prompt and the reference image into a multimodal large model, and output the next operation by the multimodal large model;
[0112] An execution module 770, configured to call a system debugging tool to execute the operation; and
[0113] A control module 780, configured to, after executing the operation, trigger the screenshot module to execute the step of obtaining the first screenshot of the user interaction interface again.
[0114] It should be noted that each of the above modules may respectively adopt the specific implementation manners of the steps in the intelligent interaction method described in the foregoing embodiments, which will not be elaborated herein.
[0115] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the intelligent interaction method described in any of the above embodiments.
[0116] Figure 8 FIG. shows a schematic hardware structure diagram of a more specific electronic device provided in this embodiment. The device may include: a processor 2010, a memory 2020, an input / output interface 2030, a communication interface 2040, and a bus 2050. Among them, the processor 2010, the memory 2020, the input / output interface 2030, and the communication interface 2040 are communicatively connected to each other inside the device through the bus 2050.
[0117] The processor 2010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0118] The memory 2020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 2020 can store the operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 2020 and are called and executed by the processor 2010.
[0119] The input / output interface 2030 is used to connect to input / output devices to achieve information input and output. Among them, the input / output devices can be configured as components in the device or externally connected to the device to provide corresponding functions. The input devices can include microphones, various sensors, etc., and the output devices can include displays, speakers, vibrators, indicator lights, etc.
[0120] The communication interface 2040 is used to connect to a communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0121] The bus 2050 includes a path for transmitting information between various components of the device (such as the processor 2010, the memory 2020, the input / output interface 2030, and the communication interface 2040).
[0122] It should be noted that although the above device only shows the processor 2010, the memory 2020, the input / output interface 2030, the communication interface 2040, and the bus 2050, in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solutions of the embodiments of this specification, and do not have to include all the components shown in the figure.
[0123] The electronic device in the above embodiment is used to implement the corresponding intelligent interaction method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0124] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the intelligent interaction method as described in any of the above embodiments.
[0125] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0126] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the task processing method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0127] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary, and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, and they are not provided in detail for the sake of brevity.
[0128] In addition, for the sake of simplicity of description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the device may be shown in block diagram form in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (i.e., these details should be completely within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0129] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0130] Embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. An intelligent interaction method, comprising: Intelligent interactive tasks that receive user input; Get the first screenshot of the user interface; Identifying at least one operable page element on the first screenshot; Assigning an operation sequence number to each of the at least one operable page elements; generating a reference image based on the first screenshot and the operation sequence number of the at least one operable page element; Generate a prompt based on the intelligent interaction task and the at least one operable page element; Input the prompt and the reference image into the multimodal large model, and the multimodal large model outputs the next operation; as well as Call the system debugging tool to execute the operation, and return to the step of obtaining the first screenshot of the user interaction interface.
2. The method according to claim 1, wherein: Acquiring a first screenshot of the user interaction interface includes: calling the system debugging tool to acquire a screenshot of the current user interaction interface as the first screenshot.
3. The method according to claim 1, wherein: Identifying at least one operable page element on the first screenshot includes: Obtaining a page layout file corresponding to the first screenshot; and The page layout file is parsed to obtain at least one operable page element on the first screenshot and its position on the first screenshot.
4. The method according to claim 3, wherein: Parsing the page layout file to obtain at least one operable page element on the first screenshot includes: Parsing the page layout file to obtain at least one page element on the first screenshot; Based on the relevant information of the at least one page element, determine whether the at least one page element is an operable page element; wherein the operable page elements include: clickable elements, slidable elements, long-pressable elements, focusable elements and text-type elements.
5. The method according to claim 3, wherein: Generating a reference image based on the first screenshot and the operation sequence number of the at least one operable page element includes: According to the position of the at least one operable page element on the first screenshot, the operation sequence number of the at least one operable page element is added to the corresponding position on the first screenshot to obtain the reference image.
6. The method according to claim 1, wherein: Generating a prompt based on the intelligent interaction task and the at least one operable page element includes: An operation manual portion for generating the prompt based on the function of the at least one operable page element and the operation sequence number; Generating a role setting part of the prompt based on a preset role template; A task design part that generates the prompt based on the task description text corresponding to the intelligent interaction task; Generate a memory portion of the prompt based on a summary of previous operations output by the multimodal large model; and The output specification part of the prompt is generated based on a preset output specification template; wherein the output of the multimodal large model defined by the output specification includes: the next operation and a summary of the previous operation.
7. The method according to claim 6, wherein: The output of the multimodal large model further includes: operation objectives; The intelligent interaction method further comprises: After calling the system debugging tool to perform the operation, obtaining a second screenshot of the user interaction interface after the operation; Inputting the first screenshot, the second screenshot and the operation target into the multimodal large model, and determining by the multimodal large model whether the execution result of the operation has completed the operation target; In response to determining that the execution result of the operation completes the operation target, recording the role of the operable page element corresponding to the operation and adding it to the operation manual part of the prompt; and In response to determining that the execution result of the operation does not complete the operation target, calling the system debugging tool to return to the previous user interaction interface, and adding a record that the execution result of the operation does not complete the operation target in the memory part of the prompt.
8. The method according to claim 1, further comprising: Initialize the operation times parameter to 0; After calling the system debugging tool to perform the operation, adding 1 to the operation count parameter; Determine whether the operation number parameter reaches a preset target operation number; In response to determining that the operation number parameter does not reach the target operation number, returning to the step of obtaining the first screenshot of the user interaction interface; as well as In response to determining that the operation number parameter has reached the target operation number, the intelligent interaction method is terminated.
9. The method according to claim 1, further comprising: Calling the system debugging tool to perform screen recording operation to obtain the target video; Obtaining a target text corresponding to the target video through speech recognition; Extracting key frames of the target video; as well as The key frame and the target text are input into the multimodal large model at the same time as the prompt and the reference image are input into the multimodal large model.
10. A method for deploying an intelligent interactive application, comprising: Deploy distributed virtual machine clusters; Marking the status of each virtual machine in the virtual machine cluster; wherein the status includes: idle state and busy state; After receiving a deployment request of the intelligent interactive application, determining a target virtual machine in an idle state; Establishing a connection with the target virtual machine; and The target virtual machine is operated to execute the intelligent interaction method as described in any one of claims 1-9.
11. The method according to claim 10, further comprising: Obtaining a pre-script corresponding to the intelligent interactive application; as well as Before acquiring the first screenshot, the target virtual machine is operated to call the system debugging tool to execute the pre-script.
12. An intelligent interactive device, comprising: An intelligent interaction task receiving module, used for receiving an intelligent interaction task input by a user; A screenshot module, used to obtain a first screenshot of the user interaction interface; A page element identification module, used to identify at least one operable page element on the first screenshot; A reference image generation module, configured to assign an operation serial number to each of the at least one operable page elements and to generate a reference image based on the first screenshot and the operation serial number of the at least one operable page element; A prompt generating module, configured to generate a prompt based on the intelligent interaction task and the at least one operable page element; An analysis module, used for inputting the prompt and the reference image into a multimodal large model, and the multimodal large model outputs a next operation; An execution module, used for calling a system debugging tool to execute the operation; as well as The control module is used to trigger the screenshot module to execute the step of obtaining the first screenshot of the user interaction interface again after executing the operation.
13. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the intelligent interaction method as described in any one of claims 1 to 9 when executing the program.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the intelligent interaction method according to any one of claims 1 to 9.
15. A computer program product, comprising computer program instructions, which, when executed on a computer, enable the computer to execute the intelligent interaction method according to any one of claims 1 to 9.
Citation Information
Cited By
Data processing method and device, electronic equipment and storage medium
CN120610770A