Execution method and electronic equipment

By leveraging the Agent capabilities of the first device and unified spatiotemporal multimodal coding technology, cross-device collaborative control of target tasks is achieved, solving the problem of collaborative control among multiple devices and realizing convenient and efficient task execution and load balancing.

CN121900969APending Publication Date: 2026-04-21LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LENOVO (BEIJING) LTD
Filing Date
2025-12-31
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, collaborative control between multiple devices has not yet been able to conveniently and efficiently meet user interaction needs, especially in scenarios involving inter-device communication and mapping, where devices without agent capabilities struggle to flexibly control target tasks.

Method used

By leveraging the Agent capabilities of the first device and utilizing unified spatiotemporal multimodal coding technology and target models, the target device and interactive operation information of the target task can be determined without needing to obtain the APP call interface of the target device, thus enabling flexible scheduling and collaborative control across devices.

Benefits of technology

It enables convenient and efficient collaborative control among multiple devices, improves the flexibility and applicability of task execution, avoids task conflicts and waste of equipment resources, and ensures load balancing and stability among devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121900969A_ABST
    Figure CN121900969A_ABST
Patent Text Reader

Abstract

The invention discloses an execution method and electronic device.The execution method is applied to first equipment and comprises the steps that based on target input, a target task corresponding to execution of the target input is determined; if it is determined that the target device executing the target task is the second device, obtaining target display data of the target device; determining target interaction operation information corresponding to execution of the target task based on the target display data; and controlling the target device to execute the target interaction operation based on the target interaction operation information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of human-computer interaction technology, and in particular to an execution method and electronic device. Background Technology

[0002] Currently, scenarios involving communication and mapping between multiple devices are becoming increasingly common, but convenient and efficient collaborative control has not yet been achieved to meet users' interaction needs in such scenarios. Summary of the Invention

[0003] The purpose of this application is to provide an execution method and an electronic device.

[0004] In a first aspect, embodiments of this application provide an execution method applied to a first device, comprising: Based on the target input, determine the target task corresponding to the target input; If the target device for performing the target task is determined to be the second device, the target display data of the target device is obtained; Based on the target display data, determine the target interactive operation information corresponding to the execution of the target task; Control the target device to perform target interactive operations based on the target interactive operation information.

[0005] In one possible implementation, determining the target interaction operation information corresponding to performing the target task based on the target display data includes: The target display data and the task information of the target task are input into the target model so that the target model outputs target interactive operation information for the target display data.

[0006] In one possible implementation, the execution method further includes: Based on the target task, determine the target application to perform the target task; Among the third devices that have the target application installed, a second device capable of performing the target task is determined.

[0007] In one possible implementation, determining a second device capable of performing the target task among the third devices on which the target application is installed includes: Determine the application resource usage status of the target application installed on each third device; Based on the resource occupancy status of each application, a second device capable of executing the target task is determined.

[0008] In one possible implementation, the execution method further includes: First display data is displayed in the first display area, and the target display data is displayed in the second display area.

[0009] In one possible implementation, the control target device performs a target interaction operation based on the target interaction operation information, including: Based on the determined target interaction operation information, construct the operation coordinate information for the second device; The operation coordinate information is sent to the second device so that the second device can simulate and execute the target interactive operation based on the operation coordinate information.

[0010] In one possible implementation, the execution method further includes: In response to the failure of the second device, obtain the information of the executed interactive operation; Based on the information of the executed interactive operations, control the fourth device to jump to the target interface; Obtain the fourth display data of the fourth device; Based on the fourth display data, determine the fourth interactive operation information corresponding to the execution of the target task; The fourth device is controlled to perform a fourth interactive operation based on the fourth interactive operation information; The target interface is the interface that appears when the second device fails to execute.

[0011] In one possible implementation, controlling the fourth device to jump to the target interface based on the executed interaction information includes: Obtain the operational coordinate mapping relationship between the second device and the fourth device; Based on the operation coordinate mapping relationship and the executed interaction operation information, the fourth device is controlled to execute a simulated target interaction operation sequence so that the fourth device jumps to the target interface; The target interaction operation sequence is the interaction operation sequence corresponding to the executed interaction operation information.

[0012] In one possible implementation, controlling the fourth device to jump to the target interface based on the executed interaction information further includes: Based on the information of the executed interactive operations, determine the page path information of the target interface; The page path information is sent to the fourth device so that the fourth device can jump to the target interface.

[0013] Secondly, embodiments of this application also provide an electronic device running a target application, the target application being able to parse at least one task based on input, the target application being used to execute: Based on the target input, determine the target task corresponding to the target input; If the target device for performing the target task is determined to be the second device, the target display data of the target device is obtained; Based on the target display data, determine the target interactive operation information corresponding to the execution of the target task; Control the target device to perform target interactive operations based on the target interactive operation information. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 A flowchart of an execution method provided in this application is shown; Figure 2 A flowchart of another execution method provided in this application is shown; Figure 3 A schematic diagram of the structure of an electronic device provided in this application is shown. Detailed Implementation

[0016] Various embodiments and features of this application are described herein with reference to the accompanying drawings.

[0017] It should be understood that various modifications can be made to the embodiments described herein. Therefore, the above description should not be considered as limiting, but merely as an example of embodiments. Other modifications within the scope and spirit of this application will be apparent to those skilled in the art.

[0018] The accompanying drawings, which are included in and form part of this specification, illustrate embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, serve to explain the principles of the present application.

[0019] These and other features of this application will become apparent from the following description of preferred forms of embodiments given as non-limiting examples, with reference to the accompanying drawings.

[0020] It should also be understood that although this application has been described with reference to some specific examples, those skilled in the art can certainly implement many other equivalent forms of this application, which have the features described in the claims and are therefore all within the scope of protection defined herein.

[0021] The above and other aspects, features and advantages of this application will become more apparent when taken in conjunction with the accompanying drawings and in view of the following detailed description.

[0022] Specific embodiments of this application are described thereafter with reference to the accompanying drawings; however, it should be understood that the claimed embodiments are merely examples of this application, which can be implemented in various ways. Well-known and / or repeated functions and structures are not described in detail to avoid unnecessary or redundant details that could obscure the application. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but merely serve as the basis and representative basis for the claims to teach those skilled in the art to use this application in a variety of substantially any suitable detailed structures.

[0023] This specification may use the phrases “in one embodiment,” “in another embodiment,” “in yet another embodiment,” or “in other embodiments,” all of which may refer to one or more of the same or different embodiments according to this application.

[0024] Explanation of terms used in the embodiments of this application: An intelligent agent is an application of artificial intelligence technology. It can be implemented based on a large model (such as a large language model). The agent's behavior is determined by the large model within the agent based on the current state and external input, and can be used to specifically solve a certain type of problem. The large model provides the decision-making foundation through learning and training on large amounts of data; the agent can also call upon tools, plugins, and knowledge bases to provide reasoning, decision-making, and execution capabilities. Simply put, the large model provides decision support, and the agent continuously optimizes the performance of its internal large model through feedback data from real-world applications (such as prompts).

[0025] Currently, when a user interacts with a terminal device, they can perform input operations on that device. The terminal device then determines the corresponding task to be executed based on the input, and generates interactive operations to perform the task based on its own display data. These interactive operations are then executed to complete the task. However, this task execution only applies to a single terminal device. In increasingly common scenarios involving multi-device communication and mapping, for collaborative control of multiple devices, some devices have agent capabilities while others do not. To enable devices without agent capabilities to also possess them, a common approach is to utilize the underlying program specifications of data or event interfaces between devices, analyze and plan the execution steps using a Large Language Model (LLM), and then execute these interfaces step by step. The advantage of this approach is high execution stability, but the functionality is relatively limited, especially for mobile app operations, where large-scale coverage is difficult, requiring individual API calls for each app.

[0026] The execution method provided in this application is applied in a scenario where multiple devices collaboratively execute a target task. By using the Agent capability of the first device, the target device executing the target task and the target interaction operation information corresponding to the target task can be determined. Without obtaining the APP call interface of the target device, the target device can be controlled to complete the response to the target input, which is highly flexible and applicable to a wide range of scenarios. Moreover, it does not require much user operation, making it convenient and efficient.

[0027] To facilitate understanding of this application, a specific implementation method provided in this application will be described in detail first.

[0028] As an example, Figure 1 A flowchart of the execution method provided in the embodiments of this application is shown, wherein the specific steps include S101-S104.

[0029] S101, Based on the target input, determine the target task corresponding to the target input.

[0030] Optionally, the execution method of this application embodiment is applied to a first device, which is any one of a plurality of devices capable of communicating and mapping with each other, and the first device is deployed with Agent capability. For example, a mobile phone, tablet computer, and laptop computer belonging to the same user can all be used as the first device; another example is the central control display screen, the passenger-side display screen, and the rear seat display screen in the same vehicle, all three of which can be used as the first device, etc.

[0031] The user performs an input operation on the first device, so that the first device can collect the input operation, that is, the first device obtains the target input. The target input can be natural language input by the user using an input device such as a touch screen, stylus, or keyboard, or natural language input by the user using an audio acquisition device included in the first device.

[0032] After obtaining the target input, the target task corresponding to the target input is determined based on the target input, such as launching the first application, closing the second application, and adjusting the operating parameters of the first device. As an example, a target application can be pre-set in the first device. After obtaining the target input, the target application is used to parse the target input to obtain at least one target task.

[0033] S102, if the target device for performing the target task is determined to be the second device, obtain the target display data of the target device.

[0034] For example, after determining the target task, the target device for executing the target task is further determined. Optionally, it is determined whether the first device can execute the target task. If the first device does not have the ability to execute the target task, the target device for executing the target task is determined to be the second device. The second device is a device that has established a communication connection with the first device, that is, the second device and the first device can communicate and map with each other. The second device can be one or more.

[0035] If the target device for performing the target task is determined to be the second device, the target display data of the target device (i.e., the second device) is obtained. The target display data is the content displayed by the target device, such as the target device's icon, running application windows, etc.

[0036] S103, determine the target interactive operation information corresponding to the target task based on the target display data.

[0037] After obtaining the target display data of the target device, the target interactive operation information corresponding to the execution of the target task is determined based on the target display data. Optionally, the target display data and the task information of the target task are input into the target model so that the target model outputs the target interactive operation information for the target display data. That is to say, the target model of this application embodiment uses visual recognition (such as processing the target display data) to determine the interactive operation information to achieve the execution of the task, and can obtain accurate target interactive operation information corresponding to the target task.

[0038] Here, the target model in this application embodiment is a machine learning model, which can perform recognition, matching and other processing on input information based on pre-configured parameters to obtain output information corresponding to the input information. In conjunction with this application embodiment, the target model recognizes the input target display data and the task information of the target task. The target display data and the task information of the target task are natural language, but can also be audio, images, tables, etc. After recognizing the input information, the target model performs voice analysis, task recognition, device information comparison, task and device matching and other processing to obtain target interactive operation information for the target display data and used to execute the target task.

[0039] The target task information includes task actions and task objects, while the target interactive operation information includes the interactive operations required to perform the target task based on the second display area of ​​the first device, and the operation objects targeted by the interactive operations.

[0040] For example, the target model is pre-trained and knows the applications installed on each device. The target model is used to construct target interaction information to perform the target task in order to complete the target task.

[0041] S104, control the target device to perform target interactive operations based on target interactive operation information.

[0042] After obtaining the target interaction operation information, the control target device executes the target interaction operation based on the target interaction operation information to perform and complete the target task. As an example, the first device can transmit the target interaction operation information to the target device, so that the target device can execute and complete the target task based on the target interaction operation information, thereby realizing the control of the target device by the first device, that is, realizing cross-device interaction operation by coordinating multiple devices.

[0043] As an example, when controlling the target device to perform a target interactive operation based on the target interactive operation information, operation coordinate information for the second device can be constructed based on the determined target interactive operation information.

[0044] Considering that multiple devices may differ across various dimensions, such as variations in device body, model, and component configuration, as well as differences in screen resolution, size, and coordinate origin, leading to variations in screen images, and different system times, the display times corresponding to the obtained display data may differ. Furthermore, multiple devices may not clearly define their own and other devices' identity and / or location information, potentially resulting in inaccurate correspondence between obtained display data and devices, or even between devices to which multiple display data belong. Therefore, the potential differences between multiple devices and the lack of clear identification can lead to inaccurate coordination and inaccurate determined interactive operation information. To address these issues, this application utilizes unified spatiotemporal multimodal coding technology to map multiple devices (i.e., a first device and at least one second device) into a unified feature space according to mapping rules. For example, the device IDs, device locations, and display data of multiple devices are mapped to the coordinate system corresponding to the feature space. Under the same coordinate system, each device is expressed in the same way, that is, the display data, identity information, and location information of different devices are aligned at the same semantic level, providing a basic guarantee for the collaborative scheduling of multiple devices. Therefore, multiple devices can be monitored and scheduled based on this coordinate system.

[0045] It is worth noting that the method of using unified spatiotemporal multimodal coding technology to uniformly plan information for multiple devices requires the combination of data preprocessing, correction, mapping calculation and other means, such as preprocessing the display data of each device (such as pixel filling, format adjustment, redundancy information deletion, etc.) and correcting the expression format of device IDs. Compared with the simple extension of the existing methods, the method of this application embodiment is more comprehensive and accurate in unifying information for multiple devices.

[0046] Optionally, when generating target interactive operation information using the target model, the device ID, device location, display data, and task information of each device in the coordinate system corresponding to the feature space can be input into the target model. The target model calculates the interactive operation information simulated for the target device, which is specific to the target device.

[0047] Based on this, after obtaining the target interaction operation information, the operation coordinate information for the second device is determined according to the aforementioned mapping rules. That is, the target interaction operation information of the second device in the coordinate system corresponding to the feature space is mapped to the operation coordinate information of the second device in the actual space. Then, the operation coordinate information is sent to the second device so that the second device can simulate and execute the target interaction operation based on the operation coordinate information; that is, it performs the target interaction operation on the target operation object to execute and complete the target task.

[0048] It is worth noting that there may be cases where multiple target interaction operations are required to perform the target task. In this case, after performing one more target interaction operation, you can continue to perform steps 103 and 104 above until the target task is completed.

[0049] This application embodiment obtains target display data from the target device through the first device to determine the corresponding target interaction operation information for executing the target task. It then controls the target device to execute the target interaction operation based on this information, achieving flexible scheduling and task allocation for multiple devices. This process can control the target device to respond to target inputs without needing to obtain the target device's APP call interface, thus improving the flexibility of task execution and expanding its applicability. Furthermore, it avoids task failures caused by task conflicts or abnormal inter-device communication, effectively achieving convenient and efficient collaborative control of multiple devices to meet users' cross-device interaction needs. It also ensures load balancing across devices, preventing waste of device resources and impacting other tasks. Additionally, the first device has Agent capabilities deployed, thus enabling other devices without Agent capabilities to reuse the first device's Agent capabilities.

[0050] Of course, this application embodiment describes the target device for executing the target task as the second device, that is, it describes the process of executing the target task across devices. However, there is a case where the target device for executing the target task is the first device. That is, after obtaining the target input and determining the target task, the target device determined for executing the target task is the first device. In this case, the target interaction operation information corresponding to the execution of the target task can be directly determined based on the display data of the first device, and the target interaction operation can be executed to complete the target task.

[0051] For example, after determining the target task corresponding to the target input, a second device for performing the target task can be further determined. Optionally, after determining the target task, a target application for performing the target task can be determined based on the target task. For example, if the target task is to open NetEase Cloud Music, then the target application is NetEase Cloud Music; as another example, if the target task is to lower the volume of audio playback, then the application currently playing the audio is the target application. As yet another example, the target application for performing the target task can also be determined based on the directional information of the target input and / or the relative position information of each device.

[0052] After identifying the target application, the applications installed on each of the multiple devices are traversed to determine the third device with the target application installed. Optionally, the processor of each device can know the applications installed on that device and form a mapping table between devices and applications. Typically, one device corresponds to multiple applications. Of course, there are also cases where one device corresponds to only one application. This application embodiment does not limit this.

[0053] As an example, a mapping relationship library is formed based on mapping tables from multiple devices. This mapping relationship library can be stored on any one of the multiple devices, such as a device with high user operation frequency or a device that is convenient for users to operate. Of course, other devices can obtain and use the mapping relationship library at any time. As another example, the mapping relationship library can be stored on each of the multiple devices, so that each device can use the mapping relationship library at any time and in a timely manner. Compared to storing the mapping relationship library on each of the multiple devices, storing it on any one of the multiple devices avoids the occupation of storage resources; compared to storing it on any one of the multiple devices, storing it on each of the multiple devices avoids the problem of the mapping relationship library being unavailable and unusable due to communication failures between devices.

[0054] Based on this, it can be determined whether to store the mapping relationship library on any one of the multiple devices, or on each of the multiple devices, depending on actual needs. Actual needs may include the storage space for each device and / or the communication methods between the multiple devices.

[0055] After identifying the target application, the third device with the target application installed is determined based on the mapping database. It's worth noting that, alternatively, after identifying the target application, the icon interface of each device can be obtained in real time, allowing the identification of the third device with the target application installed based on the icon interface of each device.

[0056] Further, after identifying the third devices with the target application installed, a second device capable of executing the target task is determined from among the third devices with the target application installed. Optionally, the current display data and / or process information of each third device is first obtained to determine the application resource occupancy status of the target application installed on each third device based on the display data and / or process information. Then, based on the application resource occupancy status, the second device capable of executing the target task is determined. The application resource occupancy status includes idle and occupied. For example, if the display data of device A indicates that the display interface includes a dialog box corresponding to the target application, then the application resource occupancy status of device A is determined to be occupied; as another example, if the process information of device B indicates that the target application is not among the currently running applications, then the application resource occupancy status of device B is determined to be idle.

[0057] This application addresses the need to determine target interaction information based on target display data of the target device. Specifically, it employs visual recognition technology to render the target display data of the target device through the application of the first device, and then determines the location corresponding to the simulated target interaction. Considering that a single application can support rendering and displaying different application interfaces at the same time, this application further determines a second device capable of executing the target task by determining whether the target application is occupied. Specifically, when rendering the target display data, the target display data can be rendered and displayed on the display screen of the first device, or it can be rendered and displayed on a simulated virtual screen.

[0058] In this embodiment, a third device whose application resource occupancy status is idle is determined to be a second device capable of executing the target task. This avoids situations where the target task cannot be executed or needs to be delayed due to using a device whose application resource occupancy status is occupied, thus providing a guarantee for executing the target task.

[0059] It is worth noting that if the application resources of the target applications installed on the third device are all occupied, the application resource usage status of the target applications installed on each third device is monitored in real time, and the third device whose application resource usage status is switched to idle is identified as the second device; it is also possible to terminate the task executed by the target application on one of the third devices to identify that third device as the second device.

[0060] If the application resources of the target application installed on the third device are all idle, the second device can be determined based on the current load of the third device and / or the communication efficiency with the first device, prioritizing the third device with low load and high communication efficiency with the first device. Alternatively, if the application resources of the target application installed on the third device are all idle, all third devices can be used as second devices. After controlling each second device to execute the target task, if at least two second devices execute successfully, the second device with the shortest execution time is determined to have completed the target task; if only one second device executes successfully, the successfully executed second device is determined to have completed the target task.

[0061] In addition, there may be only one third device with the target application installed. In this case, if the application resource of the target application on the third device is occupied, the task executed by the target application can be terminated directly, and the third device can be identified as the second device.

[0062] Optionally, after obtaining the target display data, the first display data can be displayed in a first display area, and the target display data can be displayed in a second display area. The first display data is the content currently displayed on the first device. Preferably, the area of ​​the first display area is larger than the area of ​​the second display area to avoid significant changes in the display area corresponding to the content currently displayed on the first device, which could affect the display effect and ensure a better viewing experience for the user.

[0063] This embodiment controls the second display area to display target data, allowing the user to intuitively view the simulated execution of a target task using the first device. Therefore, the area of ​​the second display area is set smaller than that of the first display area to avoid affecting the first device's display of its own data. It should be noted that if the first device currently has no task in progress, meaning it does not need to display task data, and / or if it displays its corresponding screensaver data or a specific interface that has not been updated for a certain period, then the entire display area of ​​the first device can be controlled to display the target data for the user to view.

[0064] For example, in a vehicle scenario, the screen displaying the target data can be determined based on the location corresponding to the target input, including the central control display, the passenger-side display, and the rear seat display. For instance, if the driver in the driver's seat determines the target input based on the central control display, the second display area of ​​the central control display will then be controlled to show the target data. Similarly, if a rear passenger determines the target input based on the rear seat display, the screen displaying the target data can be determined based on the left and right spatial positions.

[0065] In practical implementation, after the target device executes the target interaction operation based on the target interaction operation information, there are cases where the target task is executed successfully and cases where the target task fails. Based on this, this application embodiment also provides another execution method, wherein the specific steps are as follows: Figure 2 The flowchart shown includes S201-S205.

[0066] S201, in response to the failure of the second device, obtain the information of the executed interactive operation.

[0067] S202, based on the interactive operation information already executed, control the fourth device to jump to the target interface.

[0068] S203, Obtain the fourth display data of the fourth device.

[0069] S204, determine the fourth interactive operation information corresponding to the target task based on the fourth display data.

[0070] S205, control the fourth device to perform the fourth interactive operation based on the fourth interactive operation information.

[0071] In the event that the target task fails to execute, in response to the failure of the target device, i.e. the second device, the information on the executed interactive operations is obtained. The information on the executed interactive operations includes the interactive operations executed based on the second device and the operation objects targeted by the interactive operations executed by the device.

[0072] After obtaining information about the executed interactive operations, the system controls the fourth device to navigate to the target interface based on this information. The fourth device is one of the third devices with the target application installed, and the target interface is the interface displayed when the second device fails. Correspondingly, controlling the fourth device to navigate to the target interface means controlling the fourth device to execute the target task.

[0073] Optionally, embodiments of this application illustrate the following two control methods for controlling the fourth device to jump to the target interface: First control method When controlling the fourth device to jump to the target interface, the operation coordinate mapping relationship between the second and fourth devices in the coordinate system corresponding to the feature space can be obtained first based on the target model. Based on the operation coordinate mapping relationship, the interactive operation information for the fourth device can be obtained.

[0074] Optionally, display data for each device can be recorded using timestamps. For example, for a target device, target display data and its corresponding display time, display data after performing a target interactive operation and its corresponding display time, etc., can be recorded. Of course, there can be multiple target interactive operations, and thus multiple sets of display data after performing each target interactive operation.

[0075] Upon receiving interactive operation information for the fourth device, multiple interactive operation information for the fourth device are determined based on the operation coordinate mapping relationship between the second and fourth devices, as well as each display data of the second device and its corresponding display time. Each display data of the second device is the display data generated when the second device performs an interactive operation based on the previously executed interactive operation information. The multiple interactive operation information for the fourth device ensures that the fourth device sequentially displays the display data generated when the second device performs an interactive operation based on the previously executed interactive operation information, and the display order is the same as the display order of the second device.

[0076] For multiple interactive operation information corresponding to multiple interactive operations of the fourth device, the multiple interactive operations can be combined with the display time to form a target interactive operation sequence in chronological order. This target interactive operation sequence is the interactive operation sequence corresponding to the executed interactive operation information.

[0077] Furthermore, the fourth device is controlled to execute a simulated target interactive operation sequence, that is, the fourth device is controlled to execute the multiple interactive operations in the order in which the second device executes the multiple interactive operations, so that the fourth device can accurately jump to the target interface.

[0078] The second control method When controlling the fourth device to jump to the target interface, the page path information of the target interface can be determined based on the interactive operation information already performed. This page path information includes the page address corresponding to the target interface, through which the target interface can be directly obtained. The page address can be a network address or a local address. Based on this, after determining the page path information of the target interface, the page path information is directly sent to the fourth device so that the fourth device jumps to the target interface, that is, the fourth device obtains and displays the target interface based on the page address included in the page path information.

[0079] For example, after the fourth device navigates to the target interface, the fourth display data of the fourth device is acquired, and the fourth interactive operation information corresponding to the target task is determined based on the fourth display data. Similarly, the fourth display data and the task information of the target task can be input into the target model so that the target model outputs the fourth interactive operation information for the fourth display data. Then, the fourth device is controlled to execute the fourth interactive operation based on the fourth interactive operation information to continue executing the target task.

[0080] It should be noted that there may be cases where the fourth device fails to execute. In this case, the corresponding steps S201-S205 above can be executed again until the target task is completed. No additional operation is required from the user. The device can be switched automatically without the user's awareness to ensure the successful execution of the target task.

[0081] Taking the objective task of playing "Happy Birthday" through the NetEase Cloud Music app as an example, after device A receives user input and determines the corresponding task, it iterates through its own applications and determines that it does not have NetEase Cloud Music downloaded. Therefore, it searches for devices that have NetEase Cloud Music downloaded, finding devices B and C. Based on the application resource usage status of devices B and C, it determines that device B will execute the user's task. Device A obtains the display data of device B so that the target model can obtain interactive operation information for device B based on the display data and the task. It then controls device B to execute interactive operations based on this information. These interactive operations can be multiple, such as double-clicking the NetEase Cloud Music app icon, clicking the search box to enter edit mode, entering "Happy Birthday" to search for a playlist, and clicking on "Happy Birthday" in the playlist to play.

[0082] Of course, the above interactive operations are executed sequentially, that is, the interactive operation information corresponding to each interactive operation is sent to device B one by one.

[0083] After device B performs the interactive operation of inputting the characters "Happy Birthday," it fails to display the searched playlist interface due to a network interruption. At this point, device A, based on the interactive operation information already performed by device B, controls device C to jump to the interface corresponding to the input of the characters "Happy Birthday," and controls device C to perform the search interactive operation. After device C responds to this interactive operation and displays the searched playlist, device A, based on the currently displayed interface of device C and the user's task, further determines the interactive operation information for device C, and controls device C to perform the interactive operation based on the interactive operation information for device C, in order to continue to complete the user's task, that is, to play Happy Birthday through NetEase Cloud Music.

[0084] In addition, considering that the embodiments of this application coordinate and schedule multiple devices, and that a single device solution can only make decisions based on local information, the execution method of the embodiments of this application can build a hierarchical global decision engine on the first device. The lower layer of the hierarchical global decision engine can retain the local real-time response capability of each device, and the upper layer of the hierarchical global decision engine performs multi-device scheduling through the target model in the Agent to realize task execution and resource optimization.

[0085] When optimizing resources, given that the embodiments of this application utilize unified spatiotemporal multimodal coding technology to map information of multiple devices to a unified feature space, the upper layer can use multi-objective optimization algorithms to solve problems such as multi-device resource load balancing, network latency minimization, and task completion time optimization during the multi-device scheduling process. This ensures that the determined target device, combined with the target interaction operation information, can efficiently and accurately complete the target task without causing a burden on the device load.

[0086] Furthermore, when the target device malfunctions or its state changes (such as the example above where the target task is to play the Happy Birthday song through the NetEase Cloud Music application), a multi-objective optimization algorithm is used to automatically reconstruct the target device and target interaction operation information to ensure the stability between multiple devices and the continuity of target task execution.

[0087] The hierarchical decision-making approach of the hierarchical global decision engine not only ensures the response speed to user input and the success rate of task execution, but also guarantees the comprehensiveness and rationality of multi-device scheduling, which is superior to single-device decision-making in all dimensions.

[0088] The second aspect of this application also provides an electronic device corresponding to the execution method. Since the principle by which the electronic device solves the problem is similar to the aforementioned execution method, the implementation of the electronic device can be found in the implementation of the method, and repeated details will not be elaborated further. A schematic diagram of the electronic device can be shown below. Figure 3 As shown, it includes at least a memory 301 and a processor 302, and the processor 302 of the electronic device runs a target application, which can be a pre-built intelligent agent that can parse at least one task based on input.

[0089] Optionally, the target application in this embodiment of the application is used to execute: Based on the target input, determine the target task corresponding to the target input; If the target device for performing the target task is determined to be the second device, the target display data of the target device is obtained; Based on the target display data, determine the target interactive operation information corresponding to the execution of the target task; Control the target device to perform target interactive operations based on the target interactive operation information.

[0090] In another example, when the target application determines the target interactive operation information corresponding to the execution of the target task based on the target display data, it performs the following steps: The target display data and the task information of the target task are input into the target model so that the target model outputs target interactive operation information for the target display data.

[0091] As another example, the target application is also used to perform: Based on the target task, determine the target application to perform the target task; Among the third devices that have the target application installed, a second device capable of performing the target task is determined.

[0092] In another example, when the target application determines a second device capable of performing the target task among the third devices on which the target application is installed, it performs the following steps: Determine the application resource usage status of the target application installed on each third device; Based on the resource occupancy status of each application, a second device capable of executing the target task is determined.

[0093] As another example, the target application is also used to perform: First display data is displayed in the first display area, and the target display data is displayed in the second display area.

[0094] In another example, when the target application controls the target device to perform a target interactive operation based on the target interactive operation information, the following steps are performed: Based on the determined target interaction operation information, construct the operation coordinate information for the second device; The operation coordinate information is sent to the second device so that the second device can simulate and execute the target interactive operation based on the operation coordinate information.

[0095] As another example, the target application is also used to perform: In response to the failure of the second device, obtain the information of the executed interactive operation; Based on the information of the executed interactive operations, control the fourth device to jump to the target interface; Obtain the fourth display data of the fourth device; Based on the fourth display data, determine the fourth interactive operation information corresponding to the execution of the target task; The fourth device is controlled to perform a fourth interactive operation based on the fourth interactive operation information; The target interface is the interface that appears when the second device fails to execute.

[0096] In another example, when the target application controls the fourth device to jump to the target interface based on the previously executed interaction information, it performs the following steps: Obtain the operational coordinate mapping relationship between the second device and the fourth device; Based on the operation coordinate mapping relationship and the executed interaction operation information, the fourth device is controlled to execute a simulated target interaction operation sequence so that the fourth device jumps to the target interface; The target interaction operation sequence is the interaction operation sequence corresponding to the executed interaction operation information.

[0097] In another example, when the target application controls the fourth device to jump to the target interface based on the already performed interaction information, it also performs the following steps: Based on the information of the executed interactive operations, determine the page path information of the target interface; The page path information is sent to the fourth device so that the fourth device can jump to the target interface.

[0098] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk. Optionally, in this embodiment, the processor executes the method steps described in the above embodiments according to the program code stored in the storage medium. Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, which will not be repeated here. Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed on a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be executed in a different order than those described here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any specific hardware and software combination.

[0099] Furthermore, although exemplary embodiments have been described herein, their scope includes any and all embodiments based on this application that have equivalent elements, modifications, omissions, combinations (e.g., schemes involving intersections of various embodiments), adaptations, or alterations. Elements in the claims will be interpreted broadly based on the language used in the claims and are not limited to the examples described in this specification or during the implementation of this application, which will be interpreted as non-exclusive. Therefore, this specification and examples are intended to be considered illustrative only, and the true scope and spirit are indicated by the following claims and the full scope of their equivalents.

[0100] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of them) can be used in combination with each other. Other embodiments may be used by those skilled in the art upon reading the above description. Furthermore, in the above detailed description, various features may be grouped together to simplify the application. This should not be construed as an intention that a disclosed feature not claimed is necessary for any claim. Rather, the subject matter of this application may be less than all the features of a particular disclosed embodiment. Thus, the following claims are incorporated herein by reference as examples or embodiments, wherein each claim is an independent, separate embodiment, and these embodiments are contemplated as being possible in various combinations or arrangements. The scope of this application should be determined by reference to the appended claims and the full scope of their equivalents.

[0101] The foregoing has described in detail several embodiments of this application, but this application is not limited to these specific embodiments. Those skilled in the art can make various variations and modifications based on the concept of this application, and all such variations and modifications should fall within the scope of protection claimed in this application.

Claims

1. An execution method applied to a first device, comprising: Based on the target input, determine the target task corresponding to the target input; If the target device for performing the target task is determined to be the second device, the target display data of the target device is obtained; Based on the target display data, determine the target interactive operation information corresponding to the execution of the target task; Control the target device to perform target interactive operations based on the target interactive operation information.

2. The execution method according to claim 1, wherein determining the target interactive operation information corresponding to executing the target task based on the target display data includes: The target display data and the task information of the target task are input into the target model so that the target model outputs target interactive operation information for the target display data.

3. The execution method according to claim 1 further includes: Based on the target task, determine the target application to perform the target task; Among the third devices that have the target application installed, a second device capable of performing the target task is determined.

4. The execution method according to claim 3, wherein determining a second device capable of executing the target task among the third devices on which the target application is installed includes: Determine the application resource usage status of the target application installed on each third device; Based on the resource occupancy status of each application, a second device capable of executing the target task is determined.

5. The execution method according to claim 1, further comprising: First display data is displayed in the first display area, and the target display data is displayed in the second display area.

6. The execution method according to any one of claims 1-5, wherein the control target device performs a target interaction operation based on the target interaction operation information, comprising: Based on the determined target interaction operation information, construct the operation coordinate information for the second device; The operation coordinate information is sent to the second device so that the second device can simulate and execute the target interactive operation based on the operation coordinate information.

7. The execution method according to claim 1, further comprising: In response to the failure of the second device, obtain the information of the executed interactive operation; Based on the information of the executed interactive operations, control the fourth device to jump to the target interface; Obtain the fourth display data of the fourth device; Based on the fourth display data, determine the fourth interactive operation information corresponding to the execution of the target task; The fourth device is controlled to perform a fourth interactive operation based on the fourth interactive operation information; The target interface is the interface that appears when the second device fails to execute.

8. The execution method according to claim 7, wherein controlling the fourth device to jump to the target interface based on the executed interactive operation information includes: Obtain the operational coordinate mapping relationship between the second device and the fourth device; Based on the operation coordinate mapping relationship and the executed interaction operation information, the fourth device is controlled to execute a simulated target interaction operation sequence so that the fourth device jumps to the target interface; The target interaction operation sequence is the interaction operation sequence corresponding to the executed interaction operation information.

9. The execution method according to claim 7, wherein controlling the fourth device to jump to the target interface based on the executed interactive operation information further includes: Based on the information of the executed interactive operations, determine the page path information of the target interface; The page path information is sent to the fourth device so that the fourth device can jump to the target interface.

10. An electronic device running a target application, the target application being capable of parsing at least one task based on input, the target application being used to perform: Based on the target input, determine the target task corresponding to the target input; If the target device for performing the target task is determined to be the second device, the target display data of the target device is obtained; Based on the target display data, determine the target interactive operation information corresponding to the execution of the target task; Control the target device to perform target interactive operations based on the target interactive operation information.